EMC gets fat and flashy with Greenplum appliances

Take that, Teradata, Exadata, Netezza


For one brief shining moment, when it bought Data General a zillion years ago, EMC was a server maker, and last year's acquisition of Greenplum makes it a server vendor (of sorts) once again. More like a data analytics systems integrator, but let's not split hairs.

The original Data Computing Appliance, or DCA, was based on Sun Fire x64-based servers from Sun Microsystems. But Luke Lonergan, chief technology officer at EMC's Greenplum unit, tells El Reg that it is giving customers a choice of OEMed two-socket servers from Dell, Hewlett-Packard, and Huawei as the basis of the mainstream DCA boxes the company is now selling. Those server options are also available for two new variants of the DCA machine, equipped with either fat disks for extra capacity or solid state disks for faster data access. These are the High Capacity DCA and the High Performance DCA, using the EMC nomenclature.

You might be thinking that Greenplum could just use Vblock configurations based on the "California" Unified Computing System, given that EMC, Cisco, and VMware are all buddy-buddy in the Virtual Computing Environment Company partnership formerly known as Acadia. The Vblocks take Cisco's converged switching and servers, EMC's storage, and VMware's server virtualization and wrap them together in three different configurations that are supposed to show up preassembled and ready to load server instances on.

While this is great for server virtualization, a Vblock is not necessarily the best platform on which to run a clustered database and data analytics workloads. And so Greenplum is for the moment sticking to local disk or flash storage on server nodes in the DCA clusters and using 10 Gigabit Ethernet links between server nodes so they can munch each other's data and, if you believe all the talk, turn it into information.

EMC Greenplum Appliance

EMC's Greenplum data analytics appliance

The current DCA server nodes are kosher 2U rack-mounted boxes with two sockets and sporting Intel's six-core "Westmere-EP" Xeon 5600 processors. The base DCA box has 600GB SATA disks, and it was suitable for a wide range of workloads. The DCA has two master servers controlling the cluster and 16 segment servers for storing databases, with a total of 192 cores for chewing through data.

These segment server nodes have an aggregate of 768GB of main memory and one disk drive per core. The DCA has 36 TB of uncompressed usable capacity for a data warehouse and 144 TB with compression turned on; it can scan data at 24GB/sec and has a data load rate of 10TB/hour. The DCA can scale up to six racks in a single system, and you can buy in quarter-rack increments, according to Lonergan.

With the High Capacity DCA, Greenplum is swapping out the 600GB drives and replacing them with 2TB SATA disks spinning at 7200 RPM. Some DCA customers were looking to chew on larger data sets, and the same set of server nodes, Greenplum can now offer 124TB of usable capacity (496TB compressed). Nothing is free of course, and there is a performance hit with the larger capacity.

"This is mostly a bandwidth game," explains Lonergan. "The outer platter of a SATA drive is spinning more than enough to saturate a SATA controller."

When you do the numbers, the High Capacity DCA has a scan rate of 16GB/sec, a 33 per cent hit compared to the regular DCA, and the data load rate is cut in half, to 4.8TB/hour. You can scale this High Capacity DCA up to six racks, just like the plain DCA.

If you really want speed, then Greenplum has the new High Performance DCA, which puts 24 solid state disks into a server chassis that has four half-node, two-socket Xeon 5600s servers in the chassis. Six SSDs are allocated to each node, which has only one six-core Xeon 5600 plugged in. This yields the same one drive, one core ratio of the original DCA. The SSDs link into the server nodes over 6Gb/sec SAS channels.

The High Performance DCA has two master servers, and a total of 14 additional rack servers yielding 56 segment servers. All told, they have 1,344GB of aggregate main memory, 336 cores, and 336 SSDs. The usable capacity of this flashy DCA is actually 44TB (176TB compressed), which is more than the plain vanilla DCA. The High Performance DCA has a database scan rate of 72GB/sec - three times the regular DCA - and a data load rate of 20TB/hour - twice the vanilla cluster. This version of the machine only comes in a single rack; you can't scale it any further.

Lonergan tells El Reg that EMC doesn't expect the flashy DCA to be a big seller in terms of numbers of racks, mainly because SSDs are still twice as expensive as disks. He did not provider pricing, but Lonergan said that the target for the High Performance DCA was to yield four times the performance for twice the money. "We did much better than that," he says. (Scan rate was only 3X and load rate was only 2X the performance, as noted above. But neither of these is application performance, which is what Lonergan was referring to.)

In addition to the new iron, EMC is rolling out Greenplum Database 4.1, which the company says has better integration with Hadoop clusters as well as an extended range of analytical functions.

SAS Institute was also part of EMC's Greenplum announcements. SAS is cooking up a parallel cluster version of its analytics tools, and has run some tests putting the SAS tools on top of a DCA appliance. The SAS code is currently limited to the size of main memory of a single server, according to Lonergan, and that explains in part why SAS was so popular on Sun Microsystems Unix boxes for so many years. But big analytical jobs take a long time. In one benchmark that SAS and EMC have run, it took 27 hours to crunch some data using the SAS tools on a big (and unspecified) SMP box. On a 32-node vanilla DCA, this same analysis took 50 seconds.

SAS says the Greenplum version of its data analytics suite will be available in the fourth quarter of this year. Looks like EMC will be knocking on the doors of Sun/Oracle shops that have SAS tools running. ®

Similar topics


Other stories you might like

  • Lonestar plans to put datacenters in the Moon's lava tubes
    How? Founder tells The Register 'Robots… lots of robots'

    Imagine a future where racks of computer servers hum quietly in darkness below the surface of the Moon.

    Here is where some of the most important data is stored, to be left untouched for as long as can be. The idea sounds like something from science-fiction, but one startup that recently emerged from stealth is trying to turn it into a reality. Lonestar Data Holdings has a unique mission unlike any other cloud provider: to build datacenters on the Moon backing up the world's data.

    "It's inconceivable to me that we are keeping our most precious assets, our knowledge and our data, on Earth, where we're setting off bombs and burning things," Christopher Stott, founder and CEO of Lonestar, told The Register. "We need to put our assets in place off our planet, where we can keep it safe."

    Continue reading
  • Conti: Russian-backed rulers of Costa Rican hacktocracy?
    Also, Chinese IT admin jailed for deleting database, and the NSA promises no more backdoors

    In brief The notorious Russian-aligned Conti ransomware gang has upped the ante in its attack against Costa Rica, threatening to overthrow the government if it doesn't pay a $20 million ransom. 

    Costa Rican president Rodrigo Chaves said that the country is effectively at war with the gang, who in April infiltrated the government's computer systems, gaining a foothold in 27 agencies at various government levels. The US State Department has offered a $15 million reward leading to the capture of Conti's leaders, who it said have made more than $150 million from 1,000+ victims.

    Conti claimed this week that it has insiders in the Costa Rican government, the AP reported, warning that "We are determined to overthrow the government by means of a cyber attack, we have already shown you all the strength and power, you have introduced an emergency." 

    Continue reading
  • China-linked Twisted Panda caught spying on Russian defense R&D
    Because Beijing isn't above covert ops to accomplish its five-year goals

    Chinese cyberspies targeted two Russian defense institutes and possibly another research facility in Belarus, according to Check Point Research.

    The new campaign, dubbed Twisted Panda, is part of a larger, state-sponsored espionage operation that has been ongoing for several months, if not nearly a year, according to the security shop.

    In a technical analysis, the researchers detail the various malicious stages and payloads of the campaign that used sanctions-related phishing emails to attack Russian entities, which are part of the state-owned defense conglomerate Rostec Corporation.

    Continue reading
  • FTC signals crackdown on ed-tech harvesting kid's data
    Trade watchdog, and President, reminds that COPPA can ban ya

    The US Federal Trade Commission on Thursday said it intends to take action against educational technology companies that unlawfully collect data from children using online educational services.

    In a policy statement, the agency said, "Children should not have to needlessly hand over their data and forfeit their privacy in order to do their schoolwork or participate in remote learning, especially given the wide and increasing adoption of ed tech tools."

    The agency says it will scrutinize educational service providers to ensure that they are meeting their legal obligations under COPPA, the Children's Online Privacy Protection Act.

    Continue reading
  • Mysterious firm seeks to buy majority stake in Arm China
    Chinese joint venture's ousted CEO tries to hang on - who will get control?

    The saga surrounding Arm's joint venture in China just took another intriguing turn: a mysterious firm named Lotcap Group claims it has signed a letter of intent to buy a 51 percent stake in Arm China from existing investors in the country.

    In a Chinese-language press release posted Wednesday, Lotcap said it has formed a subsidiary, Lotcap Fund, to buy a majority stake in the joint venture. However, reporting by one newspaper suggested that the investment firm still needs the approval of one significant investor to gain 51 percent control of Arm China.

    The development comes a couple of weeks after Arm China said that its former CEO, Allen Wu, was refusing once again to step down from his position, despite the company's board voting in late April to replace Wu with two co-chief executives. SoftBank Group, which owns 49 percent of the Chinese venture, has been trying to unentangle Arm China from Wu as the Japanese tech investment giant plans for an initial public offering of the British parent company.

    Continue reading
  • SmartNICs power the cloud, are enterprise datacenters next?
    High pricing, lack of software make smartNICs a tough sell, despite offload potential

    SmartNICs have the potential to accelerate enterprise workloads, but don't expect to see them bring hyperscale-class efficiency to most datacenters anytime soon, ZK Research's Zeus Kerravala told The Register.

    SmartNICs are widely deployed in cloud and hyperscale datacenters as a means to offload input/output (I/O) intensive network, security, and storage operations from the CPU, freeing it up to run revenue generating tenant workloads. Some more advanced chips even offload the hypervisor to further separate the infrastructure management layer from the rest of the server.

    Despite relative success in the cloud and a flurry of innovation from the still-limited vendor SmartNIC ecosystem, including Mellanox (Nvidia), Intel, Marvell, and Xilinx (AMD), Kerravala argues that the use cases for enterprise datacenters are unlikely to resemble those of the major hyperscalers, at least in the near term.

    Continue reading

Biting the hand that feeds IT © 1998–2022