XPS still around?
Posted in 2010
Back in the 90's XPS touted near linear scalability. Now fast forward to 2010 and what we see is NoSQL in terms of the DW space. Putting the cost of the underlying database software aside, I wonder if anyone (HINT IBM!) has tried to do a benchmark of XPS using 8 core 1U boxes w 8TB disk per box. (Think 4 x 2TB SATA drives) This hardware configuration is COTS and cost effective w ToR switching. Meaning in 1 rack you can build a cluster of 24+ nodes w 8TB and 32GB Memory per node, w 10GBe Nics and a ToR 10GBe switch. Hardware wise, the cost would be approx 5K per box (120K total), and I think that you can get a 24 port ToR switch for 2K a port or ~50K (just to round up the numbers.) So for under 200K in hardware, running your favorite flavor of Linux (OpenSuSe, Centos, Ubuntu...) Sans the price of XPS, you'd have a test bed that would be similar to a small/medium cluster for Hadoop/HBase. One would think that IBM, which could easily fund this... could put together the POC in one of their campus offices. (Actually it would be cheaper since IBM could either loan the hardware from a pool, or provide the hardware on a 6 month lease which should be long enough to build, test, tune and then retest the configuration. Also note that it could be paid for out of *blue* dollars ...) The reason one would want to do this is because there's a myth about NoSQL. Most those who propagate the NoSQL mantra only have exposure to MySQL and maybe minimal postgress too. Looking back at XPS, the largest publicly talked about DW was 27TB running at First Union (Pre Wachovia) in their marketing department. Back then... 27TB was a lot. Assuming that data density grows in proportion to Moore's law, doubling ever 18 months that would be roughly 27PB of data. (If someone has a better calculation for estimate of data growth, please post. My assumption was that DW was done mid/ early 90's so it doubled 10 times or 1024 times. Even still, on the low side it would be 1+PBs by today's standards.) So, one has to wonder why IBM hasn't thought about doing something like this at scale? 8TB per node, 24 nodes == 192TB of raw disk. Just some food for thought. But hey! What do I know? I guess the brilliant marketing and R&D wonks within IBM's IM group are too busy trying to offer their own 'special branch of hadoop' which runs on their 32bit Linux offering. (I guess they too don't want to threaten IBM's sacred cow, ala DB2 as their DW engine of choice either....) ;-P -G