Re: 11.50.xC4 Compression with workgroup edition?
Posted in 2009
A question about whether the storage-compression feature in IDS 11.50.xC4 is available with Workgroup Edition turned into a vendor/architecture debate rather than a support thread. Oracle's Mark Townsend outlines Oracle's free index and bulk compression versus its paid Advanced (block-level) compression; IBM's Serge Rielau argues IBM/Informix's table-level global dictionary suits slowly-changing data while block dictionaries suit skewed/time-varying data handled via range partitioning. Ian Gumby counters on TCO, claiming the add-on price can't be justified against cheap disk/SSD for typical SMB-sized systems. No answer to the original Workgroup Edition licensing question is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Licensing & Editions, Migration, Import/Export & Data Conversion
> Oracle has a paid and a free version. Of course they're completely > different and as you can imagine the free is much more limited. Well - things play out a little differently in Oracle Firstly - Oracle has always had free index compression in it's Enterprise Edition. So for big OLTP environments (and also for index rich DW environments) the storage required for indexes can be significantly reduced. Then we have what we call bulk compression, also free in Enterprise Edition - if you load or move a lot of data at once, then we will compress the repeating data values. Typically this is used in DW and archiving environments, where the data, once compressed, does not change rapidly. The compression "cost" is also amortized across the load operation Both of these have been in the product for many years. Last year, we introduced Advanced Compression as an added option - it's basically continuous compression for rapidly changing data (i.e the data is initially stored uncompressed, but once a threshold in a block is reached, the data in the block is compressed, the compression dictionary is held at the block level, so the compression ratio doesn't degrade as new data values come in). It's cool because you can just alter existing tables under an OLTP environment, and over time they will compress themselves without any need to unload and load, or re-org, or do any of those things that are difficult to do on very large tables. Both use the same compression algorithms, and both use the same block based compression dictionary, which we think is a Good Thing (TM). They are however designed to solve different use cases - I wouldn't describe either as limited.
An easy way to distinguish between IBM (DB2 and IDS) compression and Oracle compression is to look at e.g. an employee table with, say names and addresses and other junk (like salaries ;-) Block compressions wins, as Mark points out, in fast changing data. Global compression wins in slow changing data. Now, how is the employee table clustered? I take a guess heer and say employeeid (you rprimary key). As new employees get hired the table grows and the employee id increases. If you have a company of say 1000 employees located in say Madrid. Chances are you will see a lot of overlap in the addresses and also in the names and salaries. Keep in mind that IBM compression is oblivious to column boundaries! Now as your company rows or people get hired fired, what are the chances that the data to compress will radically change? Pretty slim. Until you add new locations for your company most people will live around Madrid. Until you move out of Spain their street names and names will me similar and so on. So for a that sort of table a global dictionary provides minimal overhead in size while paying for it with soem amount of skew over time. Now lets take a look at block based. Here the frequent column values get stored into each block. let's say we have a 100 rows per block (pick a number) What are the chances that out of my 100 employees several live in Madrid? pretty high. What are teh chances that several live in teh same street? Pretty low. Same Salary? Even lower. Now, whenever there are dups what are the chances that they are NOT dups in the other blocks as well? Not that high because the table is not clustered around these frequent values. Any clustering is random. So you end up with repeating dictionaries with the very similar content reducing the benefits of compression by increasing the size of the dictionary. So, where does block based win? Skew over time? Not many bikinis sold in winter. Few x-mas trees in summer. Did we say changes over TIME? Oh yes.. of course that's where range partitioning comes in. These tables that change over time are typically range partitioned. And each partition has it's own dictionary. So you get the best of both worlds. A compression that adapts to changing values and isn't wasteful. Gumby sent me a private note (not sure why) about $/TB. Apparently his believe is that $ is made up out of purchase price. That of course is not true once you leave your basement office. Storage comes with cooling and maintenance costs, floor space rental etc. We have been in the compression business on the DB2 side for a bit now and it does make a lot of financial sense to customers as proven by them buying it and IBM investing even more into the technology. Cheers Serge -- Serge Rielau SQL Architect DB2 for LUW IBM Toronto Lab
On May 22, 8:13 am, Serge Rielau <srie...@ca.ibm.com> wrote: > So, where does block based win? Skew over time? Not many bikinis sold in > Gumby sent me a private note (not sure why) about $/TB. > Apparently his believe is that $ is made up out of purchase price. > That of course is not true once you leave your basement office. > Storage comes with cooling and maintenance costs, floor space rental etc. > Serge, I think I sent you a private e-mail because I was responding to an e- mail you sent. But actually where I live, we don't have a basement. And considering I have a rack siting next to me in my home office... I'm familiar with the amount of heat generated. And since you live in 'the lab' and rarely spend time on a customer presence lets look at the real world economies of scale.... You want 'slow drives' you have 2.5" SATA drives that are pretty efficient and are only relatively slow when talking about today's SCSI/ SAS drives or SSD drives. So in a 2U high rack you can pack in several TB of Raid storage and not really give off a lot of heat. Again relative to the amount of storage. But since you want to talk about heat, how much heat is given off by the CPU? Outch! Yeah that's where the real cooling issues have to happen. Don't see a lot of huge heat sinks on hard drives do we? SSDs while more expensive require less power and generate even less heat. The problem is that SSDs will degrade overtime and have a higher cost over comperable spinning disks. However, for the 120K, you can buy a lot of SSD storage. Now of course lets take a break and point out that IDS isn't positioned as a DW machine. So what's the average installation of IDS and how many systems are > 1TB in size? C'mon Serge, how many warehouse/distribution servers running EXE have 1+TB of data?, POS store systems?, telephony/cell switch systems? SMBs? Maybe Stuart's porn collection that he hosts for a 'friend' is > 1 TB, but in a real OLTP database? Not so much. And you can buy single 1+TB hard drives in 3.5" form factors. And thats the point about storage. Disks are cheap. SSDs will drop in price and within a couple of years, HP will solve the manufacturing issues and bring memistors to market. (Instant on memory in extremely high density). The cost of the energy used is also much less. Also Serge, there are ways of designing racks and the glass room such that you can reduce the costs of cooling and even let your glass room run hotter. So please spare me on the 'compression is green' because it isn't. > We have been in the compression business on the DB2 side for a bit now > and it does make a lot of financial sense to customers as proven by them > buying it and IBM investing even more into the technology. > Well you know those mainframe customers. They'll buy anything until their bean counters catch up and tell them to migrate off the pig iron and down to a distributed (LUW) box. That's where Oracle and Microsoft step in with their solutions. ;-) Of course I know of a lot of shops that are still running 5 1/4" form factor drives in their disk arrays so what does that tell you about their push for efficiency. Heck its probably cheaper for them to swap out their disk arrays with smaller more efficient disk subsystems... But back to Serge's employee table. I guess with respect to IBM, they don't want to do any compression to the table. I mean with all of the North American RIFs and EU RIFs then the India and Brazil hiring... way too much change to make compression worth while. :-) Rumor has it that there's another 5K coming in June.... But hey! What do I know? Its not like I've actually priced machines out and consider the TCO for customers, including foot prints... ;-) -G PS Serge, you really should quit while you're behind.
On May 22, 8:13 am, Serge Rielau <srie...@ca.ibm.com> wrote: > Gumby sent me a private note (not sure why) about $/TB. > Apparently his believe is that $ is made up out of purchase price. > That of course is not true once you leave your basement office. > Storage comes with cooling and maintenance costs, floor space rental etc. > > We have been in the compression business on the DB2 side for a bit now > and it does make a lot of financial sense to customers as proven by them > buying it and IBM investing even more into the technology. > > Cheers > Serge > -- > Serge Rielau > SQL Architect DB2 for LUW > IBM Toronto Lab Gee Serge, I wanted to add a couple of things about space, energy and efficiency... First here's a new article about IBM's introduction of SSDs. http://www.theregister.co.uk/2009/05/22/ibm_power_ssds/ The plus... you can get a 2.5" form factored SSD w 69GB formatted. (This is a 120GB drive but IBM is doing some wear leveling algorithm so the drive lasts longer). The negative... IBM is offering these drives at a very, very steep price. The point is that they use less energy and dissipate less heat. But if you want to price out a 2.5" SAS drive from IBM (IBM - Hard drive - 73 GB - hot-swap - 2.5" - SAS - 15000 rpm) You can get one online for $277 (USD) plus tax/shipping/etc. Now if we look at a complete system... http://www.supermicro.com/products/system/1U/1025/SYS-1025W-UR.cfm This doesn't have the Xeon 5500 series but I bet you can call your SuperMicro rep and price one out. 8 X 73 GB = 584 GB raw, so what's that in a RAID 10 configuration? ~290 GB? The cost for 8 SAS drives? Roughly $2280. You don't want SAS? IBM's 300GB 2.5" drive is ~$500 a drive so 4 drives in RAID 10 is what? 600GB of storage? Now I didn't price the hardware out, other than for the drives. But the point is that You can build this box for under 10 grand. I think this machine is limited to 32 GB of memory, so you may want to look at other options.... Linux? OpenSuSE is free. Informix? If you don't split the box in to two separate virtual machines, IDS for 8 cores is pretty damn expensive. Splitting it, Workgroup edition is what? $25K Now if you want the 'compression feature', on a split box that what an addition 60K? (I thought someone said 120K for 8 cores...) [Note: my memory is getting bad in my old age...] So Serge, here's a pizza box that can pretty much be put in to a warehouse, a store, most SMB companies and run a lot of their business requirements. This box could in fact be used for the *majority* of most customer's Database application needs. 60K for compression? Not a real price winner in today's technology. Heat and energy use? Again, not a real issue. 2 sockets each at what 80-90 watt output taking the place of much larger, older equipment? This box probably has as much horsepower, if not more, than a single frame of an SP2 did circa '95 timeframe. Want to talk about shrinking the footprint of you machine room? Since Serge wants to live in his la-la land of Markham labs, lets get back to reality. Most customers are not going to want to pay the surcharge for this feature. And that was the point I was trying to get to. IBM says that they've been able to get up to a 50% reduction in data storage savings. Of course if you put the data you want to retrieve in the index so you don't hit the underlying table... you'll have less savings since your index will get fatter and its not being compressed. Sorry, I'm not bashing IBM. I think that there is some value to this feature, just not at the current price point. Moore's law is working against you. -G PS. Unlike Serge, I am choosing the hardware for this example because it does fit an SMB customer's needs. 1 1U box split to be database server, webserver (w mail and dns, LDAP and SAMBA server thrown in).
Ian Michael Gumby wrote: > Blah, blah, blah, blah ... > > -G > > PS. Unlike Serge, I am choosing the hardware for this example because > it does fit an SMB customer's needs. 1 1U box split to be database > server, webserver (w mail and dns, LDAP and SAMBA server thrown in). So maybe, just MAYBE, IBM isn't targeting the SMB user with this feature? But hey, what do I know? It's not like I'm not perving over pictures of Lindsay Lohan in a three-way with Britney and Sharon Stone! -- Cheers, Obnoxio The Clown http://obotheclown.blogspot.com -- This message has been scanned for viruses and dangerous content by OpenProtect(http://www.openprotect.com), and is believed to be clean.
"Ian Michael Gumby" <im_gumby@hotmail.com> wrote in message news:673a4c68-687a-4317-a259-bde3d14b15a7@o20g2000vbh.googlegroups.com... >> Now of course lets take a break and point out that IDS isn't positioned as a DW machine. Er, funny you should mention that: The week before last Informix was being marketed as the "embedded database of choice". Last week it was the choice for OLTP This week .... http://www-01.ibm.com/software/data/informix/warehouse/