11.50.xC4 Compression with workgroup edition?
Posted in 2009
A user asked whether the new 11.50.xC4 data compression can be used with Workgroup Edition, since OAT only warns vaguely that extra licensing may be needed. The answer given: compression is a chargeable add-on (Informix Storage Optimization Feature) for Enterprise Edition only, listed around US$15.5K per 100 PVUs, so it isn't available to Workgroup users. The rest of the thread debates its worth: it's pattern-based (not just char data), rows stay compressed in the buffer cache, indexes aren't compressed, and gains depend heavily on data and access patterns; IBM's whitepaper and the compression estimator were cited for figures.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Stored Procedures & SPL, Licensing & Editions, Third-Party Tools & Monitoring
I can't seem to get any information on whether the new 11.50.xC4 compression feature can be used with the Workgroup edition. In OAT I see the ominous warning: "Enabling compression might require additional license agreements. Ensure that these agreements are in place before continuing." Well, wouldn't it be nice if instead the message could point us to the location of where we could out this information, or even better, simply state whether or not this particular IDS instance is licensed for this feature? Why build new features if we can't easily find out whether we are entitled to them? I've looked at the recent licensing agreements and see no mention of it, so why should OAT display such a message?
"Ron MacNeil" <macneil.ron@gmail.com> wrote in message news:mailman.688.1242824883.1831.informix-list@iiug.org... >> I can't seem to get any information on whether the new 11.50.xC4 >> compression feature can be used with the Workgroup edition. In OAT I see >> the ominous warning: "Enabling compression might require additional >> license agreements. Ensure that these agreements are in place before >> continuing." >> Well, wouldn't it be nice if instead the message could point us to the >> location of where we could out this information, or even better, simply >> state whether or not this particular IDS instance is licensed for this >> feature? Why build new features if we can't easily find out whether we >> are entitled to them? >> I've looked at the recent licensing agreements and see no mention of it, >> so why should OAT display such a message? Yes, amidst the fanfares about the compression feature IBM has been strangely reticent to mention this, hasn't it? Compression is a chargeable option available only for Enterprise Edition (Informix Storage Optimization Feature) and its list cost is '9,978 (about US$15,500 at today's exchange rate) per100 PVUs .. .the same as for DB2 I'm told. The high cost, and the availability for Enterprise only, will limit the appeal of this feature to all but the largest installations only.
On May 20, 9:45 am, "Captain Pedantic" <theharlequi...@hotmail.com> wrote: > "Ron MacNeil" <macneil....@gmail.com> wrote in message > > news:mailman.688.1242824883.1831.informix-list@iiug.org... > > >> I can't seem to get any information on whether the new 11.50.xC4 > >> compression feature can be used with the Workgroup edition. In OAT I see > >> the ominous warning: "Enabling compression might require additional > >> license agreements. Ensure that these agreements are in place before > >> continuing." > >> Well, wouldn't it be nice if instead the message could point us to the > >> location of where we could out this information, or even better, simply > >> state whether or not this particular IDS instance is licensed for this > >> feature? Why build new features if we can't easily find out whether we > >> are entitled to them? > >> I've looked at the recent licensing agreements and see no mention of it, > >> so why should OAT display such a message? > > Yes, amidst the fanfares about the compression feature IBM has been > strangely reticent to mention this, hasn't it? > > Compression is a chargeable option available only for Enterprise Edition > (Informix Storage Optimization Feature) and its list cost is £9,978 (about > US$15,500 at today's exchange rate) per100 PVUs .. .the same as for DB2 I'm > told. > > The high cost, and the availability for Enterprise only, will limit the > appeal of this feature to all but the largest installations only. Ok, Color me silly, but what's the real value? You have SATA, SAS and now SSDs. With respect to SATA, drives are slow(er) than SAS but have a high density so you have a long time to recognize any ROI from disk compression. With respect to SAS, the drives are faster than SAS but have a lower density yet still not that expensive, relatively speaking. SSDs... fastest 'disk' but has the highest cost per GB of storage. There are some other drawbacks too, but thats a different story. Here I could see some potential value for disk compression except for the fact that there is a cost associated with the compress/decompress of the data. Kind of defeats the purpose of the SSD. So what am I missing? IMHO it seems like this feature is like tits on a boar. ;-) $15.5K buys a lot of disk and the array chassis.
The benefits of compression are twofold, Gumby. One of course is the storage savings which you have treated reasonably below. The second benefit is processing speed. Because IDS compresses at the row level and maintains the data in compressed form in the buffer cache, rows are only decompressed when actually accessed from the cache. The gains from having more data in the same amount of memory and of moving more data in a single IO are often greater than the cost of decompressing the rows in order to access them. Many systems actually run faster with compressed data according to IBM's testing and the experiences of early adopters. There is a breakpoint around the number of times the average row is accessed once it is read into the cache and the percentage of rows on a compressed page that are accessed once the page is in memory that determines whether performance improves or degrades due to compression. No hard numbers that I've seen, but for most OLTP systems you know what tables have rows that are accessed many times and what tables contain data that is written once and read seldom. You also know what tables exhibit good locality of active rows. Larger tables with higher locality and relatively lower access frequency will tend to see improvement in processing speed for larger reports and more active systems. Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, May 20, 2009 at 12:51 PM, Ian Michael Gumby <im_gumby@hotmail.com>wrote: > On May 20, 9:45 am, "Captain Pedantic" <theharlequi...@hotmail.com> > wrote: > > "Ron MacNeil" <macneil....@gmail.com> wrote in message > > > > news:mailman.688.1242824883.1831.informix-list@iiug.org... > > > > >> I can't seem to get any information on whether the new 11.50.xC4 > > >> compression feature can be used with the Workgroup edition. In OAT I > see > > >> the ominous warning: "Enabling compression might require additional > > >> license agreements. Ensure that these agreements are in place before > > >> continuing." > > >> Well, wouldn't it be nice if instead the message could point us to the > > >> location of where we could out this information, or even better, > simply > > >> state whether or not this particular IDS instance is licensed for this > > >> feature? Why build new features if we can't easily find out whether > we > > >> are entitled to them? > > >> I've looked at the recent licensing agreements and see no mention of > it, > > >> so why should OAT display such a message? > > > > Yes, amidst the fanfares about the compression feature IBM has been > > strangely reticent to mention this, hasn't it? > > > > Compression is a chargeable option available only for Enterprise Edition > > (Informix Storage Optimization Feature) and its list cost is £9,978 > (about > > US$15,500 at today's exchange rate) per100 PVUs .. .the same as for DB2 > I'm > > told. > > > > The high cost, and the availability for Enterprise only, will limit the > > appeal of this feature to all but the largest installations only. > > Ok, > Color me silly, but what's the real value? > > You have SATA, SAS and now SSDs. > > With respect to SATA, drives are slow(er) than SAS but have a high > density so you have a long time to recognize any ROI from disk > compression. > With respect to SAS, the drives are faster than SAS but have a lower > density yet still not that expensive, relatively speaking. > SSDs... fastest 'disk' but has the highest cost per GB of storage. > There are some other drawbacks too, but thats a different story. Here > I could see some potential value for disk compression except for the > fact that there is a cost associated with the compress/decompress of > the data. Kind of defeats the purpose of the SSD. > > So what am I missing? IMHO it seems like this feature is like tits on > a boar. ;-) > > $15.5K buys a lot of disk and the array chassis. > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list >
Gumby, I agree! I was excited about upgrading to 11.5 for the compression. We run on a SAN at a data center and the costs over many systems could be of value! If they decide to charge for this it reduces the ROI and no longer worth the effort. Another nail in the coffin! Sounds like IBM is trying to be like Slybase....all the bells and whistles are "premium" features and they charge extra....Of course the "phantom spids" come free! peace Ian Michael Gumby <im_gumby@hotmail .com> To Sent by: informix-list@iiug.org informix-list-bou cc nces@iiug.org Subject Re: 11.50.xC4 Compression with 05/20/2009 12:56 workgroup edition? PM On May 20, 9:45 am, "Captain Pedantic" <theharlequi...@hotmail.com> wrote: > "Ron MacNeil" <macneil....@gmail.com> wrote in message > > news:mailman.688.1242824883.1831.informix-list@iiug.org... > > >> I can't seem to get any information on whether the new 11.50.xC4 > >> compression feature can be used with the Workgroup edition. In OAT I see > >> the ominous warning: "Enabling compression might require additional > >> license agreements. Ensure that these agreements are in place before > >> continuing." > >> Well, wouldn't it be nice if instead the message could point us to the > >> location of where we could out this information, or even better, simply > >> state whether or not this particular IDS instance is licensed for this > >> feature? Why build new features if we can't easily find out whether we > >> are entitled to them? > >> I've looked at the recent licensing agreements and see no mention of it, > >> so why should OAT display such a message? > > Yes, amidst the fanfares about the compression feature IBM has been > strangely reticent to mention this, hasn't it? > > Compression is a chargeable option available only for Enterprise Edition > (Informix Storage Optimization Feature) and its list cost is £9,978 (about > US$15,500 at today's exchange rate) per100 PVUs .. .the same as for DB2 I'm > told. > > The high cost, and the availability for Enterprise only, will limit the > appeal of this feature to all but the largest installations only. Ok, Color me silly, but what's the real value? You have SATA, SAS and now SSDs. With respect to SATA, drives are slow(er) than SAS but have a high density so you have a long time to recognize any ROI from disk compression. With respect to SAS, the drives are faster than SAS but have a lower density yet still not that expensive, relatively speaking. SSDs... fastest 'disk' but has the highest cost per GB of storage. There are some other drawbacks too, but thats a different story. Here I could see some potential value for disk compression except for the fact that there is a cost associated with the compress/decompress of the data. Kind of defeats the purpose of the SSD. So what am I missing? IMHO it seems like this feature is like tits on a boar. ;-) $15.5K buys a lot of disk and the array chassis. _______________________________________________ Informix-list mailing list Informix-list@iiug.org http://www.iiug.org/mailman/listinfo/informix-list
On May 20, 12:08 pm, Art Kagel <art.ka...@gmail.com> wrote: Many systems actually run faster with compressed data > according to IBM's testing and the experiences of early adopters. There is > a breakpoint around the number of times the average row is accessed once it > is read into the cache and the percentage of rows on a compressed page that > are accessed once the page is in memory that determines whether performance > improves or degrades due to compression. Hmmm. Art, I didn't think about that in terms of memory usage. However, have you recently looked at the price of RAM these days? ;-) Since IBM doesn't release any benchmarks... (Ooops! :-) ) ... , how is it that they can really guage the potential for improvement? How much do you really gain from compression? Isn't it going to be 'data dependent' ? I mean if I have an OLTP system where most of my data isn't in VARCHARs, how much compression can I expect? The less compression, the less gain in terms of memory paging. Lets take a look at a hotel reservation system. (Hyatt? Hilton? ...) You have a property name, but that gets translated to an id and most of your information is going to be non-var char data types. (property_id, room_id, room_type, checkin_id, etc ...) So in these systems, I don't see the potential value. So what am I missing? Don't get me wrong. I'm not trying to be dense.... I mean that I agree that if you can compress the page, you'll be more efficient in terms of memory usage. But then you increase your cpu load when you have to compress/decompress the fields. You seem to be implying that you can access the rows with the tables compressed. So does that mean if you're running a query and the table is compressed, what happens when one of the compressed fields is being used as a filter? Going back to the hotel example... Suppose you want to find an available room for a given weekend in New York City. Since the City field is probably going to be a varchar, it would be compressed in the table. But if I'm filtering on city MATCHES 'New York', does that work against a compressed field or does the engine have to decompress it when using it as a filter in a query? Sorry, I'm skeptical of its value. You also go on to say the following: "Larger tables with higher locality and relatively lower access frequency will tend to see improvement in processing speed for larger reports and more active systems. " I'm not sure what you mean by this. What do you mean exactly by 'higher locality'? I'm sorry, but even as I try to think up some design examples, most tend to be normalized and would yield little in value of compression. Again, I apologize for appearing dense. I just don't get it. 15.5K on top of an already expensive enterprise system isn't a good thing. I'd love to see some hard numbers, even if its not a real benchmark, but an example that could be reproduced by anyone. No hard numbers, just percentages of improvement. Does this make sense?
Ian Michael Gumby wrote: > On May 20, 12:08 pm, Art Kagel <art.ka...@gmail.com> wrote: > Many systems actually run faster with compressed data >> according to IBM's testing and the experiences of early adopters. There is >> a breakpoint around the number of times the average row is accessed once it >> is read into the cache and the percentage of rows on a compressed page that >> are accessed once the page is in memory that determines whether performance >> improves or degrades due to compression. > > Hmmm. > > Art, > > I didn't think about that in terms of memory usage. However, have you > recently looked at the price of RAM these days? ;-) > > Since IBM doesn't release any benchmarks... (Ooops! :-) ) ... , how is > it that they can really guage the potential for improvement? > > How much do you really gain from compression? Isn't it going to be > 'data dependent' ? I mean if I have an OLTP system where most of my > data isn't in VARCHARs, how much compression can I expect? The less > compression, the less gain in terms of memory paging. Since I may be considered an interested party I will try to be impartial. I'd say that only the customer can really understand the benefits of this feature. I'll try to explain why below and give some more information about how it works. First, looking at the disk price/GB is just part of the picture. Customer end up with multiple copies of data. The active instance, sometimes test and development (of course we all know this should be smaller and should not contain production data, but a reality check is in order...), logical logs, and several copies in backup images. Than you have the time of the backups and possibly other maintenance activities. I have *very serious* doubts that many customers really can put a cost on their GB (active, copies, backups etc.)... This is of course a very personal opinion. > > Lets take a look at a hotel reservation system. (Hyatt? Hilton? ...) > You have a property name, but that gets translated to an id and most > of your information is going to be non-var char data types. > (property_id, room_id, room_type, checkin_id, etc ...) So in these > systems, I don't see the potential value. So what am I missing? Maybe nothing, but it really depends on other factors. I have a customer with hundreds of tables of the same type. Most of their fields aren't chars. This instance takes about 1.5TB. On a typical table I can achieve nearly 50% of space savings. So, having a lot of CHAR fields can probably help, but it isn't a necessary condition. On a table with lot's of CHAR it can probably be easier to achieve much higher compression rates... But it has more to do with patterns than with datatypes. > > Don't get me wrong. I'm not trying to be dense.... > I mean that I agree that if you can compress the page, you'll be more > efficient in terms of memory usage. But then you increase your cpu > load when you have to compress/decompress the fields. True. That why the documentation mentions several situations where it may not be a good idea to compress the table. In particular if your system is CPU bound. Nonetheless the uncompression is fairly light... > > You seem to be implying that you can access the rows with the tables > compressed. So does that mean if you're running a query and the table > is compressed, what happens when one of the compressed fields is being > used as a filter? From the application/user point of view nothing happens. From the CPU point of view, it has to uncompress the row before processing it. It's not the same process, but it's more or less a similar situation to in-place altered tables. The engine is responsible to present a "current format" row to the layers above (for the sake of argument, assume a ALTER TABLE ADD (new_colum datatype DEFAULT ...) ) > Going back to the hotel example... Suppose you want to find an > available room for a given weekend in New York City. Since the City > field is probably going to be a varchar, it would be compressed in the > table. But if I'm filtering on city MATCHES 'New York', does that work > against a compressed field or does the engine have to decompress it > when using it as a filter in a query? > You're missing one important aspect of the compression algorithm. We don't compress "fields". We compress patterns and a pattern can span more than one field. > Sorry, I'm skeptical of its value. For once I find that pretty fair... > You also go on to say the following: > "Larger tables with higher locality and relatively lower access > frequency > will tend to see improvement in processing speed for larger reports > and more > active systems. " > > I'm not sure what you mean by this. What do you mean exactly by > 'higher locality'? > > I'm sorry, but even as I try to think up some design examples, most > tend to be normalized and would yield little in value of compression. Not necessarily true as I pointed out above. The Compression Estimator can be easily used to check the estimated compression ratios you'll get. Of course we have to consider several aspects: - The effective compression ratio - The cost savings you'll have from the compression ratios - The cost of the compression feature - The impacts (good or bad) on the system performance > > Again, I apologize for appearing dense. I just don't get it. 15.5K on > top of an already expensive enterprise system isn't a good thing. > I'd love to see some hard numbers, even if its not a real benchmark, > but an example that could be reproduced by anyone. No hard numbers, > just percentages of improvement. The whitepaper which is available on http://ibm.com/informix/compression has some examples for the TPC-C schema (compression ratios and performance) > > Does this make sense? > Your doubts? Sure. I'd say that either way we look at this it's a fairly complex subject. And as will all prices it can really depend. Finally a small note. The compression functionality introduced two other "small" features that may be very useful. Repack and shrink. But that's a bit off topic. Regards. -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently...
Darren_Jacobs@carmax.com wrote: > Gumby, > > I agree! > > I was excited about upgrading to 11.5 for the compression. We run on a SAN > at a data center and the costs over many systems could be of value! If > they decide to charge for this it reduces the ROI and no longer worth the > effort. Another nail in the coffin! Did you consider all the factors? The fact that it costs money obviously reduces ROI. It doesn't necessarily turn it negative... The effort is part of the ROI equation. IDS 11.5 has a lot of other features relatively to older releases (depending on the version you're using). AFAIK all RDBMS which have introduced compression charge for it... Oracle has a paid and a free version. Of course they're completely different and as you can imagine the free is much more limited. -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently...
On May 20, 7:02 pm, Fernando Nunes <domusonl...@gmail.com> wrote: > Darren_Jac...@carmax.com wrote: > Did you consider all the factors? The fact that it costs money obviously > reduces ROI. It doesn't necessarily turn it negative... The effort is part of > the ROI equation. > IDS 11.5 has a lot of other features relatively to older releases (depending on > the version you're using). Fernando, I think he did consider it and that's why he's hard pressed to justify the investment. The fact that its adding $15.5K to the bottom line, one has to consider other options that for 15.5K could also improve performance. Oh and of course you have to consider not just the 15.5K, but also increased maintenance costs going past year 1. Like I said, more memory or additional disk subsystems. I agree you have to look at the bigger picture and make your own decision. IMHO, I don't know how to justify this feature. Personally I think there are other features that are more important and that would show a better ROI. -G
Ian Michael Gumby wrote: > On May 20, 12:08 pm, Art Kagel <art.ka...@gmail.com> wrote: > Many systems actually run faster with compressed data >> according to IBM's testing and the experiences of early adopters. There is >> a breakpoint around the number of times the average row is accessed once it >> is read into the cache and the percentage of rows on a compressed page that >> are accessed once the page is in memory that determines whether performance >> improves or degrades due to compression. > > Hmmm. > > Art, > > I didn't think about that in terms of memory usage. However, have you > recently looked at the price of RAM these days? ;-) > > Since IBM doesn't release any benchmarks... (Ooops! :-) ) ... , how is > it that they can really guage the potential for improvement? > > How much do you really gain from compression? Isn't it going to be > 'data dependent' ? I mean if I have an OLTP system where most of my > data isn't in VARCHARs, how much compression can I expect? The less > compression, the less gain in terms of memory paging. > > Lets take a look at a hotel reservation system. (Hyatt? Hilton? ...) > You have a property name, but that gets translated to an id and most > of your information is going to be non-var char data types. > (property_id, room_id, room_type, checkin_id, etc ...) So in these > systems, I don't see the potential value. So what am I missing? The compression algorithm is looking for patterns. Common patterns leads to reductions - regardless of data type. The reservation systems that I've worked with tend to have a LOT of common data, (rates, day, etc...) Since it is pattern based, then non-var char data types are going to just as compressable as char data types. Even numeric data will compress. Probably the only thing which would not compress very well would be things such as GIF and JPEG objects. > > Don't get me wrong. I'm not trying to be dense.... > I mean that I agree that if you can compress the page, you'll be more > efficient in terms of memory usage. But then you increase your cpu > load when you have to compress/decompress the fields. > > You seem to be implying that you can access the rows with the tables > compressed. So does that mean if you're running a query and the table > is compressed, what happens when one of the compressed fields is being > used as a filter? > > Going back to the hotel example... Suppose you want to find an > available room for a given weekend in New York City. Since the City > field is probably going to be a varchar, it would be compressed in the > table. But if I'm filtering on city MATCHES 'New York', does that work > against a compressed field or does the engine have to decompress it > when using it as a filter in a query? And the indexes are not compressed - so no additional work. > > Sorry, I'm skeptical of its value. > > You also go on to say the following: > "Larger tables with higher locality and relatively lower access > frequency > will tend to see improvement in processing speed for larger reports > and more > active systems. " > > I'm not sure what you mean by this. What do you mean exactly by > 'higher locality'? > > I'm sorry, but even as I try to think up some design examples, most > tend to be normalized and would yield little in value of compression. > > Again, I apologize for appearing dense. I just don't get it. 15.5K on > top of an already expensive enterprise system isn't a good thing. > I'd love to see some hard numbers, even if its not a real benchmark, > but an example that could be reproduced by anyone. No hard numbers, > just percentages of improvement. > > Does this make sense? > >
On May 20, 8:21 pm, Madison Pruet <mpru...@verizon.net> wrote: > Ian Michael Gumby wrote: > The compression algorithm is looking for patterns. Common patterns > leads to reductions - regardless of data type. The reservation systems > that I've worked with tend to have a LOT of common data, (rates, day, > etc...) Since it is pattern based, then non-var char data types are > going to just as compressable as char data types. Even numeric data > will compress. Probably the only thing which would not compress very > well would be things such as GIF and JPEG objects. > > Ah, That explains a bit more. The trouble I'm having is that 15K per server is a lot of money. I mean if I spent 15K on upgrading my disks and maxing out my memory, will that give me better performance than I would get from compression? That's the trouble that I have. You don't have a concrete example that is reproducible by the customer. And if you're not willing to put the skin in the game on a benchmark, are you willing to put the skin in the game creating a sample database , application, and test set that you'll let a customer download and run to show a comparison between compressed and non-compressed performance? Don't get me wrong it sounds like you've got a 'loss less' compression algorithm that you apply either to the page or specifically to the row of data. Since you're not compressing the indexes, as long as you're not doing a table scan, you shouldn't see any performance hits. And since I asked, what happens in a sequential table scan? Please understand that there are some customers that can say 'hook me up' and they'll play with the feature and make their own judgment. I'm trying to understand the pros/cons before I decide to recommend this as a solution for a customer. The YMMV is a big killer when you consider the alternatives.
Comments on yours below. Overall, as we've all been saying, it's a complex subject and only YMMV can show the benefits if any. Is it worth the $$? Personally, I think that in general, no, but for specific customers it will be well worth the cost. Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, May 20, 2009 at 3:42 PM, Ian Michael Gumby <im_gumby@hotmail.com>wrote: > On May 20, 12:08 pm, Art Kagel <art.ka...@gmail.com> wrote: > Many systems actually run faster with compressed data > > according to IBM's testing and the experiences of early adopters. There > is > > a breakpoint around the number of times the average row is accessed once > it > > is read into the cache and the percentage of rows on a compressed page > that > > are accessed once the page is in memory that determines whether > performance > > improves or degrades due to compression. > > Hmmm. > > Art, > > I didn't think about that in terms of memory usage. However, have you > recently looked at the price of RAM these days? ;-) I didn't say anything about RAM. > > > Since IBM doesn't release any benchmarks... (Ooops! :-) ) ... , how is > it that they can really guage the potential for improvement? We all know that they did the benchmarks in-house. At the IIUG Conference Scott Lashley revealed that they used compression for many of the larger tables in the benchmark and yes it improved performance. > > > How much do you really gain from compression? Isn't it going to be > 'data dependent' ? I mean if I have an OLTP system where most of my > data isn't in VARCHARs, how much compression can I expect? The less > compression, the less gain in terms of memory paging. Yes, data dependent as to the percentage of compression. Character data will tend to compress 70-85% and binary data about 50% on average. Remember that the compression table is global to the dbspace not row specific, and that the entire row is compressed as a single stream of bytes, not individual columns - so repeated patterns between rows will be as important as the nature of the data within the row itself. > > > Lets take a look at a hotel reservation system. (Hyatt? Hilton? ...) > You have a property name, but that gets translated to an id and most > of your information is going to be non-var char data types. > (property_id, room_id, room_type, checkin_id, etc ...) So in these > systems, I don't see the potential value. So what am I missing? > > Don't get me wrong. I'm not trying to be dense.... > I mean that I agree that if you can compress the page, you'll be more > efficient in terms of memory usage. But then you increase your cpu > load when you have to compress/decompress the fields. > > You seem to be implying that you can access the rows with the tables > compressed. So does that mean if you're running a query and the table > is compressed, what happens when one of the compressed fields is being > used as a filter? If it is indexed then there will be no impact, index pages are not compressed. If table values have to be accessed for rows that will ultimately be rejected, then that will be expensive. Obviously a table that's used in table scans is a poor candidate for compression. > > > Going back to the hotel example... Suppose you want to find an > available room for a given weekend in New York City. Since the City > field is probably going to be a varchar, it would be compressed in the > table. But if I'm filtering on city MATCHES 'New York', does that work > against a compressed field or does the engine have to decompress it > when using it as a filter in a query? > > Sorry, I'm skeptical of its value. > > You also go on to say the following: > "Larger tables with higher locality and relatively lower access > frequency > will tend to see improvement in processing speed for larger reports > and more > active systems. " > > I'm not sure what you mean by this. What do you mean exactly by > 'higher locality'? Locality is the nearness of data rows that are being accessed by a single query. So, if the typical query will only access recently added data, that data will tend to be on only a few pages that are located physically together within one or two extents then the data needed for that query can be said to have high locality. A table that does not experience random deletions but only mass prunnings will tend to have high locality if only recent data is accessed by a query. Any order entry system (indeed most OLTP systems) will tend to exhibit high locality of data. Compression improves locality by cramming more rows onto a single page of a given size. Locality improvement is one area where compression can improve performance. Typically a row in a transaction table is accessed only a small number of times over the course of a relatively short time span. It is inserted, included in a detail report, read for a daily summary on the day it is inserted, reread for pick lists and similar processing reports then left alone until billing and periodic summaries are generated days later. This is what I mean by records that are access seldom. However, in our order entry system, all of the rows on a typical detail table page will belong to orders entered on the same or consecutive days, so if one of those rows is needed in memory, chances are very good that most of the rows on that page will be needed in memory. The reduced cost of IO by reading in more rows at a time more than offsets the CPU cost of decompressing the row a small number of times. Yes, the gap may be smaller or non-existent if the drives are ssd drives. That's the nature of YMMV. > > > I'm sorry, but even as I try to think up some design examples, most > tend to be normalized and would yield little in value of compression. > > Again, I apologize for appearing dense. I just don't get it. 15.5K on > top of an already expensive enterprise system isn't a good thing. > I'd love to see some hard numbers, even if its not a real benchmark, > but an example that could be reproduced by anyone. No hard numbers, > just percentages of improvement. > > Does this make sense? > > > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list >
"Ian Michael Gumby" <im_gumby@hotmail.com> wrote in message news:b136bbc1-8d91-4bfb-8983-ac89e3c11f97@q2g2000vbr.googlegroups.com... On May 20, 8:21 pm, Madison Pruet <mpru...@verizon.net> wrote: > Ian Michael Gumby wrote: >> The trouble I'm having is that 15K per server is a lot of money. I mean if I spent 15K on upgrading my disks and maxing out my memory, will that give me better performance than I would get from >> compression? First of all, I just looked it up on the "Informix" website; the list price is US$15,300 per PVU. Be clear now that this is per PVU. If you're running Enterprise Edition (as I understand it compression is available only for Enterprise) you are probably running with more than the equivalent of 8 Intel cores, or 400 PVUs (otherwise you'd be using Workgroup, unless you use ER or one of the other precluded functionalities). So let's say the entry level is 800 PVUs (4 quad-core Intel cpus, or 2 quad-core Sparc VII procs). That would make the cost of compression 8 x 15.3 = ~$122k per server, not the $15.5 you quote above. I tend to agree with Art: there will probably be some customers for whom compression will make sense in ROI terms. But as I said in an earlier post, at these prices and with the Enterprise-only restriction, it is very much a niche offering. Perhaps Serge can tell us if the Enterprise-only restriction applies to DB2 also?