JFS2 on AIX on IDS
Posted in 2009
Someone asked whether JFS2 on AIX is suitable for IDS chunks. Art Kagel initially argued journaled filesystems are bad for databases (extra I/O overhead, page migration, fragmentation, safety issues like ext3/ext4), but after research agreed JFS2 is lightweight (journals metadata only, no page migration) and reasonably safe. Practical advice from IBM whitepapers: align the filesystem to 2K/4K, mount with Concurrent I/O for decent speed (DIO alone can be far slower), and note IDS 11.50.xC4+ supports CIO while IDS 9/10 lack DIRECT_IO. One site running ~100 JFS2 servers saw no performance change. Art still recommends raw devices; no single definitive conclusion.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Platform-Specific Issues
Does anyone have opinions - good, bad, or ugly - on using JFS2 on AIX for IDS? Any comments would be appreciated! Thanks - Mark Xtivia Inc. www.markscranton.com
Mark Scranton (Xtivia Inc.) wrote: > Does anyone have opinions - good, bad, or ugly - on using JFS2 on AIX > for IDS? Any comments would be appreciated! > > Thanks - > Mark > Xtivia Inc. > www.markscranton.com > System crashes (someone pulls plug out of wall). JFS does "Recovery to a point in time 'A'" Engine does recovery from another point in time (before or after A). Engine may fail recovery, or may not fail - is data on disk correct??
Journaled filesystems and databases are a BAD combination! Why? Note: - When you rewrite a page in a journaled filesystem the FS does not overwrite the existing page, it writes the page back to a new location and marks the original copy for later cleanup. That means: - There is filesystem overhead eating IO bandwidth and competing with the DB. - Journaled FS's assume all IO is cached so there is no provision I know of to give higher priority to external IO versus the internal FS cleanup operations. Since IDS uses DIRECT_IO and/or O_SYNC it is competing directly with the filesystem itself for priority access to the drives/ - IDS's chunks become increasingly more fragmented and dispersed across the disk partition over time due to the page movement on rewrite. - Don't know about JFS, but some journaled filesystems are designed badly and so are unsafe (EXT3 with write-back enabled and EXT4 always write meta data identifying a modified page's new location before writing the actual new copy of the page to disk). Altogether, avoid journaled filesystems. Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, May 6, 2009 at 11:05 AM, Mark Scranton (Xtivia Inc.) < mark.scranton@gmail.com> wrote: > Does anyone have opinions - good, bad, or ugly - on using JFS2 on AIX > for IDS? Any comments would be appreciated! > > Thanks - > Mark > Xtivia Inc. > www.markscranton.com > > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list >
Art Kagel wrote: > Journaled filesystems and databases are a BAD combination! Why? Note: > > * When you rewrite a page in a journaled filesystem the FS does not > overwrite the existing page, it writes the page back to a new > location and marks the original copy for later cleanup. That means: > o There is filesystem overhead eating IO bandwidth and > competing with the DB. > o Journaled FS's assume all IO is cached so there is no > provision I know of to give higher priority to external IO > versus the internal FS cleanup operations. Since IDS uses > DIRECT_IO and/or O_SYNC it is competing directly with the > filesystem itself for priority access to the drives/ > o IDS's chunks become increasingly more fragmented and > dispersed across the disk partition over time due to the > page movement on rewrite. > * Don't know about JFS, but some journaled filesystems are designed > badly and so are unsafe (EXT3 with write-back enabled and EXT4 > always write meta data identifying a modified page's new location > before writing the actual new copy of the page to disk). > > Altogether, avoid journaled filesystems. > I'd love to see this discussion going further... And it would be nice to have some technical documentation about this. Some of your observations make me wonder... I found this: http://publib.boulder.ibm.com/infocenter/pseries/v5r3/topic/com.ibm.aix.prftungd/doc/prftungd/reorg_fs_logs_log_vols.htm Since we don't "grow" chunks, I don't know if we cause any logging... I understand the rational behind the concerns, but do they really apply to a JFS2 used only for IDS chunks using DIRECT_IO? As a side note, IDS 11.50.xC4 supports CIO on AIX. Regards. -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently...
Yes, Fernando, after I replied I went to research JFS2 in more depth. It turns out that as a Journaled FS JFS2 is very light weight and only journals the meta-data not the actual data pages. JFS2 also does not use page migration when you overwrite existing data. That would point to it as being probably the only journaled file system that one might consider for database chunks. I would want to benchmark it, but it actually doesn't sound bad. Art Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, May 6, 2009 at 6:39 PM, Fernando Nunes <domusonline@gmail.com>wrote: > Art Kagel wrote: > > Journaled filesystems and databases are a BAD combination! Why? Note: > > > > * When you rewrite a page in a journaled filesystem the FS does not > > overwrite the existing page, it writes the page back to a new > > location and marks the original copy for later cleanup. That > means: > > o There is filesystem overhead eating IO bandwidth and > > competing with the DB. > > o Journaled FS's assume all IO is cached so there is no > > provision I know of to give higher priority to external IO > > versus the internal FS cleanup operations. Since IDS uses > > DIRECT_IO and/or O_SYNC it is competing directly with the > > filesystem itself for priority access to the drives/ > > o IDS's chunks become increasingly more fragmented and > > dispersed across the disk partition over time due to the > > page movement on rewrite. > > * Don't know about JFS, but some journaled filesystems are designed > > badly and so are unsafe (EXT3 with write-back enabled and EXT4 > > always write meta data identifying a modified page's new location > > before writing the actual new copy of the page to disk). > > > > Altogether, avoid journaled filesystems. > > > > I'd love to see this discussion going further... And it would be nice to > have > some technical documentation about this. > Some of your observations make me wonder... I found this: > > > http://publib.boulder.ibm.com/infocenter/pseries/v5r3/topic/com.ibm.aix.prftungd/doc/prftungd/reorg_fs_logs_log_vols.htm > > Since we don't "grow" chunks, I don't know if we cause any logging... > I understand the rational behind the concerns, but do they really apply to > a > JFS2 used only for IDS chunks using DIRECT_IO? > > As a side note, IDS 11.50.xC4 supports CIO on AIX. > > Regards. > > -- > Fernando Nunes > Portugal > > http://informix-technology.blogspot.com > My email works... but I don't check it frequently... > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list >
Art Kagel wrote: > Yes, Fernando, after I replied I went to research JFS2 in more depth. > It turns out that as a Journaled FS JFS2 is very light weight and only > journals the meta-data not the actual data pages. JFS2 also does not > use page migration when you overwrite existing data. That would point > to it as being probably the only journaled file system that one might > consider for database chunks. I would want to benchmark it, but it > actually doesn't sound bad. > > Art > > Art S. Kagel > Oninit (www.oninit.com <http://www.oninit.com>) > IIUG Board of Directors (art@iiug.org <mailto:art@iiug.org>) > > Disclaimer: Please keep in mind that my own opinions are my own opinions > and do not reflect on my employer, Oninit, the IIUG, nor any other > organization with which I am associated either explicitly or implicitly. > Neither do those opinions reflect those of other individuals affiliated > with any entity with which I am affiliated nor those of the entities > themselves. > > > > On Wed, May 6, 2009 at 6:39 PM, Fernando Nunes <domusonline@gmail.com > <mailto:domusonline@gmail.com>> wrote: > > Art Kagel wrote: > > Journaled filesystems and databases are a BAD combination! Why? > Note: > > > > * When you rewrite a page in a journaled filesystem the FS > does not > > overwrite the existing page, it writes the page back to a new > > location and marks the original copy for later cleanup. > That means: > > o There is filesystem overhead eating IO bandwidth and > > competing with the DB. > > o Journaled FS's assume all IO is cached so there is no > > provision I know of to give higher priority to > external IO > > versus the internal FS cleanup operations. Since IDS > uses > > DIRECT_IO and/or O_SYNC it is competing directly with the > > filesystem itself for priority access to the drives/ > > o IDS's chunks become increasingly more fragmented and > > dispersed across the disk partition over time due to the > > page movement on rewrite. > > * Don't know about JFS, but some journaled filesystems are > designed > > badly and so are unsafe (EXT3 with write-back enabled and EXT4 > > always write meta data identifying a modified page's new > location > > before writing the actual new copy of the page to disk). > > > > Altogether, avoid journaled filesystems. > > > > I'd love to see this discussion going further... And it would be > nice to have > some technical documentation about this. > Some of your observations make me wonder... I found this: > > http://publib.boulder.ibm.com/infocenter/pseries/v5r3/topic/com.ibm.aix.prftungd/doc/prftungd/reorg_fs_logs_log_vols.htm > > Since we don't "grow" chunks, I don't know if we cause any logging... > I understand the rational behind the concerns, but do they really > apply to a > JFS2 used only for IDS chunks using DIRECT_IO? > > As a side note, IDS 11.50.xC4 supports CIO on AIX. > > Regards. > > -- > Fernando Nunes > Portugal > > http://informix-technology.blogspot.com > My email works... but I don't check it frequently... > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org <mailto:Informix-list@iiug.org> > http://www.iiug.org/mailman/listinfo/informix-list > > This can also help: http://www.ibm.com/developerworks/library/l-journaling-filesystems/index.html As for benchmarking, there is an IBM document about CIO/DIO etc.: http://www.ibm.com/developerworks/data/library/techarticle/dm-0408lee/ And another one: https://www-03.ibm.com/systems/resources/systems_p_os_aix_whitepapers_db_perf_aix.pdf The first URL is very good from our discussion standpoint. I would say it can lead you to revise your thoughts on journaled filesystems. Look for the metadata definition, metadata modes and how each filesystem uses it... Nevertheless I still keep a few URLs I gathered from your references to Linus Torvalds "thoughts" on ext3 and ext4 (a post on c.d.i or a mail in the IIUG mailing list from you)... I just didn't have time to really understand the concerns. It would be nice to recheck the subject with the information provided by the above URL in my mind. The fact that all these URLs are coming from ibm.com is really a coincidence. They have been poping up from a lot of Google searches. In any case, if anybody has real, preferably documented bad experiences with databases and Journaled filesystems it would be interesting to check them. What theBP wrote is the obvious concern. But can it really happen considering that: - We use D_SYNC (I hppe I'm right on this one... could check) - The fs can journal only the metadata The situation can become a bit more weird if we put the chunks on a filesystem used for common files, but this should not be the best practice... Regards -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently...
OK reading the White Paper (the 3rd link below) several things are clear: - JFS2 is a fairly light weight journaled filesystem since it only journals meta-data - It would appear that it is safe to use (unlike EXT4 and EXT3 with write-back enabled). - If you use JFS2 for IDS chunks you must either create the filesystem with 2K alignment or 4K alignment. If you use 4K alignment you must not use dbspace pagesizes that are not multiples of 4K. - You must mount the FS (or at least the directory containing the chunk files) with Concurrent IO (CIO) enabled in order to get decent performance compared to RAW. - IBM's own testing using Oracle for OLTP benchmark loads shows that using JFS2 with CIO enabled and FS alignment correct is about 8% slower than using RAW chunks and that just using Direct IO (DIO) with proper FS alignment is as much as 70% slower than RAW which is MUCH worse than the 25-40% I have tested over the years with traditional non-jounaled filesystems. >From the second White Paper which tests CIO and DIO using DB2 UDB we learn: - Using Direct IO (DIO) on JFS2 is only efficient if the filesystem is NOT enabled for large files. So here we are back to 2GB or smaller chunks. - CIO on large file enabled FSs is OK. - The test was IO bound as the disk farm had insufficient bandwidth to satisfy the tests IO demands. - Here CIO comes much closer to RAW but IB that's only because the test was IO bound. Personally I am going to continue to recommend that we avoid all journaled file systems forchunks. I still think that RAW is the best way to go and that if one must use a filesystem it should be the simplest filesystem available that supports DIO. Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, May 6, 2009 at 9:21 PM, Fernando Nunes <domusonline@gmail.com>wrote: > Art Kagel wrote: > > Yes, Fernando, after I replied I went to research JFS2 in more depth. > > It turns out that as a Journaled FS JFS2 is very light weight and only > > journals the meta-data not the actual data pages. JFS2 also does not > > use page migration when you overwrite existing data. That would point > > to it as being probably the only journaled file system that one might > > consider for database chunks. I would want to benchmark it, but it > > actually doesn't sound bad. > > > > Art > > > > Art S. Kagel > > Oninit (www.oninit.com <http://www.oninit.com>) > > IIUG Board of Directors (art@iiug.org <mailto:art@iiug.org>) > > > > Disclaimer: Please keep in mind that my own opinions are my own opinions > > and do not reflect on my employer, Oninit, the IIUG, nor any other > > organization with which I am associated either explicitly or implicitly. > > Neither do those opinions reflect those of other individuals affiliated > > with any entity with which I am affiliated nor those of the entities > > themselves. > > > > > > > > On Wed, May 6, 2009 at 6:39 PM, Fernando Nunes <domusonline@gmail.com > > <mailto:domusonline@gmail.com>> wrote: > > > > Art Kagel wrote: > > > Journaled filesystems and databases are a BAD combination! Why? > > Note: > > > > > > * When you rewrite a page in a journaled filesystem the FS > > does not > > > overwrite the existing page, it writes the page back to a > new > > > location and marks the original copy for later cleanup. > > That means: > > > o There is filesystem overhead eating IO bandwidth and > > > competing with the DB. > > > o Journaled FS's assume all IO is cached so there is no > > > provision I know of to give higher priority to > > external IO > > > versus the internal FS cleanup operations. Since IDS > > uses > > > DIRECT_IO and/or O_SYNC it is competing directly with > the > > > filesystem itself for priority access to the drives/ > > > o IDS's chunks become increasingly more fragmented and > > > dispersed across the disk partition over time due to > the > > > page movement on rewrite. > > > * Don't know about JFS, but some journaled filesystems are > > designed > > > badly and so are unsafe (EXT3 with write-back enabled and > EXT4 > > > always write meta data identifying a modified page's new > > location > > > before writing the actual new copy of the page to disk). > > > > > > Altogether, avoid journaled filesystems. > > > > > > > I'd love to see this discussion going further... And it would be > > nice to have > > some technical documentation about this. > > Some of your observations make me wonder... I found this: > > > > > http://publib.boulder.ibm.com/infocenter/pseries/v5r3/topic/com.ibm.aix.prftungd/doc/prftungd/reorg_fs_logs_log_vols.htm > > > > Since we don't "grow" chunks, I don't know if we cause any logging... > > I understand the rational behind the concerns, but do they really > > apply to a > > JFS2 used only for IDS chunks using DIRECT_IO? > > > > As a side note, IDS 11.50.xC4 supports CIO on AIX. > > > > Regards. > > > > -- > > Fernando Nunes > > Portugal > > > > http://informix-technology.blogspot.com > > My email works... but I don't check it frequently... > > _______________________________________________ > > Informix-list mailing list > > Informix-list@iiug.org <mailto:Informix-list@iiug.org> > > http://www.iiug.org/mailman/listinfo/informix-list > > > > > > This can also help: > > > http://www.ibm.com/developerworks/library/l-journaling-filesystems/index.html > > As for benchmarking, there is an IBM document about CIO/DIO etc.: > > http://www.ibm.com/developerworks/data/library/techarticle/dm-0408lee/ > > And another one: > > > https://www-03.ibm.com/systems/resources/systems_p_os_aix_whitepapers_db_perf_aix.pdf > > The first URL is very good from our discussion standpoint. I would say it > can > lead you to revise your thoughts on journaled filesystems. Look for the > metadata definition, metadata modes and how each filesystem uses it... > Nevertheless I still keep a few URLs I gathered from your references to > Linus > Torvalds "thoughts" on ext3 and ext4 (a post on c.d.i or a mail in the IIUG > mailing list from you)... I just didn't have time to really understand the > concerns. It would be nice to recheck the subject with the information > provided > by the above URL in my mind. > > The fact that all these URLs are coming from ibm.com is really a@@
We have about 100 servers using JFS2 on AIX 5.2 with Informix (versions 9 and 10). And yes folks, we are in the process of upgrading them to AIX 6.1 and Informix 11.5. Anyway, we saw no difference in Informix's performance when we started using JFS2. A few years ago we saw: https://www-03.ibm.com/systems/resources/systems_p_os_aix_whitepapers_db_perf_aix.pdf and were excited to get the performance Oracle got, but alas, to no avail. But our system administrators had their reasons so we are still glad to be "up to date". -L.S.
LIGHT SCANS wrote: > We have about 100 servers using JFS2 on AIX 5.2 with Informix > (versions 9 and 10). And yes folks, we are in the process of > upgrading them to AIX 6.1 and Informix 11.5. Anyway, we saw no > difference in Informix's performance when we started using JFS2. A > few years ago we saw: > > https://www-03.ibm.com/systems/resources/systems_p_os_aix_whitepapers_db_perf_aix.pdf > > and were excited to get the performance Oracle got, but alas, to no > avail. But our system administrators had their reasons so we are > still glad to be "up to date". > > -L.S. > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list > IDS 9 and 10 don't even support DIRECT_IO. You will probably see improvements once you go to IDS 11.5. Remember that if you want to use CIO you must go to IDS 11.50.FC4... Of course this is a pretty new version... Careful testing is recommended as always. Regards.