question on a strange IDS issue for the group -- l
Posted in 2012
Topics: Performance & Tuning, Logging & Checkpoints, Platform-Specific Issues
symptoms.. -- dramatic slowdown in checkpoint performance . -- decrease in checkpoint performance correllates almost exactly with buffer flush rate drop -- no stuck queries or long running queries/ no user or log file errors (just slow response) -- dirty pages well within Informix limits (~5000 pages/flush - same as before issues arose) environment: -- 11.70FC3 run on a Linux 2.6 VM - one RSS - 12GB memory (6GB mem for informix) -- data stored on a disk array LUN -- system was up 47 days before issue, has been rebooted steps so far: -- rebooted Informix -- rebooted VM -- changed LUN disk -- VMotioned to new VM -- moved to new ESX server results: -- nothing has shown a consistent difference in response -- running manual checkpoints on much shorter intervals gives shorter checkpoints, but no diff in flush times -- running 'dd' commands show a marked difference in performance, so all indications are that this is a system issue, not informix, but the ? arises because nothing seems to have been changed on the box, and its only running informix, so im wondering if anybody has seen this before and anything might have some ideas
First issue is there was a bug in the new readahead algorithm in 11.70 that wasn't fixed until 11.70.xC4, so the first thing I would do is to upgrade to 11.70.xC4 or .xC5 to get that fix installed. Then see how it goes. Second issue is that you are running on a VM and IO on VMWare is VERY VERY poor. Databases live and die on IO performance and checkpoints are where you see that most glaringly. Get back onto bare iron and give up on the VM environment. ESX is better but still not close to running Linux directly on the hardware. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) Blog: http://informix-myview.blogspot.com/ Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Thu, Jul 12, 2012 at 8:22 PM, JAMES ROTANTE <jamesrotante@sprintmail.com>wrote: > symptoms.. > > -- dramatic slowdown in checkpoint performance . > > -- decrease in checkpoint performance correllates almost exactly with > buffer > flush rate drop > > -- no stuck queries or long running queries/ no user or log file errors > (just > slow response) > > -- dirty pages well within Informix limits (~5000 pages/flush - same as > before > issues arose) > > environment: > -- 11.70FC3 run on a Linux 2.6 VM - one RSS - 12GB memory (6GB mem for > informix) > -- data stored on a disk array LUN > -- system was up 47 days before issue, has been rebooted > > steps so far: > -- rebooted Informix > -- rebooted VM > -- changed LUN disk > -- VMotioned to new VM > -- moved to new ESX server > > results: > -- nothing has shown a consistent difference in response > -- running manual checkpoints on much shorter intervals gives shorter > checkpoints, but no diff in flush times > -- running 'dd' commands show a marked difference in performance, so all > indications are that this is a system issue, not informix, but the ? arises > because nothing seems to have been changed on the box, and its only running > informix, so im wondering if anybody has seen this before and anything > might > have some ideas > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --bcaec5186a1654783804c4b3c879
I don't believe it is fixed in FC4 Cheers Paul -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Art Kagel Sent: Friday, July 13, 2012 5:46 AM To: ids@iiug.org Subject: Re: question on a strange IDS issue for the gr.... [27593] First issue is there was a bug in the new readahead algorithm in 11.70 that wasn't fixed until 11.70.xC4, so the first thing I would do is to upgrade to 11.70.xC4 or .xC5 to get that fix installed. Then see how it goes. Second issue is that you are running on a VM and IO on VMWare is VERY VERY poor. Databases live and die on IO performance and checkpoints are where you see that most glaringly. Get back onto bare iron and give up on the VM environment. ESX is better but still not close to running Linux directly on the hardware. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) Blog: http://informix-myview.blogspot.com/ Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Thu, Jul 12, 2012 at 8:22 PM, JAMES ROTANTE <jamesrotante@sprintmail.com>wrote: > symptoms.. > > -- dramatic slowdown in checkpoint performance . > > -- decrease in checkpoint performance correllates almost exactly with > buffer > flush rate drop > > -- no stuck queries or long running queries/ no user or log file errors > (just > slow response) > > -- dirty pages well within Informix limits (~5000 pages/flush - same as > before > issues arose) > > environment: > -- 11.70FC3 run on a Linux 2.6 VM - one RSS - 12GB memory (6GB mem for > informix) > -- data stored on a disk array LUN > -- system was up 47 days before issue, has been rebooted > > steps so far: > -- rebooted Informix > -- rebooted VM > -- changed LUN disk > -- VMotioned to new VM > -- moved to new ESX server > > results: > -- nothing has shown a consistent difference in response > -- running manual checkpoints on much shorter intervals gives shorter > checkpoints, but no diff in flush times > -- running 'dd' commands show a marked difference in performance, so all > indications are that this is a system issue, not informix, but the ? arises > because nothing seems to have been changed on the box, and its only running > informix, so im wondering if anybody has seen this before and anything > might > have some ideas > > > > **************************************************************************** *** > Forum Note: Use "Reply" to post a response in the discussion forum. > > --bcaec5186a1654783804c4b3c879 **************************************************************************** *** Forum Note: Use "Reply" to post a response in the discussion forum.
I believe there were a couple of significant issues with the RA ... IBM ended up supplying me several patched 11.70.fc3 versions to finally get them all corrected. I was informed that they are all fixed in FC5. Peter Peter Logan Senior Database Administrator Phone: 616/878-8309 From: "Paul Watson" <paul@oninit.com> To: ids@iiug.org Date: 07/13/2012 08:29 AM Subject: RE: question on a strange IDS issue for the gr.... [27595] Sent by: ids-bounces@iiug.org I don't believe it is fixed in FC4 Cheers Paul -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Art Kagel Sent: Friday, July 13, 2012 5:46 AM To: ids@iiug.org Subject: Re: question on a strange IDS issue for the gr.... [27593] First issue is there was a bug in the new readahead algorithm in 11.70 that wasn't fixed until 11.70.xC4, so the first thing I would do is to upgrade to 11.70.xC4 or .xC5 to get that fix installed. Then see how it goes. Second issue is that you are running on a VM and IO on VMWare is VERY VERY poor. Databases live and die on IO performance and checkpoints are where you see that most glaringly. Get back onto bare iron and give up on the VM environment. ESX is better but still not close to running Linux directly on the hardware. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) Blog: http://informix-myview.blogspot.com/ Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Thu, Jul 12, 2012 at 8:22 PM, JAMES ROTANTE <jamesrotante@sprintmail.com>wrote: > symptoms.. > > -- dramatic slowdown in checkpoint performance . > > -- decrease in checkpoint performance correllates almost exactly with > buffer > flush rate drop > > -- no stuck queries or long running queries/ no user or log file errors > (just > slow response) > > -- dirty pages well within Informix limits (~5000 pages/flush - same as > before > issues arose) > > environment: > -- 11.70FC3 run on a Linux 2.6 VM - one RSS - 12GB memory (6GB mem for > informix) > -- data stored on a disk array LUN > -- system was up 47 days before issue, has been rebooted > > steps so far: > -- rebooted Informix > -- rebooted VM > -- changed LUN disk > -- VMotioned to new VM > -- moved to new ESX server > > results: > -- nothing has shown a consistent difference in response > -- running manual checkpoints on much shorter intervals gives shorter > checkpoints, but no diff in flush times > -- running 'dd' commands show a marked difference in performance, so all > indications are that this is a system issue, not informix, but the ? arises > because nothing seems to have been changed on the box, and its only running > informix, so im wondering if anybody has seen this before and anything > might > have some ideas > > > > **************************************************************************** *** > Forum Note: Use "Reply" to post a response in the discussion forum. > > --bcaec5186a1654783804c4b3c879 **************************************************************************** *** Forum Note: Use "Reply" to post a response in the discussion forum. ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
thank you all very much for the help. latest actions were to turn off KAIO. this only slightly improved the flush rates, but for all practicality removed the user waits on the checkpoint. my assumption is that since informix is now managing the i/o, it can be a little smarter about it and if there's no response for a buffer flush, put the flush in a wait queue and run a user thread section. any confirmation/ correction/ improvement of that theory would be welcomed. next step: i convinced them to separate the temp dbspaces (80-90% of their i/o) to a separate disk channel. im hoping that will show a flush rate improvement during checkpoints. my theory being that most checkpoint-related flush/write activity would not be to the temp dbspaces (any thoughts/ confirmation/ correction )?