Huge checkpoint times during backup
Posted in 2000
Topics: Logging & Checkpoints
I am running Informix 7.31 on a Sun 3500 (Solaris 2.6). Starting Monday night, during the nightly level-0 backup, the checkpoint times increased from 0-1 seconds to 80-103 seconds. The backup which normally runs 1 hour ran in 4 hours. Initally I thought the backups were competing with another job, but that was not the case. On Tuesday night, I set up some cron jobs to watch the Informix engine during the backups. The same thing happened --- extreamly high checkpoints. The backup ran for 6 hours. My cron jobs verified that the backups were not competing with anything else. I did notice that the number of dirty buffers was extreamly high --- 20,000 to 40,000 dirty. The database is set to start flushing at 1% and stop and 0%. There are no other complaints about the system being slow. Nothing has changed .... prior to Monday, we had no problems. Any clues?
In article <38fdaf88.56531998@news.intersurf.net>, Jay Aymond <jaymond@communitycoffee.com> writes >I am running Informix 7.31 on a Sun 3500 (Solaris 2.6). > 7.31.UC? ^What 1,2,3,4 or 5? >Starting Monday night, during the nightly level-0 backup, the >checkpoint times increased from 0-1 seconds to 80-103 seconds. The >backup which normally runs 1 hour ran in 4 hours. > >Initally I thought the backups were competing with another job, but >that was not the case. > >On Tuesday night, I set up some cron jobs to watch the Informix engine >during the backups. The same thing happened --- extreamly high >checkpoints. The backup ran for 6 hours. My cron jobs verified that >the backups were not competing with anything else. I did notice that >the number of dirty buffers was extreamly high --- 20,000 to 40,000 >dirty. The database is set to start flushing at 1% and stop and 0%. > >There are no other complaints about the system being slow. > >Nothing has changed .... prior to Monday, we had no problems. > >Any clues? -- David Williams
Jay Aymond <jaymond@communitycoffee.com> wrote in message news:38fdaf88.56531998@news.intersurf.net... > I am running Informix 7.31 on a Sun 3500 (Solaris 2.6). > > Starting Monday night, during the nightly level-0 backup, the > checkpoint times increased from 0-1 seconds to 80-103 seconds. The > backup which normally runs 1 hour ran in 4 hours. Probably, there is some great activity when backup is running ? Archive checkpoints differ from simple checkpoints, because during archive checkpoints dirty pages are written into the vector placed in temporary dbspace. Post your logs and monitoring data for more detailed analysing. > > Initally I thought the backups were competing with another job, but > that was not the case. > > On Tuesday night, I set up some cron jobs to watch the Informix engine > during the backups. The same thing happened --- extreamly high > checkpoints. The backup ran for 6 hours. My cron jobs verified that > the backups were not competing with anything else. I did notice that > the number of dirty buffers was extreamly high --- 20,000 to 40,000 > dirty. The database is set to start flushing at 1% and stop and 0%. What is the cleaners state during flushing ? Are they all busy ? Are they starting to flush at 1 % ? How many checkpoint and LRU writes you have ? -- ------------------------------------------------- With best regards, Yuri Dovgart SAP R/3, Informix technical consultant, Informix Certified Professional, Senior System Consultant System Architecture and High Availability Systems, 'Telecominvest' company Email y_dovgart@tci.ukrtel.net ICQ 39284285
I seem to recall discussions at IWUC and in the 'advanced admin/internals' class about this situation. It has to do with 'very old' pages in the database where the 'timestamp' is close to wrapping. The 'timestamp' on the data pages is actually a 32-bit counter that is incremented every time an update occurs. During an archive, every page's timestamp is compared to the 'current' timestamp, the 'start of archive' timestamp, and a 'minimum' timestamp (which is related to the 'start of archive' timestamp -- being the timestamp that is greatest in absolute difference from the start timestamp). If a page's timestamp is greater than the current timestamp but less than the minimum, it is considered a 'very old' page, it's timestamp is updated, the page is rewritten and a special physical log record is created. If you have a lot of these during an archive, then you would see the dirty buffers and physical log activity increase significantly. The extended archive time is a known issue, but I do not think that Informix has ever fixed it. I do remember one company at IWUC had mentioned that they had written a program that examined these timestamps and could warn when such a long backup was imminent, because their backup times doubled or tripled when this occurred. At the time, Informix's response seemed to be 'yes, it is a problem, but design limitations prevent a solution in the near term'. In other words, live with it or update all your database pages regularly. Hope that helps, Doug "David Williams" <djw@smooth1.demon.co.uk> wrote in message news:U+wgyFAijj$4Eww7@smooth1.demon.co.uk... > In article <38fdaf88.56531998@news.intersurf.net>, Jay Aymond > <jaymond@communitycoffee.com> writes > >I am running Informix 7.31 on a Sun 3500 (Solaris 2.6). > > > > 7.31.UC? > ^What 1,2,3,4 or 5? > > >Starting Monday night, during the nightly level-0 backup, the > >checkpoint times increased from 0-1 seconds to 80-103 seconds. The > >backup which normally runs 1 hour ran in 4 hours. > > > >Initally I thought the backups were competing with another job, but > >that was not the case. > > > >On Tuesday night, I set up some cron jobs to watch the Informix engine > >during the backups. The same thing happened --- extreamly high > >checkpoints. The backup ran for 6 hours. My cron jobs verified that > >the backups were not competing with anything else. I did notice that > >the number of dirty buffers was extreamly high --- 20,000 to 40,000 > >dirty. The database is set to start flushing at 1% and stop and 0%. > > > >There are no other complaints about the system being slow. > > > >Nothing has changed .... prior to Monday, we had no problems. > > > >Any clues? > > -- > David Williams