Sudden Increase in Checkpoint Lengths
Posted in 2009
Topics: Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
All
Running IDS 9.4 FC6 on AIX 5.3, application written in INFORMIX-4GL
Version 7.32.FC3.
(I know it's a bit old, but it runs well (till now !!))
Around Friday lunchtime the checkpoint times on this server jumped
from between 0 and 1 second to between 30 and 60 seconds (although
many are running arounf 6 to 10 seconds, big variances). Even though
the checkpoints are fuzzy (with several thousand buffers not being
flushed) the users are being blocked for this length of time every 5
minutes (and they are not happy !!!). All disk writes are Chunk Writes
(no LRU or Fg), buffers are only getting around 6% dirty between
checkpoints with low and high set to 8%/15%. Capturing onstat -u
during a checkpoint shows around 30 user sessions blocked on the
checkpoint but no-one in a critical section
Except for an engine bounce (planned) about 4 weeks ago, the engine
has run continuously with no configuration change for over 9 months.
There have been no application changes for the past 7 days.
Only clue to the long checkpoints was a couple of 'Failed _aioreturn'
errors on each of two separate (unrelated) chunks in the online log.
The O/S logs shows these errors also, but only as isolated entries,
nothing continuous or on-going, also there are no hardware error
lights illuminated nor does the DS4000 monitoring software show any
issues.
I have run page checks on the chunks with no errors and am currently
running -cxI and -cd on each table in the DBSpaces involved.
Any suggestions where else to look?
Keith
what does sar -d 5 5 show?
"Keith Simmons" <smiley73@googlemail.com> wrote in message
news:mailman.4.1249918474.1191.informix-list@iiug.org...
> All
>
> Running IDS 9.4 FC6 on AIX 5.3, application written in INFORMIX-4GL
> Version 7.32.FC3.
> (I know it's a bit old, but it runs well (till now !!))
> Around Friday lunchtime the checkpoint times on this server jumped
> from between 0 and 1 second to between 30 and 60 seconds (although
> many are running arounf 6 to 10 seconds, big variances). Even though
> the checkpoints are fuzzy (with several thousand buffers not being
> flushed) the users are being blocked for this length of time every 5
> minutes (and they are not happy !!!). All disk writes are Chunk Writes
> (no LRU or Fg), buffers are only getting around 6% dirty between
> checkpoints with low and high set to 8%/15%. Capturing onstat -u
> during a checkpoint shows around 30 user sessions blocked on the
> checkpoint but no-one in a critical section
> Except for an engine bounce (planned) about 4 weeks ago, the engine
> has run continuously with no configuration change for over 9 months.
> There have been no application changes for the past 7 days.
> Only clue to the long checkpoints was a couple of 'Failed _aioreturn'
> errors on each of two separate (unrelated) chunks in the online log.
> The O/S logs shows these errors also, but only as isolated entries,
> nothing continuous or on-going, also there are no hardware error
> lights illuminated nor does the DS4000 monitoring software show any
> issues.
> I have run page checks on the chunks with no errors and am currently
> running -cxI and -cd on each table in the DBSpaces involved.
> Any suggestions where else to look?
>
> Keith
Keith Simmons wrote: > All > > Running IDS 9.4 FC6 on AIX 5.3, application written in INFORMIX-4GL > Version 7.32.FC3. > (I know it's a bit old, but it runs well (till now !!)) > Around Friday lunchtime the checkpoint times on this server jumped > from between 0 and 1 second to between 30 and 60 seconds (although > many are running arounf 6 to 10 seconds, big variances). Keith, are you still at the same old place using the same package? The problem I described a few days ago in relation to virtual memory also happened, on what was your sister site, at lunchtimes - an invoice run at that time plus a change in invoicing pattern due to a contract change caught us out. -- Ian Hotmail is for spammers. Real mail address is igoddard at nildram co uk
Keith Simmons wrote:
> All
>
> Running IDS 9.4 FC6 on AIX 5.3, application written in INFORMIX-4GL
> Version 7.32.FC3.
> (I know it's a bit old, but it runs well (till now !!))
> Around Friday lunchtime the checkpoint times on this server jumped
> from between 0 and 1 second to between 30 and 60 seconds (although
> many are running arounf 6 to 10 seconds, big variances). Even though
> the checkpoints are fuzzy (with several thousand buffers not being
> flushed) the users are being blocked for this length of time every 5
> minutes (and they are not happy !!!). All disk writes are Chunk Writes
> (no LRU or Fg), buffers are only getting around 6% dirty between
> checkpoints with low and high set to 8%/15%. Capturing onstat -u
> during a checkpoint shows around 30 user sessions blocked on the
> checkpoint but no-one in a critical section
> Except for an engine bounce (planned) about 4 weeks ago, the engine
> has run continuously with no configuration change for over 9 months.
> There have been no application changes for the past 7 days.
> Only clue to the long checkpoints was a couple of 'Failed _aioreturn'
> errors on each of two separate (unrelated) chunks in the online log.
> The O/S logs shows these errors also, but only as isolated entries,
> nothing continuous or on-going, also there are no hardware error
> lights illuminated nor does the DS4000 monitoring software show any
> issues.
> I have run page checks on the chunks with no errors and am currently
> running -cxI and -cd on each table in the DBSpaces involved.
> Any suggestions where else to look?
>
Any issues with disk in the syslog (or AIX's version of syslog)?
JWC