Re: Sudden Increase in Checkpoint Lengths
Posted in 2009
Topics: Backup & Restore, Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
Answers to All responses below:-
2009/8/11 Habichtsberg, Reinhard <RHabichtsberg@arz-emmendingen.de>:
> Keith.
>
> If you haven't done it already:
> set environment TRACECKPT=1; export TRACECKPT before starting IDS.> That will make the checkpoints more chatty in online.log.
>
> If they are in critical section most of the time it could be an issue with
> the btscanners.
>
> HTH.
> Reinhard.
>
>> -----Original Message-----
>> From: informix-list-bounces@iiug.org
>> [mailto:informix-list-bounces@iiug.org]On Behalf Of John Carlson
>> Sent: Tuesday, August 11, 2009 3:04 AM
>> To: informix-list@iiug.org
>> Subject: Re: Sudden Increase in Checkpoint Lengths
>>
>>
>> Keith Simmons wrote:
>> > All
>> >
>> > Running IDS 9.4 FC6 on AIX 5.3, application written in INFORMIX-4GL
>> > Version 7.32.FC3.
>> > (I know it's a bit old, but it runs well (till now !!))
>> > Around Friday lunchtime the checkpoint times on this server jumped
>> > from between 0 and 1 second to between 30 and 60 seconds (although
>> > many are running arounf 6 to 10 seconds, big variances). Even though
>> > the checkpoints are fuzzy (with several thousand buffers not being
>> > flushed) the users are being blocked for this length of time every 5
>> > minutes (and they are not happy !!!). All disk writes are
>> Chunk Writes
>> > (no LRU or Fg), buffers are only getting around 6% dirty between
>> > checkpoints with low and high set to 8%/15%. Capturing onstat -u
>> > during a checkpoint shows around 30 user sessions blocked on the
>> > checkpoint but no-one in a critical section
>> > Except for an engine bounce (planned) about 4 weeks ago, the engine
>> > has run continuously with no configuration change for over
>> 9 months.
>> > There have been no application changes for the past 7 days.
>> > Only clue to the long checkpoints was a couple of 'Failed
>> _aioreturn'
>> > errors on each of two separate (unrelated) chunks in the online log.
>> > The O/S logs shows these errors also, but only as isolated entries,
>> > nothing continuous or on-going, also there are no hardware error
>> > lights illuminated nor does the DS4000 monitoring software show any
>> > issues.
>> > I have run page checks on the chunks with no errors and am currently
>> > running -cxI and -cd on each table in the DBSpaces involved.
>> > Any suggestions where else to look?
>> >
>>
>> Any issues with disk in the syslog (or AIX's version of syslog)?
>>
>> JWC
>> _______________________________________________
>> Informix-list mailing list
>> Informix-list@iiug.org
>> http://www.iiug.org/mailman/listinfo/informix-list
>>
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
Neil
sar 5 5 while a checkpoint was in progress:-
09:07:30 device %busy avque r+w/s Kbs/s avwait avserv
Average hdisk0 2 0.0 7 34 3.8 4.0
hdisk1 2 0.0 7 34 3.5 3.9
dac0 0 0.0 173 968 39.6 10.6
dac0-utm 0 0.0 0 0 0.0 0.0
dac1 0 0.0 165 678 114.9 2.4
dac1-utm 0 0.0 0 0 0.0 0.0
hdisk2 0 0.0 0 1 0.0 6.2
hdisk3 0 0.0 0 0 0.0 1.6
hdisk4 0 0.0 0 0 0.0 0.0
hdisk5 38 0.6 105 667 51.8 13.3
hdisk6 0 0.0 2 15 0.0 5.1
hdisk7 0 0.0 0 0 0.0 3.1
hdisk8 0 1.7 43 180 385.8 0.8
hdisk9 3 0.0 62 279 0.0 1.5
hdisk10 0 0.0 3 14 0.0 5.5
hdisk11 38 0.0 112 456 0.0 2.4
hdisk12 0 0.0 0 0 0.0 0.0
hdisk13 0 0.0 0 0 0.0 0.0
hdisk14 0 0.0 0 0 0.0 0.0
hdisk15 0 0.0 0 4 0.0 3.4
hdisk16 0 0.0 0 0 0.0 2.7
hdisk17 0 0.0 0 0 0.0 0.0
hdisk18 2 0.0 5 21 0.0 4.5
hdisk19 0 0.0 0 0 0.0 2.8
hdisk20 0 0.0 0 0 0.0 0.1
cd0 0 0.0 0 0 0.0 0.0
Obnoxio
No, data-set has not grown (I have just finished a round of purging
data, didn't recover the disk space but this finished 3 weeks ago and
the issue only appeared Friday). I only have one 'set' of disks, all
same size, shape, speed and connected in the same way.
Ian
Yes, same place, same package but only two memory segments, Resident
and 1 Virtual, no change from last restart.
John
Subsequent to my initial post I have had another LVM_IO_FAIL in the
AIX errpt log. I'm just about to raise call re hardware (and possibly
LVM).
Reinhard
Not possible currently as the application is 24 x 7. It is something I
will consider next time I bounce the engine. Is there any overhead
with having this turned on (apart from more entries in the online log)
and is it a feature available in IDS9.4?
All
Archive (using ontape) times have increased by 1 hour (37 %) from
before to after Friday midday. This is pointing more and more at
hardware or O/S, calls are being raised.
Many thanks
Keith
Keith Simmons wrote: > Ian > Yes, same place, same package but only two memory segments, Resident > and 1 Virtual, no change from last restart. In that case I'd have a word with the contract managers. If, for instance, daily invoicing is switched from daily to weekly, which is entirely set in data so you'd only know about it if they told you, it can make big changes to activity. In our case, which was pre-IDS, it resulted in memory growth. The package was changed to avoid that as a result of my sending them my analysis of the problem. But there could be other funnies lurking in there and a switch of activity could upset your tuning. -- Ian Hotmail is for spammers. Real mail address is igoddard at nildram co uk
Hello Keith, All the suggestions above are good. Let me just add these commands: errpt -a | pg vmstat 5 5 The "errpt" shows hardware and software errors. Look for a change when your system started slowing down. The "vmstat" shows a current overview of O.S. performance. Run it when your system is running at its worst. The "r" should be below 10 (l like it below 4). The "pi" should be 0. The "wa" should be "0". Any other column with a high number is worth looking into. -L.S.