All of a sudden long checkpoints started
Posted in 2004
Topics: Logging & Checkpoints, Platform-Specific Issues
Can someone give any pointers what to look for and where to look for long checkpoint duration? Actually we have two production servers and past week all of a sudden there is an increase in checkpoint duration on both the servers simultaneously. Not sure what is going on and what might be causing this? Here are the information OS: HP-UX 11.0 Informix Dyamic Server 7.31UD2R1 Buffers 600000 LRU MIN DIRTY 1 LRU MAX DIRTY 2 04:34:17 Checkpoint Completed: duration was 14 seconds. 04:49:45 Checkpoint Completed: duration was 19 seconds. 05:05:25 Checkpoint Completed: duration was 30 seconds. 05:21:14 Checkpoint Completed: duration was 40 seconds. 05:36:40 Checkpoint Completed: duration was 18 seconds. 05:52:07 Checkpoint Completed: duration was 17 seconds. 06:07:46 Checkpoint Completed: duration was 30 seconds. 06:23:15 Checkpoint Completed: duration was 21 seconds. 06:38:44 Checkpoint Completed: duration was 19 seconds. 06:54:08 Checkpoint Completed: duration was 16 seconds. 07:09:36 Checkpoint Completed: duration was 18 seconds. 07:25:01 Checkpoint Completed: duration was 16 seconds. 07:40:27 Checkpoint Completed: duration was 18 seconds. 07:55:54 Checkpoint Completed: duration was 18 seconds. 08:11:25 Checkpoint Completed: duration was 22 seconds. 08:26:59 Checkpoint Completed: duration was 24 seconds. 08:42:41 Checkpoint Completed: duration was 34 seconds. 08:58:20 Checkpoint Completed: duration was 30 seconds. 09:13:56 Checkpoint Completed: duration was 27 seconds. 09:29:28 Checkpoint Completed: duration was 22 seconds. __________________________________ Do you Yahoo!? Find out what made the Top Yahoo! Searches of 2003 http://search.yahoo.com/top2003
Possibly could be:
There might be a Crit lock condition that can cause the checkpoint to wait
until it is finishes.
Year end processing causing a disk bottleneck. Therefore, you should check
'sar -d 3600 1' output for 1 hour to give an understanding of disk
activity.
Massive update/insert/delete activity and having insufficient NUMAIOVPS,
and/or having insufficient CLEANERS.
Your CKPTINTVL (when the checkpoints should occur) appear to be sporatic.
It might be that the BUFFERS to LRUS ratio is too high, so you might need
to increase the LRUS. Also, you might need to increase the size of the
PHYSDBS, because when the physical log reaches the 75% threshold a
checkpoint is triggered (ie prematurely).
When a checkpoint occurs, the PHYSDBS is just logically flushed; however,
the checkpoints are directly releated to LRU flushing. I would first check
to ensure that your system is configured to keep up with the flushing of
the lru's (your setting is 2 to 1). Here is a sample script to allow you
to look at the state of the LRU dirtiness...
#!/bin/sh
echo "\\
onstat -R:"
onstat -R | egrep -i " m |queued" | sed -e "s/%//g" |
sort -b -n -r -k 3,3 | head -11 | awk '
{
if($2=="dirty,")
printf("LRU queues are %.2f%% dirty overall, Top 10 dirty: \\
",
($1 / $5) * 100);else
printf("%s ",$3);
}
END {
printf("\\
\\
");
}'
If you get desperate, there is an undocumented feature that will allow you
to flush the LRU's without inhibiting update/insert/delete activities. It
is 'onmode -B'. I use this because we have massive updates on our system
and because the LRU cleaning is already set 1 to 0. Our checkpoints were
around 10-15 seconds, they are now around 1 sec after implementing a handy
script... I am told that the undocumented feature is difficult to
implement in HDR. Here is a sample script....
#!/bin/sh
# $INFORMIXDIR/etc/$ONCONFIG: CKPTINTVL 1800
SLEEPSECS=`expr 1800 - 300` # 5 min before normal checkpoint interval
FLUSHSECS=10 # (tune!) to the time to flush buffers
while :
do
onstat - >/dev/null 2>&1 # see if informix is online
if [ $? -eq 5 ]; then # run when informix is online
onmode -B # flushes buffers without blocking updatessleep $FLUSHSECS # sleep a few seconds
onmode -c # perform checkpoint (fast!), CKPTINTVL isreset
sleep $SLEEPSECS # wait
else # informix was not online
sleep 60 # check again in 1 minute
fi
done
I hope this helps a little.
Good luck,
-Tim
"Vineet
Mehr...." To: ids@iiug.org
<vin_us@yahoo.c cc:
om> Subject: All of a sudden long checkpoints started [2408]
Sent by:
forum.subscribe
r@iiug.org
01/05/2004
10:47 AM
Can someone give any pointers what to look for and
where to look for long checkpoint duration?
Actually we have two production servers and past week
all of a sudden there is an increase in checkpoint
duration on both the servers simultaneously.
Not sure what is going on and what might be causing
this?
Here are the information
OS: HP-UX 11.0
Informix Dyamic Server 7.31UD2R1
Buffers 600000
LRU MIN DIRTY 1
LRU MAX DIRTY 2
04:34:17 Checkpoint Completed: duration was 14
seconds.
04:49:45 Checkpoint Completed: duration was 19
seconds.
05:05:25 Checkpoint Completed: duration was 30
seconds.
05:21:14 Checkpoint Completed: duration was 40
seconds.
05:36:40 Checkpoint Completed: duration was 18
seconds.
05:52:07 Checkpoint Completed: duration was 17
seconds.
06:07:46 Checkpoint Completed: duration was 30
seconds.
06:23:15 Checkpoint Completed: duration was 21
seconds.
06:38:44 Checkpoint Completed: duration was 19
seconds.
06:54:08 Checkpoint Completed: duration was 16
seconds.
07:09:36 Checkpoint Completed: duration was 18
seconds.
07:25:01 Checkpoint Completed: duration was 16
seconds.
07:40:27 Checkpoint Completed: duration was 18
seconds.
07:55:54 Checkpoint Completed: duration was 18
seconds.
08:11:25 Checkpoint Completed: duration was 22
seconds.
08:26:59 Checkpoint Completed: duration was 24
seconds.
08:42:41 Checkpoint Completed: duration was 34
seconds.
08:58:20 Checkpoint Completed: duration was 30
seconds.
09:13:56 Checkpoint Completed: duration was 27
seconds.
09:29:28 Checkpoint Completed: duration was 22
seconds.
__________________________________
Do you Yahoo!?
Find out what made the Top Yahoo! Searches of 2003
http://search.yahoo.com/top2003