Re: checkpoint duration
Posted in 1998
Willie L. Meeks wrote:
>
> I have a serious problem with checkpoint duration which cause my posting
> process to almost double in duration because of the amount of time wasted
> on checkpoints. Here is a a sample of the duration take directly from the
> online log file.
>
> 05:32:00 Checkpoint Completed: duration was 20 seconds.
> 05:34:54 Checkpoint Completed: duration was 93 seconds.
> 05:37:02 Checkpoint Completed: duration was 50 seconds.
> The following is a listing of the parameters from the onconfig file.
>
> PHYSFILE 75000 # Physical log file size (Kbytes)
[Surgical SNIP]
> CKPTINTVL 30000 # Check point interval (in sec)
LRUS 64 # Number of LRU queues(2/20/98)
LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
[SNIP]
Long checkpoints are not your main problem, I'll deal with that last.
You have your checkpoint interval set for 500 minutes yet you are
checkpointing every 2 minutes! With the checkpoints taking 20+seconds
there is little time left to process anything. You are in the
situation where your transactions, and other problems (see below), are
delaying the actual checkpoint processing and the long checkpoints are
delaying the transactions. I was there once.
The first problem is that the Physical log is filling and triggering a
premature checkpoint (depending on version you may even run into one of
several Physical Log Overflow bugs which will crash the engine). I also
see other problems.
Are you using KAIO? If not you need to increase NUMAIOVPS to at least
the level of CLEANERS. CLEANERS currently exceed LRUS this is
unneccessary, unlike in V5.xx V7.xx does not need one CLEANER per disk
at checkpoint time to perform but it does need one CLEANER per LRU to
keep the LRUS clean between checkpoints especially with aggressive
LRU_MAX/MIN_DIRTY values like yours (and like mine also).
You have LRUS=64 this is a particularly bad value. There is a bug,
which Informix has not verified yet, that I have identified that causes
the number of bufwaits to increase by 10-100 times for certain values
of LRUS. LRUS=64 is the only absolutely verified such value although I
suspect 16 and 96 also. Change LRUS to anything else, like 65, I use
127 (128 causes a harmless bug to trigger printing dire warnings about
incorrect Physical and Logical sizing so I avoid it even though it works
fine) and run 128 buffers just like you do.
Also after increasing the Physical log size to something reasonable so
you are not spending 1/5th to 1/2 of your processing time in checkpoints
reduce the checkpoint interval to something reasonable like 300 (5
minutes) or 900.
As to why the checkpoints are so long, I suspect three things:
1) The LRUS=64 bufwaits is killing you (run onstat -p and check then
change to some other value, restart and after a time onstat -p
again. Bufwaits should be a small % of bufwrites + diskreads (both
of which need to acquire an LRU latch). My guess is that this
number is anywhere from a large % of bufwrites + diskreads to
several times that figure. This will cause your updates to hang and
delay checkpoint processing.
2) Your disk farm is slow and it just takes a lot of time to flush the
buffers. I checkpoint 200000 buffers with LRU_MIN/MAX_DIRTY set as
you have it in 3-6 seconds (average 1900 dirty buffers at checkpoint
time). After adjusting the physical log you might try
LRU_MIN/MAX_DIRTY set to 1 & 0, respectively. Are you using
singleton disk drives? A large stripe set will flush faster and
improve readahead performance also. You might try a caching
controller for your drives as another way to speed flushes to disk.
3) Any transaction in a critical section (onstat -s) will cause the
checkpoint to wait. The checkpoint waits until all tasks have left
their current critical sections after locking new task out of the
semaphores they need to enter a critical section themselves. Look
for processes that spend a long time in a critical section as
culprits for causing this and rewrite them to reduce the amount of
time they spend in critical sections.
Hope this helps. Let us know.
Art S. Kagel