Long checkpoints (LRU_xxx_DIRTY / LRUS / CLEANERS adapted)
Posted in 1999
We are having long checkpoints (around 20 seconds). The number of
dirty pages to be written at checkpoint time does not even seem to
play a major role. Our SUN Sparc Storage Arrays have enough NVRAM
anyway to handle the chunk writes at checkpoint time. I have already
decreased LRU_MIN_DIRTY and LRU_MAX_DIRTY to 2 and 1. I'm now even
trying 1 and 0, but this did not decrease the checkpoint duration. I
also already increased LRUS from 16 (which is the number of CPUVPs) to
52 and now even 75, again without success. I also increased the
number of page cleaners to 52 (-> we have about 150 mirrorred disks).
Even when there are only 1000 dirty pages, the checkpoint still takes
at least 15 seconds. I have been looking at I/O activity during a
checkpoint (with iostat -x and the ssaadm utility delivered with SUN's
Sparc Storage Arrays), but did not see any overloaded disks at all.
When comparing the output from onstat runs before and after the
checkpoint, I noticed an enormous increase in bufferreads (> 300,000)
and lockreqs (300,000). Could a high number of bufferread activity at
checkpoint time cause this ? Does anyone have another idea ? Any
help would be welcome as I am already struggling with this for months.
Thank you in advance,
Mario Opsomer
PS : We are using a SUN E6000 as db server (18 CPUs). Nothing else is
running on that server. KAIO is enabled. BUFFERS is set to 512,000.
We are running Informix 7.24.UC3 with SAP R/3 (4.0B) as application.
I increased CKPTINTVL to decrease the impact of these long checkpoints
by decreasing their frequency, but the recovery of a recent crash was
so terribly long that I have decreased CKPTINTVL back to 20 minutes.