Re: Reducing checkpoint intervals...
Posted in 1998
Narkinsky, Pat wrote:
>
> I run a very busy OLTP system. Currently, checkpoints are taking between 10
> and 20 seconds, and are happening every 5 minutes (as the result of the
> checkpoint interval being hit.).
>
> For obvious reasons, I would like to reduce this. Currently, I see a couple
> of things to try:
> * Reduce LRU_MAXDIRTY/LRU_MINDIRTY from their current values of 4/2 to
> something even lower. (My idea.) I'm leaning towards 1/0 at the moment.
> * My application vendor suggest that I reduce the number of BUFFERS to
> 50000, thereby having less memory out there, forcing my buffers to become
> dirty sonner and thereby be cleaned more frequently, etc... His stated
> reason for this is that my read cache hit rate is too high.
> Which of these alternatives would you consider more reasonable? Why?
> Anyone have a third idea? To me, reducing the number of buffers to get
> faster checkpoints seems vaguely silly. But, I've Been Wrong Before.
Your vendor's statement that your read cache% is too high is what is
just plain silly (unless it is >100% that is). I recommend trying to
reduce the LRU_MAXDIRTY/MINDIRTY to 1/0 as you suggested. Also add
a couple of AIO VPs (maybe bring it up to 3) I dunno but it just seems
like 1 AIO VP cannot be right, I know the manual says only cooked I/O
needs them but ... You are using only RAW disk right?
Another thing are you using singleton drives or RAID?? or what? How
many chunks do you have and how are they being hit during a checkpoint
(ie how many chunks contain tables that are being updated)? You are
now pushing, on average, 3000 pages or 6MB out to disk at each
checkpoint. If you are pushing this all out to a single disk that's
only 300/sec no sweat for any I/O system.
Guess, it may be that the actual checkpoint is only taking 1-4 seconds
but that the checkpoint cannot get started because some process is in
a critical section for a long time. Run onstat -s -r 1 and see if any
latches seem to last for 5-7 seconds or so. That would be your culprit
then.
Try updating all rows of several tables with no apps connected until
and there are 3000 dirty pages then force a checkpoint. Was it quick?
Hmmm! Application problems, long transactions holding up checkpoints.
Art S. Kagel