Problems configuring LRU MIN/MAX and Checkpoint Interval.
Posted in 1999
Topics: Logging & Checkpoints
I orginally had problems with checkpoint intervals causing slow respones from the database in a application that was time critical (5 seconds is too slow), even with a checkpoint interval of 30 seconds. The original recommendations I received to resolve this situation was to change LRU MIN/MAX DIRTY to 0 AND 1 respectively. At first this appeared to have solved my problem, but now I'm realizing that periodically I need to load large amounts of data into the same system (ie: 1,000,000 records), and this is very slow with the LRU MIN/MAX parameters set so low. At the Informix defaults of 50/60 records are loaded at 3000/minute (still too slow, but bearable). At 0/1 the same records load at about 600 per minute. My problem is I can't afford to wait 5 seconds for a checkpoint, and I also can't afford to take 5 times longer to load records. Currently there are 16M records in the database, and it's expected to grow at 2M a month. What if any parameters should I be looking at? I have the identical software running in HP/UX (K260/2 Processors) with Dynamic Server 7.23, HP/UX (D370/2 Processors) with Workgroup Server 7.22, SCO (PII 400/2 Processors) with Workgroup Server 7.22, and SCO (PII 400/2 Processors) with Dynamic Server 7.30. All are experiencing the same problem.
Russell Bierschbach wrote: > I orginally had problems with checkpoint intervals causing slow respones > from the database in a application that was time critical (5 seconds is too > slow), even with a checkpoint interval of 30 seconds. The original > recommendations I received to resolve this situation was to change LRU > MIN/MAX DIRTY to 0 AND 1 respectively. At first this appeared to have > solved my problem, but now I'm realizing that periodically I need to load > large amounts of data into the same system (ie: 1,000,000 records), and this > is very slow with the LRU MIN/MAX parameters set so low. At the Informix > defaults of 50/60 records are loaded at 3000/minute (still too slow, but > bearable). At 0/1 the same records load at about 600 per minute. My > problem is I can't afford to wait 5 seconds for a checkpoint, and I also > can't afford to take 5 times longer to load records. Currently there are > 16M records in the database, and it's expected to grow at 2M a month. What > if any parameters should I be looking at? > > I have the identical software running in HP/UX (K260/2 Processors) with > Dynamic Server 7.23, HP/UX (D370/2 Processors) with Workgroup Server 7.22, > SCO (PII 400/2 Processors) with Workgroup Server 7.22, and SCO (PII 400/2 > Processors) with Dynamic Server 7.30. All are experiencing the same > problem. Well, is your record loads of 600 per minute a per session limit or if you have 2 sessions doing the load do you then get 1200 a minute? If it is only per session you could split your load into 5 and start up 5 loading processes. Does the table need to be accessable during the load? If not, you could try using HPL and do express mode loads which bypass the buffer pool but again, they need exclusive access to the table and require index rebuilds when the load completes (so there are draw backs). -- ******************************************************************** * Jacques P. Renaut "I'd dazzle you with brilliance * * Informix Advanced Support if I only had the knack..." * * email: jrenaut@informix.com #include <disclamier.h> * ********************************************************************
create 2 onconfig files, one for oltp and the other for batch processing. Russell Bierschbach wrote in message <3693d38d.0@news.prismnet.com>... >I orginally had problems with checkpoint intervals causing slow respones >from the database in a application that was time critical (5 seconds is too >slow), even with a checkpoint interval of 30 seconds. The original >recommendations I received to resolve this situation was to change LRU >MIN/MAX DIRTY to 0 AND 1 respectively. At first this appeared to have >solved my problem, but now I'm realizing that periodically I need to load >large amounts of data into the same system (ie: 1,000,000 records), and this >is very slow with the LRU MIN/MAX parameters set so low. At the Informix >defaults of 50/60 records are loaded at 3000/minute (still too slow, but >bearable). At 0/1 the same records load at about 600 per minute. My >problem is I can't afford to wait 5 seconds for a checkpoint, and I also >can't afford to take 5 times longer to load records. Currently there are >16M records in the database, and it's expected to grow at 2M a month. What >if any parameters should I be looking at? > >I have the identical software running in HP/UX (K260/2 Processors) with >Dynamic Server 7.23, HP/UX (D370/2 Processors) with Workgroup Server 7.22, >SCO (PII 400/2 Processors) with Workgroup Server 7.22, and SCO (PII 400/2 >Processors) with Dynamic Server 7.30. All are experiencing the same >problem. > > >
Russell Bierschbach wrote:
>
> I orginally had problems with checkpoint intervals causing slow respones
> from the database in a application that was time critical (5 seconds is too
> slow), even with a checkpoint interval of 30 seconds. The original
> recommendations I received to resolve this situation was to change LRU
> MIN/MAX DIRTY to 0 AND 1 respectively. At first this appeared to have
> solved my problem, but now I'm realizing that periodically I need to load
> large amounts of data into the same system (ie: 1,000,000 records), and this
> is very slow with the LRU MIN/MAX parameters set so low. At the Informix
> defaults of 50/60 records are loaded at 3000/minute (still too slow, but
> bearable). At 0/1 the same records load at about 600 per minute. My
> problem is I can't afford to wait 5 seconds for a checkpoint, and I also
> can't afford to take 5 times longer to load records. Currently there are
> 16M records in the database, and it's expected to grow at 2M a month. What
> if any parameters should I be looking at?
>
> I have the identical software running in HP/UX (K260/2 Processors) with
> Dynamic Server 7.23, HP/UX (D370/2 Processors) with Workgroup Server 7.22,
> SCO (PII 400/2 Processors) with Workgroup Server 7.22, and SCO (PII 400/2
> Processors) with Dynamic Server 7.30. All are experiencing the same
> problem.
Other parameters affecting checkpoint duration are CLEANERS, LRUS,
CKPTINTVL, NUMAIOVPS, and BUFFERS. Let me address these in the order I
believe
they are affecting your operation.
CKPTINTVL - you are under the mistaken impression that shortening the
frequency of checkpoints has a major effect on checkpoint duration.
Let me disavow you of that idea. Unless the LRU_MAXDIRTY and
LRU_MINDIRTY are so high that that the vast majority of your page
flushes are Chunk Writes (onstat -F) shortening the checkpoint interval
will NOT shorten checkpoint duration. What it will do is freeze your
inserts for the same N seconds that much more often. Set CKPTINTVL to
300 seconds or more and reduce LRU_MAX/MINDIRTY back to 1 and 0
respectively. That will get you <5 second checkpoints once every five
minutes rather than every 30 seconds so you can get more work done
between the pauses that checkpoints force. If you are not seeing write
cache %ages at least in the high 80s or low 90s during those bulk loads
your CKPTINTVL is too short (cache %s will be lower if you are only
loading large rows so that only one or a few fit on a page).
BUFFERS - this is one of the largest effectors of the checkpoint
duration. More buffers simply take longer to flush. Make sure you
need all of the BUFFERS you have configured.
LRUS - with more LRU queues you will be flushing more frequently and
in smaller batches. Also more LRUS means less contention between
multiple load jobs to get access to a new buffer to write to. If your
bufwaits (onstat -p) is more than a few % of (pagreads + bufwrits) then
you are experiencing LRU contention and need to increase LRUS. N.B. -
Avoid LRUS=64 or 96 these seem to trigger a bug in the LRU rehash
algorithm that causes bufwaits to go through the roof on multiprocessor
machines. Also be sure to adjust CLEANERS if you increase LRUS.
CLEANERS - the old rules of thumb from 5.xx for setting CLEANERS no
longer hold in 7.xx. However, use LRUS as a max value for this and
NUMCPUVPS as a min value. There is no way to monitor the performance
of CLEANERS and to see whether more will help with one exception. If
there are fewer than 500 dirty buffers at checkpoint time only one
cleaner will be used to flush all dirty buffers. If >500 dirty buffers
then if a cleaner is not available when the checkpoint thread has
gathers enough buffers from the same chunk to want to launch a cleaner
(ie all are still busy) then the checkpoint thread relinquishes the CPU
VP on which it is running (always VP #1) and sleeps for a while.
Therefore while it means that queries running on CPU VP#1 will stall
until all buffers have been flushed, to minimize checkpoint time it is
important to make sure that there are enough cleaners threads allocated
that there will always be on ready to run when the checkpoint threads
is looking for one. Therefore in this situation I recommend setting
CLEANERS == LRUS (OH using the max values of these - 128 - triggers a
harmless bug that reports that your physical and logical logs are zero
length at engine startup, FYI).
BTW is it possible that the data files you are loading from on being
read from the same disk drives that the Informix data is being written
to? Big no-no. Informix data should NEVER share physical drives with
OS filesystem partitions.
Art S. Kagel