Re: Checkpoint altos
Posted in 2004
Topics: Performance & Tuning, Storage & Space Management, Server Administration, Logging & Checkpoints, Platform-Specific Issues
Hola,
usually you get long checkpoints when (almost) all the writing of
buffers to disk is done during checkpoints. To verify this do
"onstat -F" and check the numbers of "Chunk Writes" versus
"LRU Writes".
To get shorter checkpoints with less impact, you want to minimize
"Chunk Writes" and increase "LRU Writes". The main instruments
to do this are the LRU_... parameters in the $ONCONFIG file:
LRUS 8 # Number of LRU queues
LRU_MAX_DIRTY 20.000000 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 10.000000 # LRU percent dirty end cleaning limit
Depending on the size of your buffer pool (BUFFERS in your
$ONCONFIG), you need a sufficient number of LRUS so that
the individual LRU queues are not too long.
With LRU_MAX_DIRTY and LRU_MIN_DIRTY you can determine
when the writing of modified buffers to disk will start and when it
will end. The extreme is to configure these to 1.0 and 0.0 which
basically means that buffer writing to disk will happen continuously.
This can often result in better overall performance than having
long checkpoints. Especially it will eliminate the impact on users
(which get blocked for the time of a checkpoint).
If this does not bring the desired results, then you probably have to
increase your buffer pool (higher number of BUFFERS).
Saludos,
Martin
--
Martin Fuerderer
IBM Informix Development Munich, Germany
Data Management Solutions
owner-informix-list@iiug.org wrote on 04.12.2004 00:00:16:
> DATOS
> =====
> IDS: Informix 7.31 FD8
> Sistema Operativo: HP-UX 11.11 ( 64 bits )
> ICN: 00110800
>
> =========================================================
>
> Hola:
> Tengo los checkpoint demasiados altos y esto
> como entenderan IMPACTA en le performance de motor.
> Me pueden ayudar para poder determinar cual es la
> causa de estos checkpoint's altos, como puedo sacar
> informacion del motor para determinar que
> proceso esta teniendo demasiaco impacto y demasiado recursos
> del motor.
>
>
> 17:25:14 Checkpoint Completed: duration was 47 seconds.
> 17:25:14 Checkpoint loguniq 19060, logpos 0x1f1c018
>
> 17:25:46 Logical Log 19060 Complete.
> 17:25:53 Logical Log 19060 - Backup Started
> 17:26:59 Logical Log 19060 - Backup Completed
> 17:31:58 Checkpoint Completed: duration was 101 seconds.
> 17:31:58 Checkpoint loguniq 19061, logpos 0x6c5018>
> Muchas gracias por la ayuda
>
> saludos
>
sending to informix-list
Martin Fuerderer wrote:
> usually you get long checkpoints when (almost) all the writing of
> buffers to disk is done during checkpoints. To verify this do
> "onstat -F" and check the numbers of "Chunk Writes" versus
> "LRU Writes".
> To get shorter checkpoints with less impact, you want to minimize
> "Chunk Writes" and increase "LRU Writes". The main instruments
> to do this are the LRU_... parameters in the $ONCONFIG file:
>
> LRUS 8 # Number of LRU queues
> LRU_MAX_DIRTY 20.000000 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 10.000000 # LRU percent dirty end cleaning limit>
> Depending on the size of your buffer pool (BUFFERS in your
> $ONCONFIG), you need a sufficient number of LRUS so that
> the individual LRU queues are not too long.
>
> With LRU_MAX_DIRTY and LRU_MIN_DIRTY you can determine
> when the writing of modified buffers to disk will start and when it
> will end. The extreme is to configure these to 1.0 and 0.0 which
> basically means that buffer writing to disk will happen continuously.
Not usually advisable - use LRU_MIN_DIRTY of 1 and LRU_MAX_DIRTY of 2.
If you had IDS 9.40 (instead of the 7.31 you report having - thanks),
then you could use fractional percentages for LRU_MIN_DIRTY and
LRU_MAX_DIRTY, but prior to that, the values were integers (so any
fractional part is ignored).
> This can often result in better overall performance than having
> long checkpoints. Especially it will eliminate the impact on users
> (which get blocked for the time of a checkpoint).
>
> If this does not bring the desired results, then you probably have to
> increase your buffer pool (higher number of BUFFERS).
One other trick which may or may not work - consider using 'onmode -B'
to flush the buffer pool (without blocking on a checkpoint) before
using 'onmode -c' to force the checkpoint. If you do this from cron,
you can schedule your checkpoints as you need them. Note that this is
only relevant if 'onmode -B' is supported (effective) in your version
of IDS. (Note: even if 'onmode' doesn't complain, it is not totally
convincing proof that it was effective; I'd want to review the page
writes and dirty lists before and after to be sure.) Also, you should
be aware that there are those who swear that doing this can slow down
a checkpoint, though I have yet to be convinced that this can happen
in a stable production system. The argument is that if the buffer
flushing takes so long that most of the pages are dirty again before
the flush completes, then you've wasted those writes. If you grant
the premise, the conclusion follows; I don't grant the premise. In
big systems, where LRU_MIN_DIRTY and LRU_MAX_DIRTY are running at low
figures, I don't see how this can happen - in small systems where you
run with large (50/60) values, it might, but still seems a trifle
implausible. YMMV, as they say.
> owner-informix-list@iiug.org wrote on 04.12.2004 00:00:16:
>
>
>>DATOS
>>=====
>>IDS: Informix 7.31 FD8
>>Sistema Operativo: HP-UX 11.11 ( 64 bits )
>>ICN: 00110800
>>
>>=========================================================
>>
>>Hola:
>> Tengo los checkpoint demasiados altos y esto
>>como entenderan IMPACTA en le performance de motor.
>>Me pueden ayudar para poder determinar cual es la
>>causa de estos checkpoint's altos, como puedo sacar
>>informacion del motor para determinar que
>>proceso esta teniendo demasiaco impacto y demasiado recursos
>>del motor.
>>
>>
>>17:25:14 Checkpoint Completed: duration was 47 seconds.
>>17:25:14 Checkpoint loguniq 19060, logpos 0x1f1c018
>>
>>17:25:46 Logical Log 19060 Complete.
>>17:25:53 Logical Log 19060 - Backup Started
>>17:26:59 Logical Log 19060 - Backup Completed
>>17:31:58 Checkpoint Completed: duration was 101 seconds.
>>17:31:58 Checkpoint loguniq 19061, logpos 0x6c5018
--
Jonathan Leffler #include <disclaimer.h>
Email: jleffler@earthlink.net, jleffler@us.ibm.com
Guardian of DBD::Informix v2003.04 -- http://dbi.perl.org/