Re: Reducing Checkpoint Duration
Posted in 1996
This is a multi-part message in MIME format.
--------------5E08BD326B040F5D24E6AD32
Content-Type: text/plain; charset=iso-8859-2
Content-Transfer-Encoding: 7bit
Lin-Chuan Lee wrote:
>
> To reduce checkpoint duration, what are recommended strategies??
>
> I have guide to performance tuning in front of me ..
> so the standard 'by-the-book' .. lower LRU_Min/MAX stuff
> I know.. Any other advice gained through real
> experience would be greatly appreciated !
>
> TIA,
> -Lin
Hav a look at:
<a href="http://www.weideneder.de/informix/faq/checkdur.html">
www.weideneder.de
</a>
--------------5E08BD326B040F5D24E6AD32
Content-Type: text/plain; charset=iso-8859-2
Content-Transfer-Encoding: 7bit
Content-Disposition: inline; filename="lin.txt"
Hi Lin,
the duration of a checkpoint depends on the number of pages that are
to be flushed to the chunks. The maximum number of pages to flush at
the time of the checkpoint is:
max_dirty_pages = BUFFERS * LRU_MAX_DIRTY
Normally:
max_dirty_pages = BUFFERS * ( (LRU_MAX~ + MIN_DIRTY ) / 2 )
( I hope that your page cleaner threads work properly between the checkpoints. )
These pages must be written in a very short time ( 0 - 1 seconds ). First,
you have to determine the number of chunks, on which it is possible to
write concurrently. Second, you have to know the average speed of your
disks when no other process accesses the disks.
The following command tells you the speed of your chunks, if your Informix
OnLine DS uses a 2KB page size. ( "... bs=4096 count=2500" if you have a 4KB
page size ). Run this command from the Korn-Shell while OnLine is Off-Line !!!
$ time dd if=/dev/chunk1 of=/dev/null bs=2048 count=5000
Take the real time of the output and dvivide it by 10. This is the time
you need to read appr. 1MB. RAID Disks with Level 5 need a little more time
to write to the disk than to read from it. Therefore you must take the
time for reading times 1.5. Run the command above twice to ensure, that
we do not read from a cooked file or there is a cache controller! The time
spent for reading the second time should be the same as for the first time!
Now we need the number of chunks:
a) which are used for (INSERTs/UPDATEs and DELETEs)
( enter "onstat -D" at the end of the day )
b) to which of a) we can write concurrently.
( much more complicated )
If we would have to write 30MB to three disks ( 10MB to each ), we could
try to write it sequentially or parallel. If we would write it sequentially
than we must take the time for one disk times three disks. Assume the time
to write 10MB to one disk would take 5 seconds, than the time to write
30MB to three disks would take 15 seconds.
Any time below 15 seconds indicates that you can write parallel to the disks.
Use the following script:
$ cat <<eof > chk.scr
:
dd if=/dev/chunk1 of=/dev/null bs=2048 count=5000 & ;# background
dd if=/dev/chunk2 of=/dev/null bs=2048 count=5000 & ;# backgroundwait
eof
$ chmod 755 chk.scr
$ time chk.scr
I'm so sorry, but you have to run the test for each chunk combination. If
you are sure that you don't use a RAID Level 5 System, you can skip the
test for chunks on the same disk (otherwise you would have a coffee break).
Now I hope you know the number of chunks to which you can write parallel.
Notice: If it would take 5 seconds to write 10MB to one disk, and if it would
take just a little longer ( about 6 seconds ) for writing 30MB to three
disks ( 10MB to each ), than you can be sure that it's possible to write
parallel to the disks. Your disk speed would be 1sec per 5MB. (6sec/30MB)
The last step: Set the number of page cleaner threads (CLEANERS) to the
number of chunks you can parallel write to. Determine if your system supports
Asynchronous Kernel AIO ( $INFORMIXDIR/release/ONLINE_7.1 document. ). If
your system does not support Asynchronous Kernel AIO set the number of
AIO-vp's (NUMAIOVPS) to the same number as CLEANERS.
Now your checkpoint duration is:
duration = BUFFERS * LRU_MAX_DIRTY * 1sec / ( 5 MB )
or
duration = BUFFERS * LRU_MAX_DIRTY * 1sec / ( 2500 Pages )
or
if you would like a duration of 1 second
LRU_MAX_DIRTY = ( 2500 Pages ) / BUFFERS
Remember: The I/O must be balanced between the several chunks. It's very
important, that you create the chunks in the correct order.
+-- 1st disk ---+ +-- 2nd disk ---+
| | | |
| 1st chunk | | 2nd chunk |
| | | |
|---------------| |---------------|
| | | |
| 3rd chunk | | 4th chunk |
| | | |
+---------------+ +---------------+
During a checkpoint each page cleaner thread flushes the data of a special
chunk to the disk. The first page cleaner thread uses the first chunk, the
second cleaner the second chunk and so on. If you would place the second
chunk to the first disk, than you can have the same coffe break as mentioned
above.
Okay, it has become late and I think I have to go to bed now. I hope you
made sure that you have a slow performance because of checkpoint waits.
Do not simply reduce LRU_MAX_DIRTY and LRU_MIN_DIRTY! These parameters
depend on your buffer pool and on your disk speed!!!
Bye,
Stefan
**************************************************************
* SEB Weideneder / Munich
* e-mail: stefan@weideneder.de
* Fax: +49 89 3617015
**************************************************************
--------------5E08BD326B040F5D24E6AD32--