long checkpoints
Posted in 2007
Topics: Storage & Space Management, Server Administration, Logging & Checkpoints, Versions, Editions & End-of-Life
hi,
i have some problem with the long running checkpoints in IDS 7.31.UD8
on HP/UX.
every day a batch utility making hig volume of insert and updates. and
this batch works about 5-10 minutes(checkpoint times included).
during this batch load one or two (long) checkpoints occur.
(unfortunately blocking checkpoints)
17:32:53 Checkpoint Completed: duration was 160 seconds.
17:32:53 Checkpoint loguniq 1250, logpos 0x1842484
17:38:59 Checkpoint Completed: duration was 155 seconds.
17:38:59 Checkpoint loguniq 1253, logpos 0x8fa098
some of the onconfig parameters are
PHYSDBS rootdbs # Location (dbspace) of physical log
PHYSFILE 100000 # Physical log file size (Kbytes)
LOGFILES 40 # Number of logical log files
LOGSIZE 20000 # Logical log size (Kbytes)LOGSMAX 50
LOCKS 150000 # Maximum number of locks
BUFFERS 200000 # Maximum number of shared buffers
NUMAIOVPS # Number of IO vps
PHYSBUFF 64 # Physical log buffer size (Kbytes)
LOGBUFF 64 # Logical log buffer size (Kbytes)
CLEANERS 50 # Number of buffer cleaner processes
SHMBASE 0x0 # Shared memory base address
SHMVIRTSIZE 64000 # initial virtual shared memorysegment size
SHMADD 32000 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
CKPTINTVL 200 # Check point interval (in sec)
LRUS 50 # Number of LRU queues
LRU_MAX_DIRTY 60 # LRU percent dirty begin cleaninglimit
LRU_MIN_DIRTY 50 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water markpercentage
LTXEHWM 60 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 64 # Stack size (Kbytes)
when checkpoint starts i look at onstat -F output and see one or two
lines with state C. the chunk writes takes much of the checkpoint
time.
i changed the LRU_MAX_DIRTY:10, LRU_MIN_DIRTY:5 and try the batch
again. but this time due to LRU writes
the batch utility gets slower and finishes its job about at 20minutes.
i've read in some notes that LRU writes is worse than chunk writes.
so what can i do to make this batch process faster and shorten
checkpoint times?
thanks
Abdullah
On Nov 21, 8:01 am, Apostrof <abdullah.ako...@gmail.com> wrote:
> hi,
Abdullah,
You've got a lot going on. First thing I see is there should be a
checkpoint between the two you've published since you have CKPTINTVL
set to 200 or 3 mins 20 secs. Hopefully this is just because the
intervening chkpt was shorter so you didn't post it.
I agree that your settings for LRU_MIN/MAX_DIRTY are way to high for
this kind of system and I would go further even than your attempt at
moving from 50/60 to 10/5 and go all the way to 2/1.
Now, what might have happened when you reduced the LRU flush settings
to cause LRU writes (yes any significant number of these - say more
than 10 - is death to throughput)? I would guess that you need more
BUFFERS. What do your server's metrics look like? Post onstat -D,
onstat -p, and the first 20 and last 20 lines of the onstat -P outputwith the time since the stats were last zero'd (onstat -z) and I'll
try to calculate them for you.
I suspect that at least part of the problem is slow disk writes. This
wouldn't be a RAID5 setup and/or using COOKED chunks would it? RAID5
alone will slow down your updates/bulk loads SIGNIFICANTLY.
Similarly, you have only the default AIO VPs configured if you are
using COOKED filesystem chunks that will also be slowing you down.
Please post more configuration information.
Art S. Kagel
> i have some problem with the long running checkpoints in IDS 7.31.UD8
> on HP/UX.
> every day a batch utility making hig volume of insert and updates. and
> this batch works about 5-10 minutes(checkpoint times included).
> during this batch load one or two (long) checkpoints occur.
> (unfortunately blocking checkpoints)
>
> 17:32:53 Checkpoint Completed: duration was 160 seconds.
> 17:32:53 Checkpoint loguniq 1250, logpos 0x1842484
> 17:38:59 Checkpoint Completed: duration was 155 seconds.
> 17:38:59 Checkpoint loguniq 1253, logpos 0x8fa098>
> some of the onconfig parameters are
>
> PHYSDBS rootdbs # Location (dbspace) of physical log
> PHYSFILE 100000 # Physical log file size (Kbytes)
>
> LOGFILES 40 # Number of logical log files
> LOGSIZE 20000 # Logical log size (Kbytes)> LOGSMAX 50
>
> LOCKS 150000 # Maximum number of locks
> BUFFERS 200000 # Maximum number of shared buffers
> NUMAIOVPS # Number of IO vps
> PHYSBUFF 64 # Physical log buffer size (Kbytes)
> LOGBUFF 64 # Logical log buffer size (Kbytes)
> CLEANERS 50 # Number of buffer cleaner processes
> SHMBASE 0x0 # Shared memory base address
> SHMVIRTSIZE 64000 # initial virtual shared memory> segment size
> SHMADD 32000 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
> 0=>unlimited
> CKPTINTVL 200 # Check point interval (in sec)
> LRUS 50 # Number of LRU queues
> LRU_MAX_DIRTY 60 # LRU percent dirty begin cleaning> limit
> LRU_MIN_DIRTY 50 # LRU percent dirty end cleaning limit
> LTXHWM 50 # Long transaction high water mark> percentage
> LTXEHWM 60 # Long transaction high water mark
> (exclusive)
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 64 # Stack size (Kbytes)>
> when checkpoint starts i look at onstat -F output and see one or two
> lines with state C. the chunk writes takes much of the checkpoint
> time.
> i changed the LRU_MAX_DIRTY:10, LRU_MIN_DIRTY:5 and try the batch
> again. but this time due to LRU writes
> the batch utility gets slower and finishes its job about at 20minutes.
> i've read in some notes that LRU writes is worse than chunk writes.
> so what can i do to make this batch process faster and shorten
> checkpoint times?
> thanks
>
> Abdullah