onstat -g ckp - block x wait
Posted in 2010
Topics: Logging & Checkpoints
Hi,
I get in some environments situations what for me isn't clear about the
non-blocking checkpoints.. (11.50 )
Monitoring the "onstat -g ckp" , with some frequency I get somes checkpoints
what the "block time" is close of zero and the "wait time" and "# waits" go up
(in some cases, more then 10 seconds because of the # of dirty_buffer x
slow_disks).
In this situations where have "wait time" with high value, the instance appear
to work like blocking checkpoint because the user sessions appear to stop
working... what for me don't make sense. (even consider the slow disks, once
mostly of data access are in buffers)
So... what's the difference between "wait time" and "block time" ???
Check what the manual says:
Total TimeTotal checkpoint duration in seconds from request time to checkpoint
completionFlush TimeTime, in seconds, to flush bufferpoolsBlock TimeIndividual
transaction blocking time, in seconds, for that particular
checkpoint# WaitsNumber of transactions blocked waiting for checkpointCkpt
TimeTime, in seconds, for all transactions to recognize a requested
checkpointWait TimeAverage time, in seconds, transactions waited for
checkpointLong TimeLongest amount of time, in seconds, a transaction waited
for checkpoint
Original post:
Hi,
I get in some environments situations what for me isn't clear about the
non-blocking checkpoints.. (11.50 )
Monitoring the "onstat -g ckp" , with some frequency I get somes checkpoints
what the "block time" is close of zero and the "wait time" and "# waits" go up
(in some cases, more then 10 seconds because of the # of dirty_buffer x
slow_disks).
In this situations where have "wait time" with high value, the instance appear
to work like blocking checkpoint because the user sessions appear to stop
working... what for me don't make sense. (even consider the slow disks, once
mostly of data access are in buffers)
So... what's the difference between "wait time" and "block time" ???
Check what the manual says:
Total TimeTotal checkpoint duration in seconds from request time to checkpoint
completionFlush TimeTime, in seconds, to flush bufferpoolsBlock TimeIndividual
transaction blocking time, in seconds, for that particular
checkpoint# WaitsNumber of transactions blocked waiting for checkpointCkpt
TimeTime, in seconds, for all transactions to recognize a requested
checkpointWait TimeAverage time, in seconds, transactions waited for
checkpointLong TimeLongest amount of time, in seconds, a transaction waited
for checkpoint
Response:
This is a bit confusing. From the onstat -g ckp output the "block time" column
is if the system had been performing a non-blocking checkpoint, but ran out of
resources (logical log/physical log etc...) and had to switch from the
non-blocking checkpoint to a blocking checkpoint. The "wait time" column can
still be seen on a non-blocking checkpoint because the only non-blocking
portion of the checkpoint is the buffer pool flush. There is still a blocking
portion of the checkpoint, and during that portion people are prevented from
getting into a critical section (or prevente from modifying data) and that
wait time column I believe is the avg amount of time that each thread that
ended up having to block averaged waiting (so it's a total amount of time
spent waiting divided by the total number of threads that had to wait, rather
then the long wait time or something like that which is the actual longest
time any 1 single thread had to wait because of the checkpoint).
Does that make more sense?
Jacques Renaut
IBM Informix Advanced Support
APD Team
Hi Jacques,
Thanks for your answer...
Help, but I still have the doubt... for how much time really/effective happen
a blocking checkpoint ? is the block time + wait time?
As you says, isn't every time the wait time means blocking or non-blocking ...
how to know?
--- Em sex, 27/8/10, JACQUES RENAUT <jrenaut@us.ibm.com> escreveu:
De: JACQUES RENAUT <jrenaut@us.ibm.com>
Assunto: Re: onstat -g ckp - block x wait [21078]
Para: ids@iiug.org
Data: Sexta-feira, 27 de Agosto de 2010, 16:21
Original post:
Hi,
I get in some environments situations what for me isn't clear about the
non-blocking checkpoints.. (11.50 )
Monitoring the "onstat -g ckp" , with some frequency I get somes checkpoints
what the "block time" is close of zero and the "wait time" and "# waits" go up
(in some cases, more then 10 seconds because of the # of dirty_buffer x
slow_disks).
In this situations where have "wait time" with high value, the instance appear
to work like blocking checkpoint because the user sessions appear to stop
working... what for me don't make sense. (even consider the slow disks, once
mostly of data access are in buffers)
So... what's the difference between "wait time" and "block time" ???
Check what the manual says:
Total TimeTotal checkpoint duration in seconds from request time to checkpoint
completionFlush TimeTime, in seconds, to flush bufferpoolsBlock TimeIndividual
transaction blocking time, in seconds, for that particular
checkpoint# WaitsNumber of transactions blocked waiting for checkpointCkpt
TimeTime, in seconds, for all transactions to recognize a requested
checkpointWait TimeAverage time, in seconds, transactions waited for
checkpointLong TimeLongest amount of time, in seconds, a transaction waited
for checkpoint
Response:
This is a bit confusing. From the onstat -g ckp output the "block time" column
is if the system had been performing a non-blocking checkpoint, but ran out of
resources (logical log/physical log etc...) and had to switch from the
non-blocking checkpoint to a blocking checkpoint. The "wait time" column can
still be seen on a non-blocking checkpoint because the only non-blocking
portion of the checkpoint is the buffer pool flush. There is still a blocking
portion of the checkpoint, and during that portion people are prevented from
getting into a critical section (or prevente from modifying data) and that
wait time column I believe is the avg amount of time that each thread that
ended up having to block averaged waiting (so it's a total amount of time
spent waiting divided by the total number of threads that had to wait, rather
then the long wait time or something like that which is the actual longest
time any 1 single thread had to wait because of the checkpoint).
Does that make more sense?
Jacques Renaut
IBM Informix Advanced Support
APD Team
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Original post:
Hi Jacques,
Thanks for your answer...
Help, but I still have the doubt... for how much time really/effective happen
a blocking checkpoint ? is the block time + wait time?
As you says, isn't every time the wait time means blocking or non-blocking ...
how to know?
Response:
Actually, the block time and wait time are both blocking, but the causes for
them are completely different. If you are getting block time from the onstat
-g ckp output, then that means you are running out of resources before the
buffer pool can be flushed in a non-blocking manner, and the checkpoint has to
switch from a non-blocking buffer pool flush, to an old style checkpoint where
the users need to be blocked from entering critical sections (or basically
doing any writes/updates) while the remainder of the buffer pool is then
finished up being flushed.
Wait time, on the other hand, is also blocking, but it is blocking that is
happening during the blocking portion of the new non-blocking flushing of the
buffer pool checkpoint code. Basically, prior to starting the buffer pool
flushing, there is still a blocking portion of the checkpoint code where users
are still prevented from entering critical sections. Again, the only
non-blocking portion of the new checkpoint code is the period while the buffer
pool is being flushed to disk. Now, generally speaking, that should be the
largest chunk of time for a checkpoint, but it isn't the only thing that a
checkpoint needs to do, so the new non-blocking checkpoints, do not mean
people never wait. The non-blocking portion is only while the buffer pool
flushes.
If you are seeing large "wait time" column output, then that means the server
is having to spend a lot of time to complete what it needs to get done during
the blocking portion of the new checkpoint code. It can be caused by a number
of things, so without more information it would be hard to say why you are
seeing this.
But again to answer your question, both block time and wait time are blocking.
However, if you are just looking for total amount of time that is spent
blocked, you would need to add up the block time column along with the ckpt
time column, and finally the Long time column. The sum of those 3 columns
should be pretty close to the amount of time during that checkpoint where
users would have been prevented from doing write type activity, but all 3 of
those columns have different reasons for what would be causing them.
The "Wait time" column is the avg length of time any waiter had to wait during
the blocking portion of the checkpoint. It's calculated by adding up all the
time that any thread that had to wait, and dividing it by the number of
threads that had to wait, to get the avg length of time spent waiting.
Jacques Renaut
IBM Informix Advanced Support
APD team
Nice!
Thanks for this more detailed explanation.
On the manual is missing an information explained like this...
Just for curiosity.
One of cases where I see longs waits into the "wait time" is mainly because of
the high RTO where practically one or no checkpoints are executed during the
day, (mostly is forced when the archive is started).
Regards
Cesar
--- Em ter, 31/8/10, JACQUES RENAUT <jrenaut@us.ibm.com> escreveu:
De: JACQUES RENAUT <jrenaut@us.ibm.com>
Assunto: Re: onstat -g ckp - block x wait [21086]
Para: ids@iiug.org
Data: Terça-feira, 31 de Agosto de 2010, 10:13
Original post:
Hi Jacques,
Thanks for your answer...
Help, but I still have the doubt... for how much time really/effective happen
a blocking checkpoint ? is the block time + wait time?
As you says, isn't every time the wait time means blocking or non-blocking ...
how to know?
Response:
Actually, the block time and wait time are both blocking, but the causes for
them are completely different. If you are getting block time from the onstat
-g ckp output, then that means you are running out of resources before the
buffer pool can be flushed in a non-blocking manner, and the checkpoint has to
switch from a non-blocking buffer pool flush, to an old style checkpoint where
the users need to be blocked from entering critical sections (or basically
doing any writes/updates) while the remainder of the buffer pool is then
finished up being flushed.
Wait time, on the other hand, is also blocking, but it is blocking that is
happening during the blocking portion of the new non-blocking flushing of the
buffer pool checkpoint code. Basically, prior to starting the buffer pool
flushing, there is still a blocking portion of the checkpoint code where users
are still prevented from entering critical sections. Again, the only
non-blocking portion of the new checkpoint code is the period while the buffer
pool is being flushed to disk. Now, generally speaking, that should be the
largest chunk of time for a checkpoint, but it isn't the only thing that a
checkpoint needs to do, so the new non-blocking checkpoints, do not mean
people never wait. The non-blocking portion is only while the buffer pool
flushes.
If you are seeing large "wait time" column output, then that means the server
is having to spend a lot of time to complete what it needs to get done during
the blocking portion of the new checkpoint code. It can be caused by a number
of things, so without more information it would be hard to say why you are
seeing this.
But again to answer your question, both block time and wait time are blocking.
However, if you are just looking for total amount of time that is spent
blocked, you would need to add up the block time column along with the ckpt
time column, and finally the Long time column. The sum of those 3 columns
should be pretty close to the amount of time during that checkpoint where
users would have been prevented from doing write type activity, but all 3 of
those columns have different reasons for what would be causing them.
The "Wait time" column is the avg length of time any waiter had to wait during
the blocking portion of the checkpoint. It's calculated by adding up all the
time that any thread that had to wait, and dividing it by the number of
threads that had to wait, to get the avg length of time spent waiting.
Jacques Renaut
IBM Informix Advanced Support
APD team
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g