What means "nonblocked checkpoint" ?!
Posted in 2008
A user testing IDS 11.5 on Linux saw his esql/c update loop pause ~3 seconds during a checkpoint, even though online.log and 'onstat -g ckp' reported zero transaction block time, and wondered why "non-blocking checkpoints" still appeared to block. Replies explained that a manually triggered checkpoint (onmode -c) is always a blocking checkpoint, as are some others (e.g. the initial backup checkpoint) or cases with insufficient physical/logical log resources; Art Kagel also noted CPU VP contention from the checkpoint thread can stall threads. A developerWorks white paper on non-blocking checkpoints was cited. The poster accepted that his use of onmode -c explained the behaviour.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Server Administration, Logging & Checkpoints
Dear All...
I am now testing IDS 11.5 in RHEL 5, since I am interesting in
nonblocking checkpoint , I config PHYSFILE 900M , BUFFERS 600M ,
and run a esql/c ap , which do endless loop update a table ,
50,000 rows of them are updated each time , size 300 bytes each row !!!!
Because both physfile and buffers are large enough , the onstat -F
showes no chunk writes and the performance looks fine , and then
I type "onmode -c" to do the checkpoint , the esql/c ap take a break
for about 3 seconds(I printf the loop counter ... it delayed for 3 secs while
checkpoint happens),the online.log showes "duration was 4 seconds."
and "Checkpoint Statistics - Avg. Txn Block Time 0.000 ,#Txns blocked 0..",
onstat -g ckp showes also "Block Time 0" !!!
It is not what I think about "nonblocking checkpoint" , seems look like
it still block the transaction while do checkpoint , It is my
misunderstanding for this concept ?!
I'm pretty sure that a user induced checkpoint (i.e. onmode -c) is alwa=
ys a
blocking checkpoint.
-------------------------------------
Madison Pruet, STSM
IDS Replication Architect
=
"MARS CHEN" =
<mars@jsun.com> =
Sent by: =
To
ids-bounces@iiug. ids@iiug.org =
org =
cc
=
Subj=
ect
08/12/2008 08:52 What means "nonblocked =
PM checkpoint" ?! [13090] =
=
=
Please respond to =
ids@iiug.org =
=
=
Dear All...
I am now testing IDS 11.5 in RHEL 5, since I am interesting in
nonblocking checkpoint , I config PHYSFILE 900M , BUFFERS 600M ,
and run a esql/c ap , which do endless loop update a table ,
50,000 rows of them are updated each time , size 300 bytes each row !!!=
!
Because both physfile and buffers are large enough , the onstat -F
showes no chunk writes and the performance looks fine , and then
I type "onmode -c" to do the checkpoint , the esql/c ap take a break
for about 3 seconds(I printf the loop counter ... it delayed for 3 secs=
while
checkpoint happens),the online.log showes "duration was 4 seconds."
and "Checkpoint Statistics - Avg. Txn Block Time 0.000 ,#Txns blocked 0=
..",
onstat -g ckp showes also "Block Time 0" !!!
It is not what I think about "nonblocking checkpoint" , seems look like=
it still block the transaction while do checkpoint , It is my
misunderstanding for this concept ?!
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
Mars,
More likely this is a problem I have tried to point out to the engine
architects since 7.24 days.
The problem may not be that your transaction was blocked, technically it was
not. Hhow many CPU VPs do you have configured? Probably just one. While
the engine is gathering the list of pages that need to be flushed to disk
for the checkpoint, the checkpoint thread runs in CPU VP #1 without
relinquishing its hold on the VP until the list is complete. This is the 3
seconds at the beginning of the checkpoint. It's not so much that you were
blocked, but that you could not run while the checkpoint thread was hogging
the main (and for you perhaps the only) CPU VP.
In earlier versions of IDS (and to my knowledge this was never fixed - the
engineers just claimed that Fuzzy Checkpoints made it a non-issue so there
was nothing to fix) even if you had many CPU VPs, all of the active threads
running in CPU VP #1 were suspended during this first checkpoint phase and
were not released to the ready queue where they might be picked up by other
CPU VPs which are not blocked. So, any threads running in CPU VP #1 would
block during the beginning of any checkpoint - even if they were not making
updates and not in a critical section. If you are running only a single SHM
poll thread in CPU VP #1 that means that this VP is likely running most of
your active threads!
I delivered test code VERY similar to yours that demonstrated the problem
about 10 years ago. You might try the same test without the updates and see
if the session is still pausing as it did when I last tested this so long
ago.
To the engine developers: I propose the same two solutions I proposed when I
first raised this one, either:
- Release all threads running in CPU VP #1 before starting the checkpoint
so that they can run in other CPU VPs if any, or better:
- Run the checkpoint thread in the ADM or MSC VP so that they do not
block or shut out anything!
Art
On Tue, Aug 12, 2008 at 9:52 PM, MARS CHEN <mars@jsun.com> wrote:
> Dear All...
>
> I am now testing IDS 11.5 in RHEL 5, since I am interesting in
> nonblocking checkpoint , I config PHYSFILE 900M , BUFFERS 600M ,
> and run a esql/c ap , which do endless loop update a table ,
> 50,000 rows of them are updated each time , size 300 bytes each row !!!!
>
> Because both physfile and buffers are large enough , the onstat -F
> showes no chunk writes and the performance looks fine , and then
> I type "onmode -c" to do the checkpoint , the esql/c ap take a break
> for about 3 seconds(I printf the loop counter ... it delayed for 3 secs
> while
> checkpoint happens),the online.log showes "duration was 4 seconds."
> and "Checkpoint Statistics - Avg. Txn Block Time 0.000 ,#Txns blocked 0..",
>
> onstat -g ckp showes also "Block Time 0" !!!>
> It is not what I think about "nonblocking checkpoint" , seems look like
> it still block the transaction while do checkpoint , It is my
> misunderstanding for this concept ?!
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do those
opinions reflect those of other individuals affiliated with any entity with
which I am affiliated nor those of the entities themselves.
Thanks Art ....
I have IDS 11.5 running in onconfig CPU VPS num = 4 ,
actually I run IDS in a small NoteBook ,and its HD is very very slow!!
that is why checkpoint duration is 4 ~ 6 seconds ...
Another esql/c ap doing select , no update , looks like not affected
by checkpoint !!!
and then I run another esql/c ap , which do endless loop update another table,
table is small , but it affected by checkpoint , too ~~~
I use onmode -c to trigger checkpoint , Madison said that
"I'm pretty sure that a user induced checkpoint (i.e. onmode -c)
is always a blocking checkpoint. " , if that is true ,
my question is solved ~~~
Thanks all ....
onmode -c is a blocking checkpoint.There are others like the initial backup checkpoint, and you may block also
if you don't have enough resources (physical and logical logs).
I think all this is explained in the docs or some papers explaining the
features. If you don't find it I can look around.
It should not block during normal checkpoints when you hae enough resources.
It doesn't mean it will never block.
Regards.
On Wed, Aug 13, 2008 at 4:24 AM, MARS CHEN <mars@jsun.com> wrote:
> Thanks Art ....
>
> I have IDS 11.5 running in onconfig CPU VPS num = 4 ,
> actually I run IDS in a small NoteBook ,and its HD is very very slow!!
> that is why checkpoint duration is 4 ~ 6 seconds ...
>
> Another esql/c ap doing select , no update , looks like not affected
> by checkpoint !!!
>
> and then I run another esql/c ap , which do endless loop update another
> table,
> table is small , but it affected by checkpoint , too ~~~
>
> I use onmode -c to trigger checkpoint , Madison said that
> "I'm pretty sure that a user induced checkpoint (i.e. onmode -c)
> is always a blocking checkpoint. " , if that is true ,
> my question is solved ~~~
>
> Thanks all ....
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
There is a nice white paper about non-blocking checkpoints
www.ibm.com/developerworks
/db2/library/techarticle/dm-0703lashley/index.html
Also I would look at onstat -g ckp to see what it is doing.
John F. Miller III
STSM, Support Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
ids-bounces@iiug.org wrote on 08/13/2008 03:33:42 AM:
> onmode -c is a blocking checkpoint.> There are others like the initial backup checkpoint, and you may block
also
> if you don't have enough resources (physical and logical logs).
> I think all this is explained in the docs or some papers explaining the
> features. If you don't find it I can look around.
>
> It should not block during normal checkpoints when you hae enough
resources.
> It doesn't mean it will never block.
> Regards.
>
> On Wed, Aug 13, 2008 at 4:24 AM, MARS CHEN <mars@jsun.com> wrote:
>
> > Thanks Art ....
> >
> > I have IDS 11.5 running in onconfig CPU VPS num = 4 ,
> > actually I run IDS in a small NoteBook ,and its HD is very very slow!!
> > that is why checkpoint duration is 4 ~ 6 seconds ...
> >
> > Another esql/c ap doing select , no update , looks like not affected
> > by checkpoint !!!
> >
> > and then I run another esql/c ap , which do endless loop update another
> > table,
> > table is small , but it affected by checkpoint , too ~~~
> >
> > I use onmode -c to trigger checkpoint , Madison said that
> > "I'm pretty sure that a user induced checkpoint (i.e. onmode -c)
> > is always a blocking checkpoint. " , if that is true ,
> > my question is solved ~~~
> >
> > Thanks all ....
> >
> >
> >
> >
>
*******************************************************************************
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
> --
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Thanks all....
In my test case , I use "onmode -c" to trigger checkpoint ,
and I was told this would always force a blocking checkpoint ,
it is my fault that ignore this behavior ...