Checkpoint Tuning in 11.5
Posted in 2010
Dan asked whether to use 11.5's auto checkpoint tuning (RTO_SERVER_RESTART, which makes CHKPTINTVL ignored) or revert to manual settings, worried that his undersized physical log was the only thing triggering checkpoints and that enlarging it would leave just one checkpoint a day at archive time. Andrew argued that's fine: with RTO set, recovery time stays near the RTO regardless of checkpoint frequency, so enlarge the plog and force a non-blocking checkpoint before backups. Others reported bugs causing runaway checkpointing and advised testing; Art ultimately recommended disabling RTO until the feature matures and relying on CHKPTINTVL plus an adequately sized physical log. No single agreed resolution, though Dan leaned toward turning RTO off.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Triggers, Constraints & Referential Integrity, Logging & Checkpoints
I have had my nose stuck in the books and a bunch of onstats trying to figure
out if it is better to "AUTO TUNE" checkpointing in 11.5 or take the old
route. What I am really concerned about is when I execute onstat -g ckp, the
only checkpoints we have are being triggered once per day from the backup and
the remainder are from the plog being 75% full. My plog is undersized on this
system by quite a bit but I am afraid if I size it to the recommended
bufferpool * 1.1. that I will not see checkpoints other than once per day with
the backup. I do have RTO_SERVER_RESTART set so checkpoint interval is
ignored. Any thoughts? I am looking for educated opinions more than anything
else I guess.
thanx,
Dan
----- Original Message -----
From: "DAN MUELLER" <dan.mueller@trnswrks.com>
To: <ids@iiug.org>
Sent: Monday, August 16, 2010 3:21 PM
Subject: Checkpoint Tuning in 11.5 [20940]
>I have had my nose stuck in the books and a bunch of onstats trying to
>figure
> out if it is better to "AUTO TUNE" checkpointing in 11.5 or take the old
> route. What I am really concerned about is when I execute onstat -g ckp,
> the
> only checkpoints we have are being triggered once per day from the backup
> and
> the remainder are from the plog being 75% full. My plog is undersized on
> this
> system by quite a bit but I am afraid if I size it to the recommended
> bufferpool * 1.1. that I will not see checkpoints other than once per day
> with
> the backup. I do have RTO_SERVER_RESTART set so checkpoint interval is
> ignored. Any thoughts? I am looking for educated opinions more than
> anything
> else I guess.
>
> thanx,
> Dan
>
Is there anything wrong with only taking a checkpoint once per day?
If you have RTO_SERVER_RESTART set to the time you're willing to let
Informix recover from a crash then how often a checkpoint occurs shouldn't
matter. Just because it takes all day to do enough work to meet the RTO that
doesn't mean it will take a whole day for the engine to recover from an
abnormal shutdown. It should just take approximately the number of seconds
RTO_SERVER_RESTART is set to.
I say set RTO_SERVER_RESTART to something you're comfortable with, increase
that physical log (no performance penalty for a ginormous physical log) and
let the engine determine how often to execute a checkpoint based on the
workload and not the time since the last checkpoint.
I have some systems with RTO_SERVER_RESTART of 60 seconds that only
checkpoint once a day. These guys have shutdown abnormally before and the
recovery time was in line with my RTO.
One thing you might also want to take a look at. The checkpoint before an
archive is a blocking checkpoint. You might want to force a regular
nonblocking checkpoint (onmode -c) before you take a backup to reduce the
time of the nonblocking checkpoint.
Andrew
You SHOULD be seeing checkpoints every 15 minutes or so! Otherwise it may
take hours to recover after a crash! The engine has to roll forward and
rollback every transaction from the last checkpoint until the moment of the
abnormal shutdown and then rollback anything that didn't commit before the
crash.
Either use AUTO checkpointing or set CHKPTINTVL to 900 seconds. In 11.50,
with non-blocking checkpoints, there is no reason to avoid proper
checkpointing.
WHY do you have the server checkpointing only a couple of times a day in the
first place?
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Mon, Aug 16, 2010 at 4:21 PM, DAN MUELLER <dan.mueller@trnswrks.com>wrote:
> I have had my nose stuck in the books and a bunch of onstats trying to
> figure
> out if it is better to "AUTO TUNE" checkpointing in 11.5 or take the old
> route. What I am really concerned about is when I execute onstat -g ckp,
> the
> only checkpoints we have are being triggered once per day from the backup
> and
> the remainder are from the plog being 75% full. My plog is undersized on
> this
> system by quite a bit but I am afraid if I size it to the recommended
> bufferpool * 1.1. that I will not see checkpoints other than once per day
> with
> the backup. I do have RTO_SERVER_RESTART set so checkpoint interval is
> ignored. Any thoughts? I am looking for educated opinions more than
> anything
> else I guess.
>
> thanx,
> Dan
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--0016e6d274b818c69b048e03885a
Dan,
I'd suggest you test this thoroughly on a test system. I have run into a
couple of bugs where the engine gets into some sort of loop and is
constantly check pointing. I finally turned off the RTO and the auto
checkpoints ... IBM was never able to determine what was going on ....
Peter Logan
Senior Database Administrator
Phone: 616/878-8309
From:
"DAN MUELLER" <dan.mueller@trnswrks.com>
To:
ids@iiug.org
Date:
08/16/2010 04:21 PM
Subject:
Checkpoint Tuning in 11.5 [20940]
Sent by:
ids-bounces@iiug.org
I have had my nose stuck in the books and a bunch of onstats trying to
figure
out if it is better to "AUTO TUNE" checkpointing in 11.5 or take the old
route. What I am really concerned about is when I execute onstat -g ckp,
the
only checkpoints we have are being triggered once per day from the backup
and
the remainder are from the plog being 75% full. My plog is undersized on
this
system by quite a bit but I am afraid if I size it to the recommended
bufferpool * 1.1. that I will not see checkpoints other than once per day
with
the backup. I do have RTO_SERVER_RESTART set so checkpoint interval is
ignored. Any thoughts? I am looking for educated opinions more than
anything
else I guess.
thanx,
Dan
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Art,
Because RTO_SERVER_RESTART is turned on (set to 300 in my case) CKPTINTVL is
ignored as per the Admin Guide. In the last 24 hours, I have seen 14
checkpoints, 13 of which were due to my plog reaching 75% of it's capacity. My
plog is grossly undersized for this instance and if it were not, I would not
be checkpointing at all except for the archive checkpoint. AUTO_CKPTS is
enabled.
You have pointed out exactly what my question/problem is in that how can the
server recover everything it did in 24 hours and still meet the 300 second
RTO_SERVER_RESTART criterea.
I have inherited this system and am about ready to turn off RTO and let this
thing checkpoint as in pre 11 days unless I can grasp some firm understanding
of how this all works. I have gotten some responses both ways about how good
RTO really is but I do not want to find out the hard way if this thing crashes.
thanx,
dan
On Tue, Aug 17, 2010 at 1:08 PM, Peter_Logan@spartanstores.com < Peter_Logan@spartanstores.com> wrote: > Dan, > > I'd suggest you test this thoroughly on a test system. I have run into a > couple of bugs where the engine gets into some sort of loop and is > constantly check pointing. I finally turned off the RTO and the auto > checkpoints ... IBM was never able to determine what was going on .... > > Peter Logan > Senior Database Administrator > Phone: 616/878-8309 > > Depending on your version, there are at least two situation where it will do frequently checkpoints: 1) due to auto update statistcs. I have no details on this currently, but I have the idea that it's normal and caused by the way it works. This of course would only happen during a short period, and while the evaluator task was working 2) another one which I watched at a customer happens if you try to create an index online on a temporary table (the reason to do it is beyond me of course). This was a bug. Don't remember exactly the version where it happened... it coudl have been FC4 or FC5. In any case, at the time it was already fixed (FC6 or FC7 certainly, maybe earlier). But none of this seems to be related to RTO... Regards. -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently... --00c09f9c974d9f9a90048e041e64
Ahh, that was not clear. I thought you were saying you were checkpointing
once a day using CHKPTINTVL plus a couple caused by the physical log at 75%
without RTO enabled. OK, I've been recommending to clients that they
disable RTO, at least until the technology matures. As others have pointed
out, there have been problems. Continue to checkpoint as always using
CHKPTINTVL and maintain a sufficient physical log to prevent premature
checkpoints. Yes, I think you should do the same.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Tue, Aug 17, 2010 at 8:23 AM, DAN MUELLER <dan.mueller@trnswrks.com>wrote:
> Art,
>
> Because RTO_SERVER_RESTART is turned on (set to 300 in my case) CKPTINTVL
> is
> ignored as per the Admin Guide. In the last 24 hours, I have seen 14
> checkpoints, 13 of which were due to my plog reaching 75% of it's capacity.
> My
> plog is grossly undersized for this instance and if it were not, I would
> not
> be checkpointing at all except for the archive checkpoint. AUTO_CKPTS is
> enabled.
>
> You have pointed out exactly what my question/problem is in that how can
> the
> server recover everything it did in 24 hours and still meet the 300 second
> RTO_SERVER_RESTART criterea.>
> I have inherited this system and am about ready to turn off RTO and let
> this
> thing checkpoint as in pre 11 days unless I can grasp some firm
> understanding
> of how this all works. I have gotten some responses both ways about how
> good
> RTO really is but I do not want to find out the hard way if this thing
> crashes.
>
> thanx,
> dan
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--0016e64768ec8defc7048e0470f3
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g