RTO_SERVER_RESTART configuration parameter
Posted in 2012
A user asked whether to enable RTO_SERVER_RESTART instead of relying on CKPTINTVL for checkpoint control. Art Kagel asked for CKPTINTVL, LRU min/max dirty, onstat -F writes, checkpoint durations and block times, and recovery SLA. Given the reported values (CKPTINTVL 375, LRU 1/2, 1-5 second checkpoints, no SLA), he concluded the engine was tuned well and RTO_SERVER_RESTART would gain little; lowering LRU dirty fractions could shift more I/O to LRU flushes if needed. He also explained that "non-blocking" checkpoints still block briefly, so block times and affected sessions matter more than checkpoint duration.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Server Administration
hello all we currently have this parameter disabled. I am looking for feedback on experiences using this parameter versus how we are doing it, which is to use the CKPTINTVL configuration parameter. thanks tom
What is the current setting for CKPTINTVL?
What are the lru_min_dirty and lru_max_dirty settings for your buffer
caches?
What does the top section of the onstat -F report look like?
How long does a typical checkpoint take to complete?
Any significant block times reported in oncheck -g ckp?
What is your service level requirement for a return to service after a
server crash?
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Fri, Sep 14, 2012 at 4:03 PM, <tomcaml@gmail.com> wrote:
> hello all
>
>
> we currently have this parameter disabled.
> I am looking for feedback on experiences using this parameter versus how
> we are doing it, which is to use the CKPTINTVL configuration parameter.
>
> thanks
> tom
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
CKPTINTVL is 375
lru_min is 1
lru_max is 2
Fg Writes LRU Writes Chunk Writes
0 155906 423903
typical is 1 to 2 seconds (sometimes being 3, 4, or 5) - we are on 11.7.FC5 now - seems like when we were on FC4, the typical was 0 seconds and 1 second here and there
no formal SLA
thanks Art,
tom
On Friday, September 14, 2012 4:03:00 PM UTC-5, Art S. Kagel wrote:
> What is the current setting for CKPTINTVL?
> What are the lru_min_dirty and lru_max_dirty settings for your buffer caches?
> What does the top section of the onstat -F report look like?
> How long does a typical checkpoint take to complete?
>
>
> Any significant block times reported in oncheck -g ckp?
> What is your service level requirement for a return to service after a server crash?
>
> Art
>
> Art S. Kagel
> Advanced DataTools (www.advancedatatools.com)
>
>
> Blog: http://informix-myview.blogspot.com/
>
>
>
> Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves.
>
>
>
>
>
>
>
> On Fri, Sep 14, 2012 at 4:03 PM, <tom...@gmail.com> wrote:
>
>
> hello all
>
>
>
>
>
> we currently have this parameter disabled.
>
> I am looking for feedback on experiences using this parameter versus how we are doing it, which is to use the CKPTINTVL configuration parameter.
>
>
>
> thanks
>
> tom
>
> _______________________________________________
>
> Informix-list mailing list
>
> Inform...@iiug.org
>
> http://www.iiug.org/mailman/listinfo/informix-list
Sounds like the engine is running just right. If you want to bring the
checkpoint times down (and that would only be a issue if the block times
are more than 0.1 sec frequently) you can set the lru_max/min_dirty values
to fractions less than 2 and 1 and shift some more of the IO to LRU flushes.
I don't see that recovery time would be more than a minute or two, so
RTO_SERVER_RESTART will likely not gain you anything.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Tue, Sep 18, 2012 at 9:28 AM, <tomcaml@gmail.com> wrote:
> CKPTINTVL is 375>
> lru_min is 1
> lru_max is 2
>
> Fg Writes LRU Writes Chunk Writes
> 0 155906 423903
>
> typical is 1 to 2 seconds (sometimes being 3, 4, or 5) - we are on
> 11.7.FC5 now - seems like when we were on FC4, the typical was 0 seconds
> and 1 second here and there
>
> no formal SLA
>
> thanks Art,
> tom
>
>
>
> On Friday, September 14, 2012 4:03:00 PM UTC-5, Art S. Kagel wrote:
> > What is the current setting for CKPTINTVL?
> > What are the lru_min_dirty and lru_max_dirty settings for your buffer
> caches?
> > What does the top section of the onstat -F report look like?
> > How long does a typical checkpoint take to complete?
> >
> >
> > Any significant block times reported in oncheck -g ckp?
> > What is your service level requirement for a return to service after a
> server crash?
> >
> > Art
> >
> > Art S. Kagel
> > Advanced DataTools (www.advancedatatools.com)
> >
> >
> > Blog: http://informix-myview.blogspot.com/
> >
> >
> >
> > Disclaimer: Please keep in mind that my own opinions are my own opinions
> and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
> other organization with which I am associated either explicitly,
> implicitly, or by inference. Neither do those opinions reflect those of
> other individuals affiliated with any entity with which I am affiliated nor
> those of the entities themselves.
> >
> >
> >
> >
> >
> >
> >
> > On Fri, Sep 14, 2012 at 4:03 PM, <tom...@gmail.com> wrote:
> >
> >
> > hello all
> >
> >
> >
> >
> >
> > we currently have this parameter disabled.
> >
> > I am looking for feedback on experiences using this parameter versus how
> we are doing it, which is to use the CKPTINTVL configuration parameter.
> >
> >
> >
> > thanks
> >
> > tom
> >
> > _______________________________________________
> >
> > Informix-list mailing list
> >
> > Inform...@iiug.org
> >
> > http://www.iiug.org/mailman/listinfo/informix-list
>
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
Can you please elaborate on the difference between what the message log says for checkpoint duration times versus block times? Are you saying the duration might be recorded as 4 seconds but the actual block time might be anywhere from 0 seconds to 4 seconds? I have heard conflicting things regarding the "non-blocking" checkpoints - One IBM person said it was not quite the "real" case that they were non-blocking so not sure what the deal is there with that. Thanks Art! On Tuesday, September 18, 2012 8:36:21 PM UTC-5, Art S. Kagel wrote: > Sounds like the engine is running just right. If you want to bring the checkpoint times down (and that would only be a issue if the block times are more than 0.1 sec frequently) you can set the lru_max/min_dirty values to fractions less than 2 and 1 and shift some more of the IO to LRU flushes. > > > > I don't see that recovery time would be more than a minute or two, so RTO_SERVER_RESTART will likely not gain you anything. > > Art
There is a brief period during any "non-blocking" checkpoint when other sessions must be blocked out of critical sections, however, the checkpoint no longer has to wait for an exclusive critical section itself. Non-blocking checkpoints are based on the realization that the engine can recover by rolling back to the previous completed checkpoint and rolling forward again. That means that the checkpoint can "complete" without waiting for all of the dirty data buffers to be written to disk which was roughly 50% of the checkpoint duration and almost all of the blocking time during the original "full" checkpoint. Informix still has to perform a "full" checkpoint under certain conditions, includeing IB when a second checkpoint begins before the previous one's IOs completed and that will be a blocking checkpoint. That is one reason not to make the checkpoint interval too short or to make the physical log too small (risking triggering an early checkpoint when the physical log reaches 75% full). So, the things we are most concerned about in 11.xx with non-blocking checkpoints are the actual checkpoint interval, the average and maximum blocking times, and the number of sessions affected by the blocks. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) Blog: http://informix-myview.blogspot.com/ Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, Sep 19, 2012 at 8:16 AM, <tomcaml@gmail.com> wrote: > Can you please elaborate on the difference between what the message log > says for checkpoint duration times versus block times? Are you saying the > duration might be recorded as 4 seconds but the actual block time might be > anywhere from 0 seconds to 4 seconds? > I have heard conflicting things regarding the "non-blocking" checkpoints > - One IBM person said it was not quite the "real" case that they were > non-blocking so not sure what the deal is there with that. > > Thanks Art! > > > > > On Tuesday, September 18, 2012 8:36:21 PM UTC-5, Art S. Kagel wrote: > > Sounds like the engine is running just right. If you want to bring the > checkpoint times down (and that would only be a issue if the block times > are more than 0.1 sec frequently) you can set the lru_max/min_dirty values > to fractions less than 2 and 1 and shift some more of the IO to LRU flushes. > > > > > > > > I don't see that recovery time would be more than a minute or two, so > RTO_SERVER_RESTART will likely not gain you anything. > > > > Art > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list >