$ONCONFIG Parameter: RTO_SERVER_RESTART
Posted in 2009
Topics: Server Administration, Logging & Checkpoints, Platform-Specific Issues
IDS 11.50.FC5 AIX 5.3
While reviewing the $ONCONFIG parameters available on this instance, I came
across this new parameter:
## Recovery Time Objective Server Restart
RTO_SERVER_RESTART 0 # Number of Seconds, Overrides Checkpoint Interval
## Shared Memory Space Parameter
CKPTINTVL 60 # Check point interval (in sec)
Has anyone played with this parameter yet? Any downside to setting this
parameter?
Initial setting: same as the "now-defunct (overriden)" CKPTINTVL?
Thanks in advance.
Clifton
_________________________________________________________________
Bing brings you maps, menus, and reviews organized in one place.
http://www.bing.com/search?q=restaurants&form=MFESRP&publ=WLHMTAG&crea=TEXT_MFES
RP_Local_MapsMenu_Resturants_1x1
RTO_SERVER_RESTART sets the recovery time goal that the engine uses to
autonomically adjust the CKPTINTVL and LRU_MIN/MAX_DIRTY settings. It
informs the engine that in the event of a hard crash you want the instance
to be able to recover in RTO_SERVER_RESTART seconds or less. The shorter
this value the more aggressively the engine has to checkpoint and flush data
during and between checkpoints. I have a couple of clients that have set
this without adverse effect on production performance, however, I would warn
against making it too aggressive. Setting it to a 10 second recovery time
might be bad.
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
On Mon, Nov 2, 2009 at 2:36 PM, Clifton Bean <clifton_bean@hotmail.com>wrote:
> IDS 11.50.FC5 AIX 5.3
>
> While reviewing the $ONCONFIG parameters available on this instance, I came
> across this new parameter:
>
> ## Recovery Time Objective Server Restart
> RTO_SERVER_RESTART 0 # Number of Seconds, Overrides Checkpoint Interval>
> ## Shared Memory Space Parameter
> CKPTINTVL 60 # Check point interval (in sec)>
> Has anyone played with this parameter yet? Any downside to setting this
> parameter?
>
> Initial setting: same as the "now-defunct (overriden)" CKPTINTVL?
>
> Thanks in advance.
>
> Clifton
>
> _________________________________________________________________
> Bing brings you maps, menus, and reviews organized in one place.
>
>
>
http://www.bing.com/search?q=restaurants&form=MFESRP&publ=WLHMTAG&crea=TEXT_MFES
RP_Local_MapsMenu_Resturants_1x1
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--0023545bd64cb960880477696842
I set this parameter when we migrated to 11.50 and am very happy so far.
RTO_SERVER_RESTART is the number of seconds you want fast recovery to take.
IDS will make an educated guess as to how long fast recovery will take by
keeping track of things like how fast you can write to your logical and
physical logs (kept persistent via RAS_PLOG_SPEED and RAS_LLOG_SPEED in the
ONCONFIG) and how much data you are writing to these logs.
When the amount of time IDS estimates fast recovery will take reaches
RTO_SERVER_RESTART it will fire a checkpoint.
You can monitor the currently estimated time to perform fast recovery via
onstat -g ckp
</home/informix> onstat -g ckp
IBM Informix Dynamic Server Version 11.50.FC3 -- On-Line (Prim) -- Up 68 days
02:38:53 -- 3740228 Kbytes
AUTO_CKPTS=On RTO_SERVER_RESTART=90 seconds Estimated recovery time 43 seconds
The idea is that you only issue a checkpoint when you need to in order to
maintain a recovery time objective (RTO) instead of always every X secconds.
The result is fewer checkpoints during low activity and more checkpoints
during high activity and a way to "guarantee" the amount of time it will take
to recover from a crash.
You can turn this feature off (RTO_SERVER_RESTART 0) or set it to a number of
seconds between 1 minute and 30 minutes. If you need to recover quickly from a
crash, set it on the lower end (I'm currently set at 90 seconds) and you will
have more checkpoints. If you don't care if it takes a while to recover from a
crash, set it on the higher end and you will have fewer checkpoints (and
probably better throughput).
It is also a good idea to enable AUTO_LRU_TUNING with this feature. If during
a checkpoint the engine realizes it would not be able to meet the RTO time in
the event of a crash it will tune your LRUMIN/LRUMAX to flush the LRU queues
to disk more agressively in between checkpoints.
Hope this helps,
Andrew
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g