Re: Anybody seen a very fast wild cat out there?
Posted in 2007
Discussion of IDS 11 ("Cheetah") automatic checkpoints driven by a target recovery time (RTO). The poster asked how the server can predict how long recovery will take, since identical workloads recover at different speeds on different machines, and wanted the formula documented and visible via onstat/SMI. Answers: the estimate is dynamic, based on statistics gathered from the running instance (buffer cleaning rates per disk, (K)AIO rates, etc.), with recovery roughly proportional to the original work time; onstat -g iof and the new onstat -g ckp expose some of this. The exact formula was never given — the poster was pointed to IBM's Cheetah developerWorks forum, so no resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Backup & Restore, Storage & Space Management, Logging & Checkpoints
> > Also how can does it determine how much activity can be performed
> > in the time period? That must depend on how busy other processes are
> > on the host
>
> It's not so much the amount of activity, but the amount which we need to
> recover. If we sense that you want to recover the server in 5 min in
> all cases, and that means that we need to issue a checkpoint after 10
> min., then we do so. Again, since we now have non-blocking checkpoints,
> then the cost of a checkpoint is not a big expense.
>
How do you calculate the time taken to recover for a given amount
of stuff to recover?
What stuff does the formulae behind this use?
Different machines will recover the same amount of stuff in
different times!
>
>
> > and (for me) how busy the SAN is from activity from other hosts
> > (yes 2 hosts can share a disk array!).
>
> > PS SHMVIRTALLOCSEG having an alarm feature is cool!
>
> > Also ontape to directories (finally!) is cool but the ability to list
> > multiple (colon seperated, like DBSPACETEMP) directories and have
> > individual volumes to always go to the directory that is on the
> > filesystem with the most free space would be nice. I have a HUGE
> > script to
> > do this with ontape at present and it takes a LOT of testing! (See I
> > should have been in the closed beta!).
>
> > Finally the ability to limit the amount of memory used by an
> > individual session would be nice, again with an alarm). If I have a
> > server with 64Gb of RAM
> > and a sessions goes rogue I don't want it allocating >60GB of RAM!- Hide quoted text -
>
> - Show quoted text -
david@smooth1.co.uk wrote: >>> Also how can does it determine how much activity can be performed >>> in the time period? That must depend on how busy other processes are >>> on the host >> It's not so much the amount of activity, but the amount which we need to >> recover. If we sense that you want to recover the server in 5 min in >> all cases, and that means that we need to issue a checkpoint after 10 >> min., then we do so. Again, since we now have non-blocking checkpoints, >> then the cost of a checkpoint is not a big expense. >> > > How do you calculate the time taken to recover for a given amount > of stuff to recover? > > What stuff does the formulae behind this use? > > Different machines will recover the same amount of stuff in > different times! AFAIK, the calculation is based on statistics collected during engine activity, like buffer cleans per disk, (K)AIO rate etc. It's not a static value. It will be based on measures taken on the actual system and Informix instance. Regards. -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently...
david@smooth1.co.uk wrote:
>>> Also how can does it determine how much activity can be performed
>>> in the time period? That must depend on how busy other processes are
>>> on the host
>> It's not so much the amount of activity, but the amount which we need to
>> recover. If we sense that you want to recover the server in 5 min in
>> all cases, and that means that we need to issue a checkpoint after 10
>> min., then we do so. Again, since we now have non-blocking checkpoints,
>> then the cost of a checkpoint is not a big expense.
>>
>
> How do you calculate the time taken to recover for a given amount
> of stuff to recover?
>
> What stuff does the formulae behind this use?
>
> Different machines will recover the same amount of stuff in
> different times!
But generally the recovery time will be in proportion to the amount of
time that they did their original work.
>
>
>
>>
>>> and (for me) how busy the SAN is from activity from other hosts
>>> (yes 2 hosts can share a disk array!).
>>> PS SHMVIRTALLOCSEG having an alarm feature is cool!
>>> Also ontape to directories (finally!) is cool but the ability to list
>>> multiple (colon seperated, like DBSPACETEMP) directories and have
>>> individual volumes to always go to the directory that is on the
>>> filesystem with the most free space would be nice. I have a HUGE
>>> script to
>>> do this with ontape at present and it takes a LOT of testing! (See I
>>> should have been in the closed beta!).
>>> Finally the ability to limit the amount of memory used by an
>>> individual session would be nice, again with an alarm). If I have a
>>> server with 64Gb of RAM
>>> and a sessions goes rogue I don't want it allocating >60GB of RAM!- Hide quoted text -
>> - Show quoted text -
>
>
On 17 Feb, 02:28, Fernando Nunes <s...@domus.online.pt> wrote:
> d...@smooth1.co.uk wrote:
> >>> Also how can does it determine how much activity can be performed
> >>> in the time period? That must depend on how busy other processes are
> >>> on the host
> >> It's not so much the amount of activity, but the amount which we need to
> >> recover. If we sense that you want to recover the server in 5 min in
> >> all cases, and that means that we need to issue a checkpoint after 10
> >> min., then we do so. Again, since we now have non-blocking checkpoints,
> >> then the cost of a checkpoint is not a big expense.
>
> > How do you calculate the time taken to recover for a given amount
> > of stuff to recover?
>
> > What stuff does the formulae behind this use?
>
> > Different machines will recover the same amount of stuff in
> > different times!
>
> AFAIK, the calculation is based on statistics collected during engine activity, like buffer cleans per disk, (K)AIO rate etc.
> It's not a static value. It will be based on measures taken on the actual system and Informix instance.
>
Yes, can we get these in the documentation? And a way to view all
this information via onstat/SMI?
> Regards.
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...- Hide quoted text -
>
> - Show quoted text -
david@smooth1.co.uk wrote:
>
> Yes, can we get these in the documentation? And a way to view all
> this information via onstat/SMI?
Check onstat -g iof and onstat -g ckp (new option). I believe (may be wrong), that they contain much of this information.
Not sure if it's in the format you'd like...
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
On 18 Feb, 18:41, Fernando Nunes <s...@domus.online.pt> wrote:
> d...@smooth1.co.uk wrote:
>
> > Yes, can we get these in the documentation? And a way to view all
> > this information via onstat/SMI?
>
> Check onstat -g iof and onstat -g ckp (new option). I believe (may be wrong), that they contain much of this information.
> Not sure if it's in the format you'd like...
>
What this the formulae used on these figures to calculate when a
checkpoint needs to be done?
> Regards.
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...
david@smooth1.co.uk wrote:
> On 18 Feb, 18:41, Fernando Nunes <s...@domus.online.pt> wrote:
>> d...@smooth1.co.uk wrote:
>>
>>> Yes, can we get these in the documentation? And a way to view all
>>> this information via onstat/SMI?
>> Check onstat -g iof and onstat -g ckp (new option). I believe (may be wrong), that they contain much of this information.
>> Not sure if it's in the format you'd like...
>>
>
> What this the formulae used on these figures to calculate when a
> checkpoint needs to be done?
I don't have that level of detail.
You can try to ask at the Cheetah Forum at: http://www.ibm.com/developerworks/forums/dw_forum.jsp?forum=1071&cat=19
... or wait that someone answers here.
I don't know if this information will be provided, but the forum is constantly monitored by support staff. They can get you an official answer I hope.
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g