Re: Long Checkpoints
Posted in 2000
A user on Informix 7.30 UC10 under Linux saw checkpoints lasting 15-25 seconds with 2000+ dirty buffers, and asked whether writing ~4MB really should take that long or if tuning could help. Replies suggested checking CKPTINTVL versus the actual checkpoint interval, number of page cleaners and physical disks, and setting LRU max/min dirty very low (e.g. 2/1) so few buffers are dirty at checkpoint time. A debate followed over whether to let the physical log trigger checkpoints: one poster aimed to have the log hit 75% just as the interval expires to reduce total checkpoint time, while others argued checkpoint duration depends on dirty buffer count, I/O speed and CPU, and that the physical log should simply be made large enough not to trigger checkpoints. No single confirmed fix for the original poster is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Logging & Checkpoints
In article <38BE266E.C164B425@transcom.de>,
Ulrich Eckhardt <Ulrich.Eckhardt@transcom.de> wrote:
> Hi,
>
> i have a problem with long checkpoint runs on an informix 7.30 UC 10
> running on linux. (Machine is a PIII 500 with 512 MB Memory).
>
> Sometimes checkpoints takes 15 to 25 seconds. During this
> times i have approximately 2000 to 2300 dirty LRU queues.
>
> an onstat -FR (with only 205 dirty, but Fg Writes and LRU Writes
> are also 0 with 2000 dirty).
>
> Fg Writes LRU Writes Chunk Writes
> 0 0 96227
>
> address flusher state data
> 154f24c8 0 I 0 = 0X0
> states: Exit Idle Chunk Lru
>
> 4 buffer LRU queue pairs priority levels
> # f/m pair total % of length LOW MED_LOW MED_HIGH HIGH
> 0 F 9984 99.5% 9935 0 9757 169 9
> 1 m 0.5% 49 0 49 0 0
> 2 f 9991 99.4% 9936 0 9758 171 7
> 3 m 0.6% 55 0 55 0 0
> 4 f 9993 99.5% 9944 0 9802 139 3
> 5 m 0.5% 49 0 49 0 0
> 6 f 9991 99.5% 9939 0 9753 180 6
> 7 m 0.5% 52 0 52 0 0
> 205 dirty, 39959 queued, 40000 total, 65536 hash buckets, 2048 buffer
> size
> start clean at 40% (of pair total) dirty, or 4000 buffs dirty, stop at
> 20%
>
> One thing wich i will try is to decrease the physical log and change
> the LRU max-min parameters to halve the size.
>
> But writing back 2000 queues makes approximately 4 MB of data. Could
> this really take 20 seconds to write back ? or can i tune some
> other parameters to get here a better performance ?.
>
> Thanks for help
> Uli
> --
> Ulrich Eckhardt Tr@nscom
> http://people.frankfurt.netsurf.de/uli http://www.transcom.de
> Lagerstra'e 11-15 A8
> 64807 Dieburg Germany
>
What is your checkpoint interval? Also, what is the real interval
between checkpoints? You could be hitting the checkpoint interval or
getting the physical log to 75% before the LRUs are getting to 40%. How
many cleaners and physical disks do you have?
I usually set LRU max dirty to 2 and LRU min dirty to 1. The trick to
tuning checkpoints is that you want the physical log to come as close as
possible to 75% at the time that the checkpoint interval expires
(within a couple of seconds is best). If you can do that and have the
LRU queues fairly empty when it happens, your checkpoints will decrease
quite a bit.
--
# unrm /
ksh: unrm: not found
# man cpio
Sent via Deja.com http://www.deja.com/
Before you buy.
mars1972@my-deja.com wrote: > What is your checkpoint interval? Also, what is the real interval > between checkpoints? You could be hitting the checkpoint interval or > getting the physical log to 75% before the LRUs are getting to 40%. How > many cleaners and physical disks do you have? > > I usually set LRU max dirty to 2 and LRU min dirty to 1. The trick to > tuning checkpoints is that you want the physical log to come as close as > possible to 75% at the time that the checkpoint interval expires Why would you want to do this? While a 75% full Physical log does trigger a checkpoint, the duration of checkpoint should not have anything to do with what's in the physical log. The checkpoint duration is, primary, governed by the number of dirty buffers at checkpoint time, how fast your I/O subsystem is, how well you use the I/O subsystem (avoid contention) and cpu speed (buffers are organized chunk-wise). Could you explain your logic further? Rudy
In article <38BE8308.52D4B593@americasm01.nt.com>, Rudy Fernandes <rferdy@americasm01.nt.com> wrote: > mars1972@my-deja.com wrote: > > > What is your checkpoint interval? Also, what is the real interval > > between checkpoints? You could be hitting the checkpoint interval or > > getting the physical log to 75% before the LRUs are getting to 40%. How > > many cleaners and physical disks do you have? > > > > I usually set LRU max dirty to 2 and LRU min dirty to 1. The trick to > > tuning checkpoints is that you want the physical log to come as close as > > possible to 75% at the time that the checkpoint interval expires > > Why would you want to do this? While a 75% full Physical log does trigger a > checkpoint, the duration of checkpoint should not have anything to do with > what's in the physical log. The checkpoint duration is, primary, governed > by the number of dirty buffers at checkpoint time, how fast your I/O > subsystem is, how well you use the I/O subsystem (avoid contention) and cpu > speed (buffers are organized chunk-wise). > > Could you explain your logic further? > > Rudy > > Say you get your checkpoints down to 1 second. If the physical log fills every 30 seconds, in a 5 minute interval, you would be locking the database for 10 seconds total (10 checkpoints at 1 second each). If you can get the physical log to fill at 5 minutes, the you only are going to be locking the database for 1 second in a 5 minute interval. In that part, I guess I was looking at overall checkpoint waits, not just one specific checkpoint. -- # unrm / ksh: unrm: not found # man cpio Sent via Deja.com http://www.deja.com/ Before you buy.
mars1972@my-deja.com wrote: > Say you get your checkpoints down to 1 second. If the physical log > fills every 30 seconds, in a 5 minute interval, you would be locking the > database for 10 seconds total (10 checkpoints at 1 second each). If you > can get the physical log to fill at 5 minutes, the you only are going to > be locking the database for 1 second in a 5 minute interval. > > In that part, I guess I was looking at overall checkpoint waits, not > just one specific checkpoint. I see your point. In general, however, the physical log should be large enough so that it isn't even getting close to triggering checkpoints, under normal circumstances. The downside to increasing the size of the physical log is only that of space. Rudy
In article <89m51m$o0p$1@nnrp1.deja.com>, mars1972@my-deja.com writes: > > Say you get your checkpoints down to 1 second. If the physical log > fills every 30 seconds, in a 5 minute interval, you would be locking the > database for 10 seconds total (10 checkpoints at 1 second each). If you > can get the physical log to fill at 5 minutes, the you only are going to > be locking the database for 1 second in a 5 minute interval. > If you have checkpoints of 1 second every 30 seconds you get 10 seconds total checkpoint duration. But if collecting dirty pages for 5 minutes I think your I/O-system needs about these 10 seconds to write the dirty pages out. As far as I remember there was a article in Tech-notes where the author suggested triggering checkpoints through physical log size. You get more or less constant (and controllable) checkpoint duration, because you have the same number of pages to write out. What are the risk of doing so? I see no problem for the database filling up physical log. Tommi Mäkitalo