Re: Long Checkpoints
Posted in 2000
Topics: Performance & Tuning, Storage & Space Management, Triggers, Constraints & Referential Integrity, Logging & Checkpoints
Unfortunately the original message from Ulrich Eckhardt has not come through to me yet, so I hope you are reading Ulrich. :-) mars1972@my-deja.com wrote: > > In article <38BE8308.52D4B593@americasm01.nt.com>, > Rudy Fernandes <rferdy@americasm01.nt.com> wrote: > > mars1972@my-deja.com wrote: > > > > > What is your checkpoint interval? Also, what is the real interval > > > between checkpoints? You could be hitting the checkpoint interval > or > > > getting the physical log to 75% before the LRUs are getting to 40%. > How > > > many cleaners and physical disks do you have? > > > > > > I usually set LRU max dirty to 2 and LRU min dirty to 1. The trick > to > > > tuning checkpoints is that you want the physical log to come as > close as > > > possible to 75% at the time that the checkpoint interval expires > > > > Why would you want to do this? While a 75% full Physical log does > trigger a > > checkpoint, the duration of checkpoint should not have anything to do > with > > what's in the physical log. The checkpoint duration is, primary, > governed > > by the number of dirty buffers at checkpoint time, how fast your I/O > > subsystem is, how well you use the I/O subsystem (avoid contention) > and cpu > > speed (buffers are organized chunk-wise). > > > > Could you explain your logic further? > > > > Rudy > > Say you get your checkpoints down to 1 second. If the physical log > fills every 30 seconds, in a 5 minute interval, you would be locking the > database for 10 seconds total (10 checkpoints at 1 second each). If you > can get the physical log to fill at 5 minutes, the you only are going to > be locking the database for 1 second in a 5 minute interval. > > In that part, I guess I was looking at overall checkpoint waits, not > just one specific checkpoint. I don't subscribe to the idea of matching your physical log size to your average work load. What happens when your work load is a bit high? You have performance problems. I always create my physical log too big so that an increase in work shouldn't increase the frequency of checkpoints. So I waste a bit of disk, but disk is cheap compared to performance. But then I do have to tune for users who love to mix OLTP access with batch jobs. Then they wonder why there is a performance problem. x-) The key here is how frequent are your checkpoints, and what is disk activity like. It is probably that you have a disk contention or throughput problem. Have you tuned the number of page cleaners in line with the number of LRU queues and disks? Are your chunks created across disks and not down each disk in turn? Do you have high LRU contention? Cheers, -- Mark. +----------------------------------------------------------+-----------+ | Mark D. Stock mailto:mdstock@mydas.freeserve.co.uk |//////// /| | http://www.informix.com http://www.informixhandbook.com |///// / //| | http://www.iiug.org +-----------------------------------+//// / ///| | |What year 2000 bug? year 2000 bug? |/// / ////| | |year 2000 bug? year 2000 bug? year |// / /////| | |2000 bug? year 2000 bug? year 1900 |/ ////////| +----------------------+-----------------------------------+-----------+
"Mark D. Stock" wrote:
>
> Unfortunately the original message from Ulrich Eckhardt has not come
> through to me yet, so I hope you are reading Ulrich. :-)
>
[..]
> >
> > Say you get your checkpoints down to 1 second. If the physical log
> > fills every 30 seconds, in a 5 minute interval, you would be locking the
> > database for 10 seconds total (10 checkpoints at 1 second each). If you
> > can get the physical log to fill at 5 minutes, the you only are going to
> > be locking the database for 1 second in a 5 minute interval.
> >
> > In that part, I guess I was looking at overall checkpoint waits, not
> > just one specific checkpoint.
>
> I don't subscribe to the idea of matching your physical log size to your
> average work load. What happens when your work load is a bit high? You
> have performance problems.
>
> I always create my physical log too big so that an increase in work
> shouldn't increase the frequency of checkpoints. So I waste a bit of
> disk, but disk is cheap compared to performance. But then I do have to
> tune for users who love to mix OLTP access with batch jobs. Then they
> wonder why there is a performance problem. x-)
>
> The key here is how frequent are your checkpoints, and what is disk
> activity like. It is probably that you have a disk contention or
> throughput problem. Have you tuned the number of page cleaners in line
> with the number of LRU queues and disks? Are your chunks created across
> disks and not down each disk in turn? Do you have high LRU contention?
>
> Cheers,
> --
> Mark.
>
> +----------------------------------------------------------+-----------+
> | Mark D. Stock mailto:mdstock@mydas.freeserve.co.uk |//////// /|
> | http://www.informix.com http://www.informixhandbook.com |///// / //|
> | http://www.iiug.org +-----------------------------------+//// / ///|
> | |What year 2000 bug? year 2000 bug? |/// / ////|
> | |year 2000 bug? year 2000 bug? year |// / /////|
> | |2000 bug? year 2000 bug? year 1900 |/ ////////|
> +----------------------+-----------------------------------+-----------+
Hi,
i don't think that the disks are the problem. It's all on a single
raid 5 array (one physical disk but with fast harddrives).
Sar gives the following output during a period of high checkpoint
intervalls :
gretel:/var/log/sa # sar -b -s 08:41:00 -e 08:55:00 -f sa02
Linux 2.2.14 (gretel) 03/02/00
08:41:38 tps rtps wtps bread/s bwrtn/s
08:43:38 6.65 1.43 5.22 7.81 14.55
08:45:38 34.86 19.33 15.53 152.61 98.66
08:47:38 75.85 26.95 48.89 215.41 377.38
08:49:38 92.80 54.91 37.88 425.28 279.51
08:51:38 55.11 23.50 31.60 161.11 218.66
08:53:38 11.92 2.98 8.94 21.31 28.18
08:55:38 13.34 3.68 9.65 24.56 32.31
Average: 46.20 21.52 24.68 163.92 169.49
Which should be OK for our harddrives.
So i will play a bit with the LRU cleaners and the checkpoint
intervall (currently set to 10 Minutes).
Does the number of LRU queues affects the performance on a
single processor system ?
Uli
--
Ulrich Eckhardt Tr@nscom
http://people.frankfurt.netsurf.de/uli http://www.transcom.de
Lagerstraße 11-15 A8
64807 Dieburg Germany
Ulrich Eckhardt wrote:
>
> "Mark D. Stock" wrote:
> > I don't subscribe to the idea of matching your physical log size to your
> > average work load. What happens when your work load is a bit high? You
> > have performance problems.
> >
> > I always create my physical log too big so that an increase in work
> > shouldn't increase the frequency of checkpoints. So I waste a bit of
> > disk, but disk is cheap compared to performance. But then I do have to
> > tune for users who love to mix OLTP access with batch jobs. Then they
> > wonder why there is a performance problem. x-)
> >
> > The key here is how frequent are your checkpoints, and what is disk
> > activity like. It is probably that you have a disk contention or
> > throughput problem. Have you tuned the number of page cleaners in line
> > with the number of LRU queues and disks? Are your chunks created across
> > disks and not down each disk in turn? Do you have high LRU contention?
> >
> > Cheers,
> > --
> > Mark.
> >
> > +----------------------------------------------------------+-----------+
> > | Mark D. Stock mailto:mdstock@mydas.freeserve.co.uk |//////// /|
> > | http://www.informix.com http://www.informixhandbook.com |///// / //|
> > | http://www.iiug.org +-----------------------------------+//// / ///|
> > | |What year 2000 bug? year 2000 bug? |/// / ////|
> > | |year 2000 bug? year 2000 bug? year |// / /////|
> > | |2000 bug? year 2000 bug? year 1900 |/ ////////|
> > +----------------------+-----------------------------------+-----------+
>
> Hi,
>
> i don't think that the disks are the problem. It's all on a single
> raid 5 array (one physical disk but with fast harddrives).
>
> Sar gives the following output during a period of high checkpoint
> intervalls :
> gretel:/var/log/sa # sar -b -s 08:41:00 -e 08:55:00 -f sa02
> Linux 2.2.14 (gretel) 03/02/00
>
> 08:41:38 tps rtps wtps bread/s bwrtn/s
> 08:43:38 6.65 1.43 5.22 7.81 14.55
> 08:45:38 34.86 19.33 15.53 152.61 98.66
> 08:47:38 75.85 26.95 48.89 215.41 377.38
> 08:49:38 92.80 54.91 37.88 425.28 279.51
> 08:51:38 55.11 23.50 31.60 161.11 218.66
> 08:53:38 11.92 2.98 8.94 21.31 28.18
> 08:55:38 13.34 3.68 9.65 24.56 32.31
> Average: 46.20 21.52 24.68 163.92 169.49>
> Which should be OK for our harddrives.
>
> So i will play a bit with the LRU cleaners and the checkpoint
> intervall (currently set to 10 Minutes).
>
> Does the number of LRU queues affects the performance on a
> single processor system ?
>
> Uli
Hi,
it seems that lowering the LRU Max-Min values helps very much. Also
i have now only one LRU queue which gives also some improvements.
Thanks for the help
Uli
--
Ulrich Eckhardt Tr@nscom
http://people.frankfurt.netsurf.de/uli http://www.transcom.de
Lagerstraße 11-15 A8
64807 Dieburg Germany