RE: Checkpoints taking too long!!
Posted in 1999
Topics: Performance & Tuning, Server Administration, Logging & Checkpoints
One thing I have learned when it comes to tuning is that the first thing I
should do is diagnose what is going on before I start changing parameters or
anything else.
If you had small checkpoint times before is becouse at that time the engine
had little work to do. Now it doesn't.
Why ?
There may be a number of reasons:
Indexes in your main applications have grown a lot and are not doing the job
==> rebuild if they are fragmented
Some module is puting a big load on the system ==> check the load with
onperf for example
Find what has changed and if you doubled your usual daily load maybe you
need to touch some parameters like the checkpoint interval max and min
limits and LRU's etc.etc.
But if things get to tight and you feel you need more time to findout,
reduce the checkpoint interval drastically untill the checkpoint duration is
tolerable. Then continue with the investigation.
> -----Original Message-----
> From: bridget [SMTP:b.vanderleek@auckland.ac.nz]
> Sent: Tuesday, March 23, 1999 9:01 PM
> To: informix-list@iiug.org
> Subject: Checkpoints taking to long!!
>
> Hi
>
> I've run into a problem with the checkpoints on my production system -
> they used to be all around 0 seconds long with the occasional 3 second
> checkpoint - now they have increased to more like 10-15 seconds average
> with a top of 20 seconds. We are running 730UC5 on a Sequent Box with 8
>
> CPU dedicated to the database and 40 gig of mirrored space.
>
> Below are the settings in my onconfig file - I have decreased the
> max_dirty and min_dirty to see if it makes any difference, it didn't.
>
> Should I increase the number of LRU's and decrease the size of the
> logs?. Any ideas would be more than welcome.
>
>
> # Shared Memory Parameters
>
> LOCKS 3000000 # Maximum number of locks
> BUFFERS 60000 # Maximum number of shared buffers
> NUMAIOVPS # Number of IO vps
> PHYSBUFF 32 # Physical log buffer size (Kbytes)
> LOGBUFF 32 # Logical log buffer size (Kbytes)> LOGSMAX 40 # Maximum number of logical log files
> CLEANERS 8 # Number of buffer cleaner processes
> SHMBASE 0x10000000 # Shared memory base address
> SHMVIRTSIZE 144000 # initial virtual shared memory segment>
> size
> SHMADD 16000 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
> 0=>unlimited
> CKPTINTVL 300 # Check point interval (in sec)
> LRUS 16 # Number of LRU queues
> LRU_MAX_DIRTY 40 # LRU percent dirty begin cleaning limit
>
> LRU_MIN_DIRTY 35 # LRU percent dirty end cleaning limit
> LTXHWM 50 # Long transaction high water mark> percentage
> LTXEHWM 60 # Long transaction high water mark
> (exclusive)
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 32 # Stack size (Kbytes)>
>
> thanks Bridget
>
>
The load on the server hasn't increased significantly - the checkpoints are
occuring at the checkpoint intervals of 5 minutes (300). I tried to change this
to 3 minutes by forcing checkpoints (onmode -c) to see if this would make a
difference - this didn't do as much as I thought, it only dropped them by 3
seconds.
Anyway thanks to everyone for the ideas I'll continue the investigation
Daniel Bonesana wrote:
> One thing I have learned when it comes to tuning is that the first thing I
> should do is diagnose what is going on before I start changing parameters or
> anything else.
> If you had small checkpoint times before is becouse at that time the engine
> had little work to do. Now it doesn't.
> Why ?
> There may be a number of reasons:
> Indexes in your main applications have grown a lot and are not doing the job
> ==> rebuild if they are fragmented
> Some module is puting a big load on the system ==> check the load with
> onperf for example
> Find what has changed and if you doubled your usual daily load maybe you
> need to touch some parameters like the checkpoint interval max and min
> limits and LRU's etc.etc.
> But if things get to tight and you feel you need more time to findout,
> reduce the checkpoint interval drastically untill the checkpoint duration is
> tolerable. Then continue with the investigation.
> > -----Original Message-----
> > From: bridget [SMTP:b.vanderleek@auckland.ac.nz]
> > Sent: Tuesday, March 23, 1999 9:01 PM
> > To: informix-list@iiug.org
> > Subject: Checkpoints taking to long!!
> >
> > Hi
> >
> > I've run into a problem with the checkpoints on my production system -
> > they used to be all around 0 seconds long with the occasional 3 second
> > checkpoint - now they have increased to more like 10-15 seconds average
> > with a top of 20 seconds. We are running 730UC5 on a Sequent Box with 8
> >
> > CPU dedicated to the database and 40 gig of mirrored space.
> >
> > Below are the settings in my onconfig file - I have decreased the
> > max_dirty and min_dirty to see if it makes any difference, it didn't.
> >
> > Should I increase the number of LRU's and decrease the size of the
> > logs?. Any ideas would be more than welcome.
> >
> >
> > # Shared Memory Parameters
> >
> > LOCKS 3000000 # Maximum number of locks
> > BUFFERS 60000 # Maximum number of shared buffers
> > NUMAIOVPS # Number of IO vps
> > PHYSBUFF 32 # Physical log buffer size (Kbytes)
> > LOGBUFF 32 # Logical log buffer size (Kbytes)> > LOGSMAX 40 # Maximum number of logical log files
> > CLEANERS 8 # Number of buffer cleaner processes
> > SHMBASE 0x10000000 # Shared memory base address
> > SHMVIRTSIZE 144000 # initial virtual shared memory segment> >
> > size
> > SHMADD 16000 # Size of new shared memory segments
> > (Kbytes)
> > SHMTOTAL 0 # Total shared memory (Kbytes).
> > 0=>unlimited
> > CKPTINTVL 300 # Check point interval (in sec)
> > LRUS 16 # Number of LRU queues
> > LRU_MAX_DIRTY 40 # LRU percent dirty begin cleaning limit
> >
> > LRU_MIN_DIRTY 35 # LRU percent dirty end cleaning limit
> > LTXHWM 50 # Long transaction high water mark> > percentage
> > LTXEHWM 60 # Long transaction high water mark
> > (exclusive)
> > TXTIMEOUT 0x12c # Transaction timeout (in sec)
> > STACKSIZE 32 # Stack size (Kbytes)> >
> >
> > thanks Bridget
> >
> >
Bridget,
There are several options you have open to you. I would start with doing the
following:
1) watch your cleaners with an onstat -F at 5 second intervals. Keep close eye on
the frequency of LRU writes. Watch the LRU write info.
2) watch the LRU % dirty values. (i.e. onstat -R)
what you might see that as your LRU dirty % reaches the LRU_MAX_DIRTY you are also
reaching your checkpoint. If this is the case worst case you have 24Megs of data
(assuming you have a 2K page size) to sort and write to disk as part of a check
point.
To resolve this you can do one of the following:
Increase LRUS, or decrease LRU_MAX and MIN. If this is an OLTP system, and you have
enough system horse power you can dramatically reduce MAX and MIN and increase
performance globally. I have even gone as low as 2 and 5 to get checkpoints to
under a second.
With your current configuration you have to have 1500 buffers dirty in any
particular LRU to reach your MAX. Depending on data this might require a lot of
updates in 5 minutes. Each percent equates to ~38 buffers.
If you look at what a checkpoint does, the greatest amount of time required is
syncing up the BUFFER pool and disk. Focus your efforts here, and you will be
pleasently supprised.
you can also modify the checkpoint interval, but that shold be a last resort, and
done for fine tunning only. As you see the LRU's reaching the MAX then a check
point hitting, you can increase the interval by a couple seconds, and decrease
checkpoint time. USE caution here, because it takes a lot of monitoring to
determine this pattern, and each change needs adequate monitoring prior to changing
again.
monitor your LRUS for at least a day (continuously) before making these changes,
that way you can get a feel for how the database is reacting, then make changes
only after adequate monitoring again. I know this is obvious, but sometimes it is
easy to think you see a pattern that is not there, and make changed that effect you
adversly.
Hope this helps,
Tony
bridget wrote:
> The load on the server hasn't increased significantly - the checkpoints are
> occuring at the checkpoint intervals of 5 minutes (300). I tried to change this
>
> to 3 minutes by forcing checkpoints (onmode -c) to see if this would make a
> difference - this didn't do as much as I thought, it only dropped them by 3
> seconds.
>
> Anyway thanks to everyone for the ideas I'll continue the investigation
>
> Daniel Bonesana wrote:
>
> > One thing I have learned when it comes to tuning is that the first thing I
> > should do is diagnose what is going on before I start changing parameters or
> > anything else.
> > If you had small checkpoint times before is becouse at that time the engine
> > had little work to do. Now it doesn't.
> > Why ?
> > There may be a number of reasons:
> > Indexes in your main applications have grown a lot and are not doing the job
> > ==> rebuild if they are fragmented
> > Some module is puting a big load on the system ==> check the load with
> > onperf for example
> > Find what has changed and if you doubled your usual daily load maybe you
> > need to touch some parameters like the checkpoint interval max and min
> > limits and LRU's etc.etc.
> > But if things get to tight and you feel you need more time to findout,
> > reduce the checkpoint interval drastically untill the checkpoint duration is
> > tolerable. Then continue with the investigation.
> > > -----Original Message-----
> > > From: bridget [SMTP:b.vanderleek@auckland.ac.nz]
> > > Sent: Tuesday, March 23, 1999 9:01 PM
> > > To: informix-list@iiug.org
> > > Subject: Checkpoints taking to long!!
> > >
> > > Hi
> > >
> > > I've run into a problem with the checkpoints on my production system -
> > > they used to be all around 0 seconds long with the occasional 3 second
> > > checkpoint - now they have increased to more like 10-15 seconds average
> > > with a top of 20 seconds. We are running 730UC5 on a Sequent Box with 8
> > >
> > > CPU dedicated to the database and 40 gig of mirrored space.
> > >
> > > Below are the settings in my onconfig file - I have decreased the
> > > max_dirty and min_dirty to see if it makes any difference, it didn't.
> > >
> > > Should I increase the number of LRU's and decrease the size of the
> > > logs?. Any ideas would be more than welcome.
> > >
> > >
> > > # Shared Memory Parameters
> > >
> > > LOCKS 3000000 # Maximum number of locks
> > > BUFFERS 60000 # Maximum number of shared buffers
> > > NUMAIOVPS # Number of IO vps
> > > PHYSBUFF 32 # Physical log buffer size (Kbytes)
> > > LOGBUFF 32 # Logical log buffer size (Kbytes)> > > LOGSMAX 40 # Maximum number of logical log files
> > > CLEANERS 8 # Number of buffer cleaner processes
> > > SHMBASE 0x10000000 # Shared memory base address
> > > SHMVIRTSIZE 144000 # initial virtual shared memory segment> > >
> > > size
> > > SHMADD 16000 # Size of new shared memory segments
> > > (Kbytes)
> > > SHMTOTAL 0 # Total shared memory (Kbytes).
> > > 0=>unlimited
> > > CKPTINTVL 300 # Check point interval (in sec)
> > > LRUS 16 # Number of LRU queues
> > > LRU_MAX_DIRTY 40 # LRU percent dirty begin cleaning limit
> > >
> > > LRU_MIN_DIRTY 35 # LRU percent dirty end cleaning limit
> > > LTXHWM 50 # Long transaction high water mark> > > percentage
> > > LTXEHWM 60 # Long transaction high water mark
> > > (exclusive)
> > > TXTIMEOUT 0x12c # Transaction timeout (in sec)
> > > STACKSIZE 32 # Stack size (Kbytes)> > >
> > >
> > > thanks Bridget
> > >
> > >