Long checkpoints causing crashes
Posted in 2000
Topics: Performance & Tuning, Logging & Checkpoints, Platform-Specific Issues, Clustering, Grid & MACH11, Versions, Editions & End-of-Life
Hi There
Running IDS 7.30.UC7 on HP-UX 10.20, and the server is
now hanging most evenings. This seems to be due to
long checkpoints, you often see one checkpoint of 80
seconds or so, immediatley followed by a 2nd
checkpoint of 0-11 seconds or more, followed by
Informix Dynamic Server Stopped.
Fortunately, this is within a service guard cluster,
so it just hops across to another machine and starts
again. Currently looking at how to utilise the 6 CPU's
better, but does anyone have any ideas, answers, or is
this simply a risk that you have to accept with very
long checkpoints?
LRU_MIN+MAX_DIRTY have now been reduced to 0 and 1,
but I personally think this is pretty pointless, there
can't be much difference with 5 and 10, as the buffers
will only take seconds to reach either 10% or 1%
dirty.
I've heard an idea to reduce the number of buffers,
saying this will lower checkpoint times and reduce the
problem. This is from the same person who had BUFFERS
set to 30000 when I first arrived, with tons of
ovbufs, bufwaits, and pathetic read and write caching,
I have now set BUFFERS to 180000 on this 3 gig phys
memory system, still below the 20-25% recomended. Note
any advice for BUFFERS setting on a machine that
shares informix with another application? onstat -p
shows much better performance now, but could it
actually be a good idea to cripple the system with a
crap number of BUFFERS, just to avoid stopping the
system with long checkpoints?
Final note: I thought a core and shm dump may be a
good idea for tech support to analyse exactly why
informix is stopping, but I don't see how this is
possible, as it is not an assertion failure.
Dom
__________________________________________________
Do You Yahoo!?
Talk to your friends online with Yahoo! Messenger.
http://im.yahoo.com
In article <880i7j$iuo$1@news.xmission.com>,
Dominic Penfold <bigbaddom_the_first@yahoo.com> wrote:
>
> Hi There
>
> Running IDS 7.30.UC7 on HP-UX 10.20, and the server is
> now hanging most evenings. This seems to be due to
> long checkpoints, you often see one checkpoint of 80
> seconds or so, immediatley followed by a 2nd
> checkpoint of 0-11 seconds or more, followed by
> Informix Dynamic Server Stopped.
>
Ouch.
> Fortunately, this is within a service guard cluster,
> so it just hops across to another machine and starts
> again. Currently looking at how to utilise the 6 CPU's
> better, but does anyone have any ideas, answers, or is
> this simply a risk that you have to accept with very
> long checkpoints?
Hmmm... Looking at the wrong problem. You shouldn't be accepting the
long checkpoints. While a checkpoint is firing, you can't access the
database. Checkpoints should take between 0 and 2 seconds consistantly.
>
> LRU_MIN+MAX_DIRTY have now been reduced to 0 and 1,
> but I personally think this is pretty pointless, there
> can't be much difference with 5 and 10, as the buffers
> will only take seconds to reach either 10% or 1%
> dirty.
How fast they fill isn't the issue, it's when the data in them gets
flushed to disk. The less data in the LRU queue, the faster the
checkpoint will be.
>
> I've heard an idea to reduce the number of buffers,
> saying this will lower checkpoint times and reduce the
> problem. This is from the same person who had BUFFERS
> set to 30000 when I first arrived, with tons of
> ovbufs, bufwaits, and pathetic read and write caching,
> I have now set BUFFERS to 180000 on this 3 gig phys
> memory system, still below the 20-25% recomended. Note
> any advice for BUFFERS setting on a machine that
> shares informix with another application? onstat -p
> shows much better performance now, but could it
> actually be a good idea to cripple the system with a
> crap number of BUFFERS, just to avoid stopping the
> system with long checkpoints?
>
Decreasing your buffers shouldn't help your checkpoints. A good way to
figure out how much memory you can give informix is to find out how much
memory the second application takes at the most, subtract that from the
total memory on the box minus say 5-10% of total memory.
total memory = 1GB
second app takes 250 MB
5% = 50MB
Informix can have 700MB in this instance.
> Final note: I thought a core and shm dump may be a
> good idea for tech support to analyse exactly why
> informix is stopping, but I don't see how this is
> possible, as it is not an assertion failure.
>
> Dom
>
How large is your physical log? How many Cleaners do you have? How
many physical disks? How many CPUVPS? Actually, why don't you just
reply with a copy of your onconfig.
> __________________________________________________
> Do You Yahoo!?
> Talk to your friends online with Yahoo! Messenger.
> http://im.yahoo.com
>
--
# unrm /
ksh: unrm: not found
# man cpio
Sent via Deja.com http://www.deja.com/
Before you buy.
In article <880i7j$iuo$1@news.xmission.com>, Dominic Penfold <bigbaddom_the_first@yahoo.com> wrote: > > Hi There > > Running IDS 7.30.UC7 on HP-UX 10.20, and the server is > now hanging most evenings. This seems to be due to > long checkpoints, you often see one checkpoint of 80 > seconds or so, immediatley followed by a 2nd > checkpoint of 0-11 seconds or more, followed by > Informix Dynamic Server Stopped. -- SNIP -- > Dom Dom, I had a similar but much worse problem with HP-UX 11. Nearly every night at ~ 2:10 AM, for nearly 2 months the system would freeze, with one checkpoint taking as much as 3 hours to complete. By writing a shotgun script that monitors informix and OS functions, I was able to deflect blame from informix, showing that it was a victim of an IO freeze (3-hours to complete a disk I/O) originating from the OS. For us, it turned out to be a bug in some package called "predictive", which dialed a number to send info via modem. Somehow, this was tying up the system bus. When the HP support engineer disabled predictive, the problem went away. I realize your symptoms are not identical to mine. But the 'every evening" bit caught my eye. So, if: - the problem occurs every night at about the same time, and - you have predictive installed then you might look into this as a possible culprit. Else NEVER MIND ;-) BTW, 11-second checkpoints are not that unusual; if you have 300,000 buffers then even a 1%-dirty ratio means 3000 dirty buffers. Maybe you can flush this many buffers (plus mirrors, if you use them) in 2-4 seconds on an EMC disk farm (I hear they hate that term! ;-) but in general, 300 buffers/second flushed sounds pretty fast to me. Another BTW: What jobs are running during that time. (No, we don't really need to know. ;) If you can reschedule half of them at a time (e.g. to run 2 hours later) you might be able to verge in on one job that is tickling some sensitivity in the engine. HTH -- +----- Jacob Salomon - DBA JSalomon@bn.com - --------------------------+ |------------------- Bulletin Board Announcement ----------------------| | Congregants will please note that the bowl at the back of the church | | bearing the sign "For the Sick" is for monetary contributions only. | +----------------------------------------------------------------------+ Sent via Deja.com http://www.deja.com/ Before you buy.