Re: Problem with Long Checkpoints
Posted in 1999
Topics: Performance & Tuning, Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Triggers, Constraints & Referential Integrity, Logging & Checkpoints, Platform-Specific Issues
Jeffrey Screws schrieb in Nachricht <7j19sh$npv$1@news.xmission.com>...
>
>Hello,
>
>I have a problem with long checkpoint durations on some of my systems.
>Let me describe my environment.
>
>HW Environment:
>
>7 HP 9000 K460 boxes w/
>HP-UX 10.20
>768mb or 1 gig of ram on each box
>16 to 24 - 2.1 gb disks on each box depending on the number of Informix
>instances. Each instance has 10 dbspaces on only 2 mirrored pairs(4
>disks in total). I'm sure this is at least part of the problem. But
>I'm sure there is some tuning the can help?
>
>DB Environment
>
>1 to 4 instances on each box. The database is identical on each
>instance, (except for the contents of course!), and is roughly 4 gb is
>total size. Theses are by no stretch large databases! All storage is
>mirrored using HP LVM mirroring. Each instances probably has an
>average of 30 users hitting them 24x7. The application runs on the
>box also, so each user is telnetting in and running the 4GL app locally
>through shared memory connections.
>
>Here is the problem. On the best performing db instance the checkpoint
>duration is 2 or 3 seconds at every CKPTINTVL. On the worst
>performing instance the duration is usually around 8 or 9 seconds and
>sometimes up to 25 seconds. I am including all the info that I saw Art
>Kagel ask of someone a few days ago. Hope this isn't tooooooo much
>info.
[... snip ...]
>CKPTINTVL 3000 # Check point interval (in sec)
This is maybe worth a closer look.
Depending on what your users are doing, maybe it is better
to reduce CKPTINTVL to 600 or even 300 seconds. This would mean
more but hopefully shorter checkpoints.
Your users may or may not prefer shorter still stands, but more often,
instead of having to wait 25 seconds sometimes
I did not notice that any premature checkppoint occurred after 04:07:22,
triggered
by physlog more than 75% full.
But you are maybe close, because your onstat -l listing shows physlog is
62% used.
You can check this with onstat -l ~ 48-49 minutes after a checkpoint
I use this sequence
onstat -m gives me the time of last checkpointdate
... wait ... wait .... repeat date ... wait ...
onstat -l
The 62% usage of physlog in your listing means that you write 6250+ pages
during
the next checkpoint - a lot of pages for a 5400 rpm spindle.
If you have an opportunity to test this, use
$ time dd if=/dev/zero of=/dev/<a-chunk>
bs=<your-IFX_blocksize_probably_2048> count=6250
[courtesy of Mr. Stefan Weideneder, see: www.weideneder.de ]
BUT THIS DESTROYS WHAT RESIDES ON /dev/<a-chunk> !!!
So DO NOT USE a production chunk for this !!
From your onstat -l output I can see that you have definitely a lot
of activity going: notice the logfile with the 'L'-flag (= last checkpoint)
is
LogID 40976 and the one holding the 'C'-flag (current log) is 40981.
This means 23000 - 28000 pages of logging since your
checkpoint. I would try to go for a smaller CKPTINTVL
Good luck
Dick
----------------
Richard Kofler
debis Systemhaus EDVg - Vienna
posting on a private account
In article <Om_43.6822$Hm1.8930@news.chello.at>, "nicole kofler" <nikis.atelier@teleweb.at> wrote: [... snip ...] > The 62% usage of physlog in your listing means that you write 6250+ > pages during the next checkpoint - a lot of pages for a 5400 rpm > > spindle. As I understand, the flushing of the physical log does not involve reading the blocks from it and writing them to somewhere else (it contains only before-image pages), *unless* you crash and have to do a fast recovery. At checkpoint time it's just marked empty. [ ... snip ... ] > Good luck > Dick > ---------------- > Richard Kofler > debis Systemhaus EDVg - Vienna > posting on a private account > > Cheers -- Gabor Heppes IBM Global Services gaborh@au1.ibm.com Sent via Deja.com http://www.deja.com/ Share what you know. Learn what you don't.