Re: Checkpoints - a mystery
Posted in 1999
Sean Kelsey wrote:
>
> Sorry if I have already posted this once already...we are having problems
> with our news server...
>
> HP-UX 10.20 running IDS 7.24.UC5
>
> We have been having isolated problems with checkpoints - see below. During
> the lengthy checkpoints, obviously no users can perform any update routines
> or even log in to the server ( since the process of logging in also writes
> an audit trail to an activity log within the database ). Also, performing
> "onstat -u" does not show any users in a critical section holding everything
> up.
>
> I know we have a fairly large Physical Buffer file but this hardly gets
> filled...
>
> The server is not really used as a work horse and limited resource is used
> by the Informix online instance. Not a lot of work goes on this server at
> all.
>
> Do any of you guys have any ideas what is going on ???
OK here is what I see. You have 30 LRUS, and I presume several chunks,
but only one CLEANER thread allocated. With a single cleaner thread
the checkpoint thread will gather dirty pages for a single chunk then
launch the cleaner thread with the list of pages to flush. Then it
will go back to scanning. When the next list of buffers is ready if
the cleaner thread is still busy the checkpoint thread will suspend
itself to the ready queue where it will have to wait for any non-update
queries that are pending on CPU VP#1 where the checkpoint thread always
runs. You only have one CPU VP so ALL of the pending queries from the
beginning of the checkpoint have been waiting while the checkpoint
thread hogged the CPU VP. Now that the checkpoint thread is at the
bottom of the ready queue it must wait for all of them to clear the
queue ahead of it. So the checkpoint is waiting for queries.
Solution: Increase the number of CLEANERS so that the checkpoint thread
never has to wait for a cleaner to become free. I use CLEANERS=LRUS
as a minimum and CLEANERS=min(128,#chunks) as a maximum (yes Virginia I
have servers with over 230 chunks). [Actually I cheat and set:
LRUS=CLEANERS=127 on production servers (not 128 because of a harmless
but annoying bug that reports that both my logical and physical log
buffers are zero length when the engine starts up if I use 128).]
I do not see anything else glaring. If this does not work post:
onstat -FRpdDl
taken shortly before a checkpoint is due.
Art S. Kagel