Re: Extended checkpoint times
Posted in 1998
Peter Tashkoff wrote:
> I had a problem in November of checkpoints blowing out to 10 minutes or so
> during the evening like your manifestation.
> I was advised from tech support that checkpoints can blow out two ways.
> 1. When checkpoint is requested, Online freezes all sessions for users not
> in a critical section and waits for remainder to exit critical section
> before freezing their sessions.
> A critical section is most likely to be a physical write to disk.
>
> 2. If Online is carrying out some large unusual process, like a large =
> table load, this can also cause checkpoints to blow out.
> It is possible that a disk problem could be the cause. To investigate
> this possibility, during an interval when system is frozen do the
> following;
> 1. ID the sessions in critical sections.
> 2. onstat -g sessid to ID the tables being connected to.
> 3, ID the dbspace of the tables concerned.
> 4. ID the raw partition that this relates to.
> In our case the sysadmin had been messing around with disks getting the
> machine ready for a solaris upgrade. I don't know why this might have
> affected online but the problem did go away after the disk got sorted out.
> HTH
> Peter Tashkoff <tashkop@iname.com>
> >>> anonymous masquerading as Mark said >>>
> Hi,
> We are trying to investigate a series of overly long checkpoints on one
> of our informix databases, here is
> a snippet of the online.log file detailing the events:
> 21:16:55 Checkpoint Completed: duration was 76 seconds.
> 21:22:04 Checkpoint Completed: duration was 6 seconds.
> 21:27:16 Checkpoint Completed: duration was 8 seconds.
> 21:40:15 Checkpoint Completed: duration was 476 seconds.
> 21:45:29 Checkpoint Completed: duration was 10 seconds.
> 21:50:36 Checkpoint Completed: duration was 2 seconds.
> 21:55:42 Checkpoint Completed: duration was 3 seconds.
> 22:03:22 Checkpoint Completed: duration was 155 seconds.
We see this exact situation during an archive. Occassionally during an
archive a checkpoint will last from 100->800 seconds while our normal
checkpoint duration is 5-7 seconds. Once this hung forever which
turned out to be a deadlock between two threads each of which had
acquired a latch on one of the Physical Log buffers and wanted a latch
on the other which of course was already help by the other thread this
caused a deadlock which was not detected and resolved between two
threads in critical sections stalling the checkpoint indefinitely.
I have reported both of these situations but there is as yet no
resolution nor a comprehensive explanation of what is happening in the
case that you describe where the checkpoint eventually completes.
Art S. Kagel