Re: Blocked:CKPT
Posted in 1999
Topics: Storage & Space Management, Error Codes & Troubleshooting, Server Administration, Logging & Checkpoints
Thanks Mike,
but I don't have messages in message.log like:
Assert Failed: Chunk ... is being taken OFFLINE.
Who: ...
Result: ...
Action: ...
See Also: ...
which occur when online tries to put a dbspace down.
Or maybe it is a bug in 7.30.UC6. This message does not appear
while online tries to put a dbspace down?
David...
----- Message d'origine -----
De : David Kosenko <davek@summitdata.com>
' : <informix-list@iiug.org>
Envoi' : mardi 2 f'vrier 1999 22:39
Objet : Re: Blocked:CKPT
On Tue, 02 Feb 1999 13:42:21 +0100, David Andrieu
<dandrieu@mail.atmel.fr> wrote:
>
>Hi list,
>
> Online hungs yesterday on a checkpoint. Logical logs was not full,
>and I had no message in log file.
>
>The status was:
>
>Informix Dynamic Server Version 7.30.UC6 -- On-Line (CKPT REQ) -- Up 6
>days 0s
>Blocked:CKPT
This can happen if you have ONDBSPACEDOWN set to 2 (the default) in
your onconfig file. Try changing it to 0 or 1.
Dave
On Wed, 03 Feb 1999 08:26:55 +0100, David Andrieu
<dandrieu@mail.atmel.fr> wrote:
> but I don't have messages in message.log like:
>
> Assert Failed: Chunk ... is being taken OFFLINE.
> Who: ...
> Result: ...
> Action: ...
> See Also: ...>
>which occur when online tries to put a dbspace down.
>Or maybe it is a bug in 7.30.UC6. This message does not appear
>while online tries to put a dbspace down?
No, because ONDBSPACEDOWN=2 will not mark the chunk as down - that is
the whole point behind this option. It is meant for cases where a
disk becomes temporarily unavailable (say because a SCSI cable got
unplugged). In that case, marking the chunk as down would require a
restore to recover it, even if the disk isn't really bad. A value of
2 (WAIT) means don't mark it as down, keep polling the device (in case
it becomes available again), and at the next checkpoint, block all
activity (the symptom originally reported). This is a very bad option
if your chunk really does go down - much better to either let the
server go down or markthe chunk as down and keep the server up, but
not block at the next checkpoint.
Dave
In article <36b9f2c0.15964725@nntp.best.ix.netcom.com>, David Kosenko
<davek@summitdata.com> writes
>On Wed, 03 Feb 1999 08:26:55 +0100, David Andrieu
><dandrieu@mail.atmel.fr> wrote:
>> but I don't have messages in message.log like:
>>
>> Assert Failed: Chunk ... is being taken OFFLINE.
>> Who: ...
>> Result: ...
>> Action: ...
>> See Also: ...>>
>>which occur when online tries to put a dbspace down.
>>Or maybe it is a bug in 7.30.UC6. This message does not appear
>>while online tries to put a dbspace down?
>
I have also seen this under 5 when an archive is waiting for a tape
drive. It holds a latch stopping the next checkpoint from completing.
In our case once we stopped the archive Online continued again.
Although we did have an ~11300 second checkpoint!!!
>No, because ONDBSPACEDOWN=2 will not mark the chunk as down - that is
>the whole point behind this option. It is meant for cases where a
>disk becomes temporarily unavailable (say because a SCSI cable got
>unplugged). In that case, marking the chunk as down would require a
>restore to recover it, even if the disk isn't really bad. A value of
>2 (WAIT) means don't mark it as down, keep polling the device (in case
>it becomes available again), and at the next checkpoint, block all
>activity (the symptom originally reported). This is a very bad option
>if your chunk really does go down - much better to either let the
>server go down or markthe chunk as down and keep the server up, but
>not block at the next checkpoint.
>
>Dave
>
--
David Williams