Re: Blocked:CKPT
Posted in 1999
Topics: Storage & Space Management, Error Codes & Troubleshooting, Logging & Checkpoints
Hi David,
I think you ar rigt. I changed ONDBSPACEDOWN and set it
to 1 (ABORT). However, all my dbspaces are HPUX-mirrored.
So I don't understand why an I/O error could occur. On an other
server, the dbspaces are informix-mirrored and it seems more
reliable. Now, I'm going to wait for the next I/O error in
order to see which disk is bad.
I notice I have this problem since someone modified a C library
in our application (mainly written in MF Cobol). I applied this
morning the path PHSS_15389 (millicode library milli.a). Do you
think if it is possible that there is a link between C libraries
corrupted and Informix problems?
David...
----- Message d'origine -----
De : David Kosenko <davek@summitdata.com>
' : <informix-list@iiug.org>
Envoi' : jeudi 4 f'vrier 1999 20:25
Objet : Re: Blocked:CKPT
On Wed, 03 Feb 1999 08:26:55 +0100, David Andrieu
<dandrieu@mail.atmel.fr> wrote:
> but I don't have messages in message.log like:
>
> Assert Failed: Chunk ... is being taken OFFLINE.
> Who: ...
> Result: ...
> Action: ...
> See Also: ...>
>which occur when online tries to put a dbspace down.
>Or maybe it is a bug in 7.30.UC6. This message does not appear
>while online tries to put a dbspace down?
No, because ONDBSPACEDOWN=2 will not mark the chunk as down - that is
the whole point behind this option. It is meant for cases where a
disk becomes temporarily unavailable (say because a SCSI cable got
unplugged). In that case, marking the chunk as down would require a
restore to recover it, even if the disk isn't really bad. A value of
2 (WAIT) means don't mark it as down, keep polling the device (in case
it becomes available again), and at the next checkpoint, block all
activity (the symptom originally reported). This is a very bad option
if your chunk really does go down - much better to either let the
server go down or markthe chunk as down and keep the server up, but
not block at the next checkpoint.
Dave
In article <79mdfk$nua$1@news.xmission.com>, David Andrieu
<dandrieu@mail.atmel.fr> writes
>
>Hi David,
>
> I think you ar rigt. I changed ONDBSPACEDOWN and set it
>to 1 (ABORT). However, all my dbspaces are HPUX-mirrored.
>So I don't understand why an I/O error could occur. On an other
>server, the dbspaces are informix-mirrored and it seems more
>reliable. Now, I'm going to wait for the next I/O error in
>order to see which disk is bad.
>I notice I have this problem since someone modified a C library
>in our application (mainly written in MF Cobol). I applied this
>morning the path PHSS_15389 (millicode library milli.a). Do you
>think if it is possible that there is a link between C libraries
>corrupted and Informix problems?
>
>David...
>
If you are using a shared memory connection then your program could
potentially access the engines shared memory and corrupt it!
Try using TCP loopback connections instead..
>
>
>
>
>----- Message d'origine -----
>De : David Kosenko <davek@summitdata.com>
>À : <informix-list@iiug.org>
>Envoié : jeudi 4 février 1999 20:25
>Objet : Re: Blocked:CKPT
>
>
>On Wed, 03 Feb 1999 08:26:55 +0100, David Andrieu
><dandrieu@mail.atmel.fr> wrote:
>> but I don't have messages in message.log like:
>>
>> Assert Failed: Chunk ... is being taken OFFLINE.
>> Who: ...
>> Result: ...
>> Action: ...
>> See Also: ...>>
>>which occur when online tries to put a dbspace down.
>>Or maybe it is a bug in 7.30.UC6. This message does not appear
>>while online tries to put a dbspace down?
>
>No, because ONDBSPACEDOWN=2 will not mark the chunk as down - that is
>the whole point behind this option. It is meant for cases where a
>disk becomes temporarily unavailable (say because a SCSI cable got
>unplugged). In that case, marking the chunk as down would require a
>restore to recover it, even if the disk isn't really bad. A value of
>2 (WAIT) means don't mark it as down, keep polling the device (in case
>it becomes available again), and at the next checkpoint, block all
>activity (the symptom originally reported). This is a very bad option
>if your chunk really does go down - much better to either let the
>server go down or markthe chunk as down and keep the server up, but
>not block at the next checkpoint.
>
>Dave
>
>
>
--
David Williams