Bad sector in raw OnLine chunk?
Posted in 1996
What is the best way to find and mark a sector of a disk bad that is in
a raw Informix file?
On an RS6000, while running a 4GL program,
INFORMIX-4GL Version 4.10.UD3,
INFORMIX-OnLine Version 4.10.UE4,
we received the following error
> Following fatal error occured during eqjabld20:
> Date: 07/30/1996 Time: 08:51:27
> Program error at eqjabld20.4gl, line number 687
> SQL statement error number -408
> Invalid message type received from the sqlexec process
Error # -408:
-408 Invalid message type received from the sqlexec process.
This is an internal error that shows a problem in the communication
between the database engine and the library functions that call it.
Make sure that your program is at the same software level as the
database engine in use. If so, note all the circumstances and contact
Informix Technical Support.
The line where the program crashed was a simple insert:
INSERT INTO eq_desc_181
(
cnt_nbr,
proj_nbr,
equipment_cde,
comp_cde,
sub_comp_cde,
eq_desc,
line_no
)
VALUES
(
gv_cnt_nbr,
gv_proj_nbr,
w_rec.equipment_cde,
white_space,
white_space,
lar_eq_desc[i].desc,
i
)
>From the turbo-log:
08:46:22 Logical Log 96 Complete
08:48:34 Checkpoint Completed
08:50:25 Logical Log 97 Complete
08:51:27 newmode: Invalid page type of 0x0
08:51:27 bfcheck: bad page: pg_addr 0 != bp->bf_pagenum 10e215, userp = 30001cdc, pid = 11636, uid = 203
08:51:27 buffer header:
08:51:27 30388074: 00000001 00000000 30001cdc 303809b4 ........ 0...08..
08:51:27 30388084: 3037e018 30389874 30382074 00020000 07..08.t 08 t....
08:51:27 30388094: 02080000 0010e215 30575000 30001cdc ........ 0WP.0...
08:51:27 303880a4: 00000000 00000000 3038c19c 04000000 ........ 08......
08:51:27 page header:
08:51:27 30575000: 00000000 00000000 00010000 e4b82b40 ........ ......+@
08:51:27 30575010: 00000000 00000000 ........
08:51:27 slot table and stamp:
08:51:27 30575ffc: 00000000 ....
08:51:27 -- Fail Consistency Check -- bfput -- pid=11636 user=203 us=30001cdc
08:51:27 bfcheck: bad page: pg_addr 0 != bp->bf_pagenum 10e215, userp = 30001cdc, pid = 11636, uid = 203
08:51:27 buffer header:
08:51:27 30388074: 00000000 00000000 00000000 303809b4 ........ ....08..
08:51:27 30388084: 3037e018 30387c74 30000388 00030000 07..08|t 0.......
08:51:27 30388094: 03090000 0010e215 30575000 30001cdc ........ 0WP.0...
08:51:27 303880a4: 00000000 00000000 3038c19c 04000000 ........ 08......
08:51:27 page header:
08:51:27 30575000: 00000000 0033b257 00010000 e4b82b40 .....3.W ......+@
08:51:27 30575010: 00000000 00000000 ........
08:51:27 slot table and stamp:
08:51:27 30575ffc: 0033b257 .3.W
08:51:27 -- Fail Consistency Check -- dovrecord:bad data page -- pid=11636 user=203 us=30001cdc
08:51:27 ERROR: logundo(40) iserrno 105 us 0x30001cdc pid 11636
tx 0x30003ae4 loguniq 98 logpos 0x10929c
08:51:27 INFORMIX-OnLine Must ABORT
Log Error 'sprback() - logundo() FAILED' us 0x30001cdc pid 11636 us_flags 0x101
tx 0x30003ae4 tx_flags 0x222403 tx_loguniq 98 tx_logpos 0x10929c
08:51:27 -- Online Aborting -- us=30001cdc, pid=11636, uid=203
08:51:27 INFORMIX-OnLine entering ABORT mode!!!
08:51:27 -- Online Aborting -- us=30001b8c, pid=4742, uid=666
Tue Jul 30 14:28:43 1996
14:28:43 INFORMIX-OnLine Initialized -- Shared Memory Initialized
14:28:43 Physical Recovery Started
14:28:45 Checkpoint Completed
14:28:45 Physical Recovery Complete: 528 Pages Restored
14:28:51 Rollforward of log record failed, iserrno = 126
14:28:51 Log Record: log = 98, pos = 10929c, type = 40, trans = 2
14:28:53 Checkpoint Completed
14:28:53 ERROR: logundo(40) iserrno 126 us 0x30001cdc pid 0
tx 0x30003ae4 loguniq 98 logpos 0x10929c
14:28:53 INFORMIX-OnLine Must ABORT
Log Error 'rollback() - logundo() FAILED' us 0x30001cdc pid 0 us_flags 0x121
tx 0x30003ae4 tx_flags 0x202403 tx_loguniq 98 tx_logpos 0x10929c
14:28:53 -- Online Aborting -- us=30001cdc, pid=4742, uid=666
So we may have a bad section of the hard disk?
08:51:27 bfcheck: bad page: pg_addr 0 != bp->bf_pagenum 10e215, userp = 30001cdc, pid = 11636, uid = 203
Maybe the bad section is where a log was written?
14:28:51 Rollforward of log record failed, iserrno = 126
Trying to bring informix back up using tbmonitor gives the error that
the OnLine daemon is no longer running.
The log listing (tbstat -l) showed that about half the logs are full.
Informix cleared the logs with tbzero but the identical problem occured
again later, also during a write.
_____________________________________________________________________
| Colin McGrath cmm@trac3000.ueci.com |
| Raytheon Engineers & Constructors, Inc. (215) 422-4144 |
| Phila, PA, USA Standard Disclaimers Apply |
|___________________________________________________________________|