Server down - Assert failed - Index bad...
Posted in 2000
Robin's instance crashed with "Assert Failed: Page Check Error in btsearch: bad btree node" on a system catalog index (sysobjstate), plus oncheck reporting overlapping extents; oncheck -ci/-cI repairs failed with duplicate-key and bad ISAM format errors. Art Kagel advised unloading the data, dropping and rebuilding/reloading the table, suspecting something had overwritten the chunk; that's what Informix support ultimately did. Root cause was traced to an old ONCONFIG being used at shared-memory reinitialisation, putting the physical log back in rootdbs where it overwrote a system table; Art offered only a theory as to why the engine didn't detect this.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Storage & Space Management, Error Codes & Troubleshooting, Logging & Checkpoints
Our instance is down!!
Errors in the online.log
Assert Failed: Page Check Error in btsearch:bad btree node
Results: Possible inconsistencies in index
'database:"informix".sysobjstate#objdesc
When I run the oncheck -ci I get:
Validating indexes for 'database':informix.sysobjstate...
Index objdesc
Unable to open input file 's'
Unable to open input file 'c'
Could not bfget pagenum 21, iserrno 105
ISAM error: bad isam file format.ERROR: Index objdesc for 'database':informix.sysobjstate is bad.
Index objtab
Unable to open input file 's'
Unable to open input file 'c'
Index objdesc is bad. OK to repair it? y
Unable to open input file 's'
Unable to open input file 'c'
Error recreating index.
ISAM error: duplicate value for a record with unique key.
Error reading from network
Abnormal end !!!
Then of course it "dies"....
Also .... in the oncheck -pe it shows these .....
PHYSICAL LOG Pages 63 100000
ERROR: Extents overlap
db:bx.accidcmp 119 8
ERROR: Extents overlap
db:bx.foreign_sw 127 8
I can't seem to see a table named sysobjstate - should I?
So what do I do to fix this????
All help will be appreciated....
Thanks
Robin
P.S. I did run oncheck -cI then oncheck -cr, cc,ce and cD and came up with
the same info. I was unable to drop the "bad index".
Thanks again
Robin
> From: Boscia <rwboscia@worldnet.att.net>
> Organization: AT&T Worldnet
> Newsgroups: comp.databases.informix
> Date: Thu, 03 Aug 2000 04:16:47 GMT
> Subject: Server down - Assert failed - Index bad...
>
> Our instance is down!!
>
> Errors in the online.log
> Assert Failed: Page Check Error in btsearch:bad btree node>
> Results: Possible inconsistencies in index
> 'database:"informix".sysobjstate#objdesc
>
> When I run the oncheck -ci I get:
>
>
>
> Validating indexes for 'database':informix.sysobjstate...
> Index objdesc
> Unable to open input file 's'
> Unable to open input file 'c'
> Could not bfget pagenum 21, iserrno 105
> ISAM error: bad isam file format.> ERROR: Index objdesc for 'database':informix.sysobjstate is bad.
> Index objtab
> Unable to open input file 's'
> Unable to open input file 'c'
>
>
> Index objdesc is bad. OK to repair it? y
> Unable to open input file 's'
> Unable to open input file 'c'
> Error recreating index.
> ISAM error: duplicate value for a record with unique key.>
> Error reading from network
> Abnormal end !!!
>
>
> Then of course it "dies"....
>
> Also .... in the oncheck -pe it shows these .....
>
> PHYSICAL LOG Pages 63 100000
> ERROR: Extents overlap
> db:bx.accidcmp 119 8
> ERROR: Extents overlap
> db:bx.foreign_sw 127 8
>
> I can't seem to see a table named sysobjstate - should I?
>
>
> So what do I do to fix this????
>
> All help will be appreciated....
>
> Thanks
>
> Robin
>
>
>
>
>
Export our data, drop the entire table and rebuild/reload it.
Something wrote to the Informix chunk on which this table resides.
Hopefully only a few index pages were trashed. You may want to
examine those pages in od with various options to see if you can
recognize a data pattern that would tell you how the corruption
occurred. This would not be a RAID5 partition would it?
Art S. Kagel
Boscia wrote:
>
> P.S. I did run oncheck -cI then oncheck -cr, cc,ce and cD and came up with
> the same info. I was unable to drop the "bad index".
> Thanks again
> Robin
>
> > From: Boscia <rwboscia@worldnet.att.net>
> > Organization: AT&T Worldnet
> > Newsgroups: comp.databases.informix
> > Date: Thu, 03 Aug 2000 04:16:47 GMT
> > Subject: Server down - Assert failed - Index bad...
> >
> > Our instance is down!!
> >
> > Errors in the online.log
> > Assert Failed: Page Check Error in btsearch:bad btree node> >
> > Results: Possible inconsistencies in index
> > 'database:"informix".sysobjstate#objdesc
> >
> > When I run the oncheck -ci I get:
> >
> >
> >
> > Validating indexes for 'database':informix.sysobjstate...
> > Index objdesc
> > Unable to open input file 's'
> > Unable to open input file 'c'
> > Could not bfget pagenum 21, iserrno 105
> > ISAM error: bad isam file format.> > ERROR: Index objdesc for 'database':informix.sysobjstate is bad.
> > Index objtab
> > Unable to open input file 's'
> > Unable to open input file 'c'
> >
> >
> > Index objdesc is bad. OK to repair it? y
> > Unable to open input file 's'
> > Unable to open input file 'c'
> > Error recreating index.
> > ISAM error: duplicate value for a record with unique key.> >
> > Error reading from network
> > Abnormal end !!!
> >
> >
> > Then of course it "dies"....
> >
> > Also .... in the oncheck -pe it shows these .....
> >
> > PHYSICAL LOG Pages 63 100000
> > ERROR: Extents overlap
> > db:bx.accidcmp 119 8
> > ERROR: Extents overlap
> > db:bx.foreign_sw 127 8
> >
> > I can't seem to see a table named sysobjstate - should I?
> >
> >
> > So what do I do to fix this????
> >
> > All help will be appreciated....
> >
> > Thanks
> >
> > Robin
> >
> >
> >
> >
> >
Thanks for your help....
After being on the phone with Informix for an all day session, that is
exactly what we did..... Export, drop and rebuild/reload.
Informix did login to my system, however they still could not get the
problem fixed or even tell me how the corruption happened...
How do I examine the pages in od?
And No it is not RAID5.
Thanks again
Robin
> From: "Art S. Kagel" <kagel@bloomberg.net>
> Organization: Bloomberg LP
> Reply-To: kagel@bloomberg.net
> Newsgroups: comp.databases.informix
> Date: Thu, 03 Aug 2000 09:07:20 -0400
> To: Boscia <rwboscia@worldnet.att.net>
> Subject: Re: Server down - Assert failed - Index bad...
>
> Export our data, drop the entire table and rebuild/reload it.
> Something wrote to the Informix chunk on which this table resides.
> Hopefully only a few index pages were trashed. You may want to
> examine those pages in od with various options to see if you can
> recognize a data pattern that would tell you how the corruption
> occurred. This would not be a RAID5 partition would it?
>
> Art S. Kagel
>
> Boscia wrote:
>>
>> P.S. I did run oncheck -cI then oncheck -cr, cc,ce and cD and came up with
>> the same info. I was unable to drop the "bad index".
>> Thanks again
>> Robin
>>
>>> From: Boscia <rwboscia@worldnet.att.net>
>>> Organization: AT&T Worldnet
>>> Newsgroups: comp.databases.informix
>>> Date: Thu, 03 Aug 2000 04:16:47 GMT
>>> Subject: Server down - Assert failed - Index bad...
>>>
>>> Our instance is down!!
>>>
>>> Errors in the online.log
>>> Assert Failed: Page Check Error in btsearch:bad btree node>>>
>>> Results: Possible inconsistencies in index
>>> 'database:"informix".sysobjstate#objdesc
>>>
>>> When I run the oncheck -ci I get:
>>>
>>>
>>>
>>> Validating indexes for 'database':informix.sysobjstate...
>>> Index objdesc
>>> Unable to open input file 's'
>>> Unable to open input file 'c'
>>> Could not bfget pagenum 21, iserrno 105
>>> ISAM error: bad isam file format.>>> ERROR: Index objdesc for 'database':informix.sysobjstate is bad.
>>> Index objtab
>>> Unable to open input file 's'
>>> Unable to open input file 'c'
>>>
>>>
>>> Index objdesc is bad. OK to repair it? y
>>> Unable to open input file 's'
>>> Unable to open input file 'c'
>>> Error recreating index.
>>> ISAM error: duplicate value for a record with unique key.>>>
>>> Error reading from network
>>> Abnormal end !!!
>>>
>>>
>>> Then of course it "dies"....
>>>
>>> Also .... in the oncheck -pe it shows these .....
>>>
>>> PHYSICAL LOG Pages 63 100000
>>> ERROR: Extents overlap
>>> db:bx.accidcmp 119 8
>>> ERROR: Extents overlap
>>> db:bx.foreign_sw 127 8
>>>
>>> I can't seem to see a table named sysobjstate - should I?
>>>
>>>
>>> So what do I do to fix this????
>>>
>>> All help will be appreciated....
>>>
>>> Thanks
>>>
>>> Robin
>>>
>>>
>>>
>>>
>>>
Boscia wrote:
>
> Thanks for your help....
> After being on the phone with Informix for an all day session, that is
> exactly what we did..... Export, drop and rebuild/reload.
> Informix did login to my system, however they still could not get the
> problem fixed or even tell me how the corruption happened...
> How do I examine the pages in od?
dd bs=2k if=<chunkpath> skip=[# pages not corrupt prior to damage] |od -c|less
Then try od -h if the data is primarily binary anyway you are looking
for familiar text or data patterns.
Art S. Kagel
> And No it is not RAID5.
> Thanks again
> Robin
>
> > From: "Art S. Kagel" <kagel@bloomberg.net>
> > Organization: Bloomberg LP
> > Reply-To: kagel@bloomberg.net
> > Newsgroups: comp.databases.informix
> > Date: Thu, 03 Aug 2000 09:07:20 -0400
> > To: Boscia <rwboscia@worldnet.att.net>
> > Subject: Re: Server down - Assert failed - Index bad...
> >
> > Export our data, drop the entire table and rebuild/reload it.
> > Something wrote to the Informix chunk on which this table resides.
> > Hopefully only a few index pages were trashed. You may want to
> > examine those pages in od with various options to see if you can
> > recognize a data pattern that would tell you how the corruption
> > occurred. This would not be a RAID5 partition would it?
> >
> > Art S. Kagel
> >
> > Boscia wrote:
> >>
> >> P.S. I did run oncheck -cI then oncheck -cr, cc,ce and cD and came up with
> >> the same info. I was unable to drop the "bad index".
> >> Thanks again
> >> Robin
> >>
> >>> From: Boscia <rwboscia@worldnet.att.net>
> >>> Organization: AT&T Worldnet
> >>> Newsgroups: comp.databases.informix
> >>> Date: Thu, 03 Aug 2000 04:16:47 GMT
> >>> Subject: Server down - Assert failed - Index bad...
> >>>
> >>> Our instance is down!!
> >>>
> >>> Errors in the online.log
> >>> Assert Failed: Page Check Error in btsearch:bad btree node> >>>
> >>> Results: Possible inconsistencies in index
> >>> 'database:"informix".sysobjstate#objdesc
> >>>
> >>> When I run the oncheck -ci I get:
> >>>
> >>>
> >>>
> >>> Validating indexes for 'database':informix.sysobjstate...
> >>> Index objdesc
> >>> Unable to open input file 's'
> >>> Unable to open input file 'c'
> >>> Could not bfget pagenum 21, iserrno 105
> >>> ISAM error: bad isam file format.> >>> ERROR: Index objdesc for 'database':informix.sysobjstate is bad.
> >>> Index objtab
> >>> Unable to open input file 's'
> >>> Unable to open input file 'c'
> >>>
> >>>
> >>> Index objdesc is bad. OK to repair it? y
> >>> Unable to open input file 's'
> >>> Unable to open input file 'c'
> >>> Error recreating index.
> >>> ISAM error: duplicate value for a record with unique key.> >>>
> >>> Error reading from network
> >>> Abnormal end !!!
> >>>
> >>>
> >>> Then of course it "dies"....
> >>>
> >>> Also .... in the oncheck -pe it shows these .....
> >>>
> >>> PHYSICAL LOG Pages 63 100000
> >>> ERROR: Extents overlap
> >>> db:bx.accidcmp 119 8
> >>> ERROR: Extents overlap
> >>> db:bx.foreign_sw 127 8
> >>>
> >>> I can't seem to see a table named sysobjstate - should I?
> >>>
> >>>
> >>> So what do I do to fix this????
> >>>
> >>> All help will be appreciated....
> >>>
> >>> Thanks
> >>>
> >>> Robin
> >>>
> >>>
> >>>
> >>>
> >>>
Art, We are trying to come up with a "root cause" of why this happened. We know that the PHYSICAL LOGS ended up writing over a system table. From what we know, we had 2 config files with 2 different names. When we reinitialized shared memory the old config file was used and the PHYSICAL logs ended up back in the rootdbspace writing over a system table. Now, how can that happen? Shouldn't "it" be smart enough to know that there is a table there and know that there is not enough space and NOT write over it? And give an error? Or was my Assert Failed my error? Thanks again!! BTW - Informix had No clue how or why this happened.... ----------------------------------------------------------- Got questions? Get answers over the phone at Keen.com. Up to 100 minutes free! http://www.keen.com
rboscia wrote: > > Art, > We are trying to come up with a "root cause" of why this > happened. We know that the PHYSICAL LOGS ended up writing over > a system table. Ouch! > From what we know, we had 2 config files with 2 different names. > When we reinitialized shared memory the old config file was used > and the PHYSICAL logs ended up back in the rootdbspace writing > over a system table. My only guess as to how this might have happened is somehow someone muffed the physical log code so that it opens PHYSDBS and searches for the beginning of a previous physical log block which must be a contiguous block in a single chunk PHYSFILE KB long. Since you previously had the physical log in rootdbs but moved it elsewhere it is possible that the old physical log header was still on disk there, that the engine found it and used it but that since then a tablespace tablespace entry had been written somewhere further on in what had been that physical log block but the engine does not check for that possibility. > Now, how can that happen? Shouldn't "it" be smart enough to > know that there is a table there and know that there is not > enough space and NOT write over it? And give an error? > Or was my Assert Failed my error? Thanks again!! > BTW - Informix had No clue how or why this happened.... No doubt, not without looking into the source anyway. Art S. Kagel