Re: Assert Failed: Page Check Error in bfput
Posted in 2005
A 7.31.UC6 server reported an assert failure on page 0x30fcaa1 in chunk 3, and oncheck -pP showed the page was zeroed/UNKNOWN; nightly ontape archives then failed with "Archive detects that page ... is corrupt." Advice: drop the (already unloaded/renamed) corrupt table so the extents are freed and the archive ignores them, then take a fresh backup; also check whether chunk 3's device was tampered with (symlink/filesystem) or had a hardware fault, dumping the raw page with dd (use onstat -d and vxprint -ht to map chunk to device, and iseek rather than skip, which reads everything up to the offset). Upgrading off 7.31.UC6 was also suggested. No confirmation of the outcome is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Backup & Restore, Storage & Space Management, Error Codes & Troubleshooting
On 14 Sep 2005 23:27:24 -0700, "Superboer" <superboer7@planet.nl> wrote:
>> BAD PAGE 30fcaa1: pg_addr 0 != bp->bf_pagenum 30fcaa1
>
>
>tells that page 30fcaa1 is zeroed. or at least it's adress
>check your af files. you also may want to dump the page:
>
>oncheck -pP 0x3 0xfcaa1>in order to check if it's really zeroed...
>
>when you do not need the old table anymore you could try and drop
>it; however it may fail.
>you may also want to check the dev for chunk #3 maybe someone mucked it
>up
>
# oncheck -pP 0x3 0xfcaa1addr stamp nslots flag type frptr frcnt next prev
0 0 0 0 UNKNOWN 0 0 0 0
slot ptr len flg
My ontape also failed last night:
01:42:36 Assert Failed: Archive detects that page 0x30fcaa1 is corrupt.
01:42:36 Informix Dynamic Server Version 7.31.UC6
01:42:36 Who: Session(1601656, root@emco5, 7494, 2034648440)
Thread(2601857, arcbackup1, 79c1db20, 3)
File: rsarcbu.c Line: 2482
What should I do?
Thanks
Gary Quiring
I guess that archive was taken before you unloaded and fixed the data? If so then you have to now take a backup with your uncorrupted data and 7.31.UC6?? get onto to a 7.32UD release as soon as possible then look at upgrading to IDS 10
Hello Gary,
check what is wrong with chunk# 3; maybe someone added a extra sym link
to
it and it's reused for something else or a filesystem is build on it
or.... dono may the hardware has had a hickup.
You better do the above since it may repeat itself whenever you have
the problem fixed!!!!
for the bad page; if you are lucky you can try and drop the table
which got corrupted and you got unloaded.
if that fails and again if you are lucky that it's only 1 page
which is corrupted; then phone TS and ask them to patch this page to be
an empty datapage (if it was one and not a bitmap or index...)
try again and drop the table.
Also you may want to dd stuff out of chunk# 3
eq dd if=<chunk3> of=somefile bs=2k count=10 skip=1034913 count=1)
(or skip=1034900 count=100) and use od on somefile maybe this gives
you a hint on what may have corrupted it in the first place
Superboer.
Gary Quiring schreef:
> On 14 Sep 2005 23:27:24 -0700, "Superboer" <superboer7@planet.nl> wrote:
>
> >> BAD PAGE 30fcaa1: pg_addr 0 != bp->bf_pagenum 30fcaa1
> >
> >
> >tells that page 30fcaa1 is zeroed. or at least it's adress
> >check your af files. you also may want to dump the page:
> >
> >oncheck -pP 0x3 0xfcaa1> >in order to check if it's really zeroed...
> >
> >when you do not need the old table anymore you could try and drop
> >it; however it may fail.
> >you may also want to check the dev for chunk #3 maybe someone mucked it
> >up
> >
> # oncheck -pP 0x3 0xfcaa1> addr stamp nslots flag type frptr frcnt next prev
> 0 0 0 0 UNKNOWN 0 0 0 0
> slot ptr len flg
>
> My ontape also failed last night:
> 01:42:36 Assert Failed: Archive detects that page 0x30fcaa1 is corrupt.
> 01:42:36 Informix Dynamic Server Version 7.31.UC6
> 01:42:36 Who: Session(1601656, root@emco5, 7494, 2034648440)
> Thread(2601857, arcbackup1, 79c1db20, 3)
> File: rsarcbu.c Line: 2482>
> What should I do?
> Thanks
> Gary Quiring
Gary Quiring wrote:
> On 14 Sep 2005 23:27:24 -0700, "Superboer" <superboer7@planet.nl> wrote:
>
>
>>> BAD PAGE 30fcaa1: pg_addr 0 != bp->bf_pagenum 30fcaa1
>>
>>
>>tells that page 30fcaa1 is zeroed. or at least it's adress
>>check your af files. you also may want to dump the page:
>>
>>oncheck -pP 0x3 0xfcaa1>>in order to check if it's really zeroed...
>>
>>when you do not need the old table anymore you could try and drop
>>it; however it may fail.
>>you may also want to check the dev for chunk #3 maybe someone mucked it
>>up
>>
>
> # oncheck -pP 0x3 0xfcaa1> addr stamp nslots flag type frptr frcnt next prev
> 0 0 0 0 UNKNOWN 0 0 0 0
> slot ptr len flg
>
> My ontape also failed last night:
> 01:42:36 Assert Failed: Archive detects that page 0x30fcaa1 is corrupt.
> 01:42:36 Informix Dynamic Server Version 7.31.UC6
> 01:42:36 Who: Session(1601656, root@emco5, 7494, 2034648440)
> Thread(2601857, arcbackup1, 79c1db20, 3)
> File: rsarcbu.c Line: 2482>
> What should I do?
Gary, you should be able to just drop the renamed table. Once the
pages/extents are marked as free the archive will ignore them. If they are
reused then the page header of the damaged page will be rewritten and all
should be well. The only worry is that there was a physical cause for the
corruption (unlikely if only one page was affected) which as the Superboer
points out you can test with dd.
Art S. Kagel
"Gary Quiring" <gquiring@msn.com> wrote in message
news:o3nli19fobslaq904dp4f24p1rt5jdom10@4ax.com...
> On Fri, 16 Sep 2005 10:45:04 -0400, "Art S. Kagel" <kagel@bloomberg.net>
> wrote:
>
> I would like to check the chunk with dd but I am unclear how I know what
> chunk 3
> is? How do I turn that into my raw device names? They are all 2 gig
> partitions
> in Veritas Volume Manager.
>
> /dev/informix/db01 - /dev/informix/db98
onstat -d
On Fri, 16 Sep 2005 10:45:04 -0400, "Art S. Kagel" <kagel@bloomberg.net> wrote:
>Gary Quiring wrote:
>> On 14 Sep 2005 23:27:24 -0700, "Superboer" <superboer7@planet.nl> wrote:
>>
>>
>>>> BAD PAGE 30fcaa1: pg_addr 0 != bp->bf_pagenum 30fcaa1
>>>
>>>
>>>tells that page 30fcaa1 is zeroed. or at least it's adress
>>>check your af files. you also may want to dump the page:
>>>
>>>oncheck -pP 0x3 0xfcaa1>>>in order to check if it's really zeroed...
>>>
>>>when you do not need the old table anymore you could try and drop
>>>it; however it may fail.
>>>you may also want to check the dev for chunk #3 maybe someone mucked it
>>>up
>>>
>>
>> # oncheck -pP 0x3 0xfcaa1>> addr stamp nslots flag type frptr frcnt next prev
>> 0 0 0 0 UNKNOWN 0 0 0 0
>> slot ptr len flg
>>
>> My ontape also failed last night:
>> 01:42:36 Assert Failed: Archive detects that page 0x30fcaa1 is corrupt.
>> 01:42:36 Informix Dynamic Server Version 7.31.UC6
>> 01:42:36 Who: Session(1601656, root@emco5, 7494, 2034648440)
>> Thread(2601857, arcbackup1, 79c1db20, 3)
>> File: rsarcbu.c Line: 2482>>
>> What should I do?
>
>Gary, you should be able to just drop the renamed table. Once the
>pages/extents are marked as free the archive will ignore them. If they are
>reused then the page header of the damaged page will be rewritten and all
>should be well. The only worry is that there was a physical cause for the
>corruption (unlikely if only one page was affected) which as the Superboer
>points out you can test with dd.
>
>Art S. Kagel
Hi Art,
I would like to check the chunk with dd but I am unclear how I know what chunk 3
is? How do I turn that into my raw device names? They are all 2 gig partitions
in Veritas Volume Manager.
/dev/informix/db01 - /dev/informix/db98
Thanks
Gary
On Fri, 16 Sep 2005 15:55:08 +0100, "Neil Truby" <neil.truby@ardenta.com> wrote:
>"Gary Quiring" <gquiring@msn.com> wrote in message
>news:o3nli19fobslaq904dp4f24p1rt5jdom10@4ax.com...
>> On Fri, 16 Sep 2005 10:45:04 -0400, "Art S. Kagel" <kagel@bloomberg.net>
>> wrote:
>>
>> I would like to check the chunk with dd but I am unclear how I know what
>> chunk 3
>> is? How do I turn that into my raw device names? They are all 2 gig
>> partitions
>> in Veritas Volume Manager.
>>
>> /dev/informix/db01 - /dev/informix/db98
>
>onstat -d>
I tried running dd, it just hangs. Is Informix locking the data that I cannot
get to it?
Chunks
address chk/dbs offset size free bpages flags pathname
7704a210 1 1 0 200000 198893 PO- /dev/informix/rootdb
7704b100 2 2 0 1048550 798497 PO- /dev/informix/dblog1
7704b1e0 3 2 0 1048550 1048547 PO- /dev/informix/dblog2
7704b2c0 4 2 0 1048550 1048547 PO- /dev/informix/dblog3
dd if=/dev/informix/dblog2 of=gq bs=2k count=10 skip=1034913 count=1
Gary Quiring wrote: > I would like to check the chunk with dd but I am unclear how I know what chunk 3 > is? How do I turn that into my raw device names? They are all 2 gig partitions > in Veritas Volume Manager. > > /dev/informix/db01 - /dev/informix/db98 > > Thanks > Gary vxprint -ht
No skip reads all the data up to that point! Check the man pages - you want something like iseek= instead. But check the man pages I got this wrong once!