AF During Archive - Looks Familiar
Posted in 2010
A nightly archive on IDS 10.00.UC8 (Solaris 10) threw an assert warning "Archive detects that page 291:848713 is corrupt" in a blobspace chunk (arc_bpages/arc_blobspace stack), though the engine stayed up and the archive completed. Suggestions: restore/roll forward and hunt for the cause (shared chunk devices, silent RAID failure, prior crash, non-logged/buffered-log databases); IBM support explained it was likely a blob free-map page wrongly marking an unused, zeroed blob page as in use — consistent with clean oncheck -cD/-cDI results. Resolution: Mike created a new blobspace, unloaded, dropped and recreated the blob-bearing tables pointing at it (schema edited), reloaded, then dropped the suspect blobspace. Others noted chunks can also be dd'd to new devices instead of ALTERing.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Storage & Space Management
10.00.UC8
Solaris 10
I think this has been posted before - apologies if so. The search feature here
did not find what I was looking for...
We saw this last night:
Archive detects that page 291:848713 is corrupt.
Archive detects that page 291:848713 is corrupt.
5409b400: 00000000 00000000 00000000 00000000 ........ ........
5409b410 *
17:37:38
17:37:38 IBM Informix Dynamic Server Version 10.00.UC8D3a Software Serial
Number RDS#N000000
17:37:38 Assert Warning: Archive detects that page 291:848713 is corrupt.
17:37:38 Who: Session(82961, informix@rmkc00a, 7787, 53eea0b8)
Thread(87306, arcbackup1, 50373c50, 5)
File: rsarcbu.c Line: 3424
17:37:38 Stack for thread: 87306 arcbackup1
base: 0x564c9000
len: 266240
pc: 0x008458c4
tos: 0x565094b0
state: running
vp: 5
0x00844cfc (oninit)afhandler(0xdffc00, 0xc00800, 0x0, 0x1, 0x401, 0x401)
0x00844460 (oninit)afwarn_interface(0xdf7448, 0x0, 0x0, 0xd3e294, 0xd60, 0x30)
0x006b9448 (oninit)arc_bpages(0xdf7448, 0x5409b400, 0xdd3b7c, 0x2, 0x1, 0xffff)
0x006b9164 (oninit)arc_blobspace(0xdd3b80, 0x406e0, 0x7cc09, 0x123, 0x3d09ed,
0x7a13d)
0x006b6944 (oninit)arc_scanner(0xc00400, 0xa, 0xdd3b7c, 0x529cd808,
0xffffbfff, 0x406e0)
0x00821274 (oninit)startup (0xdffdb8, 0x7, 0x0, 0x4f0138d8, 0x0, 0x4f0138a0)
0x00818ef4 (oninit)mt_poll_yield(0x0, 0x0, 0x0, 0x0, 0x0, 0x0)
0x00000000 (*nosymtab*)0x0
As we can see from the stack - the "bad" page is in a blobspace chunk. The
messages and circumstances sure look like IC52728
(http://www-01.ibm.com/support/docview.wss?uid=swg1IC52728).
The instance was created via an imported restore - but that was back in July
of last year. The engine did not crash (AFWARN) and the archive succeeded but
I am wondering what the deal is. I know that I have dealt with this before and
just cannot find the email chain from the last time we ran into it - but also
after finding that bug it has me wondering...
Thanks -
MM
Here was my response when someone posted the same problem several years ago:
Restore from the last good archive, roll logs forward, take a new archive,
then
find out how that page became corrupted. Is it possible for example that
there
are multiple instances on this machine and two or more have been assigned
the
same disk partition as a chunk through different links? Is the disk
structure a
singleton drive or a RAID5 array that may have gone silently bad? Are there
any messages in the system logs about disk errors since the last good
archive
was taken?
It's still valid. I'd only add: Do you have any non-logged or buffered log
databases? If so has the system crashed in the past or been shutdown out
from under IDS without a proper server shutdown during active update
activity?
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Tue, Jun 8, 2010 at 9:06 AM, MIKE MAGIE <jmmagie@yahoo.com> wrote:
> 10.00.UC8
> Solaris 10
>
> I think this has been posted before - apologies if so. The search feature
> here
> did not find what I was looking for...
>
> We saw this last night:
>
> Archive detects that page 291:848713 is corrupt.
> Archive detects that page 291:848713 is corrupt.
> 5409b400: 00000000 00000000 00000000 00000000 ........ ........
> 5409b410 *
> 17:37:38
> 17:37:38 IBM Informix Dynamic Server Version 10.00.UC8D3a Software Serial
> Number RDS#N000000
>
> 17:37:38 Assert Warning: Archive detects that page 291:848713 is corrupt.
> 17:37:38 Who: Session(82961, informix@rmkc00a, 7787, 53eea0b8)
>
> Thread(87306, arcbackup1, 50373c50, 5)
>
> File: rsarcbu.c Line: 3424
> 17:37:38 Stack for thread: 87306 arcbackup1
>
> base: 0x564c9000
> len: 266240
>
> pc: 0x008458c4
> tos: 0x565094b0
> state: running
>
> vp: 5
>
> 0x00844cfc (oninit)afhandler(0xdffc00, 0xc00800, 0x0, 0x1, 0x401, 0x401)
> 0x00844460 (oninit)afwarn_interface(0xdf7448, 0x0, 0x0, 0xd3e294, 0xd60,
> 0x30)
> 0x006b9448 (oninit)arc_bpages(0xdf7448, 0x5409b400, 0xdd3b7c, 0x2, 0x1,
> 0xffff)
> 0x006b9164 (oninit)arc_blobspace(0xdd3b80, 0x406e0, 0x7cc09, 0x123,
> 0x3d09ed,
> 0x7a13d)
> 0x006b6944 (oninit)arc_scanner(0xc00400, 0xa, 0xdd3b7c, 0x529cd808,
> 0xffffbfff, 0x406e0)
> 0x00821274 (oninit)startup (0xdffdb8, 0x7, 0x0, 0x4f0138d8, 0x0,
> 0x4f0138a0)
> 0x00818ef4 (oninit)mt_poll_yield(0x0, 0x0, 0x0, 0x0, 0x0, 0x0)
> 0x00000000 (*nosymtab*)0x0
>
> As we can see from the stack - the "bad" page is in a blobspace chunk. The
> messages and circumstances sure look like IC52728
> (http://www-01.ibm.com/support/docview.wss?uid=swg1IC52728).
>
> The instance was created via an imported restore - but that was back in
> July
> of last year. The engine did not crash (AFWARN) and the archive succeeded
> but
> I am wondering what the deal is. I know that I have dealt with this before
> and
> just cannot find the email chain from the last time we ran into it - but
> also
> after finding that bug it has me wondering...
>
> Thanks -
>
> MM
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--000e0cd292a6cc0d17048887aaeb
No crashes ever on this box - very solid version and OS. We don't have any logged databases and don't backup logs at all. The only logs that get backed up are those that are active during the archive. I had this same issue in the past and may have discussed it with a former IBM colleague outside of the forum. I am trying to jog the memory here... We can't really restore in this case - it is a prod server and would take ~12 hours to finish and since this was not a crash we would not be allowed to restore. Thanks! Mike
Well, the error would indicate that you have a blob free map page that
believes it has a blob page in use and it needs to be archived, however, when
it goes to read the blob page, it appears as though it's all 0's. So it could
be that you have physical corruption on that blob page and it got all zero'd
out. You could confirm that by running an oncheck -cD on tables that have blob
columns sorted in that blobspace and if you have a row that's pointing to that
blobpage, you should get at the very least a blobtimestamp mismatch error.
Another alternative could be that the blob free map page is mistakenly thinks
that blob page is in use so it wants to archive it, but since it isn't really
in use, and it's just a bunch of zero's, then the archive reports some error
about it but drives on. In this 2nd case, oncheck -cD wouldn't report any
errors because it doesn't examine the blob free map page so it wouldn't ever
see that page. I don't have any particularly good reason why you would just
start to see it again if you hadn't been seeing it previously.
Jacques Renaut
IBM Informix Advanced Support
APD Team
Cool thanks - It looks like the second scenario is the case here as we onchecked both tables that have blobs in the affected blob space and the onchecks (cDI) came back clean. Gonna see what happens tonight with the archive and go from there. We may end up dropping the table and recreating it. MM
hi Mike, I was excercising a workaround for similar issue few weeks ago, reading now Art's post I'm quite sure the 'bad page' issue was a result of engine crash - it was not on production but on Development machine. Our developers like to test new things like new JDBS/JBOSSes and one of these (pretty sure it was the 2nd one) was used by developers during a project. JBOSS kept a pool of database connections to the IDS. One of the connection from this pool was assigned to a developer which needed connection to IDS. This time developer was not a problem itself - the JBOSS was not clearing the memory after the developer's sessions finished - it was assigning new and new memory segments and strangely even crashing the IDS once (this is our IDS support company opinion and not my own). The fix for that blobdbs we did - as proposed by the same company - to: 1. create a blobdbspace (called blobdbs2) in the same size as the original (with bag page error) one. 2. alter all the tables from blobdbs -> blobdbs2 (single user mode should be used) 3. ensuring there's nothing more in blobdbs then drop 'blobdbs' 4. rename blobdbs2 -> blobdbs (hic! quiescent mode needed) I'm about to re-create the blobdbs in the original place on the device and do the same excercise to have the blobdbs back in-place, didn't have time to do this though. Hope that helped Waldek (IDS11.50FC6W4, SUSE10.3)
Waldek, To most easily move the blobspace back, you can: - Bring the instance offline - dd the blobkspace chunks to the original devices/files - move the existing links to the new copies - bring the engine back online No need to do the ALTER TABLE unless you can't take the engine offline for the time it will take to copy the chunk(s). (You can do a test copy while the engine is online to establish a downtime baseline.) Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Tue, Jun 8, 2010 at 4:17 PM, WALDEMAR ZNOINSKI <waldek@znoinski.pl>wrote: > hi Mike, > I was excercising a workaround for similar issue few weeks ago, reading now > Art's post I'm quite sure the 'bad page' issue was a result of engine crash > - > it was not on production but on Development machine. Our developers like to > test new things like new JDBS/JBOSSes and one of these (pretty sure it was > the > 2nd one) was used by developers during a project. JBOSS kept a pool of > database connections to the IDS. One of the connection from this pool was > assigned to a developer which needed connection to IDS. This time developer > was not a problem itself - the JBOSS was not clearing the memory after the > developer's sessions finished - it was assigning new and new memory > segments > and strangely even crashing the IDS once (this is our IDS support company > opinion and not my own). > > The fix for that blobdbs we did - as proposed by the same company - to: > 1. create a blobdbspace (called blobdbs2) in the same size as the original > (with bag page error) one. > 2. alter all the tables from blobdbs -> blobdbs2 (single user mode should > be > used) > 3. ensuring there's nothing more in blobdbs then drop 'blobdbs' > 4. rename blobdbs2 -> blobdbs (hic! quiescent mode needed) > > I'm about to re-create the blobdbs in the original place on the device and > do > the same excercise to have the blobdbs back in-place, didn't have time to > do > this though. > > Hope that helped > Waldek (IDS11.50FC6W4, SUSE10.3) > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --000e0cdf19b496116004888ac3e9
hi Art,
that would be a way quicker than altering (the same the ontape - the block
operations - is quicker than anything sql-like) but i have dbspaces devices
linked to one physical device (SAN RAID10 array) so all blob/temp/datadbs are
on the same device.
I could use dd with 'seek' option to skip a proper amount of blocks on my
device and start where the 'original' blobdbs was but let's leave our
experiments out the users knowledge and concious - I'll try that on my some
testing machine though ;-)
That's one reason that I like to use a volume manager to subdivide large
disk arrays into many smaller devices - one per chunk. You gain a lot in
flexibility to move chunks around from physical store to physical store when
you don't have to deal with offsets.
Enjoy.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Tue, Jun 8, 2010 at 5:13 PM, WALDEMAR ZNOINSKI <waldek@znoinski.pl>wrote:
> hi Art,
> that would be a way quicker than altering (the same the ontape - the block
> operations - is quicker than anything sql-like) but i have dbspaces devices
> linked to one physical device (SAN RAID10 array) so all blob/temp/datadbs
> are
> on the same device.
> I could use dd with 'seek' option to skip a proper amount of blocks on my
> device and start where the 'original' blobdbs was but let's leave our
> experiments out the users knowledge and concious - I'll try that on my some
> testing machine though ;-)
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--000e0cd2560e5b7dc104888b8066
Here is I ended up doing:
1. Create a new empty blobspace.
2. Determine all tables with blobs in the suspect blobspace.
3. Collect dbschema for each table found in step 2.
4. Edit the schema file to change the name of the blobspace from old to new.
5. Unload the affected tables.
6. Drop, recreate (using schema file from step 4), load affected tables.
7. Drop suspect blobspace.