Re: Assert Failed: Page Check Error on 7.24.UC5
Posted in 1999
> Any body have experience with page check errors? I have had several page
> check errors in the past several weeks and would like some insight into what
> may be causing the problem. The page check errors do not occur on just one
> table, but seem to be random. I would like an explanation of what the check
> is comparing (memory and disk pages?) and why they may not be in sync.
> We are running IDS 7.24.UC5 on Solaris 2.6 with veritas 2.5
Each disk page in an Informix database has two timestamps - one starting at the
fourth byte of the page header and one at the end of the page. Each time the
page is written, the timestamps are incremented by one. When the page is read
back from disk into buffer cache, the two timestamps are compared. They are
supposed to be equal. If they're not, then the page is considered a bad page.
What can cause this? Several things, many of which I'll probably not even think
of in this reply. One example would be a system crash at the exact moment that
the page was being written. This could allow timestamp1 to be updated, but not
timestamp2. One would hope that when the instance was brought back up that fast
recovery would roll the transaction back and fix this problem automatically, but
that assumes a logged database. Have there been any recent crashes? Is this a
logged database?
Another possibility would involve another process writing to the same area of
disk. Note that this other process would probably not be another instance of
Informix, as that instance would probably have consistent timestamps, barring a
crash, of course. If you have two Informix instances both using the same area
of raw disk you would probably get a slightly different error stating that the
page address (pg_addr, the first four bytes of the page header) was not what was
expected.
Yet another possibility could involve bugs in the OS or LVM, though I think
that's less likely. Failing disk drive(s) are also a possibility
I've seen something like this when a sysadm put a UNIX filesystem on a raw disk
that Informix was using, because he didn't think anything was using it. Look at
the contents of the page in question, using "oncheck -pP chknum pgnum", see if
it looks like data that should be in your table or if it looks like something
from another source.
In the case below, the pg_addr is 0x01b07435, stored in the buffer at address
0x0b866800, so try "oncheck -pP 0x01b 0x07435". From the excerpt you've
included, I have to ask whether your table would contain something like "ACTION:
FILL 52607413"? This almost looks like something from a report or a data
entry screen.
As you've seen, anytime Informix encounters one of these errors it shuts down
the engine to prevent any further corruption of data.
>
> here is an example: 10:59:46 Checkpoint Completed: duration was 1 seconds.
> 11:00:29 Assert Failed: Page Check Error in phposition:isposition:bad page
> 11:00:29 Who: Session(75274, XXXXXXXX@MACHINE, 27306, 227242944)
> Thread(75400, sqlexec, d058ad8, 14) 11:00:29 Results: Possible> inconsistencies in 'ameritrade:"XXXXXXX".TABLE' 11:00:29 Action: Run
> 'oncheck -cDI ameritrade:"XXXXXX".TABLE' 11:00:29 See Also:
> /dump/af.26888d9c, shmem.26888d9c.0, /opt/informix/core 11:00:39 mt.c, line
> 10176, thread 125, proc id 24374, Unexpected virtual processor termination,
> pid = 24403, exit = 0xb . 11:00:39 Assert Failed: Unexpected virtual
> processor termination, pid = 24403, exit = 0xb
>
> 11:00:39 Who: Session(2, informix@, 0, 0) Thread(125, kaio, 0, 1) 11:00:39> Results: Fatal Internal Error requires system shutdown 11:00:39 Action:
> Restart OnLine 11:00:39 See Also: /dump/af.7d8da6, shmem.7d8da6.0,
> /opt/informix/core 11:00:47 The Master Daemon Died 11:00:47 invoke_alarm():
> /bin/sh -c '/opt/informix/etc/no_log.sh 5 6 "Internal Subsystem failure: '
> MT'" "The Master Daemon Died" ' 11:00:47 invoke_alarm(): mt_exec failed,
> status -1, errno 0 11:00:47 PANIC: Attempting to bring system down
>
> top of af file:
>
> 11:00:29 bfcheck: bad page: pg_stamp cd6588eb != PG_STAMP2 d63289db> buffer header
> 0b2e184c: 00000000 00000000 00000000 0b30dd48 ........ .....0.H
> 0b2e185c: 0b2afae0 0a00c694 0b34c9d0 00020000 .*...... .4......
> 0b2e186c: 00400000 01b07435 0b866800 00000000 .@....t5 ..h.....
> 0b2e187c: 00000000 00010010 d633a36d ........ .3.m
> page
> 0b866800: 01b07435 cd6588eb 00012001 05630295 ..t5.e.. .. ..c..
> 0b866810: 00000000 00000000 41435449 4f4e3a20 ........ ACTION:
> 0b866820: 20202020 20204649 4c4c2035 32363037 FI LL 52607
> 0b866830: 34313320 20202020 20202020 20202020 413
> 0b866840: 20202020 20202020 2020494e 53545255 INSTRU
>
>
> 0b866850: 4d454e54 3a202020 20202020 20202020 MENT:
>
> Thanks!
>
> Jim Neugebauer
> jneugebauer@ameritrade.com
> Informix/Oracle DBA
> Systems Engineering Group
> Ameritrade Holding Corp.
> 4211 South 102nd St.
> Omaha, NE 68127-1031
> phone 402.597.5623
> fax 402.5977799
Mark Collins
mcollins@us.dhl.com
Everybody at some level realizes that the calendar is a fairly
arbitrary thing, invented by humans, for reasons that have more to
do with how committees are structured than with anything that's
really happening in the heavens. And yet people look at the fact
that the calendar is about to turn 2000, and assume there's some
deity who thinks that the base 10 counting system is pretty darn
important, and make all sorts of predictions of doom and gloom as a
result. To me, that tells you everything you need to know about
human beings -- and a whole lot about the market for the NC.
Scott Adams, creator of _Dilbert_