chunks overlapping?
Posted in 2007
A 7.30.UC6 server on Solaris crashed with "Assert Failed: Page Check Error in undopgalter: slots overlap", then aborted again during logical recovery, prompting the poster to suspect overlapping chunks on raw slices. His own arithmetic showed the slices were big enough; Neil Truby warned that the rootdbs slice starting at cylinder 0 risks Solaris overwriting it with the VTOC, while others said the symptom is a corrupted page needed for rollback/logical recovery and advised oncheck -pr plus contacting support to bypass/truncate the transaction, or restoring from archive. The poster had already restored, and no definitive cause was established.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Storage & Space Management, Error Codes & Troubleshooting
Yesterday our system went down with the log file showing this below:
quote
16:18:33 Informix Dynamic Server Version 7.30.UC6 Software SerialNumber AAC#
J540938
16:18:35 Informix Dynamic Server Initialized -- Shared MemoryInitialized.
16:18:35 Physical Recovery Started.
16:18:35 Physical Recovery Complete: 1696 Pages Restored.
16:18:35 Logical Recovery Started.
16:18:43 Assert Failed: Page Check Error in undopgalter: slotsoverlap
16:18:43 Informix Dynamic Server Version 7.30.UC6
16:18:43 Who: Session(13, informix@sunhost, 0, 0)
Thread(25, fast_rec, 0, 1)
File: rsdebug.c Line: 943
16:18:43 Results: Possible inconsistencies in 'mydb:"".'
16:18:43 Action: Run 'oncheck -cD 1048930'
16:19:01 See Also: /tmp/af.19a863, shmem.19a863.0
16:19:10 rstrans.c, line 7788, thread 25, proc id 5953, DynamicServer must abo
rt.
16:19:10 Assert Failed: Dynamic Server must abort
16:19:10 Informix Dynamic Server Version 7.30.UC6
16:19:10 Who: Session(13, informix@sunhost, 0, 0)
Thread(25, fast_rec, 0, 1)
File: called by ASF mt_affail Line: 0
16:19:10 Results: Fatal Internal Error requires system shutdown
unquote
Now, I searched this group and some suggested that there might be some
chunk overlapping issue. I wonder if my current configuration indeed
has this problem. Upon recovery, the dbspace info shown was:
chk/dbs offset size free bpages flags pathname
1 1 4096 1000000 868971 PO- /ifxRawDvc.d/rootDbs
2 1 2048 1000000 3 PO- /ifxRawDvc.d/
dataChunk1
3 1 2048 1000000 58953 PO- /ifxRawDvc.d/
dataChunk2
4 1 2048 1000000 999997 PO- /ifxRawDvc.d/
dataChunk3
The raw devices are in fact symbolic links, info here:
lrwxrwxrwx 1 informix informix 18 Apr 5 2006 dataChunk1 -> /
dev/rdsk/c2t0d0s1
lrwxrwxrwx 1 informix informix 18 Apr 5 2006 dataChunk2 -> /
dev/rdsk/c2t0d0s2
lrwxrwxrwx 1 informix informix 18 Apr 5 2006 dataChunk3 -> /
dev/rdsk/c2t0d0s3
lrwxrwxrwx 1 root other 18 Apr 4 2006 rootDbs -> /dev/
rdsk/c2t0d0s0
Now for the partition table on the harddisk where the database reside:
Current partition table (original):
Total disk cylinders available: 52095 + 2 (reserved cylinders)
Part Tag Flag Cylinders Size Blocks
0 unassigned wm 0 - 1024 2.00GB (1025/0/0)
4198400
1 unassigned wm 1025 - 2049 2.00GB (1025/0/0)
4198400
2 unassigned wm 2050 - 3074 2.00GB (1025/0/0)
4198400
3 unassigned wm 3075 - 4099 2.00GB (1025/0/0)
4198400
4 unassigned wm 0 0
(0/0/0) 0
5 unassigned wm 0 0
(0/0/0) 0
6 unassigned wm 7051 - 32650 50.00GB (25600/0/0)
104857600
7 unassigned wm 0 0
(0/0/0) 0
Now, I did some calculations and got the following results:
1. size of a slice that is assigned to a chunk, according to the
partition table, is 4198400/2 = 2099200k
2. size of the rootdbs = (4096 + 1000000) * 2k = 2008192k
3. size of other data chunks, each = (2048 + 1000000) * 2k =
2004096k
Since 2099200 > 2008192 > 2004096, I fail to see how the chunks can
overlap each other. If any of you gurus see anything wrong in my
logic, please please tell. Many thanks in advance. Or maybe there was
no chunk overlap and the error had other causes?
"emrefan" <dksleung@hotmail.com> wrote in message news:1181799101.149939.132380@j4g2000prf.googlegroups.com... > > Part Tag Flag Cylinders Size Blocks > 0 unassigned wm 0 - 1024 2.00GB (1025/0/0) > 4198400 > 1 unassigned wm 1025 - 2049 2.00GB (1025/0/0) > 4198400 > 2 unassigned wm 2050 - 3074 2.00GB (1025/0/0) > 4198400 > 3 unassigned wm 3075 - 4099 2.00GB (1025/0/0) > 4198400 > 4 unassigned wm 0 0 > (0/0/0) 0 > 5 unassigned wm 0 0 > (0/0/0) 0 > 6 unassigned wm 7051 - 32650 50.00GB (25600/0/0) > 104857600 > 7 unassigned wm 0 0 > (0/0/0) 0 You don't give the operating system but I assume Solaris, in which case the problem is that your rootdbs has been partially over-written and you will need to restore the database server. You must not use the first couple of cylinders of a disk for a raw volume, as Solaris will overwrite them with the disk's VTOC each time it reboots.
> You don't give the operating system but I assume Solaris, in which case the > problem is that your rootdbs has been partially over-written and you will > need to restore the database server. > You must not use the first couple of cylinders of a disk for a raw volume, > as Solaris will overwrite them with the disk's VTOC each time it reboots.- Hide quoted text - Thanks Neil! I already have 4096 x 2k (2k = 1 informix page) at the start of the device reserved. But maybe that's not enough... Ok, I will have to rearrange it again, if just to be safe.
Your devs look ok to me. the s0 is a bit hairy since it starts at 0,
you could be lucky though that the vtoc did not overwrite anything.
i guess not since you would probably have other errors.
(run oncheck -pr if that displays stuff without bitching it should be
fine!!!!!!)
you should contact TS and let them bypass recovery or... or you can
restore from
archive, slot overlap in a page means a page is bad which is needed
for logical recovery.
Superboer.
On 14 jun, 07:31, emrefan <dksle...@hotmail.com> wrote:
> Yesterday our system went down with the log file showing this below:
>
> quote
>
> 16:18:33 Informix Dynamic Server Version 7.30.UC6 Software Serial> Number AAC#
> J540938
> 16:18:35 Informix Dynamic Server Initialized -- Shared Memory> Initialized.
> 16:18:35 Physical Recovery Started.
> 16:18:35 Physical Recovery Complete: 1696 Pages Restored.
> 16:18:35 Logical Recovery Started.
> 16:18:43 Assert Failed: Page Check Error in undopgalter: slots> overlap
> 16:18:43 Informix Dynamic Server Version 7.30.UC6
> 16:18:43 Who: Session(13, informix@sunhost, 0, 0)
> Thread(25, fast_rec, 0, 1)
> File: rsdebug.c Line: 943
> 16:18:43 Results: Possible inconsistencies in 'mydb:"".'
> 16:18:43 Action: Run 'oncheck -cD 1048930'
> 16:19:01 See Also: /tmp/af.19a863, shmem.19a863.0
> 16:19:10 rstrans.c, line 7788, thread 25, proc id 5953, Dynamic> Server must abo
> rt.
> 16:19:10 Assert Failed: Dynamic Server must abort
> 16:19:10 Informix Dynamic Server Version 7.30.UC6
> 16:19:10 Who: Session(13, informix@sunhost, 0, 0)
> Thread(25, fast_rec, 0, 1)
> File: called by ASF mt_affail Line: 0
> 16:19:10 Results: Fatal Internal Error requires system shutdown>
> unquote
>
> Now, I searched this group and some suggested that there might be some
> chunk overlapping issue. I wonder if my current configuration indeed
> has this problem. Upon recovery, the dbspace info shown was:
>
> chk/dbs offset size free bpages flags pathname
> 1 1 4096 1000000 868971 PO- /ifxRawDvc.d/rootDbs
> 2 1 2048 1000000 3 PO- /ifxRawDvc.d/
> dataChunk1
> 3 1 2048 1000000 58953 PO- /ifxRawDvc.d/
> dataChunk2
> 4 1 2048 1000000 999997 PO- /ifxRawDvc.d/
> dataChunk3
>
> The raw devices are in fact symbolic links, info here:
>
> lrwxrwxrwx 1 informix informix 18 Apr 5 2006 dataChunk1 -> /
> dev/rdsk/c2t0d0s1
> lrwxrwxrwx 1 informix informix 18 Apr 5 2006 dataChunk2 -> /
> dev/rdsk/c2t0d0s2
> lrwxrwxrwx 1 informix informix 18 Apr 5 2006 dataChunk3 -> /
> dev/rdsk/c2t0d0s3
> lrwxrwxrwx 1 root other 18 Apr 4 2006 rootDbs -> /dev/
> rdsk/c2t0d0s0
>
> Now for the partition table on the harddisk where the database reside:
>
> Current partition table (original):
> Total disk cylinders available: 52095 + 2 (reserved cylinders)
>
> Part Tag Flag Cylinders Size Blocks
> 0 unassigned wm 0 - 1024 2.00GB (1025/0/0)
> 4198400
> 1 unassigned wm 1025 - 2049 2.00GB (1025/0/0)
> 4198400
> 2 unassigned wm 2050 - 3074 2.00GB (1025/0/0)
> 4198400
> 3 unassigned wm 3075 - 4099 2.00GB (1025/0/0)
> 4198400
> 4 unassigned wm 0 0
> (0/0/0) 0
> 5 unassigned wm 0 0
> (0/0/0) 0
> 6 unassigned wm 7051 - 32650 50.00GB (25600/0/0)
> 104857600
> 7 unassigned wm 0 0
> (0/0/0) 0
>
> Now, I did some calculations and got the following results:
>
> 1. size of a slice that is assigned to a chunk, according to the
> partition table, is 4198400/2 = 2099200k
>
> 2. size of the rootdbs = (4096 + 1000000) * 2k = 2008192k
>
> 3. size of other data chunks, each = (2048 + 1000000) * 2k =
> 2004096k
>
> Since 2099200 > 2008192 > 2004096, I fail to see how the chunks can
> overlap each other. If any of you gurus see anything wrong in my
> logic, please please tell. Many thanks in advance. Or maybe there was
> no chunk overlap and the error had other causes?
On Jun 14, 4:37 pm, Superboer <superbo...@t-online.de> wrote:
> Your devs look ok to me. the s0 is a bit hairy since it starts at 0,
> you could be lucky though that the vtoc did not overwrite anything.
> i guess not since you would probably have other errors.
>
> (run oncheck -pr if that displays stuff without bitching it should be
> fine!!!!!!)
>
> you should contact TS and let them bypass recovery or... or you can
> restore from
> archive, slot overlap in a page means a page is bad which is needed
> for logical recovery.
>
> Superboer.
Thanks for the good advice but I can't benefit this time because the
restore had been done there and then. The error did not occur right
after a reboot and one thing suspicious was that chunk 2 might have
just been used up (only 3 pages left according to onstat) and the db
was just starting to use a new chunk (chunk 4, which has 999997 pages
out of 1 million free). I am not sure but the error could have
happened right at the moment of switching.
"Superboer" <superboer7@t-online.de> wrote in message
news:1181810274.448301.285450@i13g2000prf.googlegroups.com...
> Your devs look ok to me. the s0 is a bit hairy since it starts at 0,
> you could be lucky though that the vtoc did not overwrite anything.
> i guess not since you would probably have other errors.
>
> (run oncheck -pr if that displays stuff without bitching it should be
> fine!!!!!!)
>
> you should contact TS and let them bypass recovery or... or you can
> restore from
> archive, slot overlap in a page means a page is bad which is needed
> for logical recovery.
He's running an unsupported version.
> He's running an unsupported version. so what, he is deep shit, ask TS. he may have to negotiate payment whatever dono, just try it. Superboer.
How did the system go down BEFORE this ...
onmode -ky?shutdown -r
Would be good to see the online.log
emrefan wrote:
> Yesterday our system went down with the log file showing this below:
>
> quote
>
> 16:18:33 Informix Dynamic Server Version 7.30.UC6 Software Serial> Number AAC#
> J540938
> 16:18:35 Informix Dynamic Server Initialized -- Shared Memory> Initialized.
> 16:18:35 Physical Recovery Started.
> 16:18:35 Physical Recovery Complete: 1696 Pages Restored.
> 16:18:35 Logical Recovery Started.
> 16:18:43 Assert Failed: Page Check Error in undopgalter: slots> overlap
> 16:18:43 Informix Dynamic Server Version 7.30.UC6
> 16:18:43 Who: Session(13, informix@sunhost, 0, 0)
> Thread(25, fast_rec, 0, 1)
> File: rsdebug.c Line: 943
> 16:18:43 Results: Possible inconsistencies in 'mydb:"".'
> 16:18:43 Action: Run 'oncheck -cD 1048930'
> 16:19:01 See Also: /tmp/af.19a863, shmem.19a863.0
> 16:19:10 rstrans.c, line 7788, thread 25, proc id 5953, Dynamic> Server must abo
> rt.
> 16:19:10 Assert Failed: Dynamic Server must abort
> 16:19:10 Informix Dynamic Server Version 7.30.UC6
> 16:19:10 Who: Session(13, informix@sunhost, 0, 0)
> Thread(25, fast_rec, 0, 1)
> File: called by ASF mt_affail Line: 0
> 16:19:10 Results: Fatal Internal Error requires system shutdown>
On Jun 14, 6:36 pm, "TBP (The Big Potato)" <T...@NotHere.Co.Uk> wrote:
> How did the system go down BEFORE this ...
>
> onmode -ky?> shutdown -r
>
> Would be good to see the online.log
Well, here it is, thanks for the offer.
quote
14:45:23 Checkpoint Completed: duration was 0 seconds.
14:50:23 Checkpoint Completed: duration was 0 seconds.
14:55:23 Checkpoint Completed: duration was 0 seconds.
15:00:23 Checkpoint Completed: duration was 0 seconds.
15:02:02 Assert Failed: Page Check Error in undopgalter: slotsoverlap
15:02:02 Informix Dynamic Server Version 7.30.UC6
15:02:02 Who: Session(858985, me@newe450, 27223, 0)
Thread(874575, sqlexec, 0, 1)
File: rsdebug.c Line: 943
15:02:02 Results: Possible inconsistencies in 'mydb:"".'
15:02:02 Action: Run 'oncheck -cD 1048930'
15:02:14 See Also: /tmp/af.584f9669, shmem.584f9669.0
15:06:34 rstrans.c, line 7788, thread 874575, proc id 28320, DynamicServer must abort.
15:06:34 Assert Failed: Dynamic Server must abort
15:06:34 Informix Dynamic Server Version 7.30.UC6
15:06:34 Who: Session(858985, me@newe450, 27223, 0)
Thread(874575, sqlexec, 0, 1)
File: called by ASF mt_affail Line: 0
15:06:34 Results: Fatal Internal Error requires system shutdown
15:06:34 Action: Restart OnLine
15:06:52 See Also: /tmp/af.584f9669, shmem.584f9669.1
15:14:37 rstrans.c, line 7788, thread 874575, proc id 28320, DynamicServer must abort.
15:14:37 PANIC: Attempting to bring system down
unquote
emrefan wrote:
> On Jun 14, 6:36 pm, "TBP (The Big Potato)" <T...@NotHere.Co.Uk> wrote:
>
>>How did the system go down BEFORE this ...
>>
>>onmode -ky?>>shutdown -r
>>
>>Would be good to see the online.log
>
>
> Well, here it is, thanks for the offer.
>
> quote
>
> 14:45:23 Checkpoint Completed: duration was 0 seconds.
> 14:50:23 Checkpoint Completed: duration was 0 seconds.
> 14:55:23 Checkpoint Completed: duration was 0 seconds.
> 15:00:23 Checkpoint Completed: duration was 0 seconds.
> 15:02:02 Assert Failed: Page Check Error in undopgalter: slots> overlap
> 15:02:02 Informix Dynamic Server Version 7.30.UC6
> 15:02:02 Who: Session(858985, me@newe450, 27223, 0)
> Thread(874575, sqlexec, 0, 1)
> File: rsdebug.c Line: 943
> 15:02:02 Results: Possible inconsistencies in 'mydb:"".'
> 15:02:02 Action: Run 'oncheck -cD 1048930'
> 15:02:14 See Also: /tmp/af.584f9669, shmem.584f9669.0
> 15:06:34 rstrans.c, line 7788, thread 874575, proc id 28320, Dynamic> Server must abort.
> 15:06:34 Assert Failed: Dynamic Server must abort
> 15:06:34 Informix Dynamic Server Version 7.30.UC6
> 15:06:34 Who: Session(858985, me@newe450, 27223, 0)
> Thread(874575, sqlexec, 0, 1)
> File: called by ASF mt_affail Line: 0
> 15:06:34 Results: Fatal Internal Error requires system shutdown
> 15:06:34 Action: Restart OnLine
> 15:06:52 See Also: /tmp/af.584f9669, shmem.584f9669.1
> 15:14:37 rstrans.c, line 7788, thread 874575, proc id 28320, Dynamic> Server must abort.
> 15:14:37 PANIC: Attempting to bring system down>
> unquote
>
and now the first 20 or so lines from the top of the af file (you would be getting one each time you try and start)
I think you are running into a rollback defect, which ... you would need support to "truncate" transactions.