Re: 48 out of 58 chunks down
Posted in 1997
In article <s03aHAAgfhGzEwzc@kirzel.demon.co.uk>, Peter R Wotherspoon
<peter@kirzel.demon.co.uk> writes
>This is a plea for urgent help with a failure of one of our Informix
>instances.
>
>My company runs on an Informix software package, using Online 5.03,
>which we have been upgrading over the weekend.
>
>Because the upgrade involved many "alter table" statements, we archived,
>ran the upgrade without logging, then attempted to re-archive and
>restart logging.
>
>At this point, Informix decided to mark 48 out of 58 unmirrored chunks
>as "down". In fact, the chunks are mirrored in hardware, and are all OK
>from a UNIX point of view (mode=660 u=informix g=informix, too, before
>you ask). Chunks 1 to 10 are OK, chunks 11 to 58 are all down. That
>seems too neat: there must be some reason for the pattern.
>
Are the down chunks all on the same
a) disk partition(s) for cooked (filesystem) chunks)
b) raw device(s) for raw devices.
If b) then do a prtvtoc on the devices, possibly the VTOC on the
disk has been overwritten meaning that the whole of the disk
becomes inaccesible (include raw disk accesses).
Is the VTOC is unreadable then fdisk unusally provides a way to
restore a backup copy. If this works then
1. restore an archive
2. dbexport to a fast tape device (E.g. 4mm DAT)
3. drop and recreate the chunks with 100K larger offsets and
150K smaller sizes and dbimport back again off-tape.
DO NOT ontape/onarchive since a restore will reset the chunk
offsets/size back to the 'wrong' versions.
>As I understand it, what we need to do is correct the problem which is
>preventing Informix from accessing the chunks, then restore from the
>last archive, which is the only supported means of recovering from a
>non-mirrored chunk down. But we can't see any reason why the chunks
>(most of which worked fine before the upgrade) are not useable.
>
Correct - see above for one possible problem + solution.
Another maybe that these chunks are on a disk pack which was
not switch on when the machine was rebooted. This means Online
will fail to read the start of the chunks (as below) and flag them
as down. In this case Informix Tech Support do have a tool which
you can use to mark the chunk as Up. Then do 1-4 above and
check your data is ok afterwards.
>Can anyone help? We have contacted Informix (and agreed to pay them),
>but have heard nothing back from them yet.
>
>The critical part of the message log reads (manually transcribed):-
>
>I/O read() chunk 11, pagenum 4, pagecnt 1 --> errno = 22
Could not read the start of the chunk errno 22 = EINVAl (Invalid
argumnet) e.g. device not switched on hence UNIX does not include
it in the list of hardware devices and the /dev/dsk.... special
devices files start returning EINVAL errors (Not a valid device).
>tblspace header error:
>de8958: 00000000 00000000 00000000 00000000
>de8968: 01000000 0100000b 0400b000 00000000
Header error for the tblspace tblspace at the start of the
first chunk in a dbspace? (Skip two reserved pages + chunk free
list page and pagenum 4 = first page of tblspace tblspace!)
>(13 lines of zeroes omitted)
>de8a48: 00000000 182e0601 ed09a046 00000000
>de8a48: 588e1001 X...
>-- Fail Consistency Check -- pthdrpage:ptalloc:bad bfget -- pid=1234
pthdrpage = partition (tblspace) header page
ptalloc = partition alloc (tblspace allocation?)
bad bfget = bad buffer get?
>user=102 us=401bc0
>Cannot Open Dbspace 11
>
>This is all repeated 47 more times.
>
>Please copy any reply to "chris.webb@theso.co.uk". Thanks to anyone who
>can help. Please excuse the use of bandwidth, but we are pretty
>desperate.
Know the feeling - reply by 1am with a telephone number if you want
futher help...I've been there before...
--
David Williams