Re: HELP! IUS crashes during level 0 dump
Posted in 1998
>From: David Buttrick <buttrick@sportingnews.com>
>Date: Tue, Jan 13, 1998 11:19 EST
>Message-id: <34BB93FD.666737D6@sportingnews.com>
>
>My problem is that IUS cannot re-connect to a tempdbs after we do a
>level 0 dump.
>
>It reports this:
>
>09:09:54 mt_aio_wait(32f8b158)
>09:09:54 chunkio(6, 0x34136018, 0xe1, 32, 0x32f167c8, 34, 1) errno 22
>
>09:09:54 Assert Failed: I/O error, Primary Chunk
>'/usr/informix/chunks/tempdbs1
>' -- Offline
>09:09:54 Who: Session(1301, informix@dev, 3638, 0)
> Thread(1338, arcbackup2, 0, 1)
> File: rsbuff.c Line: 2939
>09:09:54 Results: Chunk is now unusable>
>Has anyone seen this before? If so, what was your solution?
>
>Your prompt attention to this matter is appreciated.
>
>I've noticed that there is a tremendous amount of activity in the
>tempdbs, so I am wondering if there is some way to flush the tempdbs,
>without taking the database into quiescent mode.
>
>Your prompt attention to this matter is appreciated.
>--
>David Buttrick buttrick@sportingnews.com
>WEB Developer (314) 993-7727
>The Sporting News http://www.sportingnews.com
> S E E A D I F F E R E N T G A M E
>
Dave,
I just checked out your case (692603). I don't really think that the problem
has anything to do with the fact that a backup was underway or that you had a
tempdbs. It looks like you actually had a memory corruption problem with a web
application. I suspect that the issue with the temp dbs is really a result of
this corruption.
I noticed that Randy is getting a debuggable underway for you. Let me suggest a
couple of things that might help in the isolation of the problem.
Try running with memory scribble turned on. What this will do is to mark
memory with special patterns when it is freed and/or allocated. This can be
turned on by "onmode -f 0x200".
This will NOT make the problem go away. What it will do is to cause the
failure to occur much closer at the spot where the corruption is actually
occuring.
Often with memory corruption problems, the problem is not actually discovered
untill way after the point that the problem actually occured. By causing the
failure much closer to the point of the corruption, you have a much better
chance of getting it fixed quicker.
So, "onmode -f 0x200" should make the failure more likely to occur and hence
easier to resolve.
Of course, I'm not encouraging you to attempt this during prime time. But if
you have a test machine (or test instance) that can run with this on, it might
make it easier to discover the root cause.
A couple of other "onmode -f" commands that might come in useful.
0x1 - run a check on all memory pools with each memory allocation -- This is
really slow.
0x8 - check the global memory pools once a second.
These are additive. So you could run with "onmode -f 0x201"
I'm going to email this to Randy House. The two of you might want to discuss
this approach tomorrow. Randy knows his stuff. Have him give me a call, if
you want.
Madison Pruet