Re: Oninit from cron leaves engine in permanent fast recovery
Posted in 2005
Topics: Storage & Space Management, Server Administration
Sorry ...
... Nog does not have it. In fact, although there have been numerous
suggestions in response to this posting, all had already been tried or
were in place and nothing works. If anyone is interested, this problem
has been passed to IBM and they don't know why it won't start either.
The only information I have not been able to supply IBM is an analysis
of a core dump of the instance process.
Does anyone know 1) how to create a core dump of oninit. I have tried
sending all sorts of kill signals to do this. I can easily kill the
engine parent when it is stuck in FR but no core file is produced.
If I can get a core, I think I can use dbx to produce a stack/procedure
dump so that IBM can pinpoint the failure ... which, as people have
noticed, appears to have something to do with the initialisation of the
DBSPACETEMP list.
There is a slight problem with the latter. When cron shuts down the
instance, it will not restart with cron. But the very same offline
instance can be started by running oninit from the shell. It does not
have a problem with DBSPACETEMP or any of the dbspaces.
Please come forward if you REALLY know how to produce a core file from
an oninit.
Thanks for all your efforts.
Kind regards,
Gerry
TBP wrote:
> Nog wrote:
> > If your CONSOLE param in the ONCONFIG file is set to /dev/console try
> > pointing it at a plain file instead, e.g
> > /usr/informix/<instance>.console.
> >
> > I had a similar problem earlier this year when my server console daemon
> > left town and informix just wouldn't start, it was waiting to write to
> > console and couldn't
> >
> I think Nog has it :)
Also check the release notes....under $INFORMIXDIR/release.
On AIX do you need to create a KAIO device (aiodev??) or something. And
run something like strload?
Make sure NOAGE = 0 in onconfig as it may not be supported.
Check RESIDENT = 0 in onconfig as I have seen it not supported on AIX.
Pity you are not on Solaris as you could run truss to trace system
calls made by the process..not sure
what the equivalent is on AIX 5.2. On AIX 4.2.1 there was a truss-like
tool but it only can with the
'Performance Toolbox' that the client was not willing to spend money
on.
What does the online log say when it hangs?
If it is in recovery mode that it is probably rolling back a
transaction.
Does the output from onstat -l change at all..I don't have a 9.40 here
(running IDS 10 here) but
run onstat -l twice to a file 1 minute apart. The difference should be
more than just the first line with the
uptime in it!.
What does onstat -g ath show? Run it a few times...which threads are
running?
Does onstat -p run 1 minutes apart change?
Post the onstat -p and onstat -l difference as well as onstat -g ath
output.
Also check the release notes....under $INFORMIXDIR/release.
On AIX do you need to create a KAIO device (aiodev??) or something. And
run something like strload?
Make sure NOAGE = 0 in onconfig as it may not be supported.
Check RESIDENT = 0 in onconfig as I have seen it not supported on AIX.
Pity you are not on Solaris as you could run truss to trace system
calls made by the process..not sure
what the equivalent is on AIX 5.2. On AIX 4.2.1 there was a truss-like
tool but it only can with the
'Performance Toolbox' that the client was not willing to spend money
on.
What does the online log say when it hangs?
If it is in recovery mode that it is probably rolling back a
transaction.
Does the output from onstat -l change at all..I don't have a 9.40 here
(running IDS 10 here) but
run onstat -l twice to a file 1 minute apart. The difference should be
more than just the first line with the
uptime in it!.
What does onstat -g ath show? Run it a few times...which threads are
running?
Does onstat -p run 1 minutes apart change?
Post the onstat -p and onstat -l difference as well as onstat -g ath
output.
gerry.cassidy@dsl.pipex.com wrote:
> Sorry ...
>
> ... Nog does not have it. In fact, although there have been numerous
> suggestions in response to this posting, all had already been tried or
> were in place and nothing works. If anyone is interested, this problem
> has been passed to IBM and they don't know why it won't start either.
>
> The only information I have not been able to supply IBM is an analysis
> of a core dump of the instance process.
>
What about the online.log?
That has been asked for before in this thread and yet ...
Also, what does a ps -ef output show, immediately after the "explicit"
oninit -v command is run (just do ps -ef | grep on if you want to limitthe output, would be interesting to see the oninit processes running,
along with anything else starting with "on").