Mysterious Master Daemon death
Posted in 2004
Topics: High Availability & Replication, Storage & Space Management, Error Codes & Troubleshooting, Server Administration, Clustering, Grid & MACH11
Ok, so
I'm ripping my hair out on this one now.
For the fourth time this year, and the second time
this month, our main Informix database running on
version 7.31.UD7X2/Solaris 9 crashed mysteriously with
only the lines:
14:13:33 The Master Daemon Died
14:13:34 PANIC: Attempting to bring system down
There is no Assert Failure message or any other
forboding warning above or below this message. Nor is
there anything suspicious in /var/adm/messages. It's
just that CPU VP #1 seems to kick the bucket all of a
sudden.
I'm aware that this is consistent with someone or
something external to Informix sending a kill -9
signal to CPU VP #1. I've tested that and seen it for
myself. However, the only other applications we have
on there are Veritas Cluster Server for
high-availability (because my company likes it better
than HDR), and Big Brother for monitoring purposes.
We did experience these Master Daemon deaths a couple
times even when VCS wasn't running. Really, neither I
nor anyone else here can think of anything that could
possibly be messing with Informix like this.
IBM Informix technical support, of course, insists
that this is external to Informix, and largely based
on my experience before this started happening, I'm
inclined to agree with them. Unfortunately, no one
else at my company is buying that explanation.
We've attempted to put a wrapper script around the
kill command to identify who's running it, but with
Solaris, almost all kills, even from the shell command
line, don't actually use /bin/kill, they use a kernel
call. We're not too fond of the idea of hacking the
kernel to monitor that. :)
We were also given the suggestion to truss CPU VP 1.
This might at least identify that a kill has been
issued, but this is a pretty hefty burden to place on
the db, particularly if it's weeks before this happens
again. We can run a cron that lops off the voluminous
output from this every so often to keep disk space
from filling, but just the I/O involved here could
feed back badly onto the database.
We also tried running onstat -i (interactive mode) in
a persistent window so that if this were to occur
again, at least we could take onstat commands for
further analysis. Informix told me, and I tested and
confirmed this, that even if the engine dies, the
interactive onstat command stays attached to the
shared memory, and can continue giving onstat outputs
based on that memory. You can even re-start Informix,
and the interactive onstat still looks at the memory
from the instance before it was killed. (Actually,
that's kind of creepy. Cool, but creepy.)
Here's the interesting bit, though. This last time,
when the Master Daemon died, I went back to the
persistent window running the interactive onstat, and
it didn't stay attached to shared memory. I tried to
run one command there, and it gave me the "shared
memory not initialized" message, and exited.
This seems really suspicious to me. Why would it be
in a test environment that I could kill Informix with
a kill -9, and onstat -i does indeed continue to see
the memory of the dead instance, but in this
mysterious case in production, it doesn't? Indeed,
this behavior makes it seem more likely that something
bad really did happen within Informix.
Informix has really become persona-non-grata around
here, and without any af files or other log messages
to go on, this problem has become impossible to
troubleshoot. Has *anyone* seen this behavior before,
or know anything about it? Any help here may save my
job and/or our continued use of Informix. Thanks very
much.
--John Bejarano
Informix DBA, Shutterfly
Hi John,
If somebody is killing ur oninit process you can try some of the below -
1) If u run informix thru informix id, u can try to change and run
using root. This will increase one level of security as then anybody
killing oninit will need root id or a script with setuid bit of root to
kill the same.
2) In my last company, for audit purposes , they had a utility which
would run when u login ( as root ) and write all inputs to a file owned
by root ( and u can place this file somewhere which only few people know
abt ), there are ways to get around this too, but still can help.
hope this helps
ps: I dont have the software which logs activities of all users, ask ur
sysadmin if he/she can help
Rgds
Preetinder
John Bejarano wrote:
>Ok, so I'm ripping my hair out on this one now.
>
>For the fourth time this year, and the second time
>this month, our main Informix database running on
>version 7.31.UD7X2/Solaris 9 crashed mysteriously with
>only the lines:
>
>14:13:33 The Master Daemon Died
>14:13:34 PANIC: Attempting to bring system down
>
>There is no Assert Failure message or any other
>forboding warning above or below this message. Nor is
>there anything suspicious in /var/adm/messages. It's
>just that CPU VP #1 seems to kick the bucket all of a
>sudden.
>
>I'm aware that this is consistent with someone or
>something external to Informix sending a kill -9
>signal to CPU VP #1. I've tested that and seen it for
>myself. However, the only other applications we have
>on there are Veritas Cluster Server for
>high-availability (because my company likes it better
>than HDR), and Big Brother for monitoring purposes.
>We did experience these Master Daemon deaths a couple
>times even when VCS wasn't running. Really, neither I
>nor anyone else here can think of anything that could
>possibly be messing with Informix like this.
>
>IBM Informix technical support, of course, insists
>that this is external to Informix, and largely based
>on my experience before this started happening, I'm
>inclined to agree with them. Unfortunately, no one
>else at my company is buying that explanation.
>
>We've attempted to put a wrapper script around the
>kill command to identify who's running it, but with
>Solaris, almost all kills, even from the shell command
>line, don't actually use /bin/kill, they use a kernel
>call. We're not too fond of the idea of hacking the
>kernel to monitor that. :)
>
>We were also given the suggestion to truss CPU VP 1.
>This might at least identify that a kill has been
>issued, but this is a pretty hefty burden to place on
>the db, particularly if it's weeks before this happens
>again. We can run a cron that lops off the voluminous
>output from this every so often to keep disk space
>from filling, but just the I/O involved here could
>feed back badly onto the database.
>
>We also tried running onstat -i (interactive mode) in
>a persistent window so that if this were to occur
>again, at least we could take onstat commands for
>further analysis. Informix told me, and I tested and
>confirmed this, that even if the engine dies, the
>interactive onstat command stays attached to the
>shared memory, and can continue giving onstat outputs
>based on that memory. You can even re-start Informix,
>and the interactive onstat still looks at the memory
>from the instance before it was killed. (Actually,
>that's kind of creepy. Cool, but creepy.)
>
>Here's the interesting bit, though. This last time,
>when the Master Daemon died, I went back to the
>persistent window running the interactive onstat, and
>it didn't stay attached to shared memory. I tried to
>run one command there, and it gave me the "shared
>memory not initialized" message, and exited.
>
>This seems really suspicious to me. Why would it be
>in a test environment that I could kill Informix with
>a kill -9, and onstat -i does indeed continue to see
>the memory of the dead instance, but in this
>mysterious case in production, it doesn't? Indeed,
>this behavior makes it seem more likely that something
>bad really did happen within Informix.
>
>Informix has really become persona-non-grata around
>here, and without any af files or other log messages
>to go on, this problem has become impossible to
>troubleshoot. Has *anyone* seen this behavior before,
>or know anything about it? Any help here may save my
>job and/or our continued use of Informix. Thanks very
>much.
>
>--John Bejarano
>Informix DBA, Shutterfly
>
>
>
>
>
Related threads
- Micro-pump is cool idea for future computer chips
- FW: Informix DBA
- How tu use a table on another Instance
- lock table in exclusive mode