HHHHEEEELLLLPPPP!!!! My servers have gone nuts!
Posted in 2000
Art Kagel reported that all his DG/UX (M88K) servers running Informix 7.24UC5/UC7/UC8 were crashing with repeated "Invalid Mutex Type" assert failures, hangs at <CKPT REQ>, and restart failures (semot errno=22, shared-memory pool corruption). Fixes were inconsistent: restoring the original ONCONFIG, dropping network connections to shared-memory only, or replacing hostname/service names in sqlhosts with a raw IP address and port number. Others suggested a possible network attack, checking af stack traces and kernel syslogs, mismatched INFORMIXDIR/versions, and DNS problems. No resolution is recorded; an Informix case (990681) remained open and crashes continued.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Error Codes & Troubleshooting, Server Administration, Logging & Checkpoints, Networking & sqlhosts Configuration
Hi folk,
My turn to need help. I have a very weird situation, all of my DG/UX 4.11
(Motorola M88K based) servers have crashed and refused to restart over the
last two weeks. In all cases I have been able to restart them but the
solution that works on one does not work on another (with a few exceptions).
A configuration that refused to start at 8AM starts fine 2 hours later and
a configuration that is the only one working on a second machine fails on
the first machine after a subsequent crash. Online version(s) on the
various machines? 7.24UC5, 7.24UC7, & 7.24UC8 (the last release for this
platform ever!)
That's the quick and scary and right now only my development server is down
having crashed 11 minutes after a production server that was fixed by making
a change that fixed this one on Friday. I have been online with Informix
off and on for two weeks and nothing. I will not post the onconfigs right
now because as you will see the parameters seem not to matter (with noted
exceptions). Ask for them if you like :-)
Symptoms:
First I see :
10:41:02 mt.c, line 6796, thread 15, proc id 11331, Invalid Mutex Type.
10:41:02 Assert Failed: Invalid Mutex Type
.....
10:41:02 Results: Fatal Internal Error requires system shutdown
over an over from various threads, mostly aio vp threads. Then the engine
reaches the next checkpoint and goes into <CKPT REQ> and stays there and
existing threads and new connections hang. You cannot shutdown and
onstat -g glo report shared memory changing so it cannot report. Sometimesone or more VPs physically crash. Cores and shared memory dumps show
nothing useful to me or tech support. I can kill -PIPE the oninits and
drag the engine down then remove any remaining shared memory, but on
startup I get:
10:55:01 DR: DRAUTO id 0 (Off)
10:55:32 mt.c, line 1210, thread 0, proc id 3101, semot: errno = 22
10:55:32 Assert Failed: errno = 22
....
10:55:32 Results: Fatal Internal Error requires system shutdown
OR
11:10:28 DR: DRAUTO id 0 (Off)
11:10:57 Assert Failed: Memory free block header corruption detected in
mt_shm_malloc_segid 1
11:10:57 Who: Thread(106, flush_sub(1), 0, 1)
11:10:57 Results: Pool repaired
11:10:57 Action: Please notify Informix Technical Support
then the previous errno = 22 messages stream mixed with Condition Failed:
(different threads), In (hang_thread)
messages.
On the first machine this happened (a week ago Sunday) Informix changed
just about every parameter in the onconfig file one or two at a time adding
and deleting interfaces. Finally after 12 hours and two techs in two
countries we got the engine up with 1/2 the CPU VPs and 1/4th the number of
AIO VPs but it crashed just 2 hours later. I then restored the original
ONCONFIG and started it again and it's been up for 8 days now, go figure.
Then this past Friday a second machine, same symptoms. Finally Sunday AM
we discovered that removing all the network connections and just running
shared memory worked and since on that machine all connections are local
(all users come into local middleware servers through other mechanisms) it
was OK for the nonce. (Of course we had tried that on the machine the week
before with no luck!) Meanwhile Friday evening a development machine went
down and we were able to restart by changing the hostname and services in
sqlhosts for TCP address and port number. So when a third production server
went GA-GA this AM and nothing else worked I tried the address/port trick
and voila it worked here also (which is good because this machine cannot
live without networked connections! Only to discover that the development
server had also crashed again 10 minutes after that production server
although I have been able to restart that one.
All of the affected machines had kernels built, independently, on Feb. 14,
some have been booted recently before the crash one had been online, OS &
engine, since March 6 without a peep.
Has ANYONE seen such nonsense before?
Art S. Kagel
Case Number???
"Art S. Kagel" wrote:
> Hi folk,
>
> My turn to need help. I have a very weird situation, all of my DG/UX 4.11
> (Motorola M88K based) servers have crashed and refused to restart over the
> last two weeks. In all cases I have been able to restart them but the
> solution that works on one does not work on another (with a few exceptions).
> A configuration that refused to start at 8AM starts fine 2 hours later and
> a configuration that is the only one working on a second machine fails on
> the first machine after a subsequent crash. Online version(s) on the
> various machines? 7.24UC5, 7.24UC7, & 7.24UC8 (the last release for this
> platform ever!)
>
> That's the quick and scary and right now only my development server is down
> having crashed 11 minutes after a production server that was fixed by making
> a change that fixed this one on Friday. I have been online with Informix
> off and on for two weeks and nothing. I will not post the onconfigs right
> now because as you will see the parameters seem not to matter (with noted
> exceptions). Ask for them if you like :-)
>
> Symptoms:
>
> First I see :
>
> 10:41:02 mt.c, line 6796, thread 15, proc id 11331, Invalid Mutex Type.
> 10:41:02 Assert Failed: Invalid Mutex Type
> .....
> 10:41:02 Results: Fatal Internal Error requires system shutdown>
> over an over from various threads, mostly aio vp threads. Then the engine
> reaches the next checkpoint and goes into <CKPT REQ> and stays there and
> existing threads and new connections hang. You cannot shutdown and
> onstat -g glo report shared memory changing so it cannot report. Sometimes> one or more VPs physically crash. Cores and shared memory dumps show
> nothing useful to me or tech support. I can kill -PIPE the oninits and
> drag the engine down then remove any remaining shared memory, but on
> startup I get:
>
> 10:55:01 DR: DRAUTO id 0 (Off)
> 10:55:32 mt.c, line 1210, thread 0, proc id 3101, semot: errno = 22
> 10:55:32 Assert Failed: errno = 22
> ....
> 10:55:32 Results: Fatal Internal Error requires system shutdown>
> OR
>
> 11:10:28 DR: DRAUTO id 0 (Off)
> 11:10:57 Assert Failed: Memory free block header corruption detected in
> mt_shm_malloc_segid 1
> 11:10:57 Who: Thread(106, flush_sub(1), 0, 1)
> 11:10:57 Results: Pool repaired
> 11:10:57 Action: Please notify Informix Technical Support>
> then the previous errno = 22 messages stream mixed with Condition Failed:
> (different threads), In (hang_thread)
> messages.
>
> On the first machine this happened (a week ago Sunday) Informix changed
> just about every parameter in the onconfig file one or two at a time adding
> and deleting interfaces. Finally after 12 hours and two techs in two
> countries we got the engine up with 1/2 the CPU VPs and 1/4th the number of
> AIO VPs but it crashed just 2 hours later. I then restored the original
> ONCONFIG and started it again and it's been up for 8 days now, go figure.
> Then this past Friday a second machine, same symptoms. Finally Sunday AM
> we discovered that removing all the network connections and just running
> shared memory worked and since on that machine all connections are local
> (all users come into local middleware servers through other mechanisms) it
> was OK for the nonce. (Of course we had tried that on the machine the week
> before with no luck!) Meanwhile Friday evening a development machine went
> down and we were able to restart by changing the hostname and services in
> sqlhosts for TCP address and port number. So when a third production server
> went GA-GA this AM and nothing else worked I tried the address/port trick
> and voila it worked here also (which is good because this machine cannot
> live without networked connections! Only to discover that the development
> server had also crashed again 10 minutes after that production server
> although I have been able to restart that one.
>
> All of the affected machines had kernels built, independently, on Feb. 14,
> some have been booted recently before the crash one had been online, OS &
> engine, since March 6 without a peep.
>
> Has ANYONE seen such nonsense before?
>
> Art S. Kagel
--
Madison Pruet
===========================================
Enterprise Replication Product Developement
Dallas, Texas
Informix Software
===========================================
"Art S. Kagel" wrote: > > [much trimmed] > My turn to need help. > ... Finally Sunday AM > we discovered that removing all the network connections and just running > shared memory worked and since on that machine all connections are local > it was OK for the nonce. ... > Meanwhile Friday evening a development machine went down and we were able > to restart by changing the hostname and services in sqlhosts for TCP > address and port number. So when a third production server > went GA-GA this AM and nothing else worked I tried the address/port trick > and voila it worked here also (which is good because this machine cannot > live without networked connections! > > Has ANYONE seen such nonsense before? This may be way off base but is it possible that you are being attacked through the network. That could account for it working one time and not another and for changing the address/port making it work OK. -- +----------------------------------------------------------+ David Kleppinger DKleppinger@dts.edu +----------------------------------------------------------+
CASE # 990681.
Art S. Kagel
Madison Pruet wrote:
>
> Case Number???
>
> "Art S. Kagel" wrote:
>
> > Hi folk,
> >
> > My turn to need help. I have a very weird situation, all of my DG/UX 4.11
> > (Motorola M88K based) servers have crashed and refused to restart over the
> > last two weeks. In all cases I have been able to restart them but the
> > solution that works on one does not work on another (with a few exceptions).
> > A configuration that refused to start at 8AM starts fine 2 hours later and
> > a configuration that is the only one working on a second machine fails on
> > the first machine after a subsequent crash. Online version(s) on the
> > various machines? 7.24UC5, 7.24UC7, & 7.24UC8 (the last release for this
> > platform ever!)
> >
> > That's the quick and scary and right now only my development server is down
> > having crashed 11 minutes after a production server that was fixed by making
> > a change that fixed this one on Friday. I have been online with Informix
> > off and on for two weeks and nothing. I will not post the onconfigs right
> > now because as you will see the parameters seem not to matter (with noted
> > exceptions). Ask for them if you like :-)
> >
> > Symptoms:
> >
> > First I see :
> >
> > 10:41:02 mt.c, line 6796, thread 15, proc id 11331, Invalid Mutex Type.
> > 10:41:02 Assert Failed: Invalid Mutex Type
> > .....
> > 10:41:02 Results: Fatal Internal Error requires system shutdown> >
> > over an over from various threads, mostly aio vp threads. Then the engine
> > reaches the next checkpoint and goes into <CKPT REQ> and stays there and
> > existing threads and new connections hang. You cannot shutdown and
> > onstat -g glo report shared memory changing so it cannot report. Sometimes> > one or more VPs physically crash. Cores and shared memory dumps show
> > nothing useful to me or tech support. I can kill -PIPE the oninits and
> > drag the engine down then remove any remaining shared memory, but on
> > startup I get:
> >
> > 10:55:01 DR: DRAUTO id 0 (Off)
> > 10:55:32 mt.c, line 1210, thread 0, proc id 3101, semot: errno = 22
> > 10:55:32 Assert Failed: errno = 22
> > ....
> > 10:55:32 Results: Fatal Internal Error requires system shutdown> >
> > OR
> >
> > 11:10:28 DR: DRAUTO id 0 (Off)
> > 11:10:57 Assert Failed: Memory free block header corruption detected in
> > mt_shm_malloc_segid 1
> > 11:10:57 Who: Thread(106, flush_sub(1), 0, 1)
> > 11:10:57 Results: Pool repaired
> > 11:10:57 Action: Please notify Informix Technical Support> >
> > then the previous errno = 22 messages stream mixed with Condition Failed:
> > (different threads), In (hang_thread)
> > messages.
> >
> > On the first machine this happened (a week ago Sunday) Informix changed
> > just about every parameter in the onconfig file one or two at a time adding
> > and deleting interfaces. Finally after 12 hours and two techs in two
> > countries we got the engine up with 1/2 the CPU VPs and 1/4th the number of
> > AIO VPs but it crashed just 2 hours later. I then restored the original
> > ONCONFIG and started it again and it's been up for 8 days now, go figure.
> > Then this past Friday a second machine, same symptoms. Finally Sunday AM
> > we discovered that removing all the network connections and just running
> > shared memory worked and since on that machine all connections are local
> > (all users come into local middleware servers through other mechanisms) it
> > was OK for the nonce. (Of course we had tried that on the machine the week
> > before with no luck!) Meanwhile Friday evening a development machine went
> > down and we were able to restart by changing the hostname and services in
> > sqlhosts for TCP address and port number. So when a third production server
> > went GA-GA this AM and nothing else worked I tried the address/port trick
> > and voila it worked here also (which is good because this machine cannot
> > live without networked connections! Only to discover that the development
> > server had also crashed again 10 minutes after that production server
> > although I have been able to restart that one.
> >
> > All of the affected machines had kernels built, independently, on Feb. 14,
> > some have been booted recently before the crash one had been online, OS &
> > engine, since March 6 without a peep.
> >
> > Has ANYONE seen such nonsense before?
> >
> > Art S. Kagel
>
> --
> Madison Pruet
>
> ===========================================
> Enterprise Replication Product Developement
> Dallas, Texas
> Informix Software
> ===========================================
What are the stack traces from the af files? Just function names from the atcive thread will do.. Also check UNIX kernel logs (Look in /etc/syslog.conf for names of logfiles). David.
Hi, I do not suppose this will solve the problem, but someone never knows. We have also very strange problems, with Online5, in 1998, Online started to crash occassionally, but 2-3 times in a month (and we have to restore data from archive, ever with data loss - logs were unusable). Reason: noncompatible versions of online ! we upgraded online from 5.0x to 5.0y and in some cases we left INFORMIXDIR pointing to old online dir (for some client 4ge binaries) because of language of messages. After removing old informix files and setting the variable approprietly, the problem went out. It seems that there were changes in shared memory structures or something like this .... Have not you installed any new program on all of the GA-GA (nice :-) boxes ? Michal Hajek ASK wrote: My turn to need help. I have a very weird situation, all of my DG/UX 4.11 (Motorola M88K based) servers have crashed and refused to restart over the last two weeks. In all cases I have been able to restart them but the solution that works on one does not work on another (with a few exceptions). A configuration that refused to start at 8AM starts fine 2 hours later and a configuration that is the only one working on a second machine fails on the first machine after a subsequent crash. Online version(s) on the various machines? 7.24UC5, 7.24UC7, & 7.24UC8 (the last release for this platform ever!) -- -------------------------------------------------------------- Michal Hajek mailto:hajek@nspuh.cz Sprava NIS http://www.nspuh.cz NsP Uherske Hradiste phone : voice +420 0632 529 204 Purkynova 365 fax +420 0632 551 014 686 68 Uherske Hradiste Czech Republic ICQ UIN: 14290832 --------------------------------------------------------------
Thanks to everyone who has send me suggested avenues of investigation both
here and in private. I have answered in private where required to reduce
bandwidth and because we are not discovering anything postable yet. Another
crash today that refused to come back until the host and service were
replaced with address and port#. The adventure continues...
Art S. Kagel
"Art S. Kagel" wrote:
>
> Hi folk,
>
> My turn to need help. I have a very weird situation, all of my DG/UX 4.11
> (Motorola M88K based) servers have crashed and refused to restart over the
> last two weeks. In all cases I have been able to restart them but the
> solution that works on one does not work on another (with a few exceptions).
> A configuration that refused to start at 8AM starts fine 2 hours later and
> a configuration that is the only one working on a second machine fails on
> the first machine after a subsequent crash. Online version(s) on the
> various machines? 7.24UC5, 7.24UC7, & 7.24UC8 (the last release for this
> platform ever!)
>
> That's the quick and scary and right now only my development server is down
> having crashed 11 minutes after a production server that was fixed by making
> a change that fixed this one on Friday. I have been online with Informix
> off and on for two weeks and nothing. I will not post the onconfigs right
> now because as you will see the parameters seem not to matter (with noted
> exceptions). Ask for them if you like :-)
>
> Symptoms:
>
> First I see :
>
> 10:41:02 mt.c, line 6796, thread 15, proc id 11331, Invalid Mutex Type.
> 10:41:02 Assert Failed: Invalid Mutex Type
> .....
> 10:41:02 Results: Fatal Internal Error requires system shutdown>
> over an over from various threads, mostly aio vp threads. Then the engine
> reaches the next checkpoint and goes into <CKPT REQ> and stays there and
> existing threads and new connections hang. You cannot shutdown and
> onstat -g glo report shared memory changing so it cannot report. Sometimes> one or more VPs physically crash. Cores and shared memory dumps show
> nothing useful to me or tech support. I can kill -PIPE the oninits and
> drag the engine down then remove any remaining shared memory, but on
> startup I get:
>
> 10:55:01 DR: DRAUTO id 0 (Off)
> 10:55:32 mt.c, line 1210, thread 0, proc id 3101, semot: errno = 22
> 10:55:32 Assert Failed: errno = 22
> ....
> 10:55:32 Results: Fatal Internal Error requires system shutdown>
> OR
>
> 11:10:28 DR: DRAUTO id 0 (Off)
> 11:10:57 Assert Failed: Memory free block header corruption detected in
> mt_shm_malloc_segid 1
> 11:10:57 Who: Thread(106, flush_sub(1), 0, 1)
> 11:10:57 Results: Pool repaired
> 11:10:57 Action: Please notify Informix Technical Support>
> then the previous errno = 22 messages stream mixed with Condition Failed:
> (different threads), In (hang_thread)
> messages.
>
> On the first machine this happened (a week ago Sunday) Informix changed
> just about every parameter in the onconfig file one or two at a time adding
> and deleting interfaces. Finally after 12 hours and two techs in two
> countries we got the engine up with 1/2 the CPU VPs and 1/4th the number of
> AIO VPs but it crashed just 2 hours later. I then restored the original
> ONCONFIG and started it again and it's been up for 8 days now, go figure.
> Then this past Friday a second machine, same symptoms. Finally Sunday AM
> we discovered that removing all the network connections and just running
> shared memory worked and since on that machine all connections are local
> (all users come into local middleware servers through other mechanisms) it
> was OK for the nonce. (Of course we had tried that on the machine the week
> before with no luck!) Meanwhile Friday evening a development machine went
> down and we were able to restart by changing the hostname and services in
> sqlhosts for TCP address and port number. So when a third production server
> went GA-GA this AM and nothing else worked I tried the address/port trick
> and voila it worked here also (which is good because this machine cannot
> live without networked connections! Only to discover that the development
> server had also crashed again 10 minutes after that production server
> although I have been able to restart that one.
>
> All of the affected machines had kernels built, independently, on Feb. 14,
> some have been booted recently before the crash one had been online, OS &
> engine, since March 6 without a peep.
>
> Has ANYONE seen such nonsense before?
>
> Art S. Kagel
"Art S. Kagel" wrote: > > Thanks to everyone who has send me suggested avenues of investigation both > here and in private. I have answered in private where required to reduce > bandwidth and because we are not discovering anything postable yet. Another > crash today that refused to come back until the host and service were > replaced with address and port#. The adventure continues... > Have the servers the same primary/secondary DNS server ? Or do they use any other, but all the boxes the same, net service ? Michal Hajek -- -------------------------------------------------------------- Michal Hajek mailto:hajek@nspuh.cz --------------------------------------------------------------
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g