RE: HHHHEEEELLLLPPPP!!!! My servers have gone nuts!
Posted in 2000
Very weird... Just from what you say, I wonder if there is something flaky
about the network. Maybe you can get some of the network people to put some
sniffers out there, see if they see anything. No, I don't see how or why
something in the network would cause AF's, but it looks to me like network
is your common thread in all of this...
FWIW,
Paul Mosser
-----Original Message-----
From: Art S. Kagel [mailto:kagel@bloomberg.net]
Sent: Monday, September 25, 2000 11:27 AM
To: informix-list@iiug.org
Subject: HHHHEEEELLLLPPPP!!!! My servers have gone nuts!
Hi folk,
My turn to need help. I have a very weird situation, all of my DG/UX 4.11
(Motorola M88K based) servers have crashed and refused to restart over the
last two weeks. In all cases I have been able to restart them but the
solution that works on one does not work on another (with a few exceptions).
A configuration that refused to start at 8AM starts fine 2 hours later and
a configuration that is the only one working on a second machine fails on
the first machine after a subsequent crash. Online version(s) on the
various machines? 7.24UC5, 7.24UC7, & 7.24UC8 (the last release for this
platform ever!)
That's the quick and scary and right now only my development server is down
having crashed 11 minutes after a production server that was fixed by making
a change that fixed this one on Friday. I have been online with Informix
off and on for two weeks and nothing. I will not post the onconfigs right
now because as you will see the parameters seem not to matter (with noted
exceptions). Ask for them if you like :-)
Symptoms:
First I see :
10:41:02 mt.c, line 6796, thread 15, proc id 11331, Invalid Mutex Type.
10:41:02 Assert Failed: Invalid Mutex Type
.....
10:41:02 Results: Fatal Internal Error requires system shutdown
over an over from various threads, mostly aio vp threads. Then the engine
reaches the next checkpoint and goes into <CKPT REQ> and stays there and
existing threads and new connections hang. You cannot shutdown and
onstat -g glo report shared memory changing so it cannot report. Sometimesone or more VPs physically crash. Cores and shared memory dumps show
nothing useful to me or tech support. I can kill -PIPE the oninits and
drag the engine down then remove any remaining shared memory, but on
startup I get:
10:55:01 DR: DRAUTO id 0 (Off)
10:55:32 mt.c, line 1210, thread 0, proc id 3101, semot: errno = 22
10:55:32 Assert Failed: errno = 22
....
10:55:32 Results: Fatal Internal Error requires system shutdown
OR
11:10:28 DR: DRAUTO id 0 (Off)
11:10:57 Assert Failed: Memory free block header corruption detected in
mt_shm_malloc_segid 1
11:10:57 Who: Thread(106, flush_sub(1), 0, 1)
11:10:57 Results: Pool repaired
11:10:57 Action: Please notify Informix Technical Support
then the previous errno = 22 messages stream mixed with Condition Failed:
(different threads), In (hang_thread)
messages.
On the first machine this happened (a week ago Sunday) Informix changed
just about every parameter in the onconfig file one or two at a time adding
and deleting interfaces. Finally after 12 hours and two techs in two
countries we got the engine up with 1/2 the CPU VPs and 1/4th the number of
AIO VPs but it crashed just 2 hours later. I then restored the original
ONCONFIG and started it again and it's been up for 8 days now, go figure.
Then this past Friday a second machine, same symptoms. Finally Sunday AM
we discovered that removing all the network connections and just running
shared memory worked and since on that machine all connections are local
(all users come into local middleware servers through other mechanisms) it
was OK for the nonce. (Of course we had tried that on the machine the week
before with no luck!) Meanwhile Friday evening a development machine went
down and we were able to restart by changing the hostname and services in
sqlhosts for TCP address and port number. So when a third production server
went GA-GA this AM and nothing else worked I tried the address/port trick
and voila it worked here also (which is good because this machine cannot
live without networked connections! Only to discover that the development
server had also crashed again 10 minutes after that production server
although I have been able to restart that one.
All of the affected machines had kernels built, independently, on Feb. 14,
some have been booted recently before the crash one had been online, OS &
engine, since March 6 without a peep.
Has ANYONE seen such nonsense before?
Art S. Kagel