RE: HHHHEEEELLLLPPPP!!!! My servers have gone nuts!
Posted in 2000
I'll give my $.02 worth . . .
> -----Original Message-----
> From: Art S. Kagel [mailto:kagel@bloomberg.net]
> Sent: Monday, September 25, 2000 2:27 PM
> To: informix-list@iiug.org
> Subject: HHHHEEEELLLLPPPP!!!! My servers have gone nuts!
>
>
> Hi folk,
>
> My turn to need help. I have a very weird situation, all of
> my DG/UX 4.11
> (Motorola M88K based) servers have crashed and refused to
> restart over the
> last two weeks. In all cases I have been able to restart
> them but the
> solution that works on one does not work on another (with a
> few exceptions).
> A configuration that refused to start at 8AM starts fine 2
> hours later and
> a configuration that is the only one working on a second
> machine fails on
> the first machine after a subsequent crash. Online version(s) on the
> various machines? 7.24UC5, 7.24UC7, & 7.24UC8 (the last
> release for this
> platform ever!)
>
Older versions . . . any bug-related problem with these versions? We were
on 7.24 for a bit with no problems.
> That's the quick and scary and right now only my development
> server is down
> having crashed 11 minutes after a production server that was
> fixed by making
> a change that fixed this one on Friday. I have been online
> with Informix
> off and on for two weeks and nothing. I will not post the
> onconfigs right
> now because as you will see the parameters seem not to matter
> (with noted
> exceptions). Ask for them if you like :-)
>
What type of changes did you make? ONCONFIG??
> Symptoms:
>
> First I see :
>
> 10:41:02 mt.c, line 6796, thread 15, proc id 11331, Invalid> Mutex Type.
> 10:41:02 Assert Failed: Invalid Mutex Type
> .....
> 10:41:02 Results: Fatal Internal Error requires system shutdown>
Ouch! Any OS errors in syslog around this time??
> over an over from various threads, mostly aio vp threads.
> Then the engine
> reaches the next checkpoint and goes into <CKPT REQ> and
> stays there and
> existing threads and new connections hang. You cannot shutdown and
> onstat -g glo report shared memory changing so it cannot> report. Sometimes
> one or more VPs physically crash. Cores and shared memory dumps show
> nothing useful to me or tech support. I can kill -PIPE the
> oninits and
> drag the engine down then remove any remaining shared memory, but on
> startup I get:
>
> 10:55:01 DR: DRAUTO id 0 (Off)
> 10:55:32 mt.c, line 1210, thread 0, proc id 3101, semot: errno = 22
> 10:55:32 Assert Failed: errno = 22
> ....
> 10:55:32 Results: Fatal Internal Error requires system shutdown>
> OR
>
> 11:10:28 DR: DRAUTO id 0 (Off)
> 11:10:57 Assert Failed: Memory free block header corruption> detected in
> mt_shm_malloc_segid 1
> 11:10:57 Who: Thread(106, flush_sub(1), 0, 1)
> 11:10:57 Results: Pool repaired
> 11:10:57 Action: Please notify Informix Technical Support>
> then the previous errno = 22 messages stream mixed with
> Condition Failed:
> (different threads), In (hang_thread)
> messages.
>
Errno 22 in HPUX is 'invalid argument'. What kernel changes to the OS could
have done this??
> On the first machine this happened (a week ago Sunday)
> Informix changed
> just about every parameter in the onconfig file one or two at
> a time adding
> and deleting interfaces. Finally after 12 hours and two techs in two
> countries we got the engine up with 1/2 the CPU VPs and 1/4th
> the number of
> AIO VPs but it crashed just 2 hours later. I then restored
> the original
> ONCONFIG and started it again and it's been up for 8 days
> now, go figure.
What were the basic changes in the ONCONFIG. Net vs. shm connections??
> Then this past Friday a second machine, same symptoms.
> Finally Sunday AM
> we discovered that removing all the network connections and
> just running
> shared memory worked and since on that machine all
> connections are local
> (all users come into local middleware servers through other
> mechanisms) it
> was OK for the nonce. (Of course we had tried that on the
> machine the week
> before with no luck!) Meanwhile Friday evening a development
> machine went
> down and we were able to restart by changing the hostname and
> services in
> sqlhosts for TCP address and port number. So when a third
> production server
> went GA-GA this AM and nothing else worked I tried the
> address/port trick
> and voila it worked here also (which is good because this
> machine cannot
> live without networked connections! Only to discover that
> the development
> server had also crashed again 10 minutes after that production server
> although I have been able to restart that one.
I'm not as familiar with networking, but has something changed on the
network? I've never heard of a network crashing a RDBMS, though. (At least
not Informix!)
>
> All of the affected machines had kernels built,
> independently, on Feb. 14,
> some have been booted recently before the crash one had been
> online, OS &
> engine, since March 6 without a peep.
>
> Has ANYONE seen such nonsense before?
I had to reread this message to make sure that I wasn't having a daydream.
Scary!
Let's recap. Multiple servers, multiple versions (minor versions!!) of
Informix. Same OS. What about kernel parameters . ... same or different??
Unless you've hit a major bug in the 7.24 family, I'm wondering if the OS or
network could be the culprit. But then you have the aio VP issues . . . .
Anything DGUX specific?