Strange Lockups on IDS 7.31.UD8
Posted in 2005
Topics: Installation, Setup & Upgrades, Storage & Space Management, Error Codes & Troubleshooting, Server Administration, Logging & Checkpoints, Platform-Specific Issues, Cloud, Docker & Containers, Versions, Editions & End-of-Life
Hey all,
We finally upgraded to UD8 about four months ago, after some fairly
thorough testing. However, we're having some problems that seem to be
mostly unexplainable.
Quick background on the system:
The system is a dual Athlon 1900 MP system w/ 3.5GB of RAM, and (5)
Hard Disks. The dbspaces are all stored on the first two drives, which
are RAID-1. There is a second RAID-1 array which we dump backups too,
and the last drive is a spare. We're running Redhat Linux 7.3, and a
2.4.20 kernel. The instance is about 11.5GB in size, not counting the
root or the tmp spaces. And it's growing at a quick clip (doubled in
size from June 1).
About two weeks ago, our production instance of IDS just quit
responding. onmonitor and dbaccess would time out looking for the db
instance. However, the log files simply showed a Completed log,
similar to:
07:31:58 Checkpoint Completed: duration was 0 seconds.
07:31:58 Checkpoint loguniq 9612, logpos 0x9e4018
Logs were continuing to run / pile up, even though no one could access
the database. However, they weren't showing up any faster / slower
than normal. I was unable to kill it w/ an onmode -k, either. After
waiting about 20 minutes, I did a kill -9, which showed a:
09:08:10 The Master Daemon Died
09:08:12 PANIC: Attempting to bring system down
Crossing my fingers, it came back up just fine.
This has happened TWICE in the past few weeks. The system is running,
the logs are fine, but completely unresponsive. In fact, the system
load is usually under 0.21 or 0.30 when this happens.
I'm not sure if it's related, but it's becoming more and more
frustrating, is the fact that my indexes are being continually
corrupted. It would seem that we have a bad index or two EVERY day
now. For instance, when I came in this morning, I had the following
waiting for me:
09:44:58 IBM Informix Dynamic Server Version 7.31.UD8
09:44:58 Who: Session(19, informix@corp.inventconnect.com, -1,2146684976)
Thread(154, sqlexec, 7ff2b3c0, 1)
File: rspartn.c Line: 1867
09:44:58 Results: Could not complete operation on'development:"root".invention'
09:44:58 Action: Run 'oncheck -cDI development:"root".invention'
09:45:02 See Also: /db0/tmp/af.482edd9, shmem.482edd9.0
I ran oncheck, it found the bad index (no problems w/ the table check)
and fixed it. Ran it again, no problems found. However, I did the
exact same thing last Thursday, on the same table.
About two weeks ago, I did a complete oncheck of the dbspace. Both -cD
and -cI, and neither found any problems. We're running Art's dostats
program to tune the indexes, about once every week or two.
I don't know if it helps, but an output of onstat_rau_ur.sh shows:
[root@corp azure]# ./onstat_rau_ur.sh
Read Utilization (UR): 96.697 %
Bufwaits Ratio (BR): 0.693446 %
The UR should ideally be very near 100%. The higher the better.
The BR should be below 7%. The lower the better.
When we were running under 7.31.UD2, the UR was 99.95% and the Bufwaits
was closer to 2%. Not sure what happened.
I'd be happy to post any stats / onconfig files that you might need to
help me solve this issue(s).
Thanks for your help.
--Anthony
Anthony wrote:
> Hey all,
>
> We finally upgraded to UD8 about four months ago, after some fairly
> thorough testing. However, we're having some problems that seem to be
> mostly unexplainable.
>
> Quick background on the system:
>
> The system is a dual Athlon 1900 MP system w/ 3.5GB of RAM, and (5)
> Hard Disks. The dbspaces are all stored on the first two drives, which
> are RAID-1. There is a second RAID-1 array which we dump backups too,
> and the last drive is a spare. We're running Redhat Linux 7.3, and a
> 2.4.20 kernel. The instance is about 11.5GB in size, not counting the
> root or the tmp spaces. And it's growing at a quick clip (doubled in
> size from June 1).
Not what you were asking but a RAID 10 array using four discs may be
better for you in performance terms than two RAID-1 mirrors.
> About two weeks ago, our production instance of IDS just quit
> responding. onmonitor and dbaccess would time out looking for the db
> instance. However, the log files simply showed a Completed log,
> similar to:
>
> 07:31:58 Checkpoint Completed: duration was 0 seconds.
> 07:31:58 Checkpoint loguniq 9612, logpos 0x9e4018>
> Logs were continuing to run / pile up, even though no one could access
> the database. However, they weren't showing up any faster / slower
> than normal. I was unable to kill it w/ an onmode -k, either. After
> waiting about 20 minutes, I did a kill -9, which showed a:
>
> 09:08:10 The Master Daemon Died
> 09:08:12 PANIC: Attempting to bring system down>
> Crossing my fingers, it came back up just fine.
>
> This has happened TWICE in the past few weeks. The system is running,
> the logs are fine, but completely unresponsive. In fact, the system
> load is usually under 0.21 or 0.30 when this happens.
Logical logs not being backed up??? Please post "onstat -l".
> I'm not sure if it's related, but it's becoming more and more
> frustrating, is the fact that my indexes are being continually
> corrupted. It would seem that we have a bad index or two EVERY day
> now. For instance, when I came in this morning, I had the following
> waiting for me:
>
> 09:44:58 IBM Informix Dynamic Server Version 7.31.UD8
> 09:44:58 Who: Session(19, informix@corp.inventconnect.com, -1,> 2146684976)
> Thread(154, sqlexec, 7ff2b3c0, 1)
> File: rspartn.c Line: 1867
> 09:44:58 Results: Could not complete operation on> 'development:"root".invention'
> 09:44:58 Action: Run 'oncheck -cDI development:"root".invention'
> 09:45:02 See Also: /db0/tmp/af.482edd9, shmem.482edd9.0>
> I ran oncheck, it found the bad index (no problems w/ the table check)
> and fixed it. Ran it again, no problems found. However, I did the
> exact same thing last Thursday, on the same table.
If you're suffering disc problems then you may see something in
/var/log/messages
If you have technical support (since you upgraded I am guessing you do)
you could raise a case and send them the af.482edd9 file. This may be
your best option rather than asking on this group.
> About two weeks ago, I did a complete oncheck of the dbspace. Both -cD
> and -cI, and neither found any problems. We're running Art's dostats
> program to tune the indexes, about once every week or two.
>
> I don't know if it helps, but an output of onstat_rau_ur.sh shows:
> [root@corp azure]# ./onstat_rau_ur.sh
> Read Utilization (UR): 96.697 %
> Bufwaits Ratio (BR): 0.693446 %
> The UR should ideally be very near 100%. The higher the better.
> The BR should be below 7%. The lower the better.
>
> When we were running under 7.31.UD2, the UR was 99.95% and the Bufwaits
> was closer to 2%. Not sure what happened.
What is BUFFERS set to? It would be helpful to post a full ONCONFIG
file. Also please post "onstat -p". If you have a lot more data than
previously then the read utilisation may well be falling.
> I'd be happy to post any stats / onconfig files that you might need to
> help me solve this issue(s).
>
> Thanks for your help.
>
> --Anthony
>