Informix hung itself?
Posted in 2011
Topics: Installation, Setup & Upgrades, Migration, Import/Export & Data Conversion, Cloud, Docker & Containers
HPUX B11.31 IA - BL870c
IDS 11.50.FC6
Twice last week (01/03/11) we had our Informix instance hang. I have a theory
why this happened but not sure if I'm correct.
What I know was happening at the time:
The system was running with about average levels of transaction from the user
community. Also a developer was 'copying' a larger than normal table from
production database to the test database. This transfer is done using a script
that unloads the production table then breaks this file into smaller ones of
5000 rows and then load these to the train database. I know this was the first
time the script was run on a large table since upgrading to 11.50, and that it
had been run on smaller ones since the upgrade with no problems.
We have used this script on larger tables but mostly used on smaller tables
then the one the developer did this time. So besides the large load running
everything else on the system was like any normal day.
When I was first contacted about 'slowness' on the systems I did an onstat -
and saw we had a long transaction rolling back. Didn't think much of this as
once or twice a month we get these when someone forgets a where clause and
tried to delete/update a million rows. About 10 minute later got re-contacted
and knew something unusual was happening. I tried to logon to OAT, but could
never connect to OAT. Back to the old standby of command line and started to
use onstat's to figure out what is going on. The log files showed they were
backup up and one about 2/3 full. Also saw I think it was 4 updates, 2 inserts
and a bunch of selects as what I assumed to be active running sqls. I checked
space on OS and DB all had free space. CPU and memory usage I think looked
within normal range. Everything looked ok except nothing was working. If I did
my commands correctly (very rusty with onstat), the sql that had the long
transaction was an update that was updating one row.
My theory is because the load of the table to train by the developer was
eating through log files like there was no tomorrow, we circled around on the
log files to where the other update was and started the long transaction
rollback of that and possible of the load. Somehow they locked each other up
so neither could continue. Am I grasping at straws or could this be the cause?
Since we get our Informix support through the ERP provider I could not get IBM
involved. The ERP provider resolved the hung system by killing one of the
oninit processes then doing a oninit to restart it. I was not pleased by this,
but it is what happened.
I doubt I will ever know what exactly cause the hung systems and my onstat
command knowledge is very rusty; was wondering if this happens again what
commands might I use to find the cause of the system being hung?
John Adamski
Network specialist
Graceland University
The system was running with about average levels of transaction from the user
community. Also a developer was 'copying' a larger than normal table from
production database to the test database. This transfer is done using a script
that unloads the production table then breaks this file into smaller ones of
5000 rows and then load these to the train database. I know this was the first
time the script was run on a large table since upgrading to 11.50, and that it
had been run on smaller ones since the upgrade with no problems.
I'd guess OAT was using tcpip and onstat shm. So, since shm worked, do yo have
enough threads configured for tcpip? An onstat -u might have indicated any
excessive threads being consumed in informix. I'll guess onint's were not
running 100. Any PDQ for this instance?