Logical log backup with onbar acting strangely
Posted in 2005
Hello everyone.
I'm having problem with backing up logical logs for the whole week.
Platform is IDS 9.40 UC5W2, Solaris 9, Sun V480.
Storage manager is Legato Networker.
On Monday, I've had an Informix engine hang (possibly due to a bug) but
couldn't bring it down, neither gracefully nor immediately, and the oninit
process didn't want to be killed, either.
So I rebooted the machine and started Informix again and no problem, every
piece of data was intact. Great.
Unfortunately, strange things started to happen: very often I was kicked out
of the shell with "fork failed" message. This turned out to be because of
the number of processes per user limit was hit.
After checking the process spawn rate it was obvious that logical log backup
was forking more than 8 thousand processes (namely, onbar -b -l).
15 to 20 seconds after this "multifork" thing the number of processes start
to drop and in few seconds they all die away. Logical log that fired backing
up gets backed up normally and everything seems fine until next logical log
fills up. Same thing with scheduled backups, logical log that gets backed up
in the end fires the problem again. And now the spicy detail: all of this
doesn't happen always, and I cannot recognize the rule. Lucky thing is that
Informix works normally, no problem with it, but still, it sucks to be
denied access to the system even if it's working OK.
Since this is a production machine I'm quite nervous about having problems
like this so any help would be appreciated.
Thanks everyone!
P.S. I've checked kernel parameters after reboot, no strange things there.