oninit taking 98% cpu
Posted in 2009
A Solaris site reported oninit pegging ~98% CPU after a user tried to kill oninit, the app failed over, and the standby machine died and was powered off. Respondents asked for diagnostics (onstat -c/-u/-p/-g iov, sar, ps, ONCONFIG changes, Solaris/IDS versions) and suggested the spike was normal post-failure recovery/resync activity that would subside. No stats were ever posted; the original poster later reported the secondary was still down but the primary was running fine with CPU back to normal, so no specific cause was identified.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Platform-Specific Issues
Got an application running on Solaris that uses informix as its
database.
Stupidly, one of the users decided to try and kill oninit. When this
failed, they did stop the app properly and it attempted to switch to
its standby machine.
This standby machine failed (for various reasons) and has now been
switched off. After restarting the main machine, it works OK apart
from oninit taking up so much CPU.
Any ideas?
Well how about giving us a clue, by posting some statistics about the
database server.
eg onstat-c
onstat -u
onstat -p
onstat -g iov
?
"BertieBigBollox@gmail.com" <bertiebigbollox@gmail.com> wrote in message
news:9d5aa313-c97a-478a-b10e-ac39f7f978d0@j21g2000yqe.googlegroups.com...
> Got an application running on Solaris that uses informix as its
> database.
>
> Stupidly, one of the users decided to try and kill oninit. When this
> failed, they did stop the app properly and it attempted to switch to
> its standby machine.
>
> This standby machine failed (for various reasons) and has now been
> switched off. After restarting the main machine, it works OK apart
> from oninit taking up so much CPU.
>
> Any ideas?
On Aug 21, 6:45 pm, "Neil Truby" <neil.tr...@ardenta.com> wrote:
> Well how about giving us a clue, by posting some statistics about the
> database server.
>
> eg onstat-c
> onstat -u
> onstat -p
> onstat -g iov>
> ?
>
> "BertieBigBol...@gmail.com" <bertiebigbol...@gmail.com> wrote in message
>
> news:9d5aa313-c97a-478a-b10e-ac39f7f978d0@j21g2000yqe.googlegroups.com...
>
>
>
> > Got an application running on Solaris that uses informix as its
> > database.
>
> > Stupidly, one of the users decided to try and kill oninit. When this
> > failed, they did stop the app properly and it attempted to switch to
> > its standby machine.
>
> > This standby machine failed (for various reasons) and has now been
> > switched off. After restarting the main machine, it works OK apart
> > from oninit taking up so much CPU.
>
> > Any ideas?- Hide quoted text -
>
> - Show quoted text -
A "user" has the power to kill oninit? I can just see Mr. Leffler
beating his head against the keyboard now.
I take it you are back online, fast recovery was unremarkable, and
now,you have no idea what's running hot?
A couple of quick questions that might shed some light
- Do you have any older onstats that you can compare your current
usage pattern with?
- Has the ONCONFIG changed? If the user can kill oninit, I am sure
they can hack up the onconfig?
- Any sar data? Double check that to see if this high cpu usage is
something new or have you been running like this for a while and this
event opened your eyes to it?
- What version of Solaris and Version of IDS?
- What does ps -eaux ( or the Solaris equivelant say) - Anything
waiting on resources?
HTH
Russ
Russ,
I think the OP just brought online his secondary, and let the synch complete which dropped his cpu usage.
The point being that when you have a system fail, the recovery process kicks in and you'll see a spike. If you wait long enough, the spike will drop as the system recovers.
No?
Sort of like being the ROTC soldier (Kevin Bacon) in Animal House. "Everybody remain calm! ..." only to get flattened by the trampling crowd. ;-)
-G
> From: russclancy@aol.com
> Subject: Re: oninit taking 98% cpu
> Date: Mon, 24 Aug 2009 09:23:44 -0700
> To: informix-list@iiug.org
>
> On Aug 21, 6:45 pm, "Neil Truby" <neil.tr...@ardenta.com> wrote:
> > Well how about giving us a clue, by posting some statistics about the
> > database server.
> >
> > eg onstat-c
> > onstat -u
> > onstat -p
> > onstat -g iov> >
> > ?
> >
> > "BertieBigBol...@gmail.com" <bertiebigbol...@gmail.com> wrote in message
> >
> > news:9d5aa313-c97a-478a-b10e-ac39f7f978d0@j21g2000yqe.googlegroups.com...
> >
> >
> >
> > > Got an application running on Solaris that uses informix as its
> > > database.
> >
> > > Stupidly, one of the users decided to try and kill oninit. When this
> > > failed, they did stop the app properly and it attempted to switch to
> > > its standby machine.
> >
> > > This standby machine failed (for various reasons) and has now been
> > > switched off. After restarting the main machine, it works OK apart
> > > from oninit taking up so much CPU.
> >
> > > Any ideas?- Hide quoted text -
> >
> > - Show quoted text -
>
> A "user" has the power to kill oninit? I can just see Mr. Leffler
> beating his head against the keyboard now.
>
> I take it you are back online, fast recovery was unremarkable, and
> now,you have no idea what's running hot?
> A couple of quick questions that might shed some light
> - Do you have any older onstats that you can compare your current
> usage pattern with?
> - Has the ONCONFIG changed? If the user can kill oninit, I am sure
> they can hack up the onconfig?
> - Any sar data? Double check that to see if this high cpu usage is
> something new or have you been running like this for a while and this
> event opened your eyes to it?
> - What version of Solaris and Version of IDS?
> - What does ps -eaux ( or the Solaris equivelant say) - Anything
> waiting on resources?
>
> HTH
>
> Russ
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
_________________________________________________________________
Get back to school stuff for them and cashback for you.
http://www.bing.com/cashback?form=MSHYCB&publ=WLHMTAG&crea=TEXT_MSHYCB_BackToSchool_Cashback_BTSCashback_1x1
As I read it, the secondary was "down and out" and had been switched off and the recovery stage had completed on the primary(main), because it was back online. I can see what you mean, about the flood of traffic to get instances in sync. But the CPU should not increase until the secondary is ready and recieving log traffic, and even then, should not max out the cpu. During the log shipping phase of getting HDR in sync, the secondary works hard, but this process should not peg the CPU, but that is assuming this is an HDR set up. I am not familair enough with ER to make comment on how that would impact CPU utilization.
On Aug 24, 9:56 pm, Russ <russcla...@aol.com> wrote: > As I read it, the secondary was "down and out" and had been switched > off and the recovery stage had completed on the primary(main), because > it was back online. I can see what you mean, about the flood of > traffic to get instances in sync. But the CPU should not increase > until the secondary is ready and recieving log traffic, and even then, > should not max out the cpu. During the log shipping phase of getting > HDR in sync, the secondary works hard, but this process should not peg > the CPU, but that is assuming this is an HDR set up. I am not > familair enough with ER to make comment on how that would impact CPU > utilization. Secondary is still down. Primary is now up and running fine. CPU back to normal.