Re: System Freezes
Posted in 2011
Topics: General Discussion
On Jan 11, 3:29 pm, "Rubinstein, James" <JRU...@midwestern.edu> wrote:
> I'm running IDS 11.50.FC6 on HPUX 11.31. We recently rolled out an
> internally developed web-based system (Apache perl/mod_perl) and
> immediately started noticing that our HPUX system becomes completely
> unresponsive, for 20-30 seconds at a time, many times throughout the
> day. My first thought was network problems, but we have pretty much
> ruled this out since I can connect to a twin HPUX server which seems
> fine during the outages. During these system freezes, any connections
> to the database fail and the system is completely unresponsive to the
> point that I cannot even type any commands at the shell. I have seen
> this behavior in the past when the oninit processes use a lot of CPU
> resources, but it is usually pretty easy to track these down to some bad
> SQL/report writing. In this case, I'm trying to figure out what may be
> causing the system freezes. I have top and the dbtop utility from IIUG,
> but those don't refresh during or freezes. I am also unable to type any
> onstat commands until the system comes back.
If you can't get top or onstat commands to work, then that indicates a
problem at the OS level. Now this could be some looping in the
database code, but probably not.
Are you overconfiguring the system? (i.e. having more CPUVPS than you
have CPUs?) That might cause this type of problem because of lock
inversion. This is the nasty problem where one process has a lock and
is running at a lower priority than another process which is trying to
obtain the lock. Basically you have to wait until the process trying
to aquire the lock ages enough so that it's OS priority drops. Nasty
problem.
One thing that you might consider is somthing like the following....
while [ 1 ];do
date
onstat -g glo
sleep 1done
When/if the problem reoccurs, you can check the output to see if it
looks like the cputime increased on one of the CPUVPS when the hang
occured...
>By that time, everything
> looks pretty normal with low load averages and our oninit processes at
> normal levels. I'm looking at the various system reports in OAT, but
> don't see anything that jumps out as the culprit. I'd appreciate any
> troubleshooting ideas.
James,
We had a situation that might be similar - we were running an
Informix, Apache, HPUX
and a perl web application.
Our problem had nothing to do with Informix though and everything to
do with
the apache version/configuration. What version of apache are you
running? I will go back and
look at my notes, but we were in a panic since it was a production
system, so we
threw 3 fixes in at one time and one of the three fixed the issue,
none of
them had anything to do with informix:
1) Changed configuration file for httpd.conf
2) downgraded the version of apache (we were using 2.2)
3) I'll have to check my notes, I forgot #3.
I always believed it was the verson of apache. We always use the pre-
compiled
version from HP. Once we backed off on the apache version, the CPU
stopped pegging at 100%.
We had a call open with HP, so HP should have case notes.
Might be worth a try.
-Daniel
> On Jan 11, 3:29 pm, "Rubinstein, James" <JRU...@midwestern.edu> wrote:
>
> > I'm running IDS 11.50.FC6 on HPUX 11.31. We recently rolled out an
> > internally developed web-based system (Apache perl/mod_perl) and
> > immediately started noticing that our HPUX system becomes completely
> > unresponsive, for 20-30 seconds at a time, many times throughout the
> > day. My first thought was network problems, but we have pretty much
> >By that time, everything
> > looks pretty normal with low load averages and our oninit processes at
> > normal levels. I'm looking at the various system reports in OAT, but
> > don't see anything that jumps out as the culprit. I'd appreciate any
> > troubleshooting ideas.
Are you using PRM (process resource management) ? We had similar problems when we hit the limits set up in PRM. Even if you are not using PRM I would think that problem is to be found in HP-UX. That there is nothing in the syslog that indicates this does not prove that you have not hit limits in the OS such as number of files etc. Ulf
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g