TOP Process
Posted in 2001
Dirk saw I4GL processes on HP-UX 11/IDS 9.21 burning 60-90% CPU while onstat showed them idle in 'cond wait (netnorm)' with a stale, fast-running SELECT as the last statement. Replies explained netnorm means the server is waiting on the client, so the loop is in the 4GL program, and suggested truss to see the system calls. Andrew Hamm diagnosed it as a detached session: when a user's telnet/PC connection dies during MENU/INPUT, read() returns 0 and 4GL spins endlessly. Workarounds offered: keepalive/kernel tuning, cron jobs to kill such processes, alarm/timeout functions, user education. David Williams said the bug was fixed in 4GL version 4, though Hamm reported still seeing it occasionally; no definitive fix is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Server Administration, Platform-Specific Issues, Versions, Editions & End-of-Life
HP-UX 11.00
IDS 9.21FC1
Hallo everyone, I have something that's been a problem to me for quite some
time now. If I run a "top" on HP, showing me the processes using the most
cpu power, I often find 4gi's sitting right on top, using 60 - 90 % cpu
power.
I also think these programs are not doing any selects (as seen from the
current sql statement from onstat -g ses), but something else. How can I
find out what the program is doing, and why it is using so much cpu power ?
One example I have this morning:
onstat -u | grep 15373
c000000094be4828 Y--P--- 15373 izakb tPe c000000088e55d80 0 163 0
Flags are normal. If I get the current sql statement using "onstat -g ses",
and I run this in dbaccess, it runs very quickly (sub second).
Current SQL statement :
select unique waybill_no from waybill_ref where waybill_ref.ref_no='18405'
The user says that on his side, the program seems to be hanging. Has anyone
had similar experiences on HP & IDS 9.21FC1 ?
We get this every now and then, the program will sit right on top for a
couple of minutes and then disappear again.
I don't know if this is an Informix problem or an HP problem, and I have
people thinking its the database because they can see the current sql
statement sitting there and not changing. I also get the status - cond wait
(netnorm) - when I run onstat -g ses, which I am not too familiar with,
accept that I think the program is waiting for response from another
process.
Dirk Moolman
Database Administrator
Reach Technologies
"Bravery is the capacity to perform properly even when scared half to
death."
- General Omar Bradley
A netnorm state means that the database is waiting for a request from
the user. So this means there is no select going on.
I would say this is a problem in the program. If you can, take a look
at the source code.
Hope this helps,
Will
In article <93bsu8$do3$1@news.xmission.com>,
"Dirk Moolman" <dirkm@reach.co.za> wrote:
>
> HP-UX 11.00
> IDS 9.21FC1
>
> Hallo everyone, I have something that's been a problem to me for
quite some
> time now. If I run a "top" on HP, showing me the processes using
the most
> cpu power, I often find 4gi's sitting right on top, using 60 - 90 %
cpu
> power.
>
> I also think these programs are not doing any selects (as seen from
the
> current sql statement from onstat -g ses), but something else. How
can I
> find out what the program is doing, and why it is using so much cpu
power ?
>
> One example I have this morning:
>
> onstat -u | grep 15373
> c000000094be4828 Y--P--- 15373 izakb tPe c000000088e55d800 1
> 63 0
>
> Flags are normal. If I get the current sql statement using "onstat
-g ses",
> and I run this in dbaccess, it runs very quickly (sub second).
>
> Current SQL statement :
> select unique waybill_no from waybill_ref wherewaybill_ref.ref_no='18405'
>
> The user says that on his side, the program seems to be hanging. Has
anyone
> had similar experiences on HP & IDS 9.21FC1 ?
> We get this every now and then, the program will sit right on top for
a
> couple of minutes and then disappear again.
> I don't know if this is an Informix problem or an HP problem, and I
have
> people thinking its the database because they can see the current sql
> statement sitting there and not changing. I also get the status -
cond wait
> (netnorm) - when I run onstat -g ses, which I am not too familiar
with,
> accept that I think the program is waiting for response from another
> process.
>
> Dirk Moolman
> Database Administrator
> Reach Technologies
>
> "Bravery is the capacity to perform properly even when scared half to
> death."
> - General Omar Bradley
>
>
Sent via Deja.com
http://www.deja.com/
Dirk Moolman wrote:
>
> HP-UX 11.00
> IDS 9.21FC1
>
> Hallo everyone, I have something that's been a problem to me for quite some
> time now. If I run a "top" on HP, showing me the processes using the most
> cpu power, I often find 4gi's sitting right on top, using 60 - 90 % cpu
> power.
>
> I also think these programs are not doing any selects (as seen from the
> current sql statement from onstat -g ses), but something else. How can I
> find out what the program is doing, and why it is using so much cpu power ?
>
> One example I have this morning:
>
> onstat -u | grep 15373
> c000000094be4828 Y--P--- 15373 izakb tPe c000000088e55d80 0 1> 63 0
>
> Flags are normal. If I get the current sql statement using "onstat -g ses",
> and I run this in dbaccess, it runs very quickly (sub second).
>
> Current SQL statement :
> select unique waybill_no from waybill_ref where waybill_ref.ref_no='18405'>
> The user says that on his side, the program seems to be hanging. Has anyone
> had similar experiences on HP & IDS 9.21FC1 ?
> We get this every now and then, the program will sit right on top for a
> couple of minutes and then disappear again.
> I don't know if this is an Informix problem or an HP problem, and I have
> people thinking its the database because they can see the current sql
> statement sitting there and not changing. I also get the status - cond wait
> (netnorm) - when I run onstat -g ses, which I am not too familiar with,
> accept that I think the program is waiting for response from another
> process.
I would suggest that the program is indeed hanging, for reasons as yet
unclear. Andrew Hamm says the server is doing nothing, which is
consistent with the I4GL code actually spinning its wheels.
One way which this used to appear on some systems was when the program
was disconnected from a network connection, it would get a SIGHUP
signal, and as part of processing that, would try to restore the
terminal state before exiting, but the restore operation triggered
another SIGHUP, so the program tried to restore, ... which took a long
time to run out of stack space on systems with big hunks of virtual
memory and ate up a lot of CPU time.
However, you imply that the user has not hung up, so it should be
something else that's at fault. You might get an idea from attaching
the HP equivalent of 'truss' to the running I4GL program and seeing what
it is doing. You might not see it doing any system calls at all, in
which case you've managed to create an infinite loop (not necessarily
your fault), or you might see a repeating system call sequence which
will help indicate what is actually going on in the I4GL program.
--
Yours,
Jonathan Leffler (Jonathan.Leffler@Informix.com) #include <disclaimer.h>
Guardian of DBD::Informix v1.00.PC1 -- http://www.perl.com/CPAN
"I don't suffer from insanity; I enjoy every minute of it!"
Jonathan Leffler wrote in message <3A5A163A.46A8F7C5@informix.com>... > >I would suggest that the program is indeed hanging, for reasons as yet >unclear. Andrew Hamm says the server is doing nothing, which is >consistent with the I4GL code actually spinning its wheels. > Emm, when did I say that? It's the sort of thing I would say, but I don't recall saying it in this context... I know what's going on, but first, I'd like to know why the illustrated select-statement looks very suspiciously like one of our applications! Dirk, are you one of our customers? I don't recognise your name. Please e-mail me directly with details. I'll get the answer back to the group when I understand the situation better.
Andrew Hamm wrote in message <3a5ab7f0@news.iprimus.com.au>...
>Jonathan Leffler wrote in message <3A5A163A.46A8F7C5@informix.com>...
>>
>I know what's going on, but first, I'd like to know why the illustrated
>select-statement looks very suspiciously like one of our applications!
Dirk,
>are you one of our customers? I don't recognise your name. Please e-mail me
>directly with details. I'll get the answer back to the group when I
>understand the situation better.
>
Hokay, I've examined your web site (nice) and am satisfied that there is no
support-chain issues or even hanky-panky going on. If you were from one of
our customers I should talk to you only through proper channels. I guess
it's no big surprise that people supplying similar verticals may use the
same table and column names!
Here's the answer to the 4GL-on-top problem:
1) It likes to be on top occasionally. Doesn't everyone?
2) What you are seeing is "detached" sessions. They come about (no, that's
not another sexual innuendo) when users switch off their PC's or do
something equally evil. For reasons that escape me, often a switch-happy
user doesn't leave a dead session, but sometimes it does.
Now, here's my theory about what's going on inside the 4GL runner:
1) The telnet session to the PC is rudely interrupted.
2) Meanwhile, the 4GL program is sitting in an input state (such as at a
MENU or INPUT statement - see note below)
3) Since the 4GL is asking the operating system for the "next keystroke"
from the user but the connection has gone away, the O/S returns a code to
4GL saying: cannot get another character.
4) 4GL sees the return from the read(2) system call and says "OK, that's
interesting - gimme a character from the user!"
5) Since the 4GL is asking the operating system for the "next keystroke"
from the user but the connection has gone away, the O/S returns a code to
4GL saying: cannot get another character.
6) 4GL sees the return from the read(2) system call and says "OK, that's
interesting - gimme a character from the user!"
7) Since the 4GL is asking the operating system for the "next keystroke"
from the user but the connection has gone away, the O/S returns a code to
4GL saying: cannot get another character.
8) 4GL sees the return from the read(2) system call and says "OK, that's
interesting - gimme a character from the user!"
9) Since the 4GL is asking the operating system for the "next keystroke"
from the user but the connection has gone away, the O/S returns a code to
4GL saying: cannot get another character.
10) 4GL sees the return from the read(2) system call and says "OK, that's
interesting - gimme a character from the user!"
11) Since the 4GL is asking the operating system for the "next keystroke"
from the user but the connection has gone away, the O/S returns a code to
4GL saying: cannot get another character.
12) 4GL sees the return from the read(2) system call and says "OK, that's
interesting - gimme a character from the user!"
(interrupt me when this joke becomes old)
So, basically, this can go on for the entire time slice on offer to the 4GL
process, and hence it truly chews up a lot of CPU cycles, including a
shit-load of context switches into the kernel, and it only takes one or two
of these stupid runaways to make a really noticable impact on performance.
Sometimes, I can just FEEL a 4GL runaway, and go looking for it.
NOTE: I'm not sure if one specific 4GL statement such as MENU, INPUT,
DISPLAY, PROMPT or whatever is the guilty one, or whether it's universal. I
suppose you could invent a test to isolate the problem.
It's immaterial what the last SQL statement was, although it could be
considered a piece of interesting but difficult forensic evidence. You will
find that the process hasn't done anything interesting in the engine from
the time that the big red switch was tickled on the users PC. That can be
days sometimes.
You may find that the onstat -u entry shows that the process is in a
transaction, but that's just a really bad programming technique - waiting
for the user in the middle of real database changes? OUCH! Here comes a long
transaction problerm! I'd be EXTREMELY surprised to see the process is in a
critical section (unless it's being rolled back for a long transaction;-)
I've never seen anything that shows activity in the process.
Final piece of evidence that proves it's a detached session is to look at
the named tty listed in the ps listing for that process. It may still show
the users original terminal - in which case, you will find that the /dev
entry for that terminal has a really old time stamp which proves the time of
the switch off. Usually around 5:00 pm or lunchtime. If there is still a
listed tty, take a look at all the processes on that terminal:
ps -f -tp12 # for example
and you may find either NOTHING else, or parent processes faithfully waiting
for the naughty little sod.
Depending on the O/S, the tty may be showing as ? which is another dead-set
givaway of the problem.
SO: final solutions:
1) Educate your users, and carry a wooden yardstick with you as you visit
the guilty users.
2) Investigate keepalive settings or other kernel thingies I'm unaware of
which may help. Note: the SIGHUP should be sent from the kernel but it
appears it doesn't happen or maybe it's ignored by the 4GL?
3) Hassle informix to investigate the problem and finally fix it.
4) Setup cron jobs to look for these processes and deal with them
appropriately. Exercise the usual cautions.
5) Consider adding runner functions which I call c_alarm() and c_alarmed()
(you can perhaps guess their implementation?) and explicitly program
timeouts into your 4GL code. This is probably a really really big job...
6) Hassle Informix to add timeout facilities into the 4GL language around
the MENU, INPUT etc statements. Now that I'd like to see...
__END__
Andrew Hamm
Technical Consultant
Sanderson Australia Pty Ltd
e-mail: <mailto:ahamm@sanderson.net.au>
web: <http://www.sanderson.net.au>
--
I like cats too - let's exchange recipies
In article <3a5bca8a$1@news.iprimus.com.au>, Andrew Hamm <ahamm@sanderson.net.au> writes >1) The telnet session to the PC is rudely interrupted. > >2) Meanwhile, the 4GL program is sitting in an input state (such as at a >MENU or INPUT statement - see note below) > >3) Since the 4GL is asking the operating system for the "next keystroke" >from the user but the connection has gone away, the O/S returns a code to >4GL saying: cannot get another character. >4) 4GL sees the return from the read(2) system call and says "OK, that's >interesting - gimme a character from the user!" > Yes, the read() system call returns 0 - no characters read and so 4GL loops. This was fixed in version 4! All current 4gl versions do not have this bug. -- David Williams
David Williams wrote in message ... >In article <3a5bca8a$1@news.iprimus.com.au>, Andrew Hamm ><ahamm@sanderson.net.au> writes > > Yes, the read() system call returns 0 - no characters read and so > 4GL loops. This was fixed in version 4! All current 4gl versions > do not have this bug. > Hmmm - methinks not properly, 'cos it's still seen occasionally.
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g