Re: Run away 4GL's
Posted in 1999
On Wed, 24 Feb 1999, Watson, Paul wrote: >Informix 5.10 SE >4GL 4.20 >Sun E10000 Solaris 2.5.1 (Don't ask) >450 ish concurrent users connecting using telnet from win3.11/NT Workstation >clients > >Occasionally I'm getting (a couple a day) 4gl clocking excessive CPU, note >the 4ge not the sqlexec. The PID tree from the program is fine and >everything traces back correctly to the login shell. The sqlexec is >'sleeping' and the stack trace on the 4ge indicates the user is sat in a >screen menu. The 4ge however is clocking CPU in real time, ie for every >real second the process clocks 1 second of CPU - the joys of an E10000 I >suppose ;-) Trussing the 4ge show that the process in continuously polling >stdin, a single byte read as if it's waiting for a keystoke which tallies >back to where they are in the code. > >I can not reproduce the problem via CNTL-S/Q combinations, unplugging the >PCs from the network, holding down a key continuously. I suspect that the problem is related to what happens when the PC process vanishes, and a SIGHUP is generated by the kernel. The I4GL process then goes into a tail-spin because it tries to reset the terminal settings (stty settings) to those it found at startup, but the ioctl() call that does this generates another SIGHUP because the terminal (actually a pty or pseudo-tty device) at the far end isn't there, and this goes on repeating at rather high speed for a very long time. This is certainly a problem that has afflicted a number of machine types at various times. I don't recall it being a problem on Solaris, but it is quite possible that it is the case. You could verify what's going on by running truss on one of the runaway processes (truss -p $runaway_pid) as root. If I'm right, you'll see a clear pattern in the output running over a repeat cycle of about 15 lines. >It happens to a number of different 4ges, for users on different subnets. > >Oddly sending any message to the runaway screen clears the problem. What do you mean by this? I'd hazard a guess that sending a message to the pty device leaves the pty in a non-HUP state for long enough for the I4GL process to do its ioctl() thing and then exit. I'd be curious to see an abbreviated version of the truss output as you do send the message -- if I'm at all right, of course. >It also never happened on the old E2000 Solaris 2.4 and Informix 4gl 4.14 >SE 5.06. The termcaps on both machines are the same. But the signal handling code was modified after that... Yours, Jonathan Leffler (jleffler@informix.com) #include <wish/I/was/skiing.h> Guardian of DBD::Informix v0.60 (v0.61_02) -- http://www.perl.com/CPAN Informix IDN for D4GL & Linux -- http://www.informix.com/idn