Re: Killing database engines
Posted in 1992
>From: Tony Heskett <uunet!bnr.co.uk!A.Heskett> >Message-Id: <9205111353.3781@bay.bnr.co.uk> >Subject: Killing database engines >Date: Mon, 11 May 92 14:53:19 BST >X-Informix-List-Id: <list.1163> > >Jonathan Leffler writes: > >> [ ... lots of good stuff deleted ... ] > >> Letting the engine run should not cause corruption. In the odd >> cases where a FE program dies and the engine seems to go on forever >> (typically chewing up masses of CPU time while it is doing so), you will >> have to kill the engine. But do it gently: try signals 15 (SIGTERM) and 1 >> (SIGHUP), and SIGUSR1 (which varies: 16 on System V and 30 on SunOS (BSD?)) >> first, and allow it time (say 10 seconds) to respond in each >> case, and only if all those fail should you go for the kill 9 (SIGKILL). > >Hmm, speaking as someone who's often had to pick up the pieces >after my friends' kill -9's and tbmode -z (now safe with OnLine ?), Safe, yes. Effective: probably in 4.10, more certainly in 5.00. >I'd like a $0.02. Please don't take any of this as >criticism, it ain't :-) Pas de problem. >I was told to kill -13 the backend (SIGPIPE: "the front-end's >died"), after trying kill -15 ineffectually. Is this still >reasonable, before trying anything more serious ? It's a fair cop, guv! It was a Monday morning, though. Yes, try SIGPIPE (13). You could consider SIGALRM (14) too, but it probably won't have any useful effects. The basic recipe is try anything which (a) does not produce a core dump, and (b) is not SIGKILL *before* you use SIGKILL. SIGKILL is violent, and "Violence is the last resort of the incompetent", if I remember Salvor Hardin in "Foundation" by Asimov (on whom be peace). Regrettably, violence is sometimes necessary -- but only as a last resort! >10 sec for a rollback seems a bit short: even with a fairly >small DB, it's often taken some minutes to sort itself out. Also true. The significance of 10 seconds was that you should wait longer than the time it takes to type: kill -1 14356 kill -15 14356 kill -14 14356 kill -13 14356 kill -9 14356 between sending signals. >Roughly as long as it took to get there, they used to tell me. Depends on how much user interaction there was during the transaction, of course. >So long as it looks like it's rolling back, I usually try to >leave it alone. Yes. But it can be difficult to tell whether it is rolling back, or just chewing up the CPU. It seems to be the morning for proverbs: patience is a virtues. >With OnLine, the whole DB shuts down/should shut down if an >engine terminates abnormally. So kill -9 on one of the engines >may have the same effect as a graceless shutdown. Insofar as >you want to prevent any shared-memory corruption, this seems >like a good idea: but it can come as a bit of a surprise to >your other users ... Avoid killing OnLine engines with kill -9 except as a last resort. I think that OnLine is only taken down if the sqlturbo is in a critical section and has things latched. This means it was modifying shared memory structures and tbinit has no chance of knowing how far it had got. If the engine was sitting idle waiting for something to come from the application, it can be rolled back by tbundo (aka tbinit). Yours, Jonathan Leffler (johnl@obelix.informix.com) #include <disclaimer.h>