How Does IDS Handle Dead Connections?
Posted in 2001
Topics: Java & JDBC Development
What happens internally with IDS when a connection is cut abruptly? My developers are experiencing problems after their Java servlet applications bomb, usually in the midst of a transaction. When they attempt to continue a short while later, errors messages about locked rows result, indicating that resources are still being held. How does Informix clean itself up in this case? How does the engine determine whether a connection is still being held -- what are the mechanics here? How long before it determines that an sqlexec thread is dead? Is there a way to determine how long a connection has been maintained, or find the origin (ie network address, username, etc) of the connection?
Red Valsen wrote: > What happens internally with IDS when a connection is cut abruptly? My > developers are experiencing problems after their Java servlet > applications bomb, usually in the midst of a transaction. When they > attempt to continue a short while later, errors messages about locked > rows result, indicating that resources are still being held. > > How does Informix clean itself up in this case? How does the engine > determine whether a connection is still being held -- what are the > mechanics here? How long before it determines that an sqlexec thread is > dead? Is there a way to determine how long a connection has been > maintained, or find the origin (ie network address, username, etc) of > the connection? There are different answers, depending on the method of connecting to the database. Ignoring esoteric options, you either have a shared memory connection or a network (TCP) connection. In both cases, there is no automatic mechanism for the system to detect the failure of the client immediately. AFAIK, there are a couple of mechanisms used to determine whether there is still a client on the other end of the connection. For a shared memory connection, some process rattles round all the idle connections and checks whether the process that initiated them still exists, cleaning up (rollback work, which releases locks, etc) behind the client. I'm not sure what mechanism is used with the network connections; the trouble is that the network VP will be doing a select() operation -- that's a network select, not an SQL select -- to see which file descriptors have any data ready to read, and a dead conneection won't have any data to read. It would be possible for the network connections to send an empty message and wait for a response (and get an error since the far end is no longer present), but I don't know whether that is what is used. At some point, the network connection *is* timed out and the corresponding connection rolled back, releasing locks etc. You could use the SMI tables in the SysMaster database to establish (most of) the information you seek. I don't know which ones precisely, but you can probably bash the manuals just as well as I can. -- Yours, Jonathan Leffler (Jonathan.Leffler@Informix.com) #include <disclaimer.h> Guardian of DBD::Informix v1.00.PC1 -- http://www.perl.com/CPAN "I don't suffer from insanity; I enjoy every minute of it!"
Red Valsen wrote in message <3A55EA6C.BAC5B603@yahoo.com>... >What happens internally with IDS when a connection is cut abruptly? My >developers are experiencing problems after their Java servlet >applications bomb, usually in the midst of a transaction. When they >attempt to continue a short while later, errors messages about locked >rows result, indicating that resources are still being held. > >How does Informix clean itself up in this case? How does the engine >determine whether a connection is still being held -- what are the >mechanics here? How long before it determines that an sqlexec thread is >dead? Is there a way to determine how long a connection has been >maintained, or find the origin (ie network address, username, etc) of >the connection? > I'm looking in the Administrators Guide for IDS, Version 7.3. See pages 4-17 forward, in particular, pay special attention to page 4-48 describing the k (keepalive option). Make sure it is enabled in the client and server. Even using keepalive, the nature of network activity is that it can take a while for a failed responses from a keep-alive to be definitively labelled a connection failure by the network software. This is necessary because networks can be slow (well, duhhh) and thus the lowlevel network routines need to be somewhat conservative before deciding that a connection is dead. In the meantime, if you have any way of firing off a "set isolation to dirty read" command thru your connection to the database, you will get past a hell of a lot more locks during reads - See note below. I'm very scared of the load likely to be applied to the machine if you start following suggestions to setup monitor daemons which attempt to check for dead connections. Furthermore, I can't see how you can realistically and reliably program better tests than the keepalives; you may make nastier tests, but they are likely to cause a lot of false alerts and consequently a lot of accidentally slaughtered connections. Are the connected processes down the end of a slow(ish) connection? I'm afraid I don't know what options are available for locating Java processing, but I know that if you use either TCP or shared memory processes on the same machine as the engine, the engine will generally discover the loss of the process very quickly. Ditto for connections over a LAN. Jonathan, to answer an implied question you had: a select() when applied to a closed filehandle yields immediately, and then when you attempt an actual read you get zero bytes back and thus you know it's time to wrap up that handle. This is from my memory of about 1 year ago programming a daemon which used select() to fan in and fan out processing, including connections from remote machines. NOTE: Lest you freak out that using "dirty read" isolation level is "dangerous", I can tell you that it is very easy to work with such an isolation level, and very easy to keep your data consistent. No matter what isolation level, you can paint a picture containing "sufficient" tables and states which may cause consistency problems, so the general solution is to apply sensible techniques and you can work with any isolation level.