Re: Is iswrite atomic?
Posted in 2000
On Wed, 10 May 2000, Rod Bruce wrote: >Can anyone tell me whether iswrite() is atomic, ie, if it is >interupted by a kill, will it be ok or will it possibly leave >the data and/or index file corrupt? Good question! I don't see any mention of signals in the C-ISAM documentation, probably because I don't see any reference to signal() or sigaction() in the source code either. To jump to the conclusion, the answer to your question is "No". What follows is an extended explanation of why I think that's the answer. Let's consider what iswrite() has to do: * Given an ISAM file descriptor and a pointer to a data record (and, if it is a variable length record, the length from isreclen) * Ensure the audit trail for the file is open * Write information to the audit trail * Write information to the transaction log * Write new entries in each key (index) in the index file; this might require updates to multiple nodes if the new entry causes splits. * Write the new record to the data file That's a lot of stuff going on behind the scenes! What happens will depend on the signal handling set by the application, and on the signal that is received by the C-ISAM program. Some signals cannot be prevented: SIGKILL in particular is lethal because the program cannot do anything about it. The state of the files after a SIGKILL will be indeterminate; assume they are corrupt. Don't do it! If you are talking about regular interrupt or quit signals, then what happens will depend on your interrupt handler and on the o/s. If you are neither ignoring the signal nor fielding them with your own signal handler, then the program will stop (SIG_DFL stops the program, giving a core dump if it is SIGQUIT). Again, it is simplest to assume that the files are corrupt. If the program has its own signal handler, then what happens next depends on what the signal handler does. If the signal handler does something fancy like a longjmp(), that terminates the iswrite() call at whatever random location it was in; again, you'd have to assume that the files are corrupt after this. If the signal handler returns (recommended), then it depends on whether the signal happened during a system call such as write() or in other code. If it happened during other code, everything should be OK; the signal handler returned, and the program continued inside iswrite() as if nothing had happened. If it happened during a system call (assume it is write; the others probably aren't too significant), then one of two things can happen, and which happens depends on the o/s you are using: 1. The write() call continues and completes normally. Everything is OK, and there should be no corruption. 2. The write() call stops with errno EINTR. The write is incomplete. I've not scrutinized the code, but I'd suspect that it does not retry the write() after repositioning the file pointer if it is interrupted. I may be doing the code an injustice, but this is coordinated by function that is 330-odd lines long, and has to deal with complicated things like the variable length data portion being written into a remainder node in the index file while the fixed length portion is written into the data file, and so on. The code line count doesn't include this commentary: /* ... * arg xtype = transfer type: * CK=check,RD=read,CR=create,WR=write/update,TR=truncate,DE=delete, * IN=infobytes. * * This routine is a mess. But it's an orderly mess. It's about 6 routines * all tangled up into one. Untangling them would mean 6 separate routines * with a lot of similar tricky code not easily shared as subroutines, and * a total length even greater than this routine. The tradeoffs are * complex. So far it seems that this way things are easier to * maintain, even if they are harder to understand at first. * Someday it may make sense to put the effort into better readability. ... */ ...time passes... OK, a bit more scrutiny of the code finds calls to a dxwrite() routine, which calls bfwrite(), which calls write(). If the return from write() is not the number of bytes you attempted to write, iserrio is set (to IO_IDX + IO_WRIT), iserrno is set to errno (hence to EINTR), and an error is returned. There is no attempt here to try a rewrite on EINTR. So, you know that something has gone wrong, but you don't know what the state of the files is. This appears to be used for the index file; I didn't track down what happens with the data file, but it is presumably similar. If your system does not continue or restart i/o calls after a signal handler returns, then you've probably got some corruption. POSIX.1:1996 says in section 6.4.2.2: If a write() is interrupted by a signal before it writes any data, it shall return -1 with errno set to [EINTR]. If a write() is interrupted by a signal after it successfully writes some data, either it shall return -1 with errno set to [EINTR], or it shall return the number of bytes written. FWIW, on Solaris 7, the system documentation for write(2) says: If a write() is interrupted by a signal before it writes any data, it will return -1 with errno set to EINTR. If a write() is interrupted by a signal after it successfully writes some data, it will return the number of bytes written. That is, the Solaris documentation states which of the two POSIX.1 alternatives is implemented. And note that even if the system can return some number of bytes successfully written, the system call can also return -1 if no data was written. The C-ISAM code does not deal with either case particularly well. So, to answer your question finally, iswrite() is not atomic. -- Yours, Jonathan Leffler (Jonathan.Leffler@Informix.com) #include <disclaimer.h> Guardian of DBD::Informix v1.00.PC1 -- http://www.perl.com/CPAN "I don't suffer from insanity; I enjoy every minute of it!"