RE: HDR with Synchronous replication
Posted in 2004
Revisit: The problem is if the primary fails. If the secondary fails all is as I described in the previous post. Hmm, looks like I have to eat crow about the synchonicity if the primary fails, however. DRINTERVAL -1 is billed as 'synchronous'. It just means that logical log entries are shipped immediately to the secondary. Looks like the logical log write on the primary waits for either the secondary to receive the log records or for it to crash (or be unavailable for DRTIMEOUT seconds). If this happens it writes the logical logs locally anyway (unlike two-phase commit) and the recovery mechanism I described on CDI take care of the secondary. If the primary crashes, however, before the local log write but after the secondary receives the records all is still OK (assuming a commit was sent). The problem, which as you surmise, can only happen with DRINTERVAL >=0 is if the primary fails after committing a transaction locally but before all of the logical log records have been flushed to the secondary. Such transactions would be lost. DRINTERVAL>=0 is send logical log entries as soon as possible but do not wait to write the logical log records locally. ASAP means as soon as the logical log buffer flushes to disk. This happens if a commit record for an UNBUFFERED database is written to the log buffer, when the log buffer fills, or when the DRTIMEOUT expires without a write. This is pure asynch (we use this BTW). What ever the primary wanted to send to the secondary when it was ready to commit is saved in the DRLOSTFOUND directory when the primary restarts but cannot be reapplied to the secondary. If the primary crashes before a logical log buffer containing a commit record is sent that transaction could be lost. HOWEVER, if all databases are UNBUFFERED, the local commit cannot have been logged or acknowledged either and the logical log buffered is flushed to disk and to the secondary as soon as the commit record is added to it. This is what we do and what I recommend. DRINTERVAL==0 and all databases UNBUFFERED. It is close to 100% as safe as synchronous without the possible delay in local commit that may be caused by communications delay to the secondary. Art S. Kagel sending to informix-list