Re: Informix replication doesn't....
Posted in 1994
>>>>> "Bill" == Bill Duncan <bduncan@netcom.com> writes: [last paragraph cited first:] > I know that everybody else's replication approach suffers from some if not > all of these problems, Not everybody else's. We developed a replication system (infix-synchro) for Informix OnLine 5.0x and SE 5.0x engines. It takes a totally different approach and so does not suffer from a number of problems you mentioned. (On the other hand it may suffer from different problems depending on your needs.) [...] > I was just wondering if things had changed and if anyone in the know > could shed some light on these issues. I just thought to share our experiences and compare our solution step by step. > (1) The primary and secondary site hardware must be the same, or atleast > offer identical paging structure. Our system even replicates between OnLine and SE. It does not depend on the machine architecture. > (2) The OnLine config. on both the primary and secondary sites must > be set up exactly the same. Not in our case. > (3) You can only have one secondary site per primary site, i.e. a 1:1 > correspondence between primary and secondary sites. We do not have primary and secondary sites at all. We support full peer to peer replication. > (4) Secondary sites are read-only. Our system supports "updates anywhere". However, be prepared for conflicts that can arise if the same data is updated at two or more sites before synchronisation occurs. Some of these conflicts can be resolved automatically but the nature of certain conflicts requires human intervention. An example: If a customer's postal address is changed in different ways at two sites, one of these changes must be wrong. Someone has to decide which one is correct. We provide detailed conflict reports in this case. > (5) From the secondary site, you can not determine if a failure is at the > primary site itself or a failure of the network. Our approach does not differentiate between site and network failures. We send out update messages where every single operation is acknowledged. In case of any error, the system can determine the state to continue from and restart the replication cycle. > (6) Launching the entire replication process is difficult because both sites > have to be kept closely in sync. Sorry, what do you mean by "have to be kept closely in sync"? > (7) There are no administrative tools, just a few variables that can be > queried or set to direct some of the replicate process behavior. In our system, a configuration language is used to describe the replication paramneters, including o which data is supplied to which sites, o the relationships between database tables, o one- or bidirectional distribution (separately for each data object), o the names for data objects in an end user's terminology, o collision handling methods (on the attribute level). A browser is available to display messages about errors and conflicts that occured during replication. Our replication system has an option to populate new peers with data. It is also possible to re-synchronise sites. This is needed if one site has to restore from tape and therefore contains an older snapshot of the database than its peers assume. We currently have two additional support tools: One determines the state of the system in case of any error. Another restarts an interrupted replication cycle automatically. > (8) There is no mechanism for resolving inflight transactions that do not > make it to the replicate site in the case of a hot-site backup scenario. Re- > syncing the primary site after it comes back on-line is not possible without > heavy manual intervention. Our approach is not suitable for hot site backup at all. > (9) High transaction rates on the primary replication site can cause the log > to grow to fast for replicates to keep up....the result is loss of data from > the logs. The actual transaction rate does not affect the operation of our replication system. However, our system automatically constructs shadow tables to catch the peer's data state. So what's the main difference? The current Informix approach is targeted at "hot site backup" and "read only copy for decision support". It works by replicating transaction log entries. Our approach is targeted at geographically distributed databases where updates should be possible at any site and a greater degree of fault-tolerance is needed. We use shadow tables to accomplish this. In addition, since we do not replicate transactions, our system has to know about the relationships between tables to preserve database integrity. We require the actual tables' contents to be compared to the corresponding shadow tables. This way our system determines which updates to send out. This process (albeit highly optimised) surely takes some time. The shadow entries also impose a certain space overhead (only for replicated portions of a table). On the other hand, we do not require any additional processing during normal database use since we do not have to copy transaction log entries. And - our approach is not vulnerable to lost transactions. I hope this information sheds some light on general replication issues. Although I have not already worked with Informix' 6.0 replication I get the impression that it does the things it was designed for just right. If you have different needs you might take another approach as we did. I guess currently you won't get all the options in a single package. -- Oliver Okrongli infix Software-Systeme GmbH Phone +49 531 238090 Rebenring 33 Fax +49 531 2380935 oliver@infix.de 38106 Braunschweig F.R. Germany