Enterprise Replication Problem
Posted in 2011
A test environment cloned from production backups (IDS 7.31) had one-way Enterprise Replication failures: A replicated to B and C, but neither B nor C replicated back to A, with no ATS/RIS files or online.log errors. Connectivity and sqlhosts were ruled out. Madison Pruet explained the cause: the restored servers carry production's syscdr 'progress' tables, which hold the last sent/applied Log Sequence Numbers, so incoming transactions with earlier LSNs are treated as duplicates and silently dropped. His advice was to fully remove ER on the restored servers (cdr delete server, issued both on the server being deleted and on another server) and redefine it with cdr define server, cdr change participant, etc. The poster's follow-up asking which table holds the LSNs and whether it could be edited manually went unanswered, so no confirmed outcome is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Installation, Setup & Upgrades, Versions, Editions & End-of-Life
Hello Informix-Users. OS: SLES11, SP2 We want to migrate from IDS 7.31.UD10X1 to Informix Ultimate 11.70.UC5 and in this context we set up an Testenvironment which should be a copy of our productive System. In this Testenvironment we`ve got a problem with Enterprise Replication. The Testenvironment is still running under 7.31, and we want to upgrade the testservers step by step to 11.70, but therefor ER must be running under 7.31, because we must upgrade our productive servers by running business and we want to adjust this on the testservers first. Example: In our Testenvironment we have three Servers: testserverA with a db-backup from productive server A testserverB with a db-backup from productive server B testserverC with a db-backup from productive server C The database-structure of the productive servers is equal. The online-Backups were done in running business and to different times (a few hours difference). In our productive-enviroment Enterprise Replication works without any problems, but in the test-environment we`ve got the problem, that one way of replication fails: testserver A replicates to B and C without any problems, testserver B replicates only to C and NOT to A, testserver C relpicates only to B and NOT to A. In addition there are no ats- or ris-files generated and there is no error in the online.log. Every Server owns the same replicates, same tables, same primary keys, etc. and i can not understand why the ER fails at only one replication-way. Have somebody of the experts here a good advice or a clue where i have to look or what i can do to get ER completely running. Could it be, that Informix remembers the different Backuptimestamps and that this is the barrier for a running ER? Thanks for your help. Sascha
hi Hard to say whats going on. anyway if you have not done so already ,for the 3 servers to talk to each other you need to have entries in all sqlhosts files for the servers. So in sqlhosts file for A you need to have entries for B and C for host B the sqlhosts file has to have entries for and C and A and so on. Times of backups wont be the problem only means your dbs will be out of sync . The other problem I have struck is ports not being open on the firewall for servers to talk to each other.
Hi!
The sqlhosts-file and the ports are ok.
A connect via dbaccess is possible between all servers. The replication is
partially working, here in my example between testserver B and C. Only one way
of replication is not working.
The problem must be located on an other position.
Any other suggestions?
Thanks.
Sascha
Can you show me how you defined your replicates? What was the cdr command? From: "SASCHA KURATIS" <sascha.kuratis@westfleisch.de> To: ids@iiug.org Date: 10/11/2011 08:41 AM Subject: Enterprise Replication Problem [25162] Sent by: ids-bounces@iiug.org Hello Informix-Users. OS: SLES11, SP2 We want to migrate from IDS 7.31.UD10X1 to Informix Ultimate 11.70.UC5 and in this context we set up an Testenvironment which should be a copy of our productive System. In this Testenvironment we`ve got a problem with Enterprise Replication. The Testenvironment is still running under 7.31, and we want to upgrade the testservers step by step to 11.70, but therefor ER must be running under 7.31, because we must upgrade our productive servers by running business and we want to adjust this on the testservers first. Example: In our Testenvironment we have three Servers: testserverA with a db-backup from productive server A testserverB with a db-backup from productive server B testserverC with a db-backup from productive server C The database-structure of the productive servers is equal. The online-Backups were done in running business and to different times (a few hours difference). In our productive-enviroment Enterprise Replication works without any problems, but in the test-environment we`ve got the problem, that one way of replication fails: testserver A replicates to B and C without any problems, testserver B replicates only to C and NOT to A, testserver C relpicates only to B and NOT to A. In addition there are no ats- or ris-files generated and there is no error in the online.log. Every Server owns the same replicates, same tables, same primary keys, etc. and i can not understand why the ER fails at only one replication-way. Have somebody of the experts here a good advice or a clue where i have to look or what i can do to get ER completely running. Could it be, that Informix remembers the different Backuptimestamps and that this is the barrier for a running ER? Thanks for your help. Sascha ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Hi. Here is an exemplary "cdr define repl" command: cdr define repl R_branche -Ctimestamp -Stran -A \\\\ "db01@testacdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testbcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testccdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testdcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testecdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testfcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testgcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testhcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testicdr:root.branche" \\\\ "select * from branche " \\\\ In our testenvironment we don`t define all replicates new, because we have the testservers restored on the basis of backups of our productive servers, where ER is running fine. The used backups are from the same day, but not exactly from the same time (few hours difference) and they are done in running business. Thanks for your help. Sascha
If you use backups to create an alternate node, you will have problems because the progress tables will reflect what the original source node has sent/received. You will first have to completely remove ER from the restored server and then redefine it. Don't forget that in removing ER, you must issue "cdr delete server" twice, once on the server being deleted and once on one of the other servers. From: "SASCHA KURATIS" <sascha.kuratis@westfleisch.de> To: ids@iiug.org Date: 10/13/2011 01:20 AM Subject: Re: Enterprise Replication Problem [25191] Sent by: ids-bounces@iiug.org Hi. Here is an exemplary "cdr define repl" command: cdr define repl R_branche -Ctimestamp -Stran -A \\\\ "db01@testacdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testbcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testccdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testdcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testecdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testfcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testgcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testhcdr:root.branche" \\\\ "select * from branche " \\\\ "db01@testicdr:root.branche" \\\\ "select * from branche " \\\\ In our testenvironment we don`t define all replicates new, because we have the testservers restored on the basis of backups of our productive servers, where ER is running fine. The used backups are from the same day, but not exactly from the same time (few hours difference) and they are done in running business. Thanks for your help. Sascha ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Hi. Thanks for your answers. The problem is, when we must redefine whole ER, that we do not have a testenvironment which is a 1:1 copy of our productive-environment. Could you or anyone else told us, where we find the "progress tables, which will reflect what the original source has sent/received" ? What kind of tables are these? Is there a way to edit them? Sascha
The progress tables are internal system tables used by ER and is contained in the syscdr database. From: "SASCHA KURATIS" <sascha.kuratis@westfleisch.de> To: ids@iiug.org Date: 10/13/2011 07:03 AM Subject: Re: Enterprise Replication Problem [25194] Sent by: ids-bounces@iiug.org Hi. Thanks for your answers. The problem is, when we must redefine whole ER, that we do not have a testenvironment which is a 1:1 copy of our productive-environment. Could you or anyone else told us, where we find the "progress tables, which will reflect what the original source has sent/received" ? What kind of tables are these? Is there a way to edit them? Sascha ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Hi. Thanks for your efforts, Madison. Which syscdr-tables contains this information about ER? Which are these "progress-tables" ? Is there a syscdr-table which contains information about the time, when the restore was implemented and the database and ER was inititalized first time? Sascha
The progress tables contain the LSN (Log Sequence Number) of the last transaction which was sent (on a source side) or applied (on the target side). If a transaction is received which appears to be previous to that then it is considered to be a duplicate and is rejected. If you simply copied the production to a test environment, then it will have the progress tables from the production system, which I would assume are far in advance of the other test systems. You would need to re-establish ER on that system by running "cdr delete serv" followed by "cdr define serv", "cdr change participant", etc. M.P. From: "SASCHA KURATIS" <sascha.kuratis@westfleisch.de> To: ids@iiug.org Date: 10/14/2011 02:06 AM Subject: Re: Enterprise Replication Problem [25196] Sent by: ids-bounces@iiug.org Hi. Thanks for your efforts, Madison. Which syscdr-tables contains this information about ER? Which are these "progress-tables" ? Is there a syscdr-table which contains information about the time, when the restore was implemented and the database and ER was inititalized first time? Sascha ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
In which syscdr-table are this "Log Sequence Number"-Informations? Could we manipulate this entry by way of trial? Annotation: We build our test-environment completely new with copies of our productive-environment-databases. There exists no old databases before this db-restores. Sascha