Re: HDR with Synchronous replication
Posted in 2004
Topics: High Availability & Replication, Logging & Checkpoints
There is an important thing to consider if you choose to use HDR . The secondary server must be a "Stand By" Server. By "Stand By" i mean that, no hard work like dss queries on the secondary server, because checkpoint between servers are synchronous (regardless DRINTERVAL's value). =========================================
"Francisco Roldan" <f_roldan@admcentralagricola.com.gt> wrote in message news:cbcrq3$u35$1@news.xmission.com... > > There is an important thing to consider if you choose > to use HDR . The secondary server must be a "Stand By" Server. > By "Stand By" i mean that, no hard work like dss queries > on the secondary server, because checkpoint between > servers are synchronous (regardless DRINTERVAL's value). Yes I am aware of this. For us the secondary server will most likely by only a stand by. May be some reports during off peak hours. ER is ruled out due to too much pain to set it up. This option is coming up as a possible solution. The other option we have is to use SQL Server replication, either log shipping or transactional replication. Both are not as smooth as HDR. HDR is probably the easiest to set up replication between two servers. And if we use the group concept than the clients can switch to the next server without any change. The only change required will be when the original primary is backup. By then we have to change INFORMIXSQLHOSTS to change the order of servers within the group. That's OK. The manual intervention is not required for the switch at the time of failure.
Francisco Roldan wrote: > There is an important thing to consider if you choose > to use HDR . The secondary server must be a "Stand By" Server. > By "Stand By" i mean that, no hard work like dss queries > on the secondary server, because checkpoint between > servers are synchronous (regardless DRINTERVAL's value). No! I'll explain later. > > ========================================= > From the Administration Guide for IDS 7.3 : (4354.pdf) : > - Chapter 25 "What is High Availability Data Replication": > (Page 578) > "Checkpoints Between Database Servers" > Checkpoint Between Database Servers in a High-Availability data-Replication > pair are synchronous, regardless of the value of DRINTERVAL . A checkpoint > on the primary database server completes only after it completes on the > secondary database server. > ========================================= This is "true" but requires a bit of thought... The "chekpoint" on a secondary server is not exactly the same as in the primary server. Since the secondary is read-only there aren't usually many buffers to flush. But the primary will only complete the checkpoint when one of the following happens: 1- It receives the aknoledgement from the secondary 2- 4*DRTIMEOUT time (in seconds) expires withou an answer from the secondary (in this situation it will disconnect the replication and try later) Anything else is probably a BUG (or a mistake from me... ;) ) > If you run a heavy DSS Query on secondary server, > the checkpoint's duration on primary gets really long. Not necessarily. And I can (in certain versions) cause a big problem with some simple statements... As stated above... bug. > > Fernando Nunes and me had a semi-discussion regarding > this issue last month (may), with the subject > "Extremely long checkpoints: How to find the reason ... " > He reported this issue to IBM and they told him that > this behavior is "by dessign", so I believe there is no > much intention to modify it. I don't recall to have written this. If I did, I was mistaken. As far as I recall (Google must know better than me) what I said was something like this: It certainly ISN'T by design. There was a 9.3x version which forced certain parameters (BUFFERS, CPUVPS etc.) to be the same on primary and secondary and this was considered a BUG. This clearly states that Informix (and IBM) assume the servers can be used for different things besides fault tolerance. HDR is a perfect solution to take some heavy load of DSS queries from your OLTP primary server. The situation I was facing when I participated in the thread Francisco mentioned is clarified by now. It's one of two bugs, corrected in 7.31.FD8 and 9.40x (can't be sure about 9.4 version). I could reproduce the situation and I couldn't in the patched versions. To make this perfectly clear: HDR can BY DESIGN be used for DSS queries on the secondary. This doensn't mean you will never hit a problem. But this problems are either bad configuration, bad usage or most probably bugs. The HDR mechanism is prepared to disconnect and later reconnect in case of network problems or other causes of "load". Regards. P.S.: I must apologize to Franscisco because I said I would give feedback on this case and I didn't. Too much work... or is it too less time...? :)
Francisco Roldan wrote: > There is an important thing to consider if you choose > to use HDR . The secondary server must be a "Stand By" Server. > By "Stand By" i mean that, no hard work like dss queries > on the secondary server, because checkpoint between > servers are synchronous (regardless DRINTERVAL's value). No! I'll explain later. > > ========================================= > From the Administration Guide for IDS 7.3 : (4354.pdf) : > - Chapter 25 "What is High Availability Data Replication": > (Page 578) > "Checkpoints Between Database Servers" > Checkpoint Between Database Servers in a High-Availability data-Replication > pair are synchronous, regardless of the value of DRINTERVAL . A checkpoint > on the primary database server completes only after it completes on the > secondary database server. > ========================================= This is "true" but requires a bit of thought... The "chekpoint" on a secondary server is not exactly the same as in the primary server. Since the secondary is read-only there aren't usually many buffers to flush. But the primary will only complete the checkpoint when one of the following happens: 1- It receives the aknoledgement from the secondary 2- 4*DRTIMEOUT time (in seconds) expires withou an answer from the secondary (in this situation it will disconnect the replication and try later) Anything else is probably a BUG (or a mistake from me... ;) ) > If you run a heavy DSS Query on secondary server, > the checkpoint's duration on primary gets really long. Not necessarily. And I can (in certain versions) cause a big problem with some simple statements... As stated above... bug. > > Fernando Nunes and me had a semi-discussion regarding > this issue last month (may), with the subject > "Extremely long checkpoints: How to find the reason ... " > He reported this issue to IBM and they told him that > this behavior is "by dessign", so I believe there is no > much intention to modify it. I don't recall to have written this. If I did, I was mistaken. As far as I recall (Google must know better than me) what I said was something like this: It certainly ISN'T by design. There was a 9.3x version which forced certain parameters (BUFFERS, CPUVPS etc.) to be the same on primary and secondary and this was considered a BUG. This clearly states that Informix (and IBM) assume the servers can be used for different things besides fault tolerance. HDR is a perfect solution to take some heavy load of DSS queries from your OLTP primary server. The situation I was facing when I participated in the thread Francisco mentioned is clarified by now. It's one of two bugs, corrected in 7.31.FD8 and 9.40x (can't be sure about 9.4 version). I could reproduce the situation and I couldn't in the patched versions. To make this perfectly clear: HDR can BY DESIGN be used for DSS queries on the secondary. This doensn't mean you will never hit a problem. But this problems are either bad configuration, bad usage or most probably bugs. The HDR mechanism is prepared to disconnect and later reconnect in case of network problems or other causes of "load". Regards. P.S.: I must apologize to Franscisco because I said I would give feedback on this case and I didn't. Too much work... or is it too less time...? :)
Fernando Nunes wrote: >> Fernando Nunes and me had a semi-discussion regarding >> this issue last month (may), with the subject >> "Extremely long checkpoints: How to find the reason ... " >> He reported this issue to IBM and they told him that >> this behavior is "by dessign", so I believe there is no >> much intention to modify it. > > > I don't recall to have written this. If I did, I was mistaken. > As far as I recall (Google must know better than me) what I said was > something like this: It does!: http://groups.google.pt/groups?q=comp.databases.informix+fernando+nunes+hdr&hl=pt-PT&lr=&ie=UTF-8&selm=c8ihrg%243i5%241%40terabinaries.xmission.com&rnum=3 and sorry for the double post Regards