DR: Ping Timeout
Posted in 2009
Topics: High Availability & Replication, Server Administration, Logging & Checkpoints, Networking & sqlhosts Configuration
Hello, We are frequently receiving below mentioned messages in online logs of primary and secondary instances. This problem suddenly started to occur a few days ago and before that every thing was going fine. We didn't change anything (either in onconfig file or any other setting) that might have caused this issue to arise. What could be the possible reason behind this problem? Messages on Primary Server. =========================== 05:16:24 DR: ping timeout 05:16:25 DR: Turned off on primary server 05:16:25 DR: Cannot connect to secondary server 05:16:36 DR: Primary server connected 05:16:36 DR: Secondary server needs failure recovery 05:16:37 DR: Sending log 207015 (current), size 51200 pages, 0.03 percent used 05:16:38 DR: Sending Logical Logs Completed 05:16:39 DR: Primary server operational 05:16:40 Checkpoint Completed: duration was 1 seconds. 05:16:40 Thu Nov 26 - loguniq 207015, logpos 0x10018, timestamp: 0x3dfa3c3e Interval: 54783 Messages on Secondary Server. ============================= 05:16:25 DR: Receive error 05:16:25 ASF Echo-Thread Server: asfcode = -25582: oserr = 9: errstr = : Network connection is broken. System error = 9. 05:16:25 DR_ERR set to -1 05:16:25 DR: Warning - Proxy Subsystem not terminated 05:16:26 DR: Turned off on secondary server 05:16:35 DR: Secondary server connected 05:16:36 DR: Secondary server needs failure recovery 05:16:37 DR: Failure recovery from disk in progress ... 05:16:39 B-tree scanners disabled. 05:16:40 DR: HDR secondary server operational 05:16:40 Checkpoint Completed: duration was 0 seconds. 05:16:40 Thu Nov 26 - loguniq 207015, logpos 0x10018, timestamp: 0x3dfa3c30 Interval: 54787
What's between the primary and secondary? A local network or some WAN connection? If it's a WAN, contact your WAN connectivity provider and have them check the lines for noise and throughput. Request that the provider clean or reroute the lines if that's the problem. Art Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Mon, Dec 7, 2009 at 6:50 PM, ANEES AHMAD <aanees@i2cinc.com> wrote: > Hello, > We are frequently receiving below mentioned messages in online logs of > primary > and secondary instances. This problem suddenly started to occur a few days > ago > and before that every thing was going fine. We didn't change anything > (either > in onconfig file or any other setting) that might have caused this issue to > arise. > > What could be the possible reason behind this problem? > > Messages on Primary Server. > =========================== > > 05:16:24 DR: ping timeout > 05:16:25 DR: Turned off on primary server > 05:16:25 DR: Cannot connect to secondary server > 05:16:36 DR: Primary server connected > 05:16:36 DR: Secondary server needs failure recovery > > 05:16:37 DR: Sending log 207015 (current), size 51200 pages, 0.03 percent > used > 05:16:38 DR: Sending Logical Logs Completed > 05:16:39 DR: Primary server operational > 05:16:40 Checkpoint Completed: duration was 1 seconds. > 05:16:40 Thu Nov 26 - loguniq 207015, logpos 0x10018, timestamp: 0x3dfa3c3e > Interval: 54783 > > Messages on Secondary Server. > ============================= > > 05:16:25 DR: Receive error > 05:16:25 ASF Echo-Thread Server: asfcode = -25582: oserr = 9: errstr = : > Network connection is broken. > System error = 9. > 05:16:25 DR_ERR set to -1 > 05:16:25 DR: Warning - Proxy Subsystem not terminated > 05:16:26 DR: Turned off on secondary server > 05:16:35 DR: Secondary server connected > 05:16:36 DR: Secondary server needs failure recovery > > 05:16:37 DR: Failure recovery from disk in progress ... > 05:16:39 B-tree scanners disabled. > 05:16:40 DR: HDR secondary server operational > 05:16:40 Checkpoint Completed: duration was 0 seconds. > 05:16:40 Thu Nov 26 - loguniq 207015, logpos 0x10018, timestamp: 0x3dfa3c30 > Interval: 54787 > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --0015174c3c1ccbb563047a2db95b
OS version? Informix version? I had the same thing happening on RedHat AS 3 with IDS 10.00.UC3. Never did resolve it, even running crossover cables between the servers in the datacenter (to eliminate network issues). The only thing to stop it was to run 'xtrace' (somewhere under $INFORMIXDIR) after starting the db engine on both primary and secondary. xtrace was to debug the problem per IBM support but it did a nice job of "fixing" it. Bob ----- Original Message ----- From: "Art Kagel" <art.kagel@gmail.com> To: ids@iiug.org Sent: Monday, December 7, 2009 8:45:19 PM GMT -05:00 US/Canada Eastern Subject: Re: DR: Ping Timeout [18293] What's between the primary and secondary? A local network or some WAN connection? If it's a WAN, contact your WAN connectivity provider and have them check the lines for noise and throughput. Request that the provider clean or reroute the lines if that's the problem. Art Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Mon, Dec 7, 2009 at 6:50 PM, ANEES AHMAD <aanees@i2cinc.com> wrote: > Hello, > We are frequently receiving below mentioned messages in online logs of > primary > and secondary instances. This problem suddenly started to occur a few days > ago > and before that every thing was going fine. We didn't change anything > (either > in onconfig file or any other setting) that might have caused this issue to > arise. > > What could be the possible reason behind this problem? > > Messages on Primary Server. > =========================== > > 05:16:24 DR: ping timeout > 05:16:25 DR: Turned off on primary server > 05:16:25 DR: Cannot connect to secondary server > 05:16:36 DR: Primary server connected > 05:16:36 DR: Secondary server needs failure recovery > > 05:16:37 DR: Sending log 207015 (current), size 51200 pages, 0.03 percent > used > 05:16:38 DR: Sending Logical Logs Completed > 05:16:39 DR: Primary server operational > 05:16:40 Checkpoint Completed: duration was 1 seconds. > 05:16:40 Thu Nov 26 - loguniq 207015, logpos 0x10018, timestamp: 0x3dfa3c3e > Interval: 54783 > > Messages on Secondary Server. > ============================= > > 05:16:25 DR: Receive error > 05:16:25 ASF Echo-Thread Server: asfcode = -25582: oserr = 9: errstr = : > Network connection is broken. > System error = 9. > 05:16:25 DR_ERR set to -1 > 05:16:25 DR: Warning - Proxy Subsystem not terminated > 05:16:26 DR: Turned off on secondary server > 05:16:35 DR: Secondary server connected > 05:16:36 DR: Secondary server needs failure recovery > > 05:16:37 DR: Failure recovery from disk in progress ... > 05:16:39 B-tree scanners disabled. > 05:16:40 DR: HDR secondary server operational > 05:16:40 Checkpoint Completed: duration was 0 seconds. > 05:16:40 Thu Nov 26 - loguniq 207015, logpos 0x10018, timestamp: 0x3dfa3c30 > Interval: 54787 > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --0015174c3c1ccbb563047a2db95b
Art, Both of the instances i.e. Primary and Secondary are on same LAN. One more thing that I would like to mention that we are using SunOS 5.10 Generic_125100-09 sun4v sparc SUNW,Sun-Fire-T200 with IBM Informix Dynamic Server Version 11.50.FC3W1.
Your hardware platform could be part of the problem. IDS does not play well on the SPARC T-class processors (indeed only SQL Server and Sybase run well there according to my research - Oracle and DB2 don't do any better than IDS)! You would have been much better off with an M-class or best of all an X-class machine from Sun. For example, I just read today about a recently published SAP benchmark running Oracle 10 on a Sun X-server which produced more than twice the throughput of a previous benchmark on a more powerful (at least on paper) T-server (same number of cores but at almost 2x the clock speed and with 4x the memory). Personally, I would recommend disabling the hardware threading which is what causes the performance problems for processes like IDS's Virtual Processors that use only a single OS thread (OK two threads if you count the KAIO thread). IDS's multithreaded architecture is dependent on a proprietary threading library that does not use OS threads and that the OS's scheduler does not know anything about. The hardware threads see the IDS threads as a single massive very busy thread and that's not what they are optimized to handle well. That said, however, I have not heard any specific problems with network responsiveness on these systems, so I don't know. I would contact IBM tech support and open a case so that they can work with you to diagnose the problem. Art Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) See you at the 2010 IIUG Informix Conference April 25-28, 2010 Overland Park (Kansas City), KS www.iiug.org/conf Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, Dec 9, 2009 at 12:24 PM, ANEES AHMAD <aanees@i2cinc.com> wrote: > Art, > Both of the instances i.e. Primary and Secondary are on same LAN. > > One more thing that I would like to mention that we are using SunOS 5.10 > Generic_125100-09 sun4v sparc SUNW,Sun-Fire-T200 with IBM Informix Dynamic > Server Version 11.50.FC3W1. > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --00151747beea05c3f2047a4f29f7
Related threads
- System Or Internal Error InterruptedIOException
- 25582 error on high volume of short live trans
- ASF Echo-Thread Server: asfcode = -25582 oserr = 4