Re: HDR and ER Running Together
Posted in 2012
Topics: High Availability & Replication, Networking & sqlhosts Configuration
Particularly Jacques or Madison
I've had this situation four or five times now, HDR has restarted sending
updates as soon as the secondary engine has been bounced, but this
is obviously a manual process and suspends ER for the duration of the
HDR catchup.
I will raise a PMR the next time it happens, but will need to capture the
state of both engines and get HDR restarted well before Tech Support
can respond. What commands would be best to run to capture the
required detail. I could use onstat -a on both sides but would rather
be a bit more specific if possible.
Any thoughts?
Keith
On 20 June 2012 13:58, JACQUES RENAUT <jrenaut@us.ibm.com> wrote:
> Original post:
> Madison
>
> Thanks for this explanation, it makes logical sense and ties in with
> what I see occasionally. However I got the following errors a couple
> of days ago:
> Primary:
> 13:40:25 DR: Cannot connect to secondary server
> 13:40:36 DR: Primary server connected
> 13:40:36 DR: Send error
> 13:40:36 ASF Echo-Thread Server: asfcode = -25580: oserr = 32: errstr =
> : System error occurred in network function.
> System error = 32.
> 13:40:36 DR: Failure recovery error (2)
> 13:40:37 DR: Turned off on primary server
> 13:40:37 DR: Cannot connect to secondary server
> 13:40:48 DR: Primary server connected
> 13:40:49 DR: Secondary server needs failure recovery
>
> Secondary:
> 13:40:35 DR: Received connection request from remote server when DR is
> not Off
>
> [Local type: Secondary, Current state: ?]
>
> [Remote type: Primary]
> 13:40:42 DR: ping timeout
> 13:40:42 DR: Receive error
> 13:40:42 ASF Echo-Thread Server: asfcode = -25582: oserr = 4: errstr =
> : Network connection is broken.
> System error = 4.
> 13:40:44 DR: Turned off on secondary server
> 13:40:47 DR: Secondary server connected
> 13:40:49 DR: Secondary server needs failure recovery
> 13:40:49 DR: Failure recovery from disk in progress ...
> 13:41:24 DR: Received connection request from remote server when DR is
> not Off
>
> [Local type: Secondary, Current state: ?]
>
> [Remote type: Primary]
> 13:41:36 DR: Received connection request from remote server when DR is
> not Off
>
> [Local type: Secondary, Current state: ?]
>
> HDR failed to reconnect for some 40 minutes during which time ER failed to
> send transactions. HDR only restarted when the secondary engine was
> restarted.
> Thoughts would be appreciated, but please tell me if this is better as a PMR
> rather than discussion here.
>
> Keith
>
> Response:
>
> Well, I'm not Madison, but what you've shown is not expected behavior. When
> the secondary got it's ping timeout, it should have shutdown the dr threads
> and then allowed the primary to reconnect. It appears something went wrong,
> which was preventing the secondary from handling the ping timeout properly so
> it then wasn't allowing the primary to reconnect. It should be a pmr, but the
> problem is if there wasn't data collected from when it happened, it might
need
> to happen again to figure out what is happening (assuming that it doesn't
> happen at every ping timeout).
>
> Jacques Renaut
> IBM Informix Advanced Support
> APD Team
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Original post:
Particularly Jacques or Madison
I've had this situation four or five times now, HDR has restarted sending
updates as soon as the secondary engine has been bounced, but this
is obviously a manual process and suspends ER for the duration of the
HDR catchup.
I will raise a PMR the next time it happens, but will need to capture the
state of both engines and get HDR restarted well before Tech Support
can respond. What commands would be best to run to capture the
required detail. I could use onstat -a on both sides but would rather
be a bit more specific if possible.
Any thoughts?
Keith
Response:
Well, this is what you should make sure to collect on the both servers before
you bounce the secondary.
Both:
1) onstat -g ath, onstat -g stk all, onstat -a
Just secondary:
2) Ideally if you have the disk space, I would also grab an onstat -o shared
memory dump
Then when you open the PMR you can forward on to them the onstat -g ath, -g
stk all, -a output, and MSGPATH file from both servers from around the time of
the ping timeout was reported on the primary. Then you can also let them know
you also grabbed a shared memory dump (if you can/do grab that from the
secondary) if they need that as well.
Jacques Renaut
IBM Informix Advanced Support
APD Team