Removing SDS Node servername has timed out - rem
Posted in 2009
On IDS 11.50.FC4 (HP-UX 11.23) with a primary plus two SDS secondaries, the primary's online.log repeatedly showed "Removing SDS Node ... has timed out - removing", the secondaries going disconnected, reconnects rejected, and later "Cannot send sync" messages, always around long checkpoints (SDS_TIMEOUT=60). Madison Pruet suggested poor/slow network between primary and secondary causing lag, and explained that during a checkpoint all pages must be flushed; if an SDS node hasn't received the logs that dirtied a page, SDS_TIMEOUT kicks in and the node is dropped. Nilesh asked for the secondaries' logs, UPDATABLE_SECONDARY setting and CPU VP count, but no fix or final resolution is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Logging & Checkpoints, Platform-Specific Issues, Clustering, Grid & MACH11
Hi all, informix 11.5FC4 on HP-UX 11.23 I have one primary and two SD secondary instances in my cluster. In my cluster my one of the secondary instance goes down and In the online.log , IDS report the following error: 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing 10:51:31 Logical Log 119139 - Backup Completed 10:51:31 Logical Log 119140 - Backup Started 10:51:31 SDS Server prodsds2_tcp - state is now disconnected 10:51:35 SDS: Reconnect from prodsds2_tcp rejected - server not known or needs to be restarted. 10:52:42 Logical Log 119140 - Backup Completed 10:52:42 Logical Log 119141 - Backup Started SDS_TIMEOUT is 60 secs on all the machines. What shuld be the value of this parameter? What culd be the reason for this error? thanx Poonam
I'd look at the network. If the communication between the primary and = the secondary is not good, and the secondary starts lagging too far behind,= then it will be automaticically shut down. = "POONAM TANDEL" = <poonam_tandel28@ = yahoo.co.in>? = To Sent by: ids@iiug.org = ids-bounces@iiug. = cc org = Subj= ect Removing SDS Node servername has= 05/20/2009 05:52 timed out - rem [15809] = AM = = = Please respond to = ids@iiug.org = = = Hi all, informix 11.5FC4 on HP-UX 11.23 I have one primary and two SD secondary instances in my cluster. In my cluster my one of the secondary instance goes down and In the online.log , IDS report the following error: 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing= 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing= 10:51:31 Logical Log 119139 - Backup Completed 10:51:31 Logical Log 119140 - Backup Started 10:51:31 SDS Server prodsds2_tcp - state is now disconnected 10:51:35 SDS: Reconnect from prodsds2_tcp rejected - server not known o= r needs to be restarted. 10:52:42 Logical Log 119140 - Backup Completed 10:52:42 Logical Log 119141 - Backup Started SDS_TIMEOUT is 60 secs on all the machines. What shuld be the value of this parameter? What culd be the reason for this error? thanx Poonam ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
Madison, Here are some more details on the issue. Here is the online log from the primary, NOTE: a checkpoint started 60 seconds before the prodsds2_tcp disconnected and as noted SDS_TIMEOUT is 60. Would something with the checkpoint cause one of the SDS servers to timeout? 10:50:07 Logical Log 119139 - Backup Started 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing 10:51:31 Logical Log 119139 - Backup Completed 10:51:31 Logical Log 119140 - Backup Started 10:51:31 SDS Server prodsds2_tcp - state is now disconnected 10:51:35 SDS: Reconnect from prodsds2_tcp rejected - server not known or needs to be restarted. 10:52:42 Logical Log 119140 - Backup Completed 10:52:42 Logical Log 119141 - Backup Started 10:53:28 Checkpoint Completed: duration was 177 seconds. 10:53:28 Wed May 20 - loguniq 119147, logpos 0x99ac05c, timestamp: 0x4c78a87e Interval: 4137 10:53:28 Maximum server connections 157 10:53:28 Checkpoint Statistics - Avg. Txn Block Time 0.001, # Txns blocked 41, Plog used 4112149, Llog used 2225587 Thanks, Jeff -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Madison Pruet Sent: Wednesday, May 20, 2009 8:29 AM To: ids@iiug.org Subject: Re: Removing SDS Node servername has timed out.... [15811] I'd look at the network. If the communication between the primary and = the secondary is not good, and the secondary starts lagging too far behind,= then it will be automaticically shut down. = "POONAM TANDEL" = <poonam_tandel28@ = yahoo.co.in>? = To Sent by: ids@iiug.org = ids-bounces@iiug. = cc org = Subj= ect Removing SDS Node servername has= 05/20/2009 05:52 timed out - rem [15809] = AM = = = Please respond to = ids@iiug.org = = = Hi all, informix 11.5FC4 on HP-UX 11.23 I have one primary and two SD secondary instances in my cluster. In my cluster my one of the secondary instance goes down and In the online.log , IDS report the following error: 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing= 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing= 10:51:31 Logical Log 119139 - Backup Completed 10:51:31 Logical Log 119140 - Backup Started 10:51:31 SDS Server prodsds2_tcp - state is now disconnected 10:51:35 SDS: Reconnect from prodsds2_tcp rejected - server not known o= r needs to be restarted. 10:52:42 Logical Log 119140 - Backup Completed 10:52:42 Logical Log 119141 - Backup Started SDS_TIMEOUT is 60 secs on all the machines. What shuld be the value of this parameter? What culd be the reason for this error? thanx Poonam ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. = **************************************************************************** *** Forum Note: Use "Reply" to post a response in the discussion forum.
Jeff, All pages must be flushed during the checkpoint. If one of the SDS nod= es has not received the logs which dirtied the page and we can't flush tha= t buffer, then SDS_TIMEOUT will kick in because that's going to be the on= ly way that we can flush the page. = "Jeff Filippi" = <iiug@itdataconsu = lting.com> = To Sent by: ids@iiug.org = ids-bounces@iiug. = cc org = Subj= ect RE: Removing SDS Node servername= 05/20/2009 09:47 has timed out.... [15812] = AM = = = Please respond to = ids@iiug.org = = = Madison, Here are some more details on the issue. Here is the online log from the primary, NOTE: a checkpoint started 60 seconds before the prodsds2_tcp disconnected and as noted SDS_TIMEOUT i= s 60. Would something with the checkpoint cause one of the SDS servers to timeout? 10:50:07 Logical Log 119139 - Backup Started 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing= 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing= 10:51:31 Logical Log 119139 - Backup Completed 10:51:31 Logical Log 119140 - Backup Started 10:51:31 SDS Server prodsds2_tcp - state is now disconnected 10:51:35 SDS: Reconnect from prodsds2_tcp rejected - server not known o= r needs to be restarted. 10:52:42 Logical Log 119140 - Backup Completed 10:52:42 Logical Log 119141 - Backup Started 10:53:28 Checkpoint Completed: duration was 177 seconds. 10:53:28 Wed May 20 - loguniq 119147, logpos 0x99ac05c, timestamp: 0x4c78a87e Interval: 4137 10:53:28 Maximum server connections 157 10:53:28 Checkpoint Statistics - Avg. Txn Block Time 0.001, # Txns bloc= ked 41, Plog used 4112149, Llog used 2225587 Thanks, Jeff -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Madison Pruet Sent: Wednesday, May 20, 2009 8:29 AM To: ids@iiug.org Subject: Re: Removing SDS Node servername has timed out.... [15811] I'd look at the network. If the communication between the primary and =3D= the secondary is not good, and the secondary starts lagging too far behind,= =3D then it will be automaticically shut down. =3D "POONAM TANDEL" =3D <poonam_tandel28@ =3D yahoo.co.in>? =3D To Sent by: ids@iiug.org =3D ids-bounces@iiug. =3D cc org =3D Subj=3D ect Removing SDS Node servername has=3D 05/20/2009 05:52 timed out - rem [15809] =3D AM =3D =3D =3D Please respond to =3D ids@iiug.org =3D =3D =3D Hi all, informix 11.5FC4 on HP-UX 11.23 I have one primary and two SD secondary instances in my cluster. In my cluster my one of the secondary instance goes down and In the online.log , IDS report the following error: 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing= =3D 10:51:31 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing= =3D 10:51:31 Logical Log 119139 - Backup Completed 10:51:31 Logical Log 119140 - Backup Started 10:51:31 SDS Server prodsds2_tcp - state is now disconnected 10:51:35 SDS: Reconnect from prodsds2_tcp rejected - server not known o= =3D r needs to be restarted. 10:52:42 Logical Log 119140 - Backup Completed 10:52:42 Logical Log 119141 - Backup Started SDS_TIMEOUT is 60 secs on all the machines. What shuld be the value of this parameter? What culd be the reason for this error? thanx Poonam ***********************************************************************= =3D ******** Forum Note: Use "Reply" to post a response in the discussion forum. =3D ***********************************************************************= ***** *** Forum Note: Use "Reply" to post a response in the discussion forum. ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
hi jeff,madison today morning also i faced the same problem and my both the instances went down. here is the online log form the primary,logs shows message that 'dynamically allocated 1000000 locks" and then after some time sds instances went down 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:22:35 Level 0 Archive started on els_hist_idx 12:22:43 Logical Log 119640 Complete, timestamp: 0x6e38bcd0. 12:23:28 Logical Log 119641 Complete, timestamp: 0x6e4c1889. 12:23:57 dynamically allocated 1000000 locks 12:32:52 Logical Log 119610 - Backup Started 12:33:09 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing 12:33:09 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing 12:34:10 ERROR: Removing SDS Node prodsds1_tcp has timed out - removing 12:34:10 ERROR: Removing SDS Node prodsds1_tcp has timed out - removing 12:34:16 Checkpoint Completed: duration was 128 seconds. 12:34:16 Thu May 21 - loguniq 119643, logpos 0x479613c, timestamp: 0x6eafcd15 Interval: 4446 12:34:16 Maximum server connections 212 12:34:16 Checkpoint Statistics - Avg. Txn Block Time 0.001, # Txns blocked 0, Plog used 262520, Llog used 132314 12:34:16 Level 0 Archive started on els_hist08_dat 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync 12:34:16 Encountered problem processing update on secondary : Cannot send sync
Would like to see what SDS were doing at the same time. could you copy + paste the message log of prodsds1_tcp and prodsds2_tcp little bit before and around the time when they were was shutdown. Also what's the value of UPDATEABLE_SECONDARY and number of CPU VPs. thx, Nilesh. ids-bounces@iiug.org wrote on 05/21/2009 03:35:56 AM: > [image removed] > > Re: RE: Removing SDS Node servername has timed.... [15819] > > POONAM TANDEL > > to: > > ids > > 05/21/2009 03:36 AM > > Sent by: > > ids-bounces@iiug.org > > Please respond to ids > > hi jeff,madison > > today morning also i faced the same problem and my both the instances went > down. > here is the online log form the primary,logs shows message that 'dynamically > allocated 1000000 locks" and then after some time sds instances went down > > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:22:35 Level 0 Archive started on els_hist_idx > 12:22:43 Logical Log 119640 Complete, timestamp: 0x6e38bcd0. > 12:23:28 Logical Log 119641 Complete, timestamp: 0x6e4c1889. > 12:23:57 dynamically allocated 1000000 locks > > 12:32:52 Logical Log 119610 - Backup Started > 12:33:09 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing > 12:33:09 ERROR: Removing SDS Node prodsds2_tcp has timed out - removing > 12:34:10 ERROR: Removing SDS Node prodsds1_tcp has timed out - removing > 12:34:10 ERROR: Removing SDS Node prodsds1_tcp has timed out - removing > 12:34:16 Checkpoint Completed: duration was 128 seconds. > 12:34:16 Thu May 21 - loguniq 119643, logpos 0x479613c, timestamp: 0x6eafcd15 > Interval: 4446 > > 12:34:16 Maximum server connections 212 > 12:34:16 Checkpoint Statistics - Avg. Txn Block Time 0.001, # Txns blocked 0, > Plog used 262520, Llog used 132314 > > 12:34:16 Level 0 Archive started on els_hist08_dat > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > 12:34:16 Encountered problem processing update on secondary : Cannotsend sync > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. >