Enterprise replication-unable to initialize server
Posted in 2005
Topics: High Availability & Replication, Stored Procedures & SPL, Networking & sqlhosts Configuration, Versions, Editions & End-of-Life
I have three servers: server1 IDS 9.30 server2 IDS 9.30 server3 IDS 9.40 Server 1 and 2 are currently doing update anywhere replication, I want to add server 3. This is my define script: cdr define server -c omagrs -A /tmp/ats -R /tmp/ats -I omagrs -S omaha Here's what the log says when I run it... 14:18:29 Building 'syscdr' database ... 14:18:29 'syscdr' database built successfully. 14:18:29 CDR Warning: protocol not supported by Enterprise Replication group (omagrs) (grs:ipcshm) 14:18:29 Loading Module <SPLNULL> 14:18:29 CDR queuer initialization complete 14:18:29 CDR NIF listening on asf://omagrs 14:18:32 CDR Warning: protocol not supported by Enterprise Replication group (omaha) (prod:ipcshm) 14:18:33 CDR connection to server lost, id 10, name <omaha> Reason: ASF connect error (-956) 14:18:33 CDR GC: operation sync connect failed (error 5). 14:18:33 CDR GC: synchronization failed (sync abort, shutting down CDR) 14:18:33 CDR NIF Shutdown: connections all shutdown. 14:18:33 CDR The NIF sub-component has shut down. 14:18:33 CDR shutdown complete That's confusing. Why am I trying to use ipcshm?. Here's the sqlhosts file. servername nettype hostname servicename options #############################Primary GRS Replication Loop##################### omaha group - - i=10 prod_tcp onsoctcp server1 prod_tcp g=omaha prod onipcshm server1 sqlexecd g=omaha daldrprod group - - i=70 drprod_tcp onsoctcp server2 drprod_tcp g=daldrprod drprod onipcshm server2 sqlexecd g=daldrprod omagrs group - - i=90 grs_tcp onsoctcp server3 grs_tcp g=omagrs grs onipcshm server3 sqlexecd g=omagrs INFORMIXSERVER is set to grs_tcp on server3. Can you see what is mis-configured? TIA Dave Thacker -- Dave Thacker Senior Systems Administrator Omni Hotels Reservation Center V:402-952-6535 F:402-334-8013 M: 402-981-4613 (24/7)
--0__=09BBE533DF13E4EE8f9e8a93df938690918c09BBE533DF13E4EE Content-type: multipart/alternative; Boundary="1__=09BBE533DF13E4EE8f9e8a93df938690918c09BBE533DF13E4EE" --1__=09BBE533DF13E4EE8f9e8a93df938690918c09BBE533DF13E4EE Content-type: text/plain; charset=US-ASCII Content-transfer-encoding: quoted-printable That's a warning because you have IPSCHM in the group. TCP and SHM are= not equivalent ways to get to the server since TCP uses the network and SHM= doesn't. I suspect that the connection failed on the peer node becaus= e of some mismatch. Both ends of the connection go through a bit of verification during the connect request. = "Dave Thacker " = <dthacker@omnihot = els.com> = To Sent by: ids@iiug.org = forum.subscriber@ = cc iiug.org = Subj= ect Enterprise replication-unable to= 02/06/2005 03:19 initialize server [4179] = PM = = = = = = I have three servers: server1 IDS 9.30 server2 IDS 9.30 server3 IDS 9.40 Server 1 and 2 are currently doing update anywhere replication, I want = to add server 3. This is my define script: cdr define server -c omagrs -A /tmp/ats -R /tmp/ats -I omagrs -S omaha Here's what the log says when I run it... 14:18:29 Building 'syscdr' database ... 14:18:29 'syscdr' database built successfully. 14:18:29 CDR Warning: protocol not supported by Enterprise Replication= group (omagrs) (grs:ipcshm) 14:18:29 Loading Module <SPLNULL> 14:18:29 CDR queuer initialization complete 14:18:29 CDR NIF listening on asf://omagrs 14:18:32 CDR Warning: protocol not supported by Enterprise Replication= group (omaha) (prod:ipcshm) 14:18:33 CDR connection to server lost, id 10, name <omaha> Reason: ASF connect error (-956) 14:18:33 CDR GC: operation sync connect failed (error 5). 14:18:33 CDR GC: synchronization failed (sync abort, shutting down CDR= ) 14:18:33 CDR NIF Shutdown: connections all shutdown. 14:18:33 CDR The NIF sub-component has shut down. 14:18:33 CDR shutdown complete That's confusing. Why am I trying to use ipcshm?. Here's the sqlhosts= file. servername nettype hostname servicename options #############################Primary GRS Replication Loop##################### omaha group - - i=3D10 prod_tcp onsoctcp server1 prod_tcp g=3Domah= a prod onipcshm server1 sqlexecd g=3Domah= a daldrprod group - - i=3D70 drprod_tcp onsoctcp server2 drprod_tcp g=3Ddald= rprod drprod onipcshm server2 sqlexecd g=3Ddald= rprod omagrs group - - i=3D90 grs_tcp onsoctcp server3 grs_tcp g=3Domagrs= grs onipcshm server3 sqlexecd g=3Domagrs= INFORMIXSERVER is set to grs_tcp on server3. Can you see what is mis-configured? TIA Dave Thacker -- Dave Thacker Senior Systems Administrator Omni Hotels Reservation Center V:402-952-6535 F:402-334-8013 M: 402-981-4613 (24/7) = --1__=09BBE533DF13E4EE8f9e8a93df938690918c09BBE533DF13E4EE Content-type: text/html; charset=US-ASCII Content-Disposition: inline Content-transfer-encoding: quoted-printable <html><body> <p>That's a warning because you have IPSCHM in the group. TCP and SHM = are not equivalent ways to get to the server since TCP uses the network= and SHM doesn't. I suspect that the connection failed on the peer no= de because of some mismatch. Both ends of the connection go through a = bit of verification during the connect request. <br> <br> <br> <img src=3D"cid:10__=3D09BBE533DF13E4EE8f9e8a93df938@us.ibm.com" width=3D= "16" height=3D"16" alt=3D"Inactive hide details for "Dave Thacker = " <dthacker@omnihotels.com>">"Dave Thacker " <d= thacker@omnihotels.com><br> <br> <br> <table width=3D"100%" border=3D"0" cellspacing=3D"0" cellpadding=3D"0">= <tr valign=3D"top"><td style=3D"background-image:url(cid:20__=3D09BBE53= 3DF13E4EE8f9e8a93df938@us.ibm.com); background-repeat: no-repeat; " wid= th=3D"40%"> <ul> <ul> <ul> <ul><b><font size=3D"2">"Dave Thacker " <dthacker@omnihote= ls.com></font></b><font size=3D"2"> </font><br> <font size=3D"2">Sent by: forum.subscriber@iiug.org</font> <p><font size=3D"2">02/06/2005 03:19 PM</font></ul> </ul> </ul> </ul> </td><td width=3D"60%"> <table width=3D"100%" border=3D"0" cellspacing=3D"0" cellpadding=3D"0">= <tr valign=3D"top"><td width=3D"1%" valign=3D"middle"><img src=3D"cid:3= 0__=3D09BBE533DF13E4EE8f9e8a93df938@us.ibm.com" border=3D"0" height=3D"= 1" width=3D"58" alt=3D""><br> <div align=3D"right"><font size=3D"2">To</font></div></td><td width=3D"= 100%"><img src=3D"cid:30__=3D09BBE533DF13E4EE8f9e8a93df938@us.ibm.com" = border=3D"0" height=3D"1" width=3D"1" alt=3D""><br> <font size=3D"2">ids@iiug.org</font></td></tr> <tr valign=3D"top"><td width=3D"1%" valign=3D"middle"><img src=3D"cid:3= 0__=3D09BBE533DF13E4EE8f9e8a93df938@us.ibm.com" border=3D"0" height=3D"= 1" width=3D"58" alt=3D""><br> <div align=3D"right"><font size=3D"2">cc</font></div></td><td width=3D"= 100%"><img src=3D"cid:30__=3D09BBE533DF13E4EE8f9e8a93df938@us.ibm.com" = border=3D"0" height=3D"1" width=3D"1" alt=3D""><br> </td></tr> <tr valign=3D"top"><td width=3D"1%" valign=3D"middle"><img src=3D"cid:3= 0__=3D09BBE533DF13E4EE8f9e8a93df938@us.ibm.com" border=3D"0" height=3D"= 1" width=3D"58" alt=3D""><br> <div align=3D"right"><font size=3D"2">Subject</font></div></td><td widt= h=3D"100%"><img src=3D"cid:30__=3D09BBE533DF13E4EE8f9e8a93df938@us.ibm.= com" border=3D"0" height=3D"1" width=3D"1" alt=3D""><br> <font size=3D"2">Enterprise replication-unable to initialize server [41= 79]</font></td></tr> </table> <table border=3D"0" cellspacing=3D"0" cellpadding=3D"0"> <tr valign=3D"top"><td width=3D"58"><img src=3D"cid:30__=3D09BBE533DF13= E4EE8f9e8a93df938@us.ibm.com" border=3D"0" height=3D"1" width=3D"1" alt= =3D""></td><td width=3D"336"><img src=3D"cid:30__=3D09BBE533DF13E4EE8f9= e8a93df938@us.ibm.com" border=3D"0" height=3D"1" width=3D"1" alt=3D""><= /td></tr> </table> </td></tr> </table> <br> <tt>I have three servers:<br> server1 IDS 9.30<br> server2 IDS 9.30<br> server3 IDS 9.40<br> <br> Server 1 and 2 are currently doing update anywhere replication,
Check the '/etc/hosts.equiv' file, ensure server1 and server3 are trusted. I suspect, machines not being trusted is the main reason you are seeing ' ASF connect error (-956)'. Regards, Nilesh "Dave Thacker " <dthacker@omnihotels.com> Sent by: forum.subscriber@iiug.org 02/06/2005 03:19 PM To ids@iiug.org cc Subject Enterprise replication-unable to initialize server [4179] I have three servers: server1 IDS 9.30 server2 IDS 9.30 server3 IDS 9.40 Server 1 and 2 are currently doing update anywhere replication, I want to add server 3. This is my define script: cdr define server -c omagrs -A /tmp/ats -R /tmp/ats -I omagrs -S omaha Here's what the log says when I run it... 14:18:29 Building 'syscdr' database ... 14:18:29 'syscdr' database built successfully. 14:18:29 CDR Warning: protocol not supported by Enterprise Replication group (omagrs) (grs:ipcshm) 14:18:29 Loading Module <SPLNULL> 14:18:29 CDR queuer initialization complete 14:18:29 CDR NIF listening on asf://omagrs 14:18:32 CDR Warning: protocol not supported by Enterprise Replication group (omaha) (prod:ipcshm) 14:18:33 CDR connection to server lost, id 10, name <omaha> Reason: ASF connect error (-956) 14:18:33 CDR GC: operation sync connect failed (error 5). 14:18:33 CDR GC: synchronization failed (sync abort, shutting down CDR) 14:18:33 CDR NIF Shutdown: connections all shutdown. 14:18:33 CDR The NIF sub-component has shut down. 14:18:33 CDR shutdown complete That's confusing. Why am I trying to use ipcshm?. Here's the sqlhosts file. servername nettype hostname servicename options #############################Primary GRS Replication Loop##################### omaha group - - i=10 prod_tcp onsoctcp server1 prod_tcp g=omaha prod onipcshm server1 sqlexecd g=omaha daldrprod group - - i=70 drprod_tcp onsoctcp server2 drprod_tcp g=daldrprod drprod onipcshm server2 sqlexecd g=daldrprod omagrs group - - i=90 grs_tcp onsoctcp server3 grs_tcp g=omagrs grs onipcshm server3 sqlexecd g=omagrs INFORMIXSERVER is set to grs_tcp on server3. Can you see what is mis-configured? TIA Dave Thacker -- Dave Thacker Senior Systems Administrator Omni Hotels Reservation Center V:402-952-6535 F:402-334-8013 M: 402-981-4613 (24/7)
This has been resolved: -Madison Pruet and David Williams pointed out the ipcshm entries with group id's in the sqlhosts file, these have been harmless before. but removing them did correct the ipcshm protocol error. -Nilesh Ozarkar pointed out the ASF connect error (-956). This was caused by the hosts.equiv file having fully qualified domain names and the sql hosts file having only the short domain name. Adding the short domain name to hosts.equiv resolved this. -After these were resolved we still were not able to connect, so we contacted support, who directed us to stop and start the engine on the primary instance (omaha) so that the engine would see the new sqlhosts entries. This is apparently a bug (155772). The restart fixed the problem. Thanks to all for your help. On to the next question.... DT -- Dave Thacker Senior Systems Administrator Omni Hotels Reservation Center V:402-952-6535 F:402-334-8013 M: 402-981-4613 (24/7)