Nasty, unexpected cluster problem
Posted in 2013
Topics: High Availability & Replication, Networking & sqlhosts Configuration, Clustering, Grid & MACH11
We have a MACH11 cluster defined. It has an SDS Primary, an SDS secondary, and 2 RSSs. The SLA is defined thus: CLUSTER gcgob1_db LOCAL SLA Connections Service/Protocol Rule ob1_oltp_1 234463 ob1_oltp/onsoctcp DBSERVERS=primary ob1_obbo_1 0 ob1_obbo/onsoctcp DBSERVERS=SDS,primary ob1_orep_1 0 ob1_orep/onsoctcp DBSERVERS=SDS,primary and the sqlhosts file on the CM: ob1_db group - - i=1,c=0 ob1_db1_ap onsoctcp PXCINFMDB01OBFE ob1_db1_ap g=ob1_db # SDS Pri ob1_db2_ap onsoctcp PXCINFMDB02OBFE ob1_db2_ap g=ob1_db # SDS Sec ob1_db3_ap onsoctcp PXCINFMDB03OBFE ob1_db3_ap g=ob1_db # 1st RSS ..... (it is never intended to fail over automatically to either RSS) We took one of the RSS servers, ob1_db3_ap, and made it a Standard in order to do some pack and shrink housekeeping on it in connection with another PMR. Our expectation was that making it a Standard would remove it from the cluster definition. But in fact Connection Manager started sending OLTP connections to it once it was a Standard, which wasn't good at all as we therefore had *two* - completely unconnected - OLTP servers in our cluster config! Luckily we noticed it fairly quickly but it's still caused financial and reputational damage, and time clearing up, so not great! Expected behaviour? Or a bug? Thanks Neil
It's not up to me to say...But can you show us the CM config file? CM can be used to redirect connections to independent servers... but if you only have roles in your CM config file, the new server status should not match any of those roles... Nevertheless... did you remove the RSS definition from the primary? I seem to remember that we need to do that... the simple transformation of the server role may not remove the config from the primary... and the CM can get info from it... I'd say the software could do a better job... But it may be a "blank space"... something that was not thought to be possible... a non tested scenario. Regards On Wed, Oct 16, 2013 at 5:43 PM, NEIL TRUBY <neil.truby@ardenta.com> wrote: > We have a MACH11 cluster defined. It has an SDS Primary, an SDS secondary, > and > 2 RSSs. > > The SLA is defined thus: > CLUSTER gcgob1_db LOCAL > SLA Connections Service/Protocol Rule > ob1_oltp_1 234463 ob1_oltp/onsoctcp DBSERVERS=primary > ob1_obbo_1 0 ob1_obbo/onsoctcp DBSERVERS=SDS,primary > ob1_orep_1 0 ob1_orep/onsoctcp DBSERVERS=SDS,primary > > and the sqlhosts file on the CM: > > ob1_db group - - i=1,c=0 > ob1_db1_ap onsoctcp PXCINFMDB01OBFE ob1_db1_ap g=ob1_db # SDS Pri > ob1_db2_ap onsoctcp PXCINFMDB02OBFE ob1_db2_ap g=ob1_db # SDS Sec > ob1_db3_ap onsoctcp PXCINFMDB03OBFE ob1_db3_ap g=ob1_db # 1st RSS > ...... > > (it is never intended to fail over automatically to either RSS) > > We took one of the RSS servers, ob1_db3_ap, and made it a Standard in > order to > do some pack and shrink housekeeping on it in connection with another PMR. > Our > expectation was that making it a Standard would remove it from the cluster > definition. But in fact Connection Manager started sending OLTP > connections to > it once it was a Standard, which wasn't good at all as we therefore had > *two* > - completely unconnected - OLTP servers in our cluster config! > > Luckily we noticed it fairly quickly but it's still caused financial and > reputational damage, and time clearing up, so not great! > > Expected behaviour? Or a bug? > > Thanks > Neil > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently... --e89a8f234635ade7d904e8def22d
I didn't do the work but I am guessing we did. From the SDS Primary: 08:09:19 Maximum server connections 1858 08:09:19 Checkpoint Statistics - Avg. Txn Block Time 0.001, # Txns blocked 1, Plog used 2926, Llog used 3258 08:10:14 Error receiving a buffer from RSS ob1_db3_ha - shutting down 08:10:14 RSS ob1_db3_ha deleted 08:10:14 Unexpected RSS Authentication Failure during Initialization 08:10:14 Unexpected Initial ACK during CloneSendInt 08:10:16 Unexpected RSS Authentication Failure during Initialization 08:10:16 Unexpected Initial ACK during CloneSendInt 08:10:17 Unexpected RSS Authentication Failure during Initialization 08:10:17 Unexpected Initial ACK during CloneSendInt 08:10:18 Unexpected RSS Authentication Failure during Initialization 08:10:18 Unexpected Initial ACK during CloneSendInt I sent you the config file separately, although part of it was in the original message. cheers! N