Connection Manager Cluster
Posted in 2009
A user running two Connection Manager instances (oncmsm 3.50.FC4 / IDS 11.50.FC4) grouped in sqlhosts reported two issues. First, clients kept working after the CM was shut down; Art Kagel explained that's expected — the CM only hands the client the best server name, after which the client connects directly to IDS, so it isn't in the data path. Second, after restarting the database only one CM re-registered; the other logged "not active, waiting to detect primary". IBM's Jacques Renaut reproduced it and found it only occurred once the FOC line was added to the CM config, calling it a likely defect he would investigate — no fix or workaround is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Error Codes & Troubleshooting, Networking & sqlhosts Configuration, Clustering, Grid & MACH11, Java & JDBC Development
Hi,
We are facing a couple of problems with Connection manager cluster. We are
runing two Connection Manager instances on different machines.
And both are grouped in the sqlhosts file of Connection Manager. like
g_oltp group - - e=oltp2
oltp ontlitcp 192.168.0.16 9991 g=g_oltp
oltp2 ontlitcp 192.168.0.14 9991 g=g_oltp
Problem 1:
Our application acquired the connections to database through Connection
Manager. We can insert data through that setup.
Later we shutdown the Connection manager, still we are able to insert data.
How it comes?
Problem 2:
Both Connection Managers are running and one of them is serving the requests
90for inser etc).
We restarted the database. Database is up again but one of the Connection
manager is unable to detect the primary again.
Here are some extracts from logs:
Connection manager 1 log (which was stable, we can see in the log that both
connection manager are added "CM name"):
16:33:49 Connection Manager successfully connected to ids_net_rss1_11
16:33:50 Arbitrator FOC string = HDR+RSS,120
16:33:50 Arbitrator setting primary name = ids_rss1_11 [cmsm_arb.c:518]
16:33:50 FOC[0] = HDR
16:33:50 FOC[1] = RSS
16:33:50 FOC timeout = 120
16:33:50 get ids_net_rss1_11 CM_ADM event 8:3 cm_test11@16
[cmsm_server.ec:1138]
16:33:50 Arbitrator received CM event, subtype=3, cm=cm_test11@16
[cmsm_arb.c:533]
16:33:50 Arbitrator reinitialized CM names [cmsm_arb.c:479]
16:33:50 Arbitrator added CM name = cm_test11@16 [cmsm_arb.c:418]
16:33:54 SLA oltp redirect SQLI client from 192.168.0.58 to ids_net_rss1_11
192.168.0.14.6031 [cmsm_main.c:1279]
16:34:37 get ids_net_pri11 CM_ADM event 8:4 cm_test11@14 [cmsm_server.ec:1138]
16:34:37 Arbitrator received CM event, subtype=4, cm=cm_test11@14
[cmsm_arb.c:533]
16:34:39 get ids_net_rep11 CM_ADM event 8:4 cm_test11@14 [cmsm_server.ec:1138]
16:34:39 Arbitrator received CM event, subtype=4, cm=cm_test11@14
[cmsm_arb.c:533]
16:34:49 get ids_net_pri11 CM_ADM event 8:3 cm_test11@14 [cmsm_server.ec:1138]
16:34:49 get ids_net_rss1_11 CM_ADM event 8:3 cm_test11@14
[cmsm_server.ec:1138]
16:34:49 Arbitrator received CM event, subtype=3, cm=cm_test11@14
[cmsm_arb.c:533]
16:34:49 Arbitrator reinitialized CM names [cmsm_arb.c:479]
16:34:49 Arbitrator added CM name = cm_test11@16 [cmsm_arb.c:418]
16:34:49 Arbitrator added CM name = cm_test11@14 [cmsm_arb.c:418]
16:34:50 get ids_net_rep11 CM_ADM event 8:3 cm_test11@14 [cmsm_server.ec:1138]
16:35:06 SLA oltp redirect SQLI client from 192.168.33.220 to ids_net_rss1_11
192.168.0.14.6031 [cmsm_main.c:1279]
Connection Manager 2 log (here we can see that after database restart it is
unable to detect primary again)
15:13:26 FOC timeout = 120
15:13:26 get ids_net_rep11 CM_ADM event 8:3 cm_test11@14 [cmsm_server.ec:1138]
15:13:26 Connection Manager started successfully
15:16:19 get ids_net_rss1_11 SRV_ADM event 16:3 4 [cmsm_server.ec:1167]
15:16:23 fetch sysrepstats_cursor SQLCODE = (-25582,9,) [cmsm_server.ec:1085]
15:16:23 unregister event failed, sqlcode = (-1811,0,) [cmsm_server.ec:942]
15:16:23 Connection Manager disconnected from ids_net_rss1_11
15:16:23 Arbitrator detected down server, svr=ids_net_rss1_11, type=65536
[cmsm_arb.c:923]
15:16:23 Arbitrator = cm_test11@14 not active, waiting to detect primary
[cmsm_arb.c:954]
15:16:23 get ids_net_pri11 CLUST_CHG event 1:6 ids_net_rss1_11
[cmsm_server.ec:1097]
15:16:23 get ids_net_rep11 CLUST_CHG event 1:4 ids_net_rss1_11|Primary|
[cmsm_server.ec:1097]
15:16:24 Arbitrator = cm_test11@14 not active, waiting to detect primary
[cmsm_arb.c:954]
15:16:25 Arbitrator = cm_test11@14 not active, waiting to detect primary
[cmsm_arb.c:954]
15:16:26 Arbitrator = cm_test11@14 not active, waiting to detect primary
[cmsm_arb.c:954]
15:16:27 Arbitrator = cm_test11@14 not active, waiting to detect primary
[cmsm_arb.c:954]
15:16:28 Arbitrator = cm_test11@14 not active, waiting to detect primary
[cmsm_arb.c:954]
Database log (onstat -m, here we can see that one of the Connection manager is
registered but not the other one.
14:39:10 SCHAPI: Started 2 dbWorker threads.
14:39:11 CM:Connection manager cm_test11@16 registered with the server
Can some help for both problems. We are planning to use the Connection Manager
on production.
Note: We are using Java MI Pool in our application that requests the
connections through Connection manager.
regards,
Kamran
Connection manager ONLY negotiates for the best connection and passes that
servername back to the client application which then connects directly with
the selected instance of IDS. The client does not maintain an active
connection to the connection manager and the connection manager does not
participate in the communications between client and server once it has
supplied the selected servername.
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
On Wed, May 20, 2009 at 6:12 AM, KAMRAN HAQ <khaq@i2cinc.com> wrote:
> Hi,
> We are facing a couple of problems with Connection manager cluster. We are
> runing two Connection Manager instances on different machines.
> And both are grouped in the sqlhosts file of Connection Manager. like
>
> g_oltp group - - e=oltp2
> oltp ontlitcp 192.168.0.16 9991 g=g_oltp
> oltp2 ontlitcp 192.168.0.14 9991 g=g_oltp
>
> Problem 1:
>
> Our application acquired the connections to database through Connection
> Manager. We can insert data through that setup.
>
> Later we shutdown the Connection manager, still we are able to insert data.
> How it comes?
>
> Problem 2:
>
> Both Connection Managers are running and one of them is serving the
> requests
> 90for inser etc).
>
> We restarted the database. Database is up again but one of the Connection
> manager is unable to detect the primary again.
>
> Here are some extracts from logs:
>
> Connection manager 1 log (which was stable, we can see in the log that both
> connection manager are added "CM name"):
>
> 16:33:49 Connection Manager successfully connected to ids_net_rss1_11
> 16:33:50 Arbitrator FOC string = HDR+RSS,120
> 16:33:50 Arbitrator setting primary name = ids_rss1_11 [cmsm_arb.c:518]
> 16:33:50 FOC[0] = HDR
> 16:33:50 FOC[1] = RSS
> 16:33:50 FOC timeout = 120
> 16:33:50 get ids_net_rss1_11 CM_ADM event 8:3 cm_test11@16
> [cmsm_server.ec:1138]
> 16:33:50 Arbitrator received CM event, subtype=3, cm=cm_test11@16
> [cmsm_arb.c:533]
> 16:33:50 Arbitrator reinitialized CM names [cmsm_arb.c:479]
> 16:33:50 Arbitrator added CM name = cm_test11@16 [cmsm_arb.c:418]
> 16:33:54 SLA oltp redirect SQLI client from 192.168.0.58 to ids_net_rss1_11
> 192.168.0.14.6031 [cmsm_main.c:1279]
> 16:34:37 get ids_net_pri11 CM_ADM event 8:4 cm_test11@14 [
> cmsm_server.ec:1138]
> 16:34:37 Arbitrator received CM event, subtype=4, cm=cm_test11@14
> [cmsm_arb.c:533]
> 16:34:39 get ids_net_rep11 CM_ADM event 8:4 cm_test11@14 [
> cmsm_server.ec:1138]
> 16:34:39 Arbitrator received CM event, subtype=4, cm=cm_test11@14
> [cmsm_arb.c:533]
> 16:34:49 get ids_net_pri11 CM_ADM event 8:3 cm_test11@14 [
> cmsm_server.ec:1138]
> 16:34:49 get ids_net_rss1_11 CM_ADM event 8:3 cm_test11@14
> [cmsm_server.ec:1138]
> 16:34:49 Arbitrator received CM event, subtype=3, cm=cm_test11@14
> [cmsm_arb.c:533]
> 16:34:49 Arbitrator reinitialized CM names [cmsm_arb.c:479]
> 16:34:49 Arbitrator added CM name = cm_test11@16 [cmsm_arb.c:418]
> 16:34:49 Arbitrator added CM name = cm_test11@14 [cmsm_arb.c:418]
> 16:34:50 get ids_net_rep11 CM_ADM event 8:3 cm_test11@14 [
> cmsm_server.ec:1138]
> 16:35:06 SLA oltp redirect SQLI client from 192.168.33.220 to
> ids_net_rss1_11
> 192.168.0.14.6031 [cmsm_main.c:1279]
>
> Connection Manager 2 log (here we can see that after database restart it is
> unable to detect primary again)
>
> 15:13:26 FOC timeout = 120
> 15:13:26 get ids_net_rep11 CM_ADM event 8:3 cm_test11@14 [
> cmsm_server.ec:1138]
> 15:13:26 Connection Manager started successfully
> 15:16:19 get ids_net_rss1_11 SRV_ADM event 16:3 4 [cmsm_server.ec:1167]
> 15:16:23 fetch sysrepstats_cursor SQLCODE = (-25582,9,) [
> cmsm_server.ec:1085]
> 15:16:23 unregister event failed, sqlcode = (-1811,0,) [cmsm_server.ec:942
> ]
> 15:16:23 Connection Manager disconnected from ids_net_rss1_11
> 15:16:23 Arbitrator detected down server, svr=ids_net_rss1_11, type=65536
> [cmsm_arb.c:923]
> 15:16:23 Arbitrator = cm_test11@14 not active, waiting to detect primary
> [cmsm_arb.c:954]
> 15:16:23 get ids_net_pri11 CLUST_CHG event 1:6 ids_net_rss1_11
> [cmsm_server.ec:1097]
> 15:16:23 get ids_net_rep11 CLUST_CHG event 1:4 ids_net_rss1_11|Primary|
> [cmsm_server.ec:1097]
> 15:16:24 Arbitrator = cm_test11@14 not active, waiting to detect primary
> [cmsm_arb.c:954]
> 15:16:25 Arbitrator = cm_test11@14 not active, waiting to detect primary
> [cmsm_arb.c:954]
> 15:16:26 Arbitrator = cm_test11@14 not active, waiting to detect primary
> [cmsm_arb.c:954]
> 15:16:27 Arbitrator = cm_test11@14 not active, waiting to detect primary
> [cmsm_arb.c:954]
> 15:16:28 Arbitrator = cm_test11@14 not active, waiting to detect primary
> [cmsm_arb.c:954]
>
> Database log (onstat -m, here we can see that one of the Connection manager
> is
> registered but not the other one.
>
> 14:39:10 SCHAPI: Started 2 dbWorker threads.
> 14:39:11 CM:Connection manager cm_test11@16 registered with the server
>
> Can some help for both problems. We are planning to use the Connection
> Manager
> on production.
> Note: We are using Java MI Pool in our application that requests the
> connections through Connection manager.
>
> regards,
> Kamran
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--001636c5a5011fca72046a5726f7
Ok so that's why in case 1 application is communicating with database even the Connection Manager is down. And what goes wrong in case 2 (problem 2) where one of two Connection Managers become unstable and fails to detect primary node(instance restarted). And if Connection Manager detect the better link, handover to application and sleeps then why the log (Debug=1) of Connection Manager prints line for each request like: 16:33:54 SLA oltp redirect SQLI client from 192.168.0.58 to ids_net_rss1_11 192.168.0.14.6031 [cmsm_main.c:1279] Sorry if I am not getting the point. regards, Kamran
>Problem 2: > >Both Connection Managers are running and one of them is serving the requests 90for inser etc). > >We restarted the database. Database is up again but one of the Connection manager is unable to detect the primary again. I did a quick test and set up 2 connection manager configs and created a group like you had shown and bounced my database server and both of my connection managers successfully reconnected when the server comes back online. Could you post up your config files so I can compare them to what I used that appeared to be working? Also maybe a oncmsm -version as well? Jacques Renaut IBM IDS APD team
Enviornment script for Connection manager
----------------------------------------------------------------------
INFORMIXDIR=/export/informix11-50FC4 ;export INFORMIXDIR
INFORMIXSERVER=g_mach11 ;export INFORMIXSERVER
PATH=/usr/sbin:/usr/bin:/usr/local/bin:/usr/sfw/bin:/sbin:/usr/sfw/sparc-sun-solaris2.10/bin/:/usr/local/ssl/bin:$INFORMIXDIR/bin:$INFORMIXDIR/etc ; export
PATH
INFORMIXSQLHOSTS=/export/informix11-50FC4/etc/cm11_sqlhosts ;exportINFORMIXSQLHOSTS
----------------------------------------------------------------------
sqlhosts for Connection Manager
----------------------------------------------------------------------
g_mach11 group - - e=ids_net_rss2_11
ids_net_pri11 ontlitcp 192.168.0.15 sql_net_pri11 g=g_mach11
ids_net_rep11 ontlitcp 192.168.0.16 sql_net_rep11 g=g_mach11
ids_net_rss1_11 ontlitcp 192.168.0.14 sql_net_rss1_11 g=g_mach11
g_oltp group - - e=oltp
oltp2 ontlitcp 192.168.0.14 9991 g=g_oltp
oltp ontlitcp 192.168.0.16 9991 g=g_oltp
----------------------------------------------------------------------
Connection Manager configuration file
----------------------------------------------------------------------
NAME cm_test11@14
# SLA oltp=PRIMARY
SLA oltp2=PRIMARY
FOC HDR+RSS,120
SLA_WORKERS 16
DEBUG 1
LOGFILE /export/informix11-50FC4/etc/conmgr11.log
----------------------------------------------------------------------
# oncmsm -version
----------------------------------------------------------------------
Program Name: oncmsm
Build Version: 3.50.FC4
Build Number: C4
Build Host: apris
Build OS: SunOS-sparc 5.9
Build Date: Fri Apr 3 06:24:40 CDT 2009
GLS Version: glslib-4.50.FC5
----------------------------------------------------------------------
DRAUTO value is "3" in ONCONFIG file. IDS version is 11.5 FC4.
Each of Connection Manager instance works but, when database restarts one of
them lost.
regards,
Kamran
Hi, in addition to my previous post. We tperformed several tests and noted that if we have two Connection Manager instances say CM1 and CM2. If CM1 was started first and CM2 later, then after restart of database instance CM1 get registered but CM2 fails to detect database instance again and vice versa.
>Hi, >in addition to my previous post. We tperformed several tests and noted that if we have two Connection >Manager instances say CM1 and CM2. If CM1 was started first and CM2 later, then after restart of database >instance CM1 get registered but CM2 fails to detect database instance again and vice versa. Ok I can also get this behavior with the set up you appear to be using. It appeared to work and both connection managers appeared to reconnect after a server bounce before I added the FOC line in my config file. I would be inclined to think this is problem/defect. I'll look into this further. Jacques Renaut IBM IDS APD team