Wakeup RSS Server
Posted in 2015
A 11.70.FC7 site lost the network link between a primary and its RSS secondary; after the SMX thread timed out the RSS stayed 'Disconnected' and did not reconnect on its own, later ending in an assertion failure/engine panic. Art Kagel suggested re-issuing 'onmode -d RSS <primary>' on the secondary to kick the connection, though others reported that for RSS only restarting the secondary, dropping/re-adding the RSS definition, or restarting the primary works. IBM support pointed to known APARs (notably IC85187, where the primary logs a ping timeout but doesn't properly turn HDR off, plus IC89002, IC95683, IC93948), so upgrading to a fixed fixpack was the suggested remedy; no single confirmed fix for the original poster is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: General Discussion
Hi, We are running Informix 11.70.FC7 and have one RSS server connected to the primary. There was a networking issue between the primary and RSS for about a half hour. During this time the RSS was up and healthy though it just couldn't communicate with the primary. How do I "wakeup" the communication so that RSS continues talking to the primary without bouncing the RSS node? It didn't pick up automatically. Here is the last message log from the RSS node. 10:43:50 SMX thread is exiting because the timeout period of 10 seconds has elapsed. Use the IFX_SMX_TIMEOUT environment variable to set the timeout period. 10:43:52 RSS: Lost connection to test1 Here is the last message from the primary about the RSS node: 10:43:51 SMX thread is exiting because the timeout period of 10 seconds has elapsed. Use the IFX_SMX_TIMEOUT environment variable to set the timeout period. 10:43:52 Error receiving a buffer from RSS testrss - shutting down 10:43:55 RSS Server testrss - state is now disconnected The primary still thinks the secondary is down: Local server type: Primary Index page logging status: Enabled Index page logging was enabled at: 2015/01/08 17:06:41 Number of RSS servers: 1 RSS Server information: RSS Server control block: 0x0 RSS server name: testrss RSS server status: Inactive RSS connection status: Disconnected --f46d043c7e14d6d09c050e5a3fdc
Try running the onmode -d RSS <primary> again on the secondary.
Art
Art S. Kagel, President and Principal Consultant
ASK Database Management
www.askdbmgt.com
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on the IIUG, nor any other organization with which I am
associated either explicitly, implicitly, or by inference. Neither do
those opinions reflect those of other individuals affiliated with any
entity with which I am affiliated nor those of the entities themselves.
On Thu, Feb 5, 2015 at 6:01 PM, Informix DBA <in4mixdba@gmail.com> wrote:
> Hi,
>
> We are running Informix 11.70.FC7 and have one RSS server connected to the
> primary. There was a networking issue between the primary and RSS for
> about a half hour. During this time the RSS was up and healthy though it
> just couldn't communicate with the primary. How do I "wakeup" the
> communication so that RSS continues talking to the primary without bouncing
> the RSS node? It didn't pick up automatically.
>
> Here is the last message log from the RSS node.
> 10:43:50 SMX thread is exiting because the timeout period of 10 seconds
> has elapsed. Use
> the IFX_SMX_TIMEOUT environment variable to set the timeout period.
> 10:43:52 RSS: Lost connection to test1
>
> Here is the last message from the primary about the RSS node:
> 10:43:51 SMX thread is exiting because the timeout period of 10 seconds
> has elapsed. Use
> the IFX_SMX_TIMEOUT environment variable to set the timeout period.
> 10:43:52 Error receiving a buffer from RSS testrss - shutting down
> 10:43:55 RSS Server testrss - state is now disconnected
>
> The primary still thinks the secondary is down:
> Local server type: Primary
> Index page logging status: Enabled
> Index page logging was enabled at: 2015/01/08 17:06:41
> Number of RSS servers: 1
>
> RSS Server information:
>
> RSS Server control block: 0x0
> RSS server name: testrss
> RSS server status: Inactive
> RSS connection status: Disconnected
>
> --f46d043c7e14d6d09c050e5a3fdc
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--089e0158b7e47b1bb7050e5aa036
Thank you Art. Too late to try it as the Informix Engine crashed while I
was waiting. I will try it next time if it happens again.
10:43:50 SMX thread is exiting because the timeout period of 10 seconds
has elapsed. Use
the IFX_SMX_TIMEOUT environment variable to set the timeout period.
10:43:52 RSS: Lost connection to test1
12:10:04 stack trace for pid 3744 written to /informix_dump/af.497a3eb
12:10:04 Assert Failed: No Exception Handler
12:10:04 IBM Informix Dynamic Server Version 11.70.FC7W2
12:10:04 Who: Session(32, informix@ifxtest, 0, 24bb1c040)
Thread(175, RSS_Recv, 24bad69b8, 12)
File: mtex.c Line: 490
12:10:04 Results: Exception Caught. Type: MT_EX_OS, Context: mem
12:10:04 Action: Please notify IBM Informix Technical Support.
12:10:04 See Also: /informix_dump/af.497a3eb, shmem.497a3eb.0
12:11:55 mtex.c, line 490, thread 175, proc id 3744, No Exception Handler.
12:11:55 Fatal error in ADM VP at mt.c:14045
12:11:55 Unexpected virtual processor termination, pid = 3744, exit = 0x100
12:11:55 PANIC: Attempting to bring system down
--Dave
On Thu, Feb 5, 2015 at 12:28 PM, Art Kagel <art.kagel@gmail.com> wrote:
> Try running the onmode -d RSS <primary> again on the secondary.
>
> Art
>
> Art S. Kagel, President and Principal Consultant
> ASK Database Management
> www.askdbmgt.com
>
> Blog: http://informix-myview.blogspot.com/
>
> Disclaimer: Please keep in mind that my own opinions are my own opinions
> and do not reflect on the IIUG, nor any other organization with which I am
> associated either explicitly, implicitly, or by inference. Neither do
> those opinions reflect those of other individuals affiliated with any
> entity with which I am affiliated nor those of the entities themselves.
>
> On Thu, Feb 5, 2015 at 6:01 PM, Informix DBA <in4mixdba@gmail.com> wrote:
>
> > Hi,
> >
> > We are running Informix 11.70.FC7 and have one RSS server connected to
> the
> > primary. There was a networking issue between the primary and RSS for
> > about a half hour. During this time the RSS was up and healthy though it
> > just couldn't communicate with the primary. How do I "wakeup" the
> > communication so that RSS continues talking to the primary without
> bouncing
> > the RSS node? It didn't pick up automatically.
> >
> > Here is the last message log from the RSS node.
> > 10:43:50 SMX thread is exiting because the timeout period of 10 seconds
> > has elapsed. Use
> > the IFX_SMX_TIMEOUT environment variable to set the timeout period.
> > 10:43:52 RSS: Lost connection to test1
> >
> > Here is the last message from the primary about the RSS node:
> > 10:43:51 SMX thread is exiting because the timeout period of 10 seconds
> > has elapsed. Use
> > the IFX_SMX_TIMEOUT environment variable to set the timeout period.
> > 10:43:52 Error receiving a buffer from RSS testrss - shutting down
> > 10:43:55 RSS Server testrss - state is now disconnected
> >
> > The primary still thinks the secondary is down:
> > Local server type: Primary
> > Index page logging status: Enabled
> > Index page logging was enabled at: 2015/01/08 17:06:41
> > Number of RSS servers: 1
> >
> > RSS Server information:
> >
> > RSS Server control block: 0x0
> > RSS server name: testrss
> > RSS server status: Inactive
> > RSS connection status: Disconnected
> >
> > --f46d043c7e14d6d09c050e5a3fdc
> >
> >
> >
> >
>
>
*******************************************************************************
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
> --089e0158b7e47b1bb7050e5aa036
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--047d7b5d27986d4f55050e5b69cc
We get this a lot although our experience is we don't get an AF - the the thread exiting. Any network outage for an extended period of time leaves the RSS in an un-connected state. If we don't catch it early enough the logs wrap. And we have about 40 sites running RSS servers so this is something that causes us pain. Is this just an accepted behaviour or is there something we can 'tune' to avoid this?
The connection should have re-established itself. If it did not, then = we need to know the stacks of all of the threads on the secondary. What i= s the version of the server? Has tech support been contacted? From: "STUART STEPHENS" <stuart.stephens@monitorsoft.com> To: ids@iiug.org Date: 02/09/2015 03:19 AM Subject: Re: Wakeup RSS Server [34578] Sent by: ids-bounces@iiug.org We get this a lot although our experience is we don't get an AF - the t= he thread exiting. Any network outage for an extended period of time leave= s the RSS in an un-connected state. If we don't catch it early enough the log= s wrap. And we have about 40 sites running RSS servers so this is something tha= t causes us pain. Is this just an accepted behaviour or is there something we can 'tune' = to avoid this? ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
I too get this (not very often, smx thread exiting posted to the msg log ) on a HDR pair 11.70.FC7 due to network issue or latency. About half the time I need to reboot the secondary, the other half the primary too. Each time I opened a PMR and on the last event I was advised to go open up a feature request in order to get HDR functional once the smx thread exits. Needless to say this behavior is a pain point. Categorizing this as a feature request is kind of taking a cavalier attitude towards something that requires an outage to correct. Mark
Original post: I too get this (not very often, smx thread exiting posted to the msg log ) on a HDR pair 11.70.FC7 due to network issue or latency. About half the time I need to reboot the secondary, the other half the primary too. Each time I opened a PMR and on the last event I was advised to go open up a feature request in order to get HDR functional once the smx thread exits. Needless to say this behavior is a pain point. Categorizing this as a feature request is kind of taking a cavalier attitude towards something that requires an outage to correct. Mark Response: Are you using updatable secondaries or not? Well first off, if you are, there has already been at least 1 (or maybe 2) APAR's concerning the co-ordination of smx time-outs and HDR time-outs, additionally I believe there is 1 APAR that should an smx time-out happen but a HDR one doesn't (so no ping timeout), 1 of the threads on the secondary would go and re-start the smx/proxy sub-systems. I don't have APAR numbers off the top of my head. If you are not using UPDATABLE_SECONDARY, I'm not aware of any issue caused by having the smx connection between the primary and secondary happen. The only thread that uses the smx connection in the non-updatable case is the ping thread, and it's only trying to determine if your secondary is doing a checkpoint or not, and I don't recall if that not working that that would cause problems. So I'm curious, what problems are you having in your HDR connection when you have the smx timeouts? Jacques Renaut IBM Informix Advanced Support APD Team
Jacques, It is not an updateable HDR secondary, just RO. On one of the latest occurrences the secondary recorded a ping timeout followed by -28852 which implies a network error. It then followed with DR_ERR set to -1 , DR Turned off on secondary server. The primary just recorded a simple DR: ping timeout and did not make an attempt to re-establish like it normally does. I then bounced the secondary where then the primary recorded DR: Received connection request from remote server when DR is not off. However the secondary could not get past fast recovery to the point where it posts the usual ":DR HDR Secondary server is operational". I had to restart the primary to get HDR operational again. So this latest event did not have the SMX thread is exiting msg as it did in the previous event (about three months ago ). Mark
Original post: Jacques, It is not an updateable HDR secondary, just RO. On one of the latest occurrences the secondary recorded a ping timeout followed by -28852 which implies a network error. It then followed with DR_ERR set to -1 , DR Turned off on secondary server. The primary just recorded a simple DR: ping timeout and did not make an attempt to re-establish like it normally does. I then bounced the secondary where then the primary recorded DR: Received connection request from remote server when DR is not off. However the secondary could not get past fast recovery to the point where it posts the usual ":DR HDR Secondary server is operational". I had to restart the primary to get HDR operational again. So this latest event did not have the SMX thread is exiting msg as it did in the previous event (about three months ago ). Mark Response: Pretty sure what you are describing is a known APAR. Not sure if it is supposed to be fixed in the version you are running or not, but I definitely recall an APAR where a primary can record a ping timeout but not tear down HDR properly, and in that case, it will then reject connection requests from the secondary and report that it doesn't believe had HDR been turned off. It sounds like APAR IC85187. If it's not that, then I think there was another similar APAR that I can't find at the moment...but to be sure obviously I'd need details about different thread states when the problem happened (particularly on the primary since it's the one that seemingly thinks HDR is not turned off even though it reported the ping timeout, which should be what triggers it to turn HDR off). Jacques Renaut IBM Informix Advanced Support APD Team
Jacques, Thanks...good information since the client is on 11.70.FC4 (not FC7 as I previously stated) so this looks to be a good fit, just have to check again to see if smx thread exiting msg was buried somewhere in the msg log. At least we know now an upgrade will probably fix this issue. Mark http://www-01.ibm.com/support/docview.wss?uid=swg1IC85187
Original post: Jacques, Thanks...good information since the client is on 11.70.FC4 (not FC7 as I previously stated) so this looks to be a good fit, just have to check again to see if smx thread exiting msg was buried somewhere in the msg log. At least we know now an upgrade will probably fix this issue. Mark http://www-01.ibm.com/support/docview.wss?uid=swg1IC85187 Response: Mark, I'm not entirely certain that you have to have the smx exit message before the ping time-out message on the primary (item #1 in the apar doc you listed), but if you see items 2 and 3 then I believe that's a better indicator of having hit the problem. Jacques
Stuart,
What version are you running? We had some problems but now work reliably.
Art suggested 'kicking it' by re-running the 'onmode -d' command on the
secondary. I still think of this method that it can often work with HDR, but I
have never seen it do anything useful with RSS. If your RSS is not working in
my experience the only useful support actions are:
- Restart the secondary.
- Delete the RSS definition from both servers and re-add.
- Restart the primary.
Someone has mentioned defect IC85187 further down the thread but I know of
some others, all of which I have seen and raised PMRs about :) Without these
fixes RSS is not reliable. Do you have fixes for all of these?
IC89002 Primary SDS smxrcv thread getting Assert Failed trying to decrypt when
ENCRYPT_SMX = 2 and SMX_COMPRESS = 5
Fixed in 11.70.FC7W1
IC95683 RSS_SEND THREAD MAY BE LEFT IN CONDITION WAIT AFTER RSS ERROR
fixed in 11.70.FC7W1, presumably 11.70.FC8 too.
IC93948 ASSERT FAILED: INVALID MUTEX TYPE POSSIBLE DURING RSS RECONNECTION
ATTEMPT WHEN DELAY_APPLY IS USED
Fixed in 11.70.FC7W3
Ben.