Error re-Starting SDS Secondary server
Posted in 2011
Topics: High Availability & Replication, Error Codes & Troubleshooting, Logging & Checkpoints, Clustering, Grid & MACH11
IDS 11.50 FC8X4
Linux
I have an SDS secondary server that I started up last week. Today I made a
configuration change and stopped/ restarted the server. When it attempted to
restart, I got the following "Logical Log File not found." error.
08:36:02 Logical Log File not found.
08:36:02 oninit: Fatal error in shared memory initialization
08:36:02 IBM Informix Dynamic Server Stopped.
The primary server had this error
08:36:17 SMX thread is exiting because of network error code -25582
The weird thing is that the secondary server was connected and working just a
few minutes earlier.
I have also seen this error on my test servers... both primary and SDS
secondary - I restarted everything, making the primary server standard, in
order to get the server started.
We are using raw spaces, on a SAN, shared/accessed by all the servers in the
SDS cluster.
Any ideas why this is happening?
Thanks
laurie
Laurie Gustin
Database Administrator
Dept of Technology Services
Dept of Public Safety
lgustin@utah.gov
801-965-4410
Are you sure that the logs are on a shared device?
The way that SDS starts up is to perform a checkpoint on the primary and
then use the LSN of the checkpoint as the starting point for the SDS node.
So unless you are wrapping the logs in less than say a second, then I'm a
bit confused as to how the SDS node would have not found the logs.
SDS can be forced to restart if you bounce the primary without
transitioning the role of primary to one of the SDS nodes. The reason is
that when you restart the primary, then the buffer pool will not be
consistent with the primary.
The best practice would be
1) run onmode -d make primary on one of the SDS nodes that you want to keep
active. This should automatically shutdown the old SDS node.
2) make your changes
3) bring up the old primary as an SDS secondary.
4) Transition the rolw of primary back to the first server by running
onmode -d make primary
M.P.
From: "Laurie Gustin" <lgustin@utah.gov>
To: ids@iiug.org
Date: 06/06/2011 08:44 AM
Subject: Error re-Starting SDS Secondary server [23926]
Sent by: ids-bounces@iiug.org
IDS 11.50 FC8X4
Linux
I have an SDS secondary server that I started up last week. Today I made a
configuration change and stopped/ restarted the server. When it attempted
to
restart, I got the following "Logical Log File not found." error.
08:36:02 Logical Log File not found.
08:36:02 oninit: Fatal error in shared memory initialization
08:36:02 IBM Informix Dynamic Server Stopped.
The primary server had this error
08:36:17 SMX thread is exiting because of network error code -25582
The weird thing is that the secondary server was connected and working just
a
few minutes earlier.
I have also seen this error on my test servers... both primary and SDS
secondary - I restarted everything, making the primary server standard, in
order to get the server started.
We are using raw spaces, on a SAN, shared/accessed by all the servers in
the
SDS cluster.
Any ideas why this is happening?
Thanks
laurie
Laurie Gustin
Database Administrator
Dept of Technology Services
Dept of Public Safety
lgustin@utah.gov
801-965-4410
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Logs are on a shared device. The secondary server started fine the first time.
There is alot of activity on the primary server, Im thinking maybe I will wait
until it slows down a bit and try it again. When I had this problem with my
test servers, even the primary server had the Logical Log File not found
error. I had to shut everything down, bring the primary up as a standard
server, then take the steps to restart the secondary.
The change I was making on the secondary was to allow it to update. The
primary is still running fine.
Laurie
>>> "Madison Pruet" <mpruet@us.ibm.com> 6/6/2011 9:15 AM >>>
Are you sure that the logs are on a shared device?
The way that SDS starts up is to perform a checkpoint on the primary and
then use the LSN of the checkpoint as the starting point for the SDS node.
So unless you are wrapping the logs in less than say a second, then I'm a
bit confused as to how the SDS node would have not found the logs.
SDS can be forced to restart if you bounce the primary without
transitioning the role of primary to one of the SDS nodes. The reason is
that when you restart the primary, then the buffer pool will not be
consistent with the primary.
The best practice would be
1) run onmode -d make primary on one of the SDS nodes that you want to keep
active. This should automatically shutdown the old SDS node.
2) make your changes
3) bring up the old primary as an SDS secondary.
4) Transition the rolw of primary back to the first server by running
onmode -d make primary
M.P.
From: "Laurie Gustin" <lgustin@utah.gov>
To: ids@iiug.org
Date: 06/06/2011 08:44 AM
Subject: Error re-Starting SDS Secondary server [23926]
Sent by: ids-bounces@iiug.org
IDS 11.50 FC8X4
Linux
I have an SDS secondary server that I started up last week. Today I made a
configuration change and stopped/ restarted the server. When it attempted
to
restart, I got the following "Logical Log File not found." error.
08:36:02 Logical Log File not found.
08:36:02 oninit: Fatal error in shared memory initialization
08:36:02 IBM Informix Dynamic Server Stopped.
The primary server had this error
08:36:17 SMX thread is exiting because of network error code -25582
The weird thing is that the secondary server was connected and working just
a
few minutes earlier.
I have also seen this error on my test servers... both primary and SDS
secondary - I restarted everything, making the primary server standard, in
order to get the server started.
We are using raw spaces, on a SAN, shared/accessed by all the servers in
the
SDS cluster.
Any ideas why this is happening?
Thanks
laurie
Laurie Gustin
Database Administrator
Dept of Technology Services
Dept of Public Safety
lgustin@utah.gov
801-965-4410
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Related threads
- System Or Internal Error InterruptedIOException
- 25582 error on high volume of short live trans
- ASF Echo-Thread Server: asfcode = -25582 oserr = 4