HDR secondary crashed
Posted in 2012
Topics: High Availability & Replication, Error Codes & Troubleshooting, Server Administration, Platform-Specific Issues, Java & JDBC Development, Cloud, Docker & Containers, Versions, Editions & End-of-Life
Hi all, we yesterday had a strange situation with a HDR pair running 11.70FC4GE (Linux 64). The primary machine died because of a kernel trap, and the HDR secondary was turned to be standard to be adressable by our XA Java app. Immediately the instance was crashing because of a failure I would say in logfile recovery. Maybe the died instance was in the middle of sending a log to the secondary (my assume only). I opened a case with our reseller, but would like to hear if anybody else was in such a situation We were forced to restore the (relatively small, but important) instance. And that should be the thing to prevent with HDR. (Yes, we should search for the Linux problem, but the IDS problem is more severe ...) Anybody seen something like that before ? We have some HDR setups running, most with 11.10FC3, and are in the process of upgrading to 11.70, but since this was happening, I am forced to not trust the release. 20:57:08 log_get: get_logfile_by_id() failed 20:57:08 log_get: cannot read loguniq 15867 logpos 0x5044 20:57:08 logm_read: cannot read loguniq 15867 logpos 0x5044 20:57:08 logerr('relock() - logread failed') 20:57:08 tx 0x6acd60e8, tx_flags 0x8462b 20:57:08 tx_loguniq 15867, tx_logpos 0x5044 20:57:08 20:57:08 IBM Informix Dynamic Server Version 11.70.FC4GE Software Serial Number AAA#B000000 20:57:08 Assert Failed: Dynamic Server must abort 20:57:08 Who: Session(6, informix@keeper06, 22995, 0x6acce8f8) Thread(28, onmode_mon, 6ac91538, 1) File: rslog.c Line: 3661 20:57:08 Results: Dynamic Server must abort 20:57:08 Action: Reinitialize shared memory 20:57:08 Raw hex dump of stack located in /opt/informix/tmp/af.404b084.rawstk 20:57:08 Stack for thread: 28 onmode_mon base: 0x000000006cbe3000 len: 69632 pc: 0x000000000129aa26 tos: 0x000000006cbf1260 state: running vp: 1 Thanks for your thoughts, Marcus Haarmann
Hi all,
found the reason:
There was a lost XA transaction in the system, which was not finished.
Secondary wanted to get the logfile for this transaction from primary and
refused to go Online.
We were able to kill the old transaction (which was not locking records and
thus was not found
by our checks) with onmode -H.
Marcus
----- Ursprüngliche Mail -----
Von: "Marcus Haarmann" <marcus.haarmann@midoco.de>
An: ids@iiug.org
Gesendet: Freitag, 30. März 2012 12:23:41
Betreff: HDR secondary crashed [26616]
Hi all,
we yesterday had a strange situation with a HDR pair running 11.70FC4GE (Linux
64).
The primary machine died because of a kernel trap, and the HDR secondary was
turned
to be standard to be adressable by our XA Java app.
Immediately the instance was crashing because of a failure I would say in
logfile recovery.
Maybe the died instance was in the middle of sending a log to the secondary
(my assume only).
I opened a case with our reseller, but would like to hear if anybody else was
in such a situation
We were forced to restore the (relatively small, but important) instance. And
that should be the
thing to prevent with HDR.
(Yes, we should search for the Linux problem, but the IDS problem is more
severe ...)
Anybody seen something like that before ?
We have some HDR setups running, most with 11.10FC3, and are in the process of
upgrading to
11.70, but since this was happening, I am forced to not trust the release.
20:57:08 log_get: get_logfile_by_id() failed
20:57:08 log_get: cannot read loguniq 15867 logpos 0x5044
20:57:08 logm_read: cannot read loguniq 15867 logpos 0x5044
20:57:08 logerr('relock() - logread failed')
20:57:08 tx 0x6acd60e8, tx_flags 0x8462b
20:57:08 tx_loguniq 15867, tx_logpos 0x5044
20:57:08
20:57:08 IBM Informix Dynamic Server Version 11.70.FC4GE Software Serial
Number AAA#B000000
20:57:08 Assert Failed: Dynamic Server must abort
20:57:08 Who: Session(6, informix@keeper06, 22995, 0x6acce8f8)
Thread(28, onmode_mon, 6ac91538, 1)
File: rslog.c Line: 3661
20:57:08 Results: Dynamic Server must abort
20:57:08 Action: Reinitialize shared memory
20:57:08 Raw hex dump of stack located in /opt/informix/tmp/af.404b084.rawstk
20:57:08 Stack for thread: 28 onmode_mon
base: 0x000000006cbe3000
len: 69632
pc: 0x000000000129aa26
tos: 0x000000006cbf1260
state: running
vp: 1
Thanks for your thoughts,
Marcus Haarmann
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.