FW: Assert failure when restarting replication.
Posted in 2010
Topics: High Availability & Replication, Error Codes & Troubleshooting, Server Administration, Logging & Checkpoints
One of my team members is receiving the following. Any thoughts:
Informix: IBM Informix Dynamic Server Version 9.40.FC3
O/S: SunOS srapdsa20 5.8 Generic_117350-27 sun4u sparc SUNW,Netra-T12
Scenario: I do an onmode -d standard on both servers to stop replicaton,
I shut down
the secondary and moved the links to a different filesystem intending to
restore there
but abandoned the idea. Moved the links back to original location and
attempted
to restart replication, get an assert failure on the secondary with
following text:
14:30:27 DR: Failure recovery from disk in progress ...
14:30:28 Logical Recovery Started.
14:30:28 30 recovery worker threads will be started.
14:30:29 Start Logical Recovery - Start Log 239543, End Log ?
14:30:29 Starting Log Position - 239543 0x37de018
14:30:29 Clearing the physical and logical logs has started
14:31:25 ret_errno = 0
14:31:25 dr_debug_err = -5
14:31:25
14:31:25 IBM Informix Dynamic Server Version 9.40.FC3 Software
Serial Number AAA#B000000
14:31:25 Assert Failed: DR: Log Record Apply Thread Exited Abnormally.
Internal Error.
A restart of the database server shall be required to
correct
this problem.
14:31:25 Who: Session(26, informix@, 0, 2abe9b218)
Thread(77, dr_secapply, 2abe6c898, 1)
File: rshdr.c Line: 5518
14:31:25 Results: Dynamic Server must abort
14:31:25 Action: Reinitialize shared memory
14:31:25 Stack for thread: 77 dr_secapply
base: 0x00000002ad2d4000
len: 36864
pc: 0x0000000100872c74
tos: 0x00000002ad2db9c1
state: running
vp: 1
0x100872c74 oninit :: afstack + 0x34 sp=0x2ad2dc1c0(0xad134d68,
0xad134f88, 0xee6080, 0x0, 0xef0768, 0xad2dc750)
0x100871cd8 oninit :: afhandler + 0xaf8 sp=0x2ad2dc490
delta_sp=720(0xcc51e8, 0xacec5d68, 0xad189fd8, 0x601, 0x1, 0xefb448)
0x1008711d0 oninit :: afcrash_interface + 0x4c sp=0x2ad2dcbb0
delta_sp=1824(0xad2d3d78, 0xad189fd8, 0xacec5d68, 0xe01d88, 0x158e,
0xad2dce48)
0x10067243c oninit :: dr_secondary_apply + 0x1544 sp=0x2ad2dcc70
delta_sp=192(0xcc2fc0, 0xad189fd8, 0xad2d3d78, 0x2d7c, 0xffffffff, 0x37)
0x100848570 oninit :: startup + 0xfc sp=0x2ad2dce50
delta_sp=480(0xefb448, 0x0, 0xed0d38, 0xab25ab50, 0xac8f4028,
0xab25aad8)
14:31:25 See Also: /tmp/crash/af.4352f0c, shmem.4352f0c.0
Colin Fellows
Sears Holdings Corporation
B2-335B (847)286-3926
Colin.Fellows@searshc.com
I suspect what happened is you did the onmode -d standard on your secondary,
and then at that point you had 2 servers that were writable. Then, when you
onmode -ky the secondary to bring it down, that makes it write it's own newcheckpoint record which has the possibility of making that old secondary think
it's done new write activity since HDR had been torn down. At this point, your
original primary also probably continued and did new checkpoint activity and
so the old secondary and primary are now out of sync and can't be reconnected
without doing a restore from one of them, as both servers have made
modifications since the last checkpoint that HDR knew about. When you do
onmode -d standard you need to be very careful that you only do write activityon 1 of the servers (being able to recovery from this point depends on only 1
server having made changes since the HDR connection was broken so that the
other server can recover from the the 1 server that did the new work)
otherwise when you attempt to reconnect the 2 servers, both of the servers
think the other server should recover themself, which leads to the HDR pair
being out of sync and having to be restarted via a restore.
Jacques Renaut
IBM Informix Advanced Support
APD Team