Re: HDR unusable with 9.40
Posted in 2004
Hi Alexey,
this is another silly problem that I found in 9.40.UC2 and reproduces
in 9.30.FC5.
I opened a tech support case with UK IBM, and after more than 10 days
waiting,( I sent them a test case ), they answered that they
reproduced the problem under 9.30.UC2 but the didn't reproduce the
problem under 9.30.UC3 so they refused to open a new bug and they
didn't give me a bug number.
This only happens if there is no activity on the primary server and
also see my workaround below.
I didn't have time to test it under 9.40.UC3 or later
hope this helps
You will see all the info I sent to IBM below:
New secondary crashes after fail over procedure
Testing HDR fail over and following the steps from Informix admin
manual new secondary crashes a few seconds after it changes its state
to secondary server. If you try to bring it back up using oninit it
will crash again with the same error.
I found a workaround for this problem, ( see below ), but if you don't
want to use the workaround the only way to re-create hdr is applying a
new physical restore.
I tested the same procedure under 9.30.FC5 Solaris 8 and 9.40.UC2
Linux redhat 9 ( kernel 2.4.20-8 )
The procedure I followed was:
1. Take primary instance (prdhdr) off line. Go first to
quiescent mode and then offline.
2. Execute hdrmkpri.sh prhdr on secondary server (secondary).
3. Execute hdrmksec.sh sechdr on primary server ( primary )
4. Bring new primary server backup using onint -v
which is the same than
Instance A (currently Primary) Instance B (currently
Secondary)
------------------------------
--------------------------------
1] onmode -ky (server should be up)
2] hdrmkpri.sh
<primary_server_name>
3] hdrmksec.sh <secondary_server_name>
(now a Secondary server) 4] oninit (now a Primary
server)
Secondary Server online message log
09:04:01 Loading Module <BUILTINNULL>
09:04:06 Dynamically allocated new virtual shared memory segment
(size 8192KB)
09:04:06 IBM Informix Dynamic Server Version 9.40.UC2 SoftwareSerial Number AAA#B000000
09:04:06 IBM Informix Dynamic Server Initialized -- Shared MemoryInitialized.
09:04:06 DR: Reservation of the last logical log for log backupturned on
09:04:06 Data replication type and state information reset. To start
DR, use the 'onmode -d' command and wait for the pair to be
operational,
before shutting down the database server
09:04:06 Physical Recovery Started at Page (1:1784).
09:04:06 Physical Recovery Complete: 0 Pages Examined, 0 Pages
Restored.
09:04:06 Dataskip is now OFF for all dbspaces
09:04:06 Restartable Restore has been ENABLED
09:04:06 Recovery Mode
09:04:08 DR: Reservation of the last logical log for log backupturned off
09:04:08 DR: new type = secondary, primary server name = net940
09:04:08 DR: Trying to connect to primary server = net940
09:04:08 DR: Cannot connect to primary server
09:04:08 DR: Turned off on secondary server
[informix@appst archive]$ sh: line 1:
/usr/informix/etc/alarmprogram.sh: No such file or directory
sh: line 1: /usr/informix/etc/alarmprogram.sh: No such file or
directory
[informix@appst archive]$ onstat -m
shared memory not initialized for INFORMIXSERVER 'shmsec940'
Message Log File: /informix/940/online.log
Thread(19, dr_secapply, 4620d478, 1)
File: rshdr.c Line: 5497
09:05:01 Results: Dynamic Server must abort
09:05:01 Action: Reinitialize shared memory
09:05:01 stack trace for pid 4789 written to /tmp/af.3fb8b2d
09:05:01 See Also: /tmp/af.3fb8b2d, shmem.3fb8b2d.0
09:05:01 Process exited with return code 127: /bin/sh /bin/sh -c/usr/informix/etc/alarmprogram.sh 3 15 "Data Replication failure."
"DR: Turned off on secondary server
09:05:05 rshdr.c, line 5497, thread 19, proc id 4789, DR: Log RecordApply Thread Exited Abnormally. Internal Error.
A restart of the database server shall be required to
correct
this problem.
.
09:05:05 invoke_alarm(): /bin/sh -c'/usr/informix/etc/alarmprogram.sh 5 6 "Internal Subsystem failure:
'MT'" "rshdr.c, line 5497, thread 19, proc id 4789, DR: Log Record
Apply Thread Exited Abnormally. Internal Error.
A restart of the database server shall be required to
correct
this problem.
." '
09:05:05 invoke_alarm(): mt_exec failed, status 32512, errno 0
09:05:05 The Master Daemon Died
09:05:05 invoke_alarm(): /bin/sh -c'/usr/informix/etc/alarmprogram.sh 5 6 "Internal Subsystem failure:
'MT'" "The Master Daemon Died" '
09:05:05 invoke_alarm(): mt_exec failed, status 32512, errno 0
09:05:05 PANIC: Attempting to bring system down
[informix@appst archive]$
Primary server online message log
Message Log File: /informix/940/online.log
09:05:02 DR: Failure recovery error (2)
09:05:02 Process exited with return code 127: /bin/sh /bin/sh -c/usr/informix/etc/alarmprogram.sh 3 15 "Data Replication failure."
"DR: Local and Remote server type a
09:05:04 Physical Recovery Started at Page (1:1784).
09:05:04 Physical Recovery Complete: 0 Pages Examined, 0 Pages
Restored.
09:05:04 DR: Turned off on primary server
09:05:04 Logical Recovery Started.
09:05:04 10 recovery worker threads will be started.
09:05:04 Process exited with return code 127: /bin/sh /bin/sh -c/usr/informix/etc/alarmprogram.sh 3 15 "Data Replication failure."
"DR: Turned off on primary server"
09:05:04 DR: Cannot connect to secondary server
09:05:04 Process exited with return code 127: /bin/sh /bin/sh -c/usr/informix/etc/alarmprogram.sh 3 15 "Data Replication failure."
"DR: Cannot connect to secondary se
09:05:08 Logical Recovery has reached the transaction cleanup phase.
09:05:08 Logical Recovery Complete. 0 Committed, 0 Rolled Back, 0 Open, 0 Bad Locks
09:05:09 Dataskip is now OFF for all dbspaces
09:05:09 Checkpoint Completed: duration was 0 seconds.
09:05:09 Checkpoint loguniq 155, logpos 0x9018, timestamp: 10928966
09:05:09 Maximum server connections 0
09:05:09 On-Line Mode
[informix@st archive]$ onstat -g dri
IBM Informix Dynamic Server Version 9.40.UC2 -- On-Line (Prim) --
Up 00:00:26 -- 49744 Kbytes
Data Replication:
Type State Paired server Last DR CKPT (id/pg)
primary off netsec940 155 / 7
DRINTERVAL 30
DRTIMEOUT 30
DRLOSTFOUND /usr/informix/etc/dr.lostfound
By the way I found a workaround for this, if you follow this procedure
secondary server and HDR will stay up and running.
1. Take primary instance (prdhdr) off line. Go first to quiescent
mode and then offline.
2. Execute hdrmkpri.sh prhdr on secondary server (secondary).
3. Bring the instance back up on secondary using oninit -v.
4. Execute onmode -l on sechdr to switch to the next logical log file.
5. Execute onmode -c on sechdr to write a new checkpo