Recovering a primary and secindary using SAN techniques (quite long)
Posted in 2008
Topics: High Availability & Replication, Backup & Restore, Storage & Space Management, Stored Procedures & SPL, Server Administration, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
IDS 10.0FC8W2 on HP-UX 11.31:
We've set a up DR system for a customer.
Their system - a Primary database and an HDR pair - run at site A. The SAN
LUNs containing the primary database (but not the temp dbspaces) are in a
consistency group and are replicated syncronously to a SAN at Site B.
The recovery process at Site B, in the event of failure, is thus:
1. Split the syncronisation between Site A and Site B (if a test; of course
ina real disaster it may already be split)
2. Copy the replicated disks at Site B to two sets of BCVs. The DR primary
is pointed at one set of BCVs; the HDR secondary at the other.
3. Start the primary database. It fast-recovers, and off it goes.
4. Use an physical external rcovery on the secondary. When this is complete
use onmode -d on both the primary and secondary to establish HDR.
It works fine. We've done the same at a Solaris site. But there is one
issue recovering the HDR secondary:
We issue the command onbar -r -p -e and get:
Wed Nov 12 16:57:58 2008
16:57:58 Event alarms enabled. ALARMPROG =
'/opt/informix/10.0/etc/alarmprogram.sh'
16:57:58 Booting Language <c> from module <>
16:57:58 Loading Module <CNULL>
16:57:58 Booting Language <builtin> from module <>
16:57:58 Loading Module <BUILTINNULL>
16:58:05 DR: DRAUTO is 0 (Off)
16:58:05 IBM Informix Dynamic Server Version 10.00.FC8W2 Software SerialNumber AAA#B000000
16:58:07 IBM Informix Dynamic Server Initialized -- Shared MemoryInitialized.
16:58:07 DR: Reservation of the last logical log for log backup turned on
16:58:07 Data replication type and state information reset. To start DR,use
the 'onmode -d' command and wait for the pair to be operational,
before shutting down the database server
16:58:07 Dataskip is now OFF for all dbspaces
16:58:07 Restartable Restore has been ENABLED
16:58:07 Recovery Mode
16:58:08 The external backup for the root chunk (chunk number 1) is not
valid.
16:58:09 IBM Informix Dynamic Server Stopped.
When we try it with ontape (ontape -r -p -e) we also got the "The external
backup for the root chunk (chunk number 1) is not valid." message, but this
time Informix stayed up, and we were able to run the recovery through to
successful completion.
Wed Nov 12 17:25:05 2008
17:25:05 Event alarms enabled. ALARMPROG =
'/opt/informix/10.0/etc/alarmprogram.sh'
17:25:05 Booting Language <c> from module <>
17:25:05 Loading Module <CNULL>
17:25:05 Booting Language <builtin> from module <>
17:25:05 Loading Module <BUILTINNULL>
17:25:12 DR: DRAUTO is 0 (Off)
17:25:12 IBM Informix Dynamic Server Version 10.00.FC8W2 Software SerialNumber AAA#B000000
17:25:14 IBM Informix Dynamic Server Initialized -- Shared MemoryInitialized.
17:25:14 DR: Reservation of the last logical log for log backup turned on
17:25:14 Data replication type and state information reset. To start DR,use
the 'onmode -d' command and wait for the pair to be operational,
before shutting down the database server
17:25:14 Dataskip is now OFF for all dbspaces
17:25:14 Restartable Restore has been ENABLED
17:25:14 Recovery Mode
So why does onbar fail, and ontape work but give the same failure message?
IBM Tech Support tells us, under PMR 57109,019,866, that the external rstore
fails because:
"The chunk 1 is not valid is telling us that although there is no
involvement of the root chunk at all but the reason for the
error message is that when an onmode -c block/unblock is used to
obtain an external backup, it blocks the database server and the
system takes a checkpoint and suspends all update transactions.
In this situation for every dbspace first chunk if you do an
onmode -c block / onmode -c unblock some recovery information iswritten in page 0 and if that page is overwritten that
information is invalid and the engine complains. "
... which is of course true - there is no onmode -c block because these
disks have been unceremoniously split from their masters either by a
disaster or by a command when DR testing.
But, it works, and I'm not sure that there's a reason why it shouldn't. My
concern is that the ontape give-an-error-message-but-carry-on anyway
loophole may be tightened in later releases, as per onbar.
I suppose it doesn't matter that much; even if we can't recover the HDR
secondary in this way because of future restriction, we can re-establish it
with an ontape archive piped into an ontape restore. It just delays the
recovery time of the HDR instance.
Any thoughts?
rgds
Neil
Neil Truby wrote:
>
> But, it works, and I'm not sure that there's a reason why it shouldn't. My
> concern is that the ontape give-an-error-message-but-carry-on anyway
> loophole may be tightened in later releases, as per onbar.
>
> I suppose it doesn't matter that much; even if we can't recover the HDR
> secondary in this way because of future restriction, we can re-establish it
> with an ontape archive piped into an ontape restore. It just delays the
> recovery time of the HDR instance.
>
> Any thoughts?
Don't depend on it. I'm sure that the loophole will be tightened up.
--
Cheers,
Obnoxio The Clown
http://obotheclown.blogspot.com
"Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message
news:mailman.354.1227049464.874.informix-list@iiug.org...
> Neil Truby wrote:
>>
>> But, it works, and I'm not sure that there's a reason why it shouldn't.
>> My concern is that the ontape give-an-error-message-but-carry-on anyway
>> loophole may be tightened in later releases, as per onbar.
>>
>> I suppose it doesn't matter that much; even if we can't recover the HDR
>> secondary in this way because of future restriction, we can re-establish
>> it with an ontape archive piped into an ontape restore. It just delays
>> the recovery time of the HDR instance.
>>
>> Any thoughts?
>
> Don't depend on it. I'm sure that the loophole will be tightened up.
Yes, that's what I thought. It's a very useful loophole though!
Hello Neil,
oninit -r
may be the one you look for; may also be a closed loophole oneday????
Superboer