HDR Logical Recovery Problems
Posted in 2003
Topics: High Availability & Replication, Backup & Restore, Storage & Space Management, Logging & Checkpoints, Versions, Editions & End-of-Life
Hi!
We are running HDR in a testenvironment with IDS 9.30FC3X3 on HP Tru64 5.1B.
We got errors from HDR:
teszi06 ( primary ):
04:17:30 DR: ping timeout
04:17:30 DR: Receive error
04:17:30 Checkpoint Completed: duration was 66 seconds.
04:17:30 Checkpoint loguniq 4890, logpos 0x51d018
04:17:30 Maximum server connections 21
04:17:32 DR: Turned off on primary server
04:17:32 DR: Cannot connect to secondary server
04:17:42 DR: Primary server connected
04:17:42 DR: Receive error
teszr06 ( secondary ):
04:06:23 Maximum server connections 20
04:11:22 Checkpoint Completed: duration was 0 seconds.
04:11:22 Checkpoint loguniq 4889, logpos 0xf9b018
04:11:22 Maximum server connections 20
04:17:18 DR: ping timeout
04:17:42 DR: Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
04:17:54 DR: Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
04:18:06 DR: Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
04:18:18 DR: Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
Can anyone explain what this error messages mean?
I decided to restart the secondary instance.
The primary server got connected.
status on the secondary:
09:10:21 Physical Recovery Complete: 1406 Pages Examined 1406 Pages Restored.
09:10:22 Dataskip is now OFF for all dbspaces
09:10:22 Restartable Restore has been DISABLED
09:10:22 Recovery Mode
09:10:23 DR: Failure recovery from disk in progress ...
09:10:25 Logical Recovery Started.
09:10:25 10 recovery worker threads will be started.
09:10:25 Start Logical Recovery - Start Log 4889, End Log ?
09:10:25 Starting Log Position - 4889 0xf9b018
tevbis01.health.local / root:
tevbis01.health.local / root:
tevbis01.health.local / root:onstat -
Informix Dynamic Server Version 9.30.FC3X3 -- Fast Recovery (Sec) -- Up
00:12:33 -- 524288 Kbytes
Blocked:CKPT
tevbis01.health.local / root:onstat -g dri
Informix Dynamic Server Version 9.30.FC3X3 -- Fast Recovery (Sec) -- Up
00:12:43 -- 524288 Kbytes
Blocked:CKPT
Data Replication:
Type State Paired server Last DR CKPT (id/pg)
secondary off teszi06 4889 / 3995
DRINTERVAL 30
DRTIMEOUT 30
DRLOSTFOUND /usr/informix/teszi06/etc/dr.lostfound
In this state the secondary hangs.
I've analysed the situation.
The next logical log on the secondary is not backed up!
20d16a248 37 U-B---- 4886 4284bd 5000 234 4.68
20d16a298 38 U-B---- 4887 429845 5000 5000 100.00
20d16a2e8 39 U-B---- 4888 42abcd 5000 5000 100.00
20d16a338 40 U---C-L 4889 42bf55 5000 5000 100.00
20d16a388 41 U------ 4841 42d2dd 5000 4583 91.66
20d16a3d8 42 U-B---- 4842 42e665 5000 609 12.18
20d16a428 43 U-B---- 4843 42f9ed 5000 5000 100.00
On the primary:
onstat -l
.....
20d16a2e8 39 U-B---- 4888 42abcd 5000 5000 100.00
20d16a338 40 U-B---- 4889 42bf55 5000 5000 100.00
20d16a388 41 U-B---- 4890 42d2dd 5000 5000 100.00
20d16a3d8 42 U---C-L 4891 42e665 5000 4018 80.36
20d16a428 43 U-B---- 4843 42f9ed 5000 5000 100.00
This situation came from a test when the secondary was a standard server, and
then the primary came back online, the secondary became again secondary.
How can I free the log.log 4841 on the secondary?
Or how can I finish logical recoery without restore a level 0 backup?
Thanks & regards
Wolfgang
Wolfgang:
The problem is that while the secondary was in standalone mode it wrote to that
logical log and it was never backed up. The best option would be to take a
level 0 archive of the primary, shutdown the secondary and perform a physical
restore of the archive, then reestablish it as secondary. If this had been a
normal primary failed, secondary takes over situation you would have restarted
the downed primary in secondary mode to let it catch up then reversed the
relationship after the synchronization was complete. Since this was apparently
a test and you did not want the maintain the changes made on the secondary
while
it was in standalone mode, you should have performed the restore initially.
Art S. Kagel
----- Original Message -----
From: Wolfgang Neumar <wolfgang.neumar@hp.com>
At: 11/28 6:00
> Hi!
>
> We are running HDR in a testenvironment with IDS 9.30FC3X3 on HP Tru64 5.1B.
> We got errors from HDR:
>
> teszi06 ( primary ):
>
> 04:17:30 DR: ping timeout
> 04:17:30 DR: Receive error
> 04:17:30 Checkpoint Completed: duration was 66 seconds.
> 04:17:30 Checkpoint loguniq 4890, logpos 0x51d018
>
> 04:17:30 Maximum server connections 21
> 04:17:32 DR: Turned off on primary server
> 04:17:32 DR: Cannot connect to secondary server
> 04:17:42 DR: Primary server connected
> 04:17:42 DR: Receive error
>
> teszr06 ( secondary ):
>
> 04:06:23 Maximum server connections 20
> 04:11:22 Checkpoint Completed: duration was 0 seconds.
> 04:11:22 Checkpoint loguniq 4889, logpos 0xf9b018
>
> 04:11:22 Maximum server connections 20
> 04:17:18 DR: ping timeout
> 04:17:42 DR: Received connection request from remote server when DR is not
Off
> [Local type: Secondary, Current state: FAILED]
> [Remote type: Primary]
>
> 04:17:54 DR: Received connection request from remote server when DR is not
Off
> [Local type: Secondary, Current state: FAILED]
> [Remote type: Primary]
>
> 04:18:06 DR: Received connection request from remote server when DR is not
Off
> [Local type: Secondary, Current state: FAILED]
> [Remote type: Primary]
>
> 04:18:18 DR: Received connection request from remote server when DR is not
Off
> [Local type: Secondary, Current state: FAILED]
> [Remote type: Primary]
>
>
> Can anyone explain what this error messages mean?
>
> I decided to restart the secondary instance.
>
> The primary server got connected.
>
> status on the secondary:
>
> 09:10:21 Physical Recovery Complete: 1406 Pages Examined 1406 Pages Restored.
>
> 09:10:22 Dataskip is now OFF for all dbspaces
> 09:10:22 Restartable Restore has been DISABLED
> 09:10:22 Recovery Mode
> 09:10:23 DR: Failure recovery from disk in progress ...
> 09:10:25 Logical Recovery Started.
> 09:10:25 10 recovery worker threads will be started.
> 09:10:25 Start Logical Recovery - Start Log 4889, End Log ?
> 09:10:25 Starting Log Position - 4889 0xf9b018
> tevbis01.health.local / root:
> tevbis01.health.local / root:
> tevbis01.health.local / root:onstat -
>
> Informix Dynamic Server Version 9.30.FC3X3 -- Fast Recovery (Sec) -- Up
> 00:12:33 -- 524288 Kbytes
> Blocked:CKPT
>
> tevbis01.health.local / root:onstat -g dri
>
> Informix Dynamic Server Version 9.30.FC3X3 -- Fast Recovery (Sec) -- Up
> 00:12:43 -- 524288 Kbytes
> Blocked:CKPT
>
> Data Replication:
> Type State Paired server Last DR CKPT (id/pg)
> secondary off teszi06 4889 / 3995
>
> DRINTERVAL 30
> DRTIMEOUT 30
> DRLOSTFOUND /usr/informix/teszi06/etc/dr.lostfound>
> In this state the secondary hangs.
>
> I've analysed the situation.
> The next logical log on the secondary is not backed up!
>
> 20d16a248 37 U-B---- 4886 4284bd 5000 234
4.68
> 20d16a298 38 U-B---- 4887 429845 5000 5000
100.00
> 20d16a2e8 39 U-B---- 4888 42abcd 5000 5000
100.00
> 20d16a338 40 U---C-L 4889 42bf55 5000 5000
100.00
> 20d16a388 41 U------ 4841 42d2dd 5000 4583
91.66
> 20d16a3d8 42 U-B---- 4842 42e665 5000 609
12.18
> 20d16a428 43 U-B---- 4843 42f9ed 5000 5000
100.00
>
>
>
> On the primary:
>
> onstat -l
> ....
> 20d16a2e8 39 U-B---- 4888 42abcd 5000 5000
100.00
> 20d16a338 40 U-B---- 4889 42bf55 5000 5000
100.00
> 20d16a388 41 U-B---- 4890 42d2dd 5000 5000
100.00> 20d16a3d8 42 U---C-L 4891 42e665 5000 4018
80.36
> 20d16a428 43 U-B---- 4843 42f9ed 5000 5000
100.00
>
> This situation came from a test when the secondary was a standard server, and
> then the primary came back online, the secondary became again secondary.
>
> How can I free the log.log 4841 on the secondary?
> Or how can I finish logical recoery without restore a level 0 backup?
>
>
>
> Thanks & regards
> Wolfgang
>
>
Please
read documentation again, when a secondary is switched to standard, you have
to restore it with a backup from primary!
only when primary switched to standard and afterwards back to primary and the
secondary is secondary all the time and all needed logocal logs are available
on primary the data will be transferd to secondary. otherwise you have to
setup up a new hdr!
cu sven
>>This situation came from a test when the secondary was a standard server,
and then the primary came back online, the secondary became again secondary.
-----Ursprüngliche Nachricht-----
Von: WOLFGANG NEUMAR [mailto:wolfgang.neumar@hp.com]
Gesendet: Freitag, 28. November 2003 11:15
An: ids@iiug.org
Betreff: HDR Logical Recovery Problems [2267]
Hi!
We are running HDR in a testenvironment with IDS 9.30FC3X3 on HP Tru64 5.1B.
We got errors from HDR:
teszi06 ( primary ):
04:17:30 DR: ping timeout
04:17:30 DR: Receive error
04:17:30 Checkpoint Completed: duration was 66 seconds.
04:17:30 Checkpoint loguniq 4890, logpos 0x51d018
04:17:30 Maximum server connections 21
04:17:32 DR: Turned off on primary server
04:17:32 DR: Cannot connect to secondary server
04:17:42 DR: Primary server connected
04:17:42 DR: Receive error
teszr06 ( secondary ):
04:06:23 Maximum server connections 20
04:11:22 Checkpoint Completed: duration was 0 seconds.
04:11:22 Checkpoint loguniq 4889, logpos 0xf9b018
04:11:22 Maximum server connections 20
04:17:18 DR: ping timeout
04:17:42 DR: Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
04:17:54 DR: Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
04:18:06 DR: Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
04:18:18 DR: Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
Can anyone explain what this error messages mean?
I decided to restart the secondary instance.
The primary server got connected.
status on the secondary:
09:10:21 Physical Recovery Complete: 1406 Pages Examined 1406 Pages Restored.
09:10:22 Dataskip is now OFF for all dbspaces
09:10:22 Restartable Restore has been DISABLED
09:10:22 Recovery Mode
09:10:23 DR: Failure recovery from disk in progress ...
09:10:25 Logical Recovery Started.
09:10:25 10 recovery worker threads will be started.
09:10:25 Start Logical Recovery - Start Log 4889, End Log ?
09:10:25 Starting Log Position - 4889 0xf9b018
tevbis01.health.local / root:
tevbis01.health.local / root:
tevbis01.health.local / root:onstat -
Informix Dynamic Server Version 9.30.FC3X3 -- Fast Recovery (Sec) -- Up
00:12:33 -- 524288 Kbytes
Blocked:CKPT
tevbis01.health.local / root:onstat -g dri
Informix Dynamic Server Version 9.30.FC3X3 -- Fast Recovery (Sec) -- Up
00:12:43 -- 524288 Kbytes
Blocked:CKPT
Data Replication:
Type State Paired server Last DR CKPT (id/pg)
secondary off teszi06 4889 / 3995
DRINTERVAL 30
DRTIMEOUT 30
DRLOSTFOUND /usr/informix/teszi06/etc/dr.lostfound
In this state the secondary hangs.
I've analysed the situation.
The next logical log on the secondary is not backed up!
20d16a248 37 U-B---- 4886 4284bd 5000 234 4.68
20d16a298 38 U-B---- 4887 429845 5000 5000 100.00
20d16a2e8 39 U-B---- 4888 42abcd 5000 5000 100.00
20d16a338 40 U---C-L 4889 42bf55 5000 5000 100.00
20d16a388 41 U------ 4841 42d2dd 5000 4583 91.66
20d16a3d8 42 U-B---- 4842 42e665 5000 609 12.18
20d16a428 43 U-B---- 4843 42f9ed 5000 5000 100.00
On the primary:
onstat -l
....
20d16a2e8 39 U-B---- 4888 42abcd 5000 5000 100.00
20d16a338 40 U-B---- 4889 42bf55 5000 5000 100.00
20d16a388 41 U-B---- 4890 42d2dd 5000 5000 100.00
20d16a3d8 42 U---C-L 4891 42e665 5000 4018 80.36
20d16a428 43 U-B---- 4843 42f9ed 5000 5000 100.00
This situation came from a test when the secondary was a standard server, and
then the primary came back online, the secondary became again secondary.
How can I free the log.log 4841 on the secondary?
Or how can I finish logical recoery without restore a level 0 backup?
Thanks & regards
Wolfgang
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g