HDR problem with Windows
Posted in 2006
Topics: High Availability & Replication, Backup & Restore, Installation, Setup & Upgrades, Storage & Space Management, Server Administration, Logging & Checkpoints, Versions, Editions & End-of-Life
Hi, I wonder if anyone could help. We're having a problem with HDR with
two identical machines running Windows 2003 and IDS 10.00.TC3. We cannot
get these machines to replicate. I have an open case with OpenPSL for
some weeks who have escalated it to IBM and I am hoping I will get a
solution by that route, but the customer and us are now desperate for a
solution.
Here is the log from the primary server showing a level 0 archive being
taken, this is restored on the secondary and replication initiated.
16:09:24 Level 0 Archive started on rootdbs, maindbs, physdbs, logsdbs,rawdbs, mediadbs, sbspace
16:11:09 Archive on rootdbs, maindbs, physdbs, logsdbs, rawdbs,mediadbs, sbspace Completed.
16:19:45 Checkpoint Completed: duration was 2 seconds.
16:19:45 Checkpoint loguniq 1117, logpos 0xa5e018, timestamp: 0x6f30165c
16:19:45 Maximum server connections 21
16:29:45 Checkpoint Completed: duration was 0 seconds.
16:29:45 Checkpoint loguniq 1117, logpos 0xb35018, timestamp: 0x6f306186
16:29:45 Maximum server connections 21
16:39:45 Checkpoint Completed: duration was 0 seconds.
16:39:45 Checkpoint loguniq 1117, logpos 0xc0a018, timestamp: 0x6f30ac85
16:39:45 Maximum server connections 21
16:49:45 Checkpoint Completed: duration was 0 seconds.
16:49:45 Checkpoint loguniq 1117, logpos 0xcdd018, timestamp: 0x6f30f6e0
16:49:45 Maximum server connections 21
16:58:30 (53) Unable to initiate asynchronous read operation
16:58:30 DR: Cannot connect to secondary server
16:58:30 DR: Turned off on primary server
The secondary's log isn't that interesting. We just get:
16:58:26 DR: Reservation of the last logical log for log backup turned off
16:58:26 DR: new type = secondary, primary server name = ol_server1
16:58:27 DR: Trying to connect to primary server = ol_server1
I think the important part is the "(53) Unable to initiate asynchronous
read operation". We ought to be "experts" at setting up HDR by now as
we've been doing it for years on several platforms for a lot of
customers and have never seen this message. The sequence of events is as
follows:
On primary: run "onmode -d standard", perform level 0 backup with
"ontape -s -L 0".
On secondary: restore tape using "ontape -p", don't restore level 1
archive. Issue command: "onmode -d secondary ol_server1"
On primary: issue command "onmode -d primary ol_server2". It is shortly
after this that the asychronous read operation fails.
Any ideas? I can supply onconfig, "onstat -a", "onstat -g all" and
environment details from registry and ol_server1.cmd file - all of these
are with IBM. At the moment I have nothing to go on and my only option
is to upgrade to 10.00.TC4 which is disruptive for the customer.
Many thanks in advance, Ben.
Hi Ben,
I think the 53 error is being returned from the OS. Windows error 53 is
"The network path was not found". Is as far as you know connectivity
good between the machines (connect via dbaccess from one to the other etc)?
Out of curiosity, what is the root dbspace location on each machine?
Guy
Ben Thompson wrote:
> Hi, I wonder if anyone could help. We're having a problem with HDR with
> two identical machines running Windows 2003 and IDS 10.00.TC3. We cannot
> get these machines to replicate. I have an open case with OpenPSL for
> some weeks who have escalated it to IBM and I am hoping I will get a
> solution by that route, but the customer and us are now desperate for a
> solution.
>
> Here is the log from the primary server showing a level 0 archive being
> taken, this is restored on the secondary and replication initiated.
>
> 16:09:24 Level 0 Archive started on rootdbs, maindbs, physdbs, logsdbs,> rawdbs, mediadbs, sbspace
> 16:11:09 Archive on rootdbs, maindbs, physdbs, logsdbs, rawdbs,> mediadbs, sbspace Completed.
> 16:19:45 Checkpoint Completed: duration was 2 seconds.
> 16:19:45 Checkpoint loguniq 1117, logpos 0xa5e018, timestamp: 0x6f30165c
>
> 16:19:45 Maximum server connections 21
> 16:29:45 Checkpoint Completed: duration was 0 seconds.
> 16:29:45 Checkpoint loguniq 1117, logpos 0xb35018, timestamp: 0x6f306186
>
> 16:29:45 Maximum server connections 21
> 16:39:45 Checkpoint Completed: duration was 0 seconds.
> 16:39:45 Checkpoint loguniq 1117, logpos 0xc0a018, timestamp: 0x6f30ac85
>
> 16:39:45 Maximum server connections 21
> 16:49:45 Checkpoint Completed: duration was 0 seconds.
> 16:49:45 Checkpoint loguniq 1117, logpos 0xcdd018, timestamp: 0x6f30f6e0
>
> 16:49:45 Maximum server connections 21
> 16:58:30 (53) Unable to initiate asynchronous read operation
> 16:58:30 DR: Cannot connect to secondary server
> 16:58:30 DR: Turned off on primary server>
> The secondary's log isn't that interesting. We just get:
>
> 16:58:26 DR: Reservation of the last logical log for log backup turned off
> 16:58:26 DR: new type = secondary, primary server name = ol_server1
> 16:58:27 DR: Trying to connect to primary server = ol_server1>
> I think the important part is the "(53) Unable to initiate asynchronous
> read operation". We ought to be "experts" at setting up HDR by now as
> we've been doing it for years on several platforms for a lot of
> customers and have never seen this message. The sequence of events is as
> follows:
>
> On primary: run "onmode -d standard", perform level 0 backup with
> "ontape -s -L 0".
> On secondary: restore tape using "ontape -p", don't restore level 1
> archive. Issue command: "onmode -d secondary ol_server1"
> On primary: issue command "onmode -d primary ol_server2". It is shortly
> after this that the asychronous read operation fails.
>
> Any ideas? I can supply onconfig, "onstat -a", "onstat -g all" and
> environment details from registry and ol_server1.cmd file - all of these
> are with IBM. At the moment I have nothing to go on and my only option
> is to upgrade to 10.00.TC4 which is disruptive for the customer.
>
> Many thanks in advance, Ben.
Guy Bowerman wrote:
> I think the 53 error is being returned from the OS. Windows error 53 is
> "The network path was not found". Is as far as you know connectivity
> good between the machines (connect via dbaccess from one to the other etc)?
Just to add that I think this is a red herring. I have seen other
numbers such as 26 or 28 before the asynchronous read operation error.
Ben.