Re: CKPT REQ stops ontape -r recovery
Posted in 2006
Thanks to Norma and Martin for their efforts here. But it turns out
the problem was not as I perceived and presented it.
The state of CKPT REQ is apparently a common phenomenon during
recovery. When I later reran the recovery (after some fixups) it got
into that state a while after starting the recovery and stayed that way
for over 24 hours while it ran the 4.6TB recovery. So that was a red
herring. (BTW, I am revolted at red herring, AKA matjes, an ethnic
delicacy in my neighborhood. ;-)
The problem had more to do with the specific way my environment does
ontape backups, creating dozens to hundreds of files on huge disk
volumes, and creating symbolic links to them in a single directory.
One of the symlinks had gotten corrupted, pointing to a file that did
not exist. I found the correct file, fixed the symlink and retried.
The recovery went off without a hitch.
Again, thanks for trying. At least we now know not to panic at the
CKPT REQ state during a recovery.
-- J
P.S. I have written a new perl script, to posted to IIUG in the near
future. It consoldates the chunk I/O statistics from onstat -g iof and
presents them as dbspace I/O stats. The embedded shell script I
included in my original post can work only in my environment due to the
naming conventions we use here.
Beau Nanaz wrote:
> Greetings, Family.
>
> I am trying to "clone" a server from the ontape backup set of another.
> It was going great for about 11 hours and then seems to have stopped
> all actvity. From the online.log:
>
> 01:20:38 Maximum server connections 0
> 01:32:39 Checkpoint Completed: duration was 0 seconds.
> 01:32:39 Checkpoint loguniq 494406, logpos 0xe0d018, timestamp: -2141117962
>
> 01:32:39 Maximum server connections 0
> 01:44:52 Checkpoint Completed: duration was 0 seconds.
> 01:44:52 Checkpoint loguniq 494406, logpos 0xe0d018, timestamp:> -2141117565
>
> 01:44:52 Maximum server connections 0
> ---------------> That's the last activity at 1:44 AM - it stops checkpointing after that.
>
> I have a script to monitor the total output by dbspace by filtering the
> output of onstat -g iof. It shows me that there is no chunk I/O taking
> place. For the first few hours of the recovery process I was runnning
> this script every few minutes and watching the I/O count rise at a nice
> clip.
>
> Here's the last couple of lines of my output:
>
> ---- DBSpace totalops dskread dskwrite io/s
> ...
> DBSP thddb_tscon041dbs 589820 0 589820 4.80
> DBSP thddb_tscon034dbs 589820 0 589820 4.80
>
> Total IO-Stats: 33318739 122 33318617 275.30
>
> That total I/O count, 33,318,739 has not changed since I noticed the
> CKPT REQ about an hour ago.
>
> I found a old thread from 2003 that sounded similar to this but that
> one ended when the poster realized there was I/O activity on the
> chunks. No such luck here.
>
> I have tried "onmode -c unblock", to no avail. I did look into the
> undocumentd "onmode -O" but it gave a dire warning about marking
> chunks/dbspaces down and requiring a recovery. That would be a useless
> exercise for me so I answered N and exited.
>
> Any other ideas? You have my rapt attention! ;-(
>
> Thanks much.
>
> -- J.S.
>
> PS.
>
> Here's my script. I will see about posting it to the IIUG library when
> I get a round tuitt.
> (Environment-specific script snipped)