CKPT REQ stops ontape -r recovery
Posted in 2006
Topics: Backup & Restore, Storage & Space Management, Stored Procedures & SPL, Server Administration, Logging & Checkpoints
Greetings, Family.
I am trying to "clone" a server from the ontape backup set of another.
It was going great for about 11 hours and then seems to have stopped
all actvity. From the online.log:
01:20:38 Maximum server connections 0
01:32:39 Checkpoint Completed: duration was 0 seconds.
01:32:39 Checkpoint loguniq 494406, logpos 0xe0d018, timestamp:-2141117962
01:32:39 Maximum server connections 0
01:44:52 Checkpoint Completed: duration was 0 seconds.
01:44:52 Checkpoint loguniq 494406, logpos 0xe0d018, timestamp:-2141117565
01:44:52 Maximum server connections 0
---------------That's the last activity at 1:44 AM - it stops checkpointing after
that.
I have a script to monitor the total output by dbspace by filtering the
output of onstat -g iof. It shows me that there is no chunk I/O taking
place. For the first few hours of the recovery process I was runnning
this script every few minutes and watching the I/O count rise at a nice
clip.
Here's the last couple of lines of my output:
---- DBSpace totalops dskread dskwrite io/s
...
DBSP thddb_tscon041dbs 589820 0 589820 4.80
DBSP thddb_tscon034dbs 589820 0 589820 4.80
Total IO-Stats: 33318739 122 33318617 275.30
That total I/O count, 33,318,739 has not changed since I noticed the
CKPT REQ about an hour ago.
I found a old thread from 2003 that sounded similar to this but that
one ended when the poster realized there was I/O activity on the
chunks. No such luck here.
I have tried "onmode -c unblock", to no avail. I did look into the
undocumentd "onmode -O" but it gave a dire warning about marking
chunks/dbspaces down and requiring a recovery. That would be a useless
exercise for me so I answered N and exited.
Any other ideas? You have my rapt attention! ;-(
Thanks much.
-- J.S.
PS.
Here's my script. I will see about posting it to the IIUG library when
I get a round tuitt.
#!/usr/bin/ksh
# monitorDBspaceIO.sh
#
onstat -g iof|tail +6 |gawk '
BEGIN {
total_io = 0
iopersec = 0.0
}
$3 != 0 {
#print
total_io += $3
dskread += $4
dskwrite += $5
iopersec += $6
split($2, q_chunk, ".")
cur_dbspace = q_chunk[2]
dbspace_total_io[cur_dbspace] += $3
dbspace_dskread[cur_dbspace] += $4
dbspace_dskwrite[cur_dbspace] += $5
dbspace_iopersec[cur_dbspace] += $6
}
END {
printf("---- DBSpace totalops dskread dskwrite io/s\\n")
for (dbsname in dbspace_total_io)
{
if (length(dbsname) == 0)
continue
#printf ("dbsname is <%s>\\n", dbsname)
printf("DBSP %s %d %d %d %7.2f\\n", dbsname,
dbspace_total_io[dbsname], dbspace_dskread[dbsname],
dbspace_dskwrite[dbsname], dbspace_iopersec[dbsname])
}
printf ("\\nTotal IO-Stats: %d %d %d %7.2f\\n", total_io, dskread,
dskwrite, iopersec)
}
'
BTW, the above output had been piped through "beautify-unl.sh -db",
which is available at IIUG.
Hi,
"onstat -g ath" shows in what state all of the threads are.
There might be some interesting info in this output.
Next, it might be useful to get a stack trace from a few
threads of interest. This can be done using "onstat -g stk <tid>"
for threads that are not in state "running". For threads in state
"running" the stack trace from an attached debugger is more
reliable.
After (or before) collecting this info open a "case" with IBM
Informix Tech Support so that they can find out whether this
is the symptom of a known problem ...
Of course version information is essential, as always.
Regards,
Martin
--
Martin Fuerderer
IBM Informix Development Munich, Germany
Information Management
Informix URLs list: http://home.arcor.de/mfu1/informix/urls.html
IBM Information On Demand Global Conference
October 15-20, 2006, Anaheim, California
see http://www.ibm.com/events/informationondemand
informix-list-bounces@iiug.org wrote on 17.09.2006 05:09:03:
> Greetings, Family.
>
> I am trying to "clone" a server from the ontape backup set of another.
> It was going great for about 11 hours and then seems to have stopped
> all actvity. From the online.log:
>
> 01:20:38 Maximum server connections 0
> 01:32:39 Checkpoint Completed: duration was 0 seconds.
> 01:32:39 Checkpoint loguniq 494406, logpos 0xe0d018, timestamp:> -2141117962
>
> 01:32:39 Maximum server connections 0
> 01:44:52 Checkpoint Completed: duration was 0 seconds.
> 01:44:52 Checkpoint loguniq 494406, logpos 0xe0d018, timestamp:> -2141117565
>
> 01:44:52 Maximum server connections 0
> ---------------> That's the last activity at 1:44 AM - it stops checkpointing after
> that.
>
> I have a script to monitor the total output by dbspace by filtering the
> output of onstat -g iof. It shows me that there is no chunk I/O taking
> place. For the first few hours of the recovery process I was runnning
> this script every few minutes and watching the I/O count rise at a nice
> clip.
>
> Here's the last couple of lines of my output:
>
> ---- DBSpace totalops dskread dskwrite io/s
> ...
> DBSP thddb_tscon041dbs 589820 0 589820 4.80
> DBSP thddb_tscon034dbs 589820 0 589820 4.80
>
> Total IO-Stats: 33318739 122 33318617 275.30
>
> That total I/O count, 33,318,739 has not changed since I noticed the
> CKPT REQ about an hour ago.
>
> I found a old thread from 2003 that sounded similar to this but that
> one ended when the poster realized there was I/O activity on the
> chunks. No such luck here.
>
> I have tried "onmode -c unblock", to no avail. I did look into the
> undocumentd "onmode -O" but it gave a dire warning about marking
> chunks/dbspaces down and requiring a recovery. That would be a useless
> exercise for me so I answered N and exited.
>
> Any other ideas? You have my rapt attention! ;-(
>
> Thanks much.
>
> -- J.S.
>
> PS.
>
> Here's my script. I will see about posting it to the IIUG library when
> I get a round tuitt.
>
> #!/usr/bin/ksh
> # monitorDBspaceIO.sh
> #
> onstat -g iof|tail +6 |> gawk '
> BEGIN {
> total_io = 0
> iopersec = 0.0
> }
> $3 != 0 {
> #print
> total_io += $3
> dskread += $4
> dskwrite += $5
> iopersec += $6
> split($2, q_chunk, ".")
> cur_dbspace = q_chunk[2]
> dbspace_total_io[cur_dbspace] += $3
> dbspace_dskread[cur_dbspace] += $4
> dbspace_dskwrite[cur_dbspace] += $5
> dbspace_iopersec[cur_dbspace] += $6
> }
> END {
> printf("---- DBSpace totalops dskread dskwrite io/s\\n")
> for (dbsname in dbspace_total_io)
> {
> if (length(dbsname) == 0)
> continue
> #printf ("dbsname is <%s>\\n", dbsname)
> printf("DBSP %s %d %d %d %7.2f\\n", dbsname,
> dbspace_total_io[dbsname], dbspace_dskread[dbsname],
> dbspace_dskwrite[dbsname], dbspace_iopersec[dbsname])
> }
> printf ("\\nTotal IO-Stats: %d %d %d %7.2f\\n", total_io, dskread,
> dskwrite, iopersec)
> }
> '
> BTW, the above output had been piped through "beautify-unl.sh -db",
> which is available at IIUG.
>
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g