ontape restore taking time
Posted in 2005
Topics: Backup & Restore, Storage & Space Management, Logging & Checkpoints, Versions, Editions & End-of-Life
Hi,
I am working on HP_UX 11i and IDS 9.30.
When i am restoring the data using ontape the server is in Fast recovery for
long time. Its hanging at Cleaning the physical logs for a long time. I am
attaching the message from log here.
11:40:38 Checkpoint Completed: duration was 0 seconds.
11:40:38 Checkpoint loguniq 60, logpos 0x7983018
11:40:38 Maximum server connections 0
11:42:56 Checkpoint Completed: duration was 0 seconds.
11:42:56 Checkpoint loguniq 60, logpos 0x7983018
11:42:56 Maximum server connections 0
11:42:56 Checkpoint Completed: duration was 0 seconds.
11:42:56 Checkpoint loguniq 60, logpos 0x7983018
11:42:56 Maximum server connections 0
11:42:58 Physical Restore of scl_rootdbs, scldbs1, scldbs2, scldbs3, scldbs4,
scldbs5, scldbs6, scldbs7, scldbs8, scldbs9, sc
ldbs10, scllocaldbs1, scldbs11, scldbs12, scldbs Completed.
11:42:58 Checkpoint Completed: duration was 0 seconds.
11:42:58 Checkpoint loguniq 60, logpos 0x7983018
11:42:58 Maximum server connections 0
11:47:30 No logical log restore will be performed.
11:47:30 Preparing Physical Log for Fast Recovery ...
11:47:30 Clearing the physical and logical logs has started
12:28:21 Cleared 1013 MB of the physical and logical logs in 2450 seconds
12:28:21 Physical Recovery Started at Page(1:31210).
12:28:21 Physical Recovery Complete: 0 Pages Examined 0 Pages Restored.
12:28:21 Logical Recovery Started.
12:28:21 10 recovery worker threads will be started.
12:28:21 Checkpoint Completed: duration was 0 seconds.
12:28:21 Checkpoint loguniq 60, logpos 0x7983018
12:28:21 Maximum server connections 0
12:28:23 Logical Recovery has reached the transaction cleanup phase.
12:28:23 Logical Recovery Complete.
0 Committed, 0 Rolled Back, 0 Open, 0 Bad Locks
12:28:24 Bringing system to Quiescent Mode with no Logical Restore.
Please provide me some input where i can perform the resotore faster.
Bye
Hi,
this is a known "problem".
I'm not quite sure about HP-UX, but I once did a thorough
investigation of this on Solaris and some research on Linux.
For more info on the results please see the following text.
If you want to verify my results on your HP-UX platform,
you can use the system utility "tusc", which approximately
does the same as "truss" on Solaris (but probably uses
different commandline options).
Regards,
Martin
--
Martin Fuerderer
IBM Informix Development Munich, Germany
Information Management
------------------------------------------------------------------------------
Measurements using "truss" showed, that >98% of the time needed for log
clearing is spent in the system call pwrite(). While this system call
performs a lot slower than the normal write(), it is necessary to use
pwrite() because of the threading in IDS.
The performance of write() and pwrite() is further influenced by the
use of the O_SYNC flag (in open()). It guarantees that the call from
write()/pwrite() does not return before the data is handed to the
disk hardware (or controller) to ensure data integrity on disk.
[ write() performs the operation at the position where the file pointer
currently is located. When using write() to put data to a specific
location in the file, then a seek() needs to be done beforehand to
position the file pointer correctly for the following write() call.
With seek() and write() being separate system calls, there is no
guarantee in a threaded application like IDS that after a seek() the
file pointer will be at the specified location at the time of write().
A different thread can have done another operation on the same file that
has meanwhile moved the file pointer.
pwrite() solves this problem by combining the seek() and write()
operations into a single operation that cannot be interrupted by
a different thread. That way the integrity of the file pointer
between the seek() and write() inside the pwrite() is guaranteed. ]
Therefore, the IDS disk throughput during disk clearing cannot be
compared to performance of other utilities (e.g. "dd") which need
neither pwrite() nor the O_SYNC flag. Usually such utilities just do
write() and can get the full benefit of OS buffering capabilities.
On Solaris it has been observed that using pwrite() for cooked files
is particularly slow. There is a considerable performance gain when
using pwrite() on raw devices (even without utilizing KAIO).
A performance gain has also been observed on other platforms, but
generally not as much as on Solaris. Furthermore, the different timings
can depend extremely on the implementation of the file system.
E.g. it is known that on the first generation Reiser FS often
used on Linux, pwrite() performs extremely bad (much worse than on
Solaris). However, on a different file system implementation the
performance can be much better.
------------------------------------------------------------------------------
forum.subscriber@iiug.org wrote on 25.01.2005 08:46:14:
> Hi,
> I am working on HP_UX 11i and IDS 9.30.
>
> When i am restoring the data using ontape the server is in Fast recovery
for long time. Its hanging at Cleaning the physical logs for a long time.
I am attaching the message from log here.
>
> 11:40:38 Checkpoint Completed: duration was 0 seconds.
> 11:40:38 Checkpoint loguniq 60, logpos 0x7983018
>
> 11:40:38 Maximum server connections 0
> 11:42:56 Checkpoint Completed: duration was 0 seconds.
> 11:42:56 Checkpoint loguniq 60, logpos 0x7983018
>
> 11:42:56 Maximum server connections 0
> 11:42:56 Checkpoint Completed: duration was 0 seconds.
> 11:42:56 Checkpoint loguniq 60, logpos 0x7983018
>
> 11:42:56 Maximum server connections 0
> 11:42:58 Physical Restore of scl_rootdbs, scldbs1, scldbs2, scldbs3,
scldbs4, scldbs5, scldbs6, scldbs7, scldbs8, scldbs9, sc
> ldbs10, scllocaldbs1, scldbs11, scldbs12, scldbs Completed.
> 11:42:58 Checkpoint Completed: duration was 0 seconds.
> 11:42:58 Checkpoint loguniq 60, logpos 0x7983018
>
> 11:42:58 Maximum server connections 0
> 11:47:30 No logical log restore will be performed.
> 11:47:30 Preparing Physical Log for Fast Recovery ...
> 11:47:30 Clearing the physical and logical logs has started
> 12:28:21 Cleared 1013 MB of the physical and logical logs in 2450
seconds
> 12:28:21 Physical Recovery Started at Page(1:31210).
> 12:28:21 Physical Recovery Complete: 0 Pages Examined 0 Pages Restored.
>
> 12:28:21 Logical Recovery Started.
> 12:28:21 10 recovery worker threads will be started.
> 12:28:21 Checkpoint Completed: duration was 0 seconds.
> 12:28:21 Checkpoint loguniq 60, logpos 0x7983018
>
> 12:28:21 Maximum server connections 0
> 12:28:23 Logical Recovery has reached the transaction cleanup phase.
> 12:28:23 Logical Recovery Complete.
> 0 Committed, 0 Rolled Back, 0 Open, 0 Bad Locks
>
> 12:28:24 Bringing system to Quiescent Mode with no Logical Restore.
>
>
> Please provide me some input where i can perform the resotore faster.
>
> Bye
>
>
>