Very slow onbar restore
Posted in 2004
Topics: Backup & Restore, Platform-Specific Issues, Versions, Editions & End-of-Life
IDS 7.31 on HP-UX 11.11
The customer did a whole back-up onbar -b -w -L 0 on Saturday. Took about
six hours. Netbackup 3.4.1 is used to backup to a tape device attached to
another server on the network. Now they are doing a restore
(onbar -r -p -w). Although it's running OK, it is so sllllooooowwwwww. For
instance, if I zero out the stats, I see from onstat -D that the writes
clock up at perhaps a few thousand a minute: about one-tenth of the rate I
would expect.
Using sar shows fast (sub 4ms) response times on the target disks.
onstat -p also shows periods of inactivity, followed by a mass of writes,followed by more inactivity. onstat -g ses shows the ontape thread
constantly waiting on condition. there is nothing useful in the onbar,
Informix or system syslogs. I've added another cpu vp (it's using kaio), to
no avail.
All the above points towards a systemic problem with the tape stacker. But
it's not a network one: the only network card configured on the target
server shows instant response times to a ping to the server to which the
stacker is connected. The customer has previously tested restores on
identical hardware/OS/Informix/ Netbackup and they have flown through at the
expected rates.
The customer has also executed a file system restore using the same stacker
(not the samne device obviously) with the expected fast restore times. The
customer also reckons that she started the restore yesterday and it went at
normal speeds until it failed with a tape positioning error.
Any ideas? Suggestions gratefully received.
IBM support call 403279 refers.
thanks
Neil
Neil Truby wrote:
> IDS 7.31 on HP-UX 11.11
>
> The customer did a whole back-up onbar -b -w -L 0 on Saturday. Took about
> six hours. Netbackup 3.4.1 is used to backup to a tape device attached to
> another server on the network. Now they are doing a restore
> (onbar -r -p -w). Although it's running OK, it is so sllllooooowwwwww. For
> instance, if I zero out the stats, I see from onstat -D that the writes
> clock up at perhaps a few thousand a minute: about one-tenth of the rate I
> would expect.
snip
>
> Any ideas? Suggestions gratefully received.
>
>
>
Smells of hardware in the drive, or just possibly a dead tape. Either
way it's doing lots of retries to read the data.
Depending how far they are in to the restore, I'd do a quick try of the
prevoious level 0 - if it's fast it's the tape, slow it's the drive.
I also saw this once with a bad drive ( just an ontape ) took 10x longer
that I expected to restore, and ...........
the instance would not come up!
Neil,
NetBackup 3.4_1 is a OLD, non-supported version of NetBackup, and
there have been some performance improvements in newer vesions. The
latest version is 5.0 MP1. That said, try the following:
1. Make sure that BAR_DEBUG is not set in the onconfig file. If it is,
remove it and restart the backup. BAR_DEBUG is used for debugging and
causes performance slowdowns.
2. Set BAR_MAX_BACKUP to zero or a high number in the onconfig file.
The BAR_MAX_BACKUP onconfig parameter specifies the maximum number of
parallel child processes that will be forked for each dbspace backup,
one per dbspace, depending on the amount of shared memory that one has
available. Generally more is better. Setting BAR_MAX_BACKUP to zero
means an unlimited number of processes can be forked.
3. Make sure that the backup is run in parallel by not using the "-w"
flag. But also make sure that you have a corresponding set of storage
units in your class as BAR_MAX_BACKUP. Informix backups by default run
in parallel mode. Using the "-w" flag will cause a backup to run in
serial mode, which is safer but slower.
Another way backups can be run in serial mode is if BAR_MAX_BACKUP is
set to 1. Make sure that BAR_MAX_BACKUP is set to 0, or a high number.
4. If you have a bptm log
("/usr/openv/netbackup/logs/bptm/log.<date>") on your media server and
have VERBOSE=5 in your /usr/openv/netbackup/bp.conf file set you'll
see messages similar to the following:
"write_data: waited for full buffer 2134 times, delayed 9846 times"
"read_data: waited for empty buffer 100 times, delayed 5456 times"
"waiting for full buffer" is NetBackup waiting for information from
onbar. "waiting for empty buffer" is the tape library taking
information from NetBackup. Each 10,000 delays is 5 minutes. This will
help you focus on which side (or both) you need to concentrate your
efforts on.
5. Make sure that Informix and Veritas have adequate kernel resources.
Since both Informix and Veritas user memory, CPU and I/O resources one
must ensure that both products have adequate kernel resources to run
both. Informix has kernel recommendations for it's product in the
release notes at "$INFORMIXDIR/release/en_us/0333/IDS_7.3". Veritas
has kernel recommendations at "http://support.veritas.com/docs/264967"
6. Increase the BAR_NB_XPORT_COUNT and BAR_XFER_BUF_SIZE onconfig
parameters until an "onstat -g -stq" shows diminishing returns.
The BAR_NB_XPORT_COUNT Informix parameter specifies the number of data
buffers that each onbar_d process can use to exchange data with the
database server. The BAR_XFER_BUFSIZE Informix parameter specifies the
size of each transfer buffer. In general, more of each means faster
transfer of data, and these parameters have a direct effect on the
backup and restore performance. The default values of both are:
BAR_NB_XPORT_COUNT: 10
BAR_XFER_BUF_SIZE: 31 for 2K page ports, 15 for 4K page ports (see
your release notes for the page size of your instance)
Please note that if the BAR_XFER_BUFSIZE is changed, all previous
Informix backups cannot be restored.
Hope this information helps.
Brice Avila
Minneapolis, Minnesota
"Neil Truby" <neil.truby@ardenta.com> wrote in message news:<2hjtk3Fdofe8U1@uni-berlin.de>...
> IDS 7.31 on HP-UX 11.11
>
> The customer did a whole back-up onbar -b -w -L 0 on Saturday. Took about
> six hours. Netbackup 3.4.1 is used to backup to a tape device attached to
> another server on the network. Now they are doing a restore
> (onbar -r -p -w). Although it's running OK, it is so sllllooooowwwwww. For
> instance, if I zero out the stats, I see from onstat -D that the writes
> clock up at perhaps a few thousand a minute: about one-tenth of the rate I
> would expect.
>
> Using sar shows fast (sub 4ms) response times on the target disks.
> onstat -p also shows periods of inactivity, followed by a mass of writes,> followed by more inactivity. onstat -g ses shows the ontape thread
> constantly waiting on condition. there is nothing useful in the onbar,
> Informix or system syslogs. I've added another cpu vp (it's using kaio), to
> no avail.
>
> All the above points towards a systemic problem with the tape stacker. But
> it's not a network one: the only network card configured on the target
> server shows instant response times to a ping to the server to which the
> stacker is connected. The customer has previously tested restores on
> identical hardware/OS/Informix/ Netbackup and they have flown through at the
> expected rates.
>
> The customer has also executed a file system restore using the same stacker
> (not the samne device obviously) with the expected fast restore times. The
> customer also reckons that she started the restore yesterday and it went at
> normal speeds until it failed with a tape positioning error.
>
> Any ideas? Suggestions gratefully received.
>
> IBM support call 403279 refers.
>
> thanks
> Neil
Related threads
- IDS 10 table-level restore
- Informix Development Webinar December 11, 2007
- ontape -p/r with changed ROOTPATH
- Migrate from HP PA-RISC to HP ITANIUM by ontape