Re: Very slow onbar restore
Posted in 2004
I've now received the following back fro the customer.
Does it strike any chords with anyone?
Sorry it's so long - executive summary: it's because we do -w (whole)
backups.
thanks as ever
Neil
"We do our onbar backups with the -w flag, because one of our reasons for
doing them is for disaster recovery offsite, when we will never have logs
available to do a logical restore as well. This seems to back up the
dbspaces in informix dbspace ID number; however, when you restore, it does
so in alphabetical order of dbspace name once it's taken care of the
critical dbspaces. This means that the restore is looking for files in a
different order from the one in which they were backed up to tape.
Previously, while inefficient, this hasn't been a major issue since the
backups have continued at the same speed no matter how often the tape has
had to be repositioned. With this particular restore however, any tape
repositioning to access the next file seemed to make it potentially possible
for the restore speed to slow to a crawl where it only updates in short
bursts, or restore at a normal data rate. We saw it start fast, reposition
and go slow, reposition again and return to fast. We also saw it start slow
on a couple of occasions, although we mostly abandoned those before it got
far enough to need to do an out of order dbspace and reposition. Due to the
fact that we wanted to recover to a point in time as close as possible to
the time the database became corrupted we didn't experiment with too many
different backup tapes, but the restores we tried were on two different
tapes (one wholly on the first tape, the second partly on that and partly on
the other tape), the same thing happened on all three of the tape drives in
that robot, and the same dbspace from the same backup was observed to
restore either fast or slow, seemingly at random. There was no consistency
about what happened, other than that the speed change always seemed to come
when a tape reposition was required. Once running at a certain speed, it
would continue at that speed for as long as there were consecutive dbspaces
in restore order on the tape.
In the end we simply let a restore attempt run through. It took just under
30 hours, when the expected restore time would be 4-8 hours. At that, I
think we were lucky that some of the biggest dbspaces avoided getting hit by
the slow restore phenomenon. It probably made all the difference between it
finishing on Saturday night or Monday.
Do you know of any way to persuade onbar to back up and restore in the same
order, bearing in mind that we need the backup in question to be able to
restore without the need for a logical log restore as well? It seems very
bizarre that it doesn't, but none of the various experiments I tried last
week (changing the value of BAR_MAX_BACKUP, specifying the dbspaces in the
right order in a file, etc) were very helpful. This is a shame, as if we'd
been able to restore in backup order I strongly suspect that it would have
run through with no problems, as soon as we got a restore attempt to start
at the correct speed."