Re: Online 5.0 archive/tbcheck issues
Posted in 2007
Topics: Storage & Space Management, Server Administration, Migration, Import/Export & Data Conversion
On Jun 25, 4:27 pm, barries <barries...@hotmail.com> wrote:
> I'm looking for assistance with Online version 5.0.5. I need to know
> if there is a way to detect problems at the system level of an Online
> 5 instance besides using tbcheck.
>
> Informix version: Online version 5.0.5 UC1
> OS: SCO Unix System 5 Ver. 3.2
>
> The whole story:
> One of my clients provides software and servers to a number of their
> clients. Their application is running on Online 5.05.UC1 on SCO
> boxes. Upgrading is not an option; they are locked into this version
> due to regulatory constraints. The application database never exceeds
> about 500 MB; they provide a process to remove old records as the
> database grows so that their clients never have to be concerned with
> adding disk space or other administration tasks.
>
> The problem is this: They've run into an issue with several of their
> biggest clients where informix tbtape archives stop working properly.
> Rather than taking an hour as expected, the archive takes just 5-10
> minutes, then when my client tries to restore their archive to his
> server, the restore completes in just 5-10 minutes but doesn't really
> restore anything and the database is "broken" (tbchecks against the
> resulting database report errors). Tbtape reports no errors; the
> normal archive started and completed messages appear in the online
> log. No errors appear in the online log of the restored instance,
> either.
>
> I've had him run tbcheck -cr, -cc, and -ce on the originating
> instance, and all come back with no errors. There are no tbcheck data
> or index errors in the databases of the originating instances and the
> application works without error. The only symptom (and problem) is
> that an archive won't work.
>
> Because the problem is apparently at the system level of the instance,
> a workaround is to dbexport the data, reinitialize the instance, then
> reload the data. This does correct the problem. However, it would be
> preferable to have a way of recognizing that the problem exists
> without having to rely on spotting a bad archive. In some cases,
> tbstat -d and tbcheck -pe return negative page counts, which makes it
> obvious. But in most cases there is nothing at all to indicate that
> something is wrong.
>
> So, back to my original question: Is there a way to detect that there
> is something wrong at the instance level besides using tbcheck? Since
> this is version 5, there is no sysmaster database to query.
>
> Thank you in advance for any assistance you can give me,
>
> Barrie Shaw
> Xtivia, Inc.
Tbtape in OL5.xx DID NOT WORK PROPERLY in any release prior to 5.07!
It would miss backing up pages when the server was busy. In addition,
IB you're running into something else that was also fixed at about the
same time. The page timestamps have wrapped past 2^31 and gone
negative. These older versions of tbtape did not properly handle the
condition and got confused.
Regulatory problems or not, your client is going to have to find a way
to start upgrading his clients to OL5.20+, there are MANY reasons to
do that.
Highlights:
- tbtape broken prior to 5.07, restores are unreliable unless then
engine is not being updated while the archive is being made!
- istar/inet (ie TCP connections) considerably slower in 5.05 and
earlier than in 5.07 and later (in 5.06 it was just broken altogether
- oops!)
- Y2K compliance added in 5.10 and later (though obviously that's not
an issue here)
- Very old pages not handled properly.
Art S. Kagel
Art S. Kagel wrote:
>
> Tbtape in OL5.xx DID NOT WORK PROPERLY in any release prior to 5.07!
> It would miss backing up pages when the server was busy. In addition,
> IB you're running into something else that was also fixed at about the
> same time. The page timestamps have wrapped past 2^31 and gone
> negative. These older versions of tbtape did not properly handle the
> condition and got confused.
>
Oh the flashbacks - I was the poor sod that found that bug.
IIRC it was actually a bug in the SCO icc 'enhanced C compiler'
Advice from tech support at the time
'Take your DB offline every night and dd the chunks to tape by hand
until we can ship you a fix'
Our db was much too big to dbexport/import as disk was expensive then.
Took about 4 days as I recall...
--
Clive
On Jun 26, 11:34 am, Clive Eisen <c...@serendipita.com> wrote:
> Art S. Kagel wrote:
>
> > Tbtape in OL5.xx DID NOT WORK PROPERLY in any release prior to 5.07!
> > It would miss backing up pages when the server was busy. In addition,
> > IB you're running into something else that was also fixed at about the
> > same time. The page timestamps have wrapped past 2^31 and gone
> > negative. These older versions of tbtape did not properly handle the
> > condition and got confused.
>
> Oh the flashbacks - I was the poor sod that found that bug.
> IIRC it was actually a bug in the SCO icc 'enhanced C compiler'
>
> Advice from tech support at the time
>
> 'Take your DB offline every night and dd the chunks to tape by hand
> until we can ship you a fix'
>
> Our db was much too big to dbexport/import as disk was expensive then.
>
> Took about 4 days as I recall...
>
> --
> Clive
You and I must have hit it at about the same time then. We'd been
doing test restores to the same test machine using the same disks over
and over so it always tbchecked out fine because any pages missing
from a particular archive were restored from the original test archive
or one of the previous restore tests. Sigh. One day tried to restore
to a different machine and there were holes in the data. literally
pages missing in the middle of an extent. Reported it and they said,
OH, yeah, that's a bug and we'll have a patch for your 5.07 release in
a few days. It's scheduled to be fixed in 5.08. Right 5.08, not
5.07! Darn I'm getting old.
Art S. Kagel
Art S. Kagel wrote: > > You and I must have hit it at about the same time then. We'd been > doing test restores to the same test machine using the same disks over > and over so it always tbchecked out fine because any pages missing > from a particular archive were restored from the original test archive > or one of the previous restore tests. Sigh. One day tried to restore > to a different machine and there were holes in the data. literally > pages missing in the middle of an extent. Reported it and they said, > OH, yeah, that's a bug and we'll have a patch for your 5.07 release in > a few days. It's scheduled to be fixed in 5.08. Right 5.08, not > 5.07! Darn I'm getting old. You and me both Art. Almost exactly the same - test restore every night to alternate machine - never noticed anything wrong until the alternate machine lost a disk and I could no longer restore after replacing the disk. -- Clive