RE: OnBar In Production on IDS 7.24.UC1
Posted in 1999
Topics: Backup & Restore, Platform-Specific Issues, Versions, Editions & End-of-Life
In two separate environments we are running With 7.23.UC and 7.24.UC3 on AIX 4.3.2 and 4.1.5, and we don't have any problems backing up through Onbar to ADSM. We occasionally have issues with connection problems, but lately we have been running smoothly. What type of problems are you hitting?? JP > -----Original Message----- > From: dhayes@senco.com [SMTP:dhayes@senco.com] > Sent: Wednesday, February 10, 1999 5:26 PM > To: informix-list@iiug.org > Subject: OnBar In Production on IDS 7.24.UC1 > > I am running AIX 4.3.2 with IDS 7.24.UC1 and ADSM. We have experienced > some of the know bugs. How many sites are in production with IDS > 7.24.UC1? Can you share with me the test scenario that you ran before > OnBar went into production? > > Thanks, > Doug Hayes > Database Administrator, Senco Fasteners > dhayes@senco.com
Sorry my original message lacked detail. It was written out of
frustration. The following is the configuration I am working with:
IDS 7.24.UC1
AIX 4.3.2
ADSM Server Version 3, Release 1, Level 1.5
ADSM Client Version 2, Release 1, Level 0.6
The current testing has resulted in 3 problems that might be of
interest for OnBar users. Two of the 3 problems have work a around.
1. When performing backups I would receive the following error message
multiple times in the bar_act.log:
1999-01-19 21:49:32 43090 16884 ERROR: Unable to open connection to
server: could not fork server connection.
1999-01-19 21:49:32 16884 12948 The ON-Bar process 43090 exited with
a problem (exit code 130 (0x82), signal -1).
This problem was caused by the BAR_MAX_BACKUP variable being set to 0
or 'unlimited'. I was able to determine through testing that a value
of 10 worked. The limitation was not at the ADSM server, but appeared
to be an IDS/ADSM XBSA API limitation. I was able to run 2 backups for
2 instances on the same box. These 2 instances generated 20 threads or
connections to ADSM. The WATCH OUT here is that onbar -b DOES NOT
RETURN AN ERROR CODE != 0 when these CRITICAL ERRORS occur. It has
been suggested that I increase the BAR_RETRY variable, but I assume
failures would be still be logged making automated failure
notification difficult.
2. When performing a restore that requires a log to be salvaged, onbar
writes 2 entries into the ixbar file. The first entry in the file is
for a failed log backup while the second entry is for a successful
backup. If a second restore is performed after the completion of the
first, oncheck -ce indicates overlapping extents. Removing this failed
entry appears to be a correct work around. Below are examples of the
salvaged log entries:
csadmshm 404 L 0 866 0 0 20451203 1999-02-09 21:13:43 1
csadmshm 405 L 0 0 0 0 20451205 1999-02-09 21:14:51 1 FAILED BACKUP
csadmshm 405 L 0 871 0 0 20451205 1999-02-09 23:28:06 1SUCCESSFUL
3. This issue is not resolved yet. A high percentage of the time
oncheck -cr indicates address and page errors with the log dbspaceafter a restore.
Log page error: invalid address. Log number 115, addr 0
Log page error: invalid page type. Log number 115, addr 0
Log page error: invalid address. Log number 0, addr 0
Log page error: invalid page type. Log number 0, addr 0
I have been able to prove that all the appropriate data changes are in
the database after the restore. I have also been able to eliminate
this error in log file 115 by cycling a new log into log file 115. The
error for log number 0 was removed by dropping the appropriate log
file. The oncfg.<server> file gave me clues as to which log file(s)
might have a problem. Tech support is trying to reproduce this
problem. I am trying to determine how critical this is.
What kind of restore testing has been performed at other sites? I have
performed multiple test scenarios on 2 different servers to get this
far. My test scenario consists of 1 Level 0 and 2 Level 1's. Three log
files separate the archives. Each log file contains at least 1
transaction and 1 checkpoint. Six restores are performed back to back
in the scenario. The first restore point is the highest logid. Each
successive restore point is to a lower logid. The last restore point
is the logid where the level 0 archive ended.
Any comments or information would be appreciated,
Doug Hayes
On Thu, 11 Feb 1999 15:12:02 -0500, "Viegas, John"
<ViegasJP@panasonic.com> wrote:
>
>In two separate environments we are running With 7.23.UC and 7.24.UC3 on AIX
>4.3.2 and 4.1.5, and we don't have any problems backing up through Onbar to
>ADSM. We occasionally have issues with connection problems, but lately we
>have been running smoothly.
>What type of problems are you hitting??
>
>JP
>
Doug Hayes
Database Administrator, Senco Fasteners
dhayes@senco.com