Onbar & Logical Logs backup
Posted in 2000
Topics: Backup & Restore, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
We are using IDS 7.31 UC6 with Sun E10K Solaris V 2.7
We set up onbar using Veritas Netbackup, which we use successfully on many
other machines.
However, we are running into a problem that I cannot find the answer to (I'm
a newbie).
I've looked in all the Informix docs that I have and have had no luck on
this from Informix.
We have LTAPEDEV set to a file system and we have
ALARMPROGRAM set to /usr/informix/etc/log_full.sh
We have 47 logs sized at 5000 each and 1 at 2000
and only 1 database that is approx 1970 MB. That database does use
unbuffered logging.
The logical logs work as they should and back up like they should all day,
with our onbar backing up continuously. At approximately 10:00 PM, a
process starts that inserts data into one of the db tables.
The logs continue to back up for awhile (looked like 89 logs backed up last
time from restart ) and then we get a message
Logical Log Files are Full --Backup is Needed.
We then have to stop and restart the engine. At that time, we get
Warning: Checkpoint appears stalled and may not complete before the
database server shuts down.
Once we shut down and restart - of course- everything is fine until the next
evening - same time, same place.
We got very tired of this, so we changed LTAPEDEV to /dev/null, which of
course we know does not work and of course we get errors and of course onbar
is not working to back up the logical logs. So, of course I want to change
it back, but I am not sure what to look at, at this point.
We are not getting a Long Transaction alarm so I am assuming that is not our
problem.
Any help would be greatly appreciated.
And sorry for sounding so "newbie-like".
Thanks
Robin
Robin,
Question:
What does onstat -l show about the logs when the message appears? I
would be curious as to where the 'L' and 'C' flags are, relative to the
unbacked logs. Also, does either /tmp/bar_act.log or your storage manager
log show any anomaly at that time or since the last successful log backup
prior to that time?
Doug
"Robin Boscia" <rboscia@att.com> wrote in message
news:8qbib7$kjj4@kcweb01.netnews.att.com...
> We are using IDS 7.31 UC6 with Sun E10K Solaris V 2.7
>
> We set up onbar using Veritas Netbackup, which we use successfully on many
> other machines.
> However, we are running into a problem that I cannot find the answer to
(I'm
> a newbie).
> I've looked in all the Informix docs that I have and have had no luck on
> this from Informix.
>
> We have LTAPEDEV set to a file system and we have
> ALARMPROGRAM set to /usr/informix/etc/log_full.sh
> We have 47 logs sized at 5000 each and 1 at 2000
> and only 1 database that is approx 1970 MB. That database does use
> unbuffered logging.
>
> The logical logs work as they should and back up like they should all day,
> with our onbar backing up continuously. At approximately 10:00 PM, a
> process starts that inserts data into one of the db tables.
> The logs continue to back up for awhile (looked like 89 logs backed up
last
> time from restart ) and then we get a message
> Logical Log Files are Full --Backup is Needed.
>
> We then have to stop and restart the engine. At that time, we get
> Warning: Checkpoint appears stalled and may not complete before the
> database server shuts down.
> Once we shut down and restart - of course- everything is fine until the
next
> evening - same time, same place.
> We got very tired of this, so we changed LTAPEDEV to /dev/null, which of
> course we know does not work and of course we get errors and of course
onbar> is not working to back up the logical logs. So, of course I want to change
> it back, but I am not sure what to look at, at this point.
>
> We are not getting a Long Transaction alarm so I am assuming that is not
our
> problem.
>
> Any help would be greatly appreciated.
> And sorry for sounding so "newbie-like".
>
> Thanks
> Robin
>
>
>
>
>
Doug,
LBU_PRESERVE is set to 0
Our checkpoint interval is set to 300
and I am still looking at all the other things.
Thanks for your help
Robin
Robin,
It looks like, from the bar_act.log, that OnBar THOUGHT there was a log
backup already running when another tried to start (maybe a orphan process,
or a backup waiting on device?). As far as the logs go, if the first were
UL and the last were UC (in unique log number order), then you probably DID
have a long-running transaction, IMO. What are the long transaction
configuration values, and do you have LBUPRESERVE(?) set to 1? Remember,
the long transaction check is only made at logical log switch time.
Again, looking at bar_act.log, what time (and which log) was the last
successful backup prior to the crash? Compare that to the message file for
the same time frame, looking to see when log files filled up.
Another nasty thought occurs to me, since it was complaining about
checkpoint stall during 'shutdown'. Perhaps the system raised a checkpoint
request (what is your checkpoint interval?) and because the 10PM process was
in critical code, could not complete the request.
I would take a LONG, HARD look at that 10PM application and see what it is
doing and whether it can be broken up into reasonably sized update
increments, with appropriate commits. I would also consider greatly
increasing the number of logical logs (and how large is your physlog space,
BTW?) We run a fairly intensive OLTP system and have 2G of logical logs, in
200 logs of 10M each, and sometimes we'll fill a log every couple of
minutes.
Of course, it goes without saying that you want to get off of /dev/null as
soon as possible, right? Or are you backing up the system EVERY night and
requiring your users to keep a copy of all their postings for at least 3
days? At least, if you crash and lose 2 out of 3 backup volumes, your could
still restore and let your users re-enter 3 days worth of data...... after
which they would probably arrange a necktie party (western style, that is).
HTH,
Doug
-----Original Message-----
From: Boscia, Robin W, NBSO [mailto:rboscia@att.com]
Sent: Thursday, September 21, 2000 2:17 PM
To: 'Agnew, Doug'
Subject: RE: Onbar & Logical Logs backup
Doug,
I unfortunately did not make a copy of the onstat -l at the time, however
if I remember correctly all the logs were status U with one as the UC and
one as the UL. Prior to that however, all were UB and the current and the
last checkpoint. However my memory may not be correct. And we really do
not want to re-create the problem unless we have to.... Like I said right
now, we show them going to /dev/null ...
Here is some of the information from the bar_act.log file right about the
time it happened.
2000-09-08 23:19:09 29165 29164 /usr/informix/bin/onbar_d -l
2000-09-08 23:19:09 29165 29164 WARNING: A log backup is already running.
Can't start an
other.
2000-09-08 23:20:50 26387 26386 WARNING: The logical logs are full.
ON-Bar may encounter
errors writing to the sysutils database.
2000-09-08 23:20:51 26387 26386 Completed backup logical log 6813.
2000-09-08 23:20:51 26387 26386 Begin backup logical log 6814.
2000-09-08 23:20:52 29413 29412 /usr/informix/bin/onbar_d -l
2000-09-08 23:20:53 29413 29412 WARNING: A log backup is already running.
Can't start an
other.
2000-09-08 23:21:39 26387 26386 WARNING: The logical logs are full.
ON-Bar may encounter
errors writing to the sysutils database.
Any help would be greatly appreciated.
Thx
Robin
-----Original Message-----
From: Agnew, Doug [mailto:dagnew@charlottepipe.com]
Sent: Thursday, September 21, 2000 6:20 AM
To: Boscia, Robin W, NBSO
Subject: Re: Onbar & Logical Logs backup
Robin,
Question:
What does onstat -l show about the logs when the message appears? I
would be curious as to where the 'L' and 'C' flags are, relative to the
unbacked logs. Also, does either /tmp/bar_act.log or your storage manager
log show any anomaly at that time or since the last successful log backup
prior to that time?
Doug
">
>
>
>
>
"Doug Agnew" <dagnew@charlottepipe.com> wrote in message
news:Y%ny5.10343$Ku3.63023@news4.atl...
> Robin,
>
> Question:
> What does onstat -l show about the logs when the message appears? I
> would be curious as to where the 'L' and 'C' flags are, relative to the
> unbacked logs. Also, does either /tmp/bar_act.log or your storage manager
> log show any anomaly at that time or since the last successful log backup
> prior to that time?
>
> Doug
>
>
> "Robin Boscia" <rboscia@att.com> wrote in message
> news:8qbib7$kjj4@kcweb01.netnews.att.com...
> > We are using IDS 7.31 UC6 with Sun E10K Solaris V 2.7
> >
> > We set up onbar using Veritas Netbackup, which we use successfully on
many
> > other machines.
> > However, we are running into a problem that I cannot find the answer to
> (I'm
> > a newbie).
> > I've looked in all the Informix docs that I have and have had no luck
on
> > this from Informix.
> >
> > We have LTAPEDEV set to a file system and we have
> > ALARMPROGRAM set to /usr/informix/etc/log_full.sh
> > We have 47 logs sized at 5000 each and 1 at 2000
> > and only 1 database that is approx 1970 MB. That database does use
> > unbuffered logging.
> >
> > The logical logs work as they should and back up like they should all
day,
> > with our onbar backing up continuously. At approximately 10:00 PM, a
> > process starts that inserts data into one of the db tables.
> > The logs continue to back up for awhile (looked like 89 logs backed up
> last
> > time from restart ) and then we get a message
> > Logical Log Files are Full --Backup is Needed.
> >
> > We then have to stop and restart the engine. At that time, we get
> > Warning: Checkpoint appears stalled and may not complete before the
> > database server shuts down.
> > Once we shut down and restart - of course- everything is fine until the
> next
> > evening - same time, same place.
> > We got very tired of this, so we changed LTAPEDEV to /dev/null, which of
> > course we know does not work and of course we get errors and of course
> onbar> > is not working to back up the logical logs. So, of course I want to
change
> > it back, but I am not sure what to look at, at this point.
> >
> > We are not getting a Long Transaction alarm so I am assuming that is not
> our
> > problem.
> >
> > Any help would be greatly appreciated.
> > And sorry for sounding so "newbie-like".
> >
> > Thanks
> > Robin
> >
> >
> >
> >
> >
>
>
Doug,
Well I did the LBU_PRESERVE to 1 and at the "witching hour", I got an XBSA
error and the logs backup completely stopped. I could not get to back up at
all so I again changed the LTAPEDEV to /dev/null and restarted IDS. When I
came in on Monday, I s/w the System Administrator and he said that there had
been a bad tape in the machine where NetBackup was working. He still thinks
there may be a bad disk, but for some reason, he doesn't know which one and
he says he'll have to see it fail... (I don't know, but then I don't know a
lot about the backup machine either)...
Anyhow, I still seem to think that we need more logs and bigger as our
online.log showed that the logs were filling up every 20 seconds....
Thanks for your help (I may need it again when we get the disk issue
solved!) and any more info you can think of would be greatly appreciated....
Robin
> From: "Doug Agnew" <dagnew@charlottepipe.com>
> Reply-To: "Doug Agnew" <dagnew@charlottepipe.com>
> Newsgroups: comp.databases.informix
> Date: Thu, 21 Sep 2000 09:24:17 -0400
> Subject: Re: Onbar & Logical Logs backup
>
> Robin,
>
> Question:
> What does onstat -l show about the logs when the message appears? I
> would be curious as to where the 'L' and 'C' flags are, relative to the
> unbacked logs. Also, does either /tmp/bar_act.log or your storage manager
> log show any anomaly at that time or since the last successful log backup
> prior to that time?
>
> Doug
>
>
> "Robin Boscia" <rboscia@att.com> wrote in message
> news:8qbib7$kjj4@kcweb01.netnews.att.com...
>> We are using IDS 7.31 UC6 with Sun E10K Solaris V 2.7
>>
>> We set up onbar using Veritas Netbackup, which we use successfully on many
>> other machines.
>> However, we are running into a problem that I cannot find the answer to
> (I'm
>> a newbie).
>> I've looked in all the Informix docs that I have and have had no luck on
>> this from Informix.
>>
>> We have LTAPEDEV set to a file system and we have
>> ALARMPROGRAM set to /usr/informix/etc/log_full.sh
>> We have 47 logs sized at 5000 each and 1 at 2000
>> and only 1 database that is approx 1970 MB. That database does use
>> unbuffered logging.
>>
>> The logical logs work as they should and back up like they should all day,
>> with our onbar backing up continuously. At approximately 10:00 PM, a
>> process starts that inserts data into one of the db tables.
>> The logs continue to back up for awhile (looked like 89 logs backed up
> last
>> time from restart ) and then we get a message
>> Logical Log Files are Full --Backup is Needed.
>>
>> We then have to stop and restart the engine. At that time, we get
>> Warning: Checkpoint appears stalled and may not complete before the
>> database server shuts down.
>> Once we shut down and restart - of course- everything is fine until the
> next
>> evening - same time, same place.
>> We got very tired of this, so we changed LTAPEDEV to /dev/null, which of
>> course we know does not work and of course we get errors and of course
> onbar>> is not working to back up the logical logs. So, of course I want to change
>> it back, but I am not sure what to look at, at this point.
>>
>> We are not getting a Long Transaction alarm so I am assuming that is not
> our
>> problem.
>>
>> Any help would be greatly appreciated.
>> And sorry for sounding so "newbie-like".
>>
>> Thanks
>> Robin
>>
>>
>>
>>
>>
>
>