Re: Database server hosed, no back-ups
Posted in 2004
Topics: Error Codes & Troubleshooting, Logging & Checkpoints, Migration, Import/Export & Data Conversion, Platform-Specific Issues
Wouldn't it be better to recover a consistent system even if two days of
work are lost?
I think that this is prefereible, in fat some years agi the same (almos
exactly happen to us: logs to /dev/null, weekend (end of month), no
recovery plan at all... The only different was that we certainly have a
support contract but all that was told to us is that we must recover
from a restore tape.
In the long term some people, includoing us, had to work more everything
was well.
Chucho!
Neil Truby wrote:
> Art
>
> I've explained the point about the logical recovery to them. They believe
> that, because they know which 30 tables were running in the single
> transaction that was running at the time of failure, it would be preferable
> to rebuild those 30 to remove the logical inconsistency, than rebuild all
> 300 from 2-day-old ASCII unloads.
>
> Reference the RAID: I didn't ask but I did have /var/adm/syslog/syslog.log
> checked and there were no SCSI errors ....
>
> cheers
> Neil
>
> "Art S. Kagel" <kagel@bloomberg.net> wrote in message
> news:pan.2004.01.12.17.04.34.187881.12806@bloomberg.net...
>
>>On Mon, 12 Jan 2004 16:48:11 -0500, Neil Truby wrote:
>>
>>No help for it Neil. Support could truncate the logical logs so the
>
> engine
>
>>can come online, but the damage is to the data pages not the log pages so
>
> the
>
>>data's still hosed. The log is complaining that it could not rollback a
>>delete (well ... many deletes) because of bad page header/trailer matching
>>caused by disk corruption. Sounds like besides a complete lack of
>>understanding of how archives work, the site is also using RAID5 and has
>
> some
>
>>very old disks that have now gone flaky on them (I'd put some money on
>
> it).
>
>>Art S. Kagel
>>
>>
>>
>>>7.30 UC3 on HP-UX 10.20
>>>
>>>A site, not yet a customer, has today had a large single-transaction
>
> update
>
>>>crash, and now his database server is inaccessible (see log below).
>>>
>>>The only have one tape, and they constantly use it to archive at night,
>
> than
>
>>>back up the logs in the morning. So the tape with the most recent
>
> archive,
>
>>>Friday night's, has been over-written by Saturday's logs(!)
>>>
>>>They unloaded all 300 tables on Friday night, so have the option to
>
> re-build
>
>>>the database from that. The only alternative I can think of is to get
>
> Tech
>
>>>Support (and they have no current support contract) and patch the
>
> relevant
>
>>>area to prevent logical recovery. Of course, their data will be
>
> logically
>
>>>inconsistent, but they seem quite exceoted by this option.
>>>
>>>Any other flashes of genius?
>>>
>>>cheers
>>>Neil
>>>
>>>Mon Jan 12 20:25:44 2004
>>>
>>>20:25:44 Event alarms enabled. ALARMPROG =>
> '/usr/informix/etc/no_log.sh'
>
>>>20:25:51 DR: DRAUTO is 0 (Off)
>>>20:25:51 Informix Dynamic Server Version 7.30.UC3A Software Serial>
> Number
>
>>>AAD#J234612
>>>20:25:51 Informix Dynamic Server Initialized -- Shared Memory>
> Initialized.
>
>>>20:25:52 Physical Recovery Started.
>>>20:25:52 Physical Recovery Complete: 1174 Pages Restored. 20:25:52>
> Logical
>
>>>Recovery Started.
>>>20:26:17 Assert Failed: Page Check Error in bfput 20:26:17 Informix
>>>Dynamic Server Version 7.30.UC3A 20:26:17 Who: Session(11,>
> root@sta154, 0,
>
>>>0)
>>> Thread(37, fast_rec, 0, 4)
>>> File: rsdebug.c Line: 998
>>>20:26:17 Results: Possible inconsistencies in 'bha:"".' 20:26:17>
> Action:
>
>>>Run 'oncheck -cDI bha:"".' 20:27:44 See Also: /tmp/af.2502e8,
>>>shmem.2502e8.0 20:27:44 Error writing '/tmp/shmem.2502e8.0' errno = 28
>>>20:27:44 Assert Failed: Rollback error 172 20:27:44 Informix Dynamic
>>>Server Version 7.30.UC3A 20:27:44 Who: Session(11, root@sta154, 0, 0)
>>> Thread(37, fast_rec, 0, 4)
>>> File: rstrans.c Line: 2115
>>>20:27:44 Results: Log record (OLDRSAM:DELITEM) in log 47852, offset>>>0xd4018 was not rolled back
>>>20:27:44 Action: Use 'onlog' to view the transaction and repair>
> manually.
>
>>>20:29:15 See Also: /tmp/af.2502e8, shmem.2502e8.1 20:29:15 Error>
> writing
>
>>>'/tmp/shmem.2502e8.1' errno = 28 20:29:15 Assert Failed: Rollback error
>
> 172
>
>>>20:29:15 Informix Dynamic Server Version 7.30.UC3A 20:29:15 Who:
>>>Session(11, root@sta154, 0, 0)
>>> Thread(37, fast_rec, 0, 4)
>>> File: rsextlog.c Line: 1412
>>>20:29:15 Results: Log record (OLDRSAM:DELITEM) in log 47852, offset>>>0xd4018 was not rolled back
>>>20:29:15 Action: Use 'onlog' to view the transaction and repair>
> manually.
>
>>>20:30:48 See Also: /tmp/af.2502e8, shmem.2502e8.2 20:30:48 Error>
> writing
>
>>>'/tmp/shmem.2502e8.2' errno = 28 20:30:48 Assert Failed: Dynamic Server
>>>must abort 20:30:48 Informix Dynamic Server Version 7.30.UC3A 20:30:48
>
> Who:
>
>>>Session(11, root@sta154, 0, 0)
>>> Thread(37, fast_rec, 0, 4)
>>> File: rslog.c Line: 3152
>>>20:30:48 Results: Dynamic Server must abort 20:30:48 Action:>>>Reinitialize shared memory 20:30:48 stack trace for pid 10183 written
>
> to
>
>>>/tmp/af.2502e8 20:30:48 See Also: /tmp/af.2502e8, shmem.2502e8.3
>
> 20:30:48
>
>>>Error writing '/tmp/shmem.2502e8.3' errno = 28 20:30:48 rslog.c, line
>
> 3152,
>
>>>thread 37, proc id 10183, Dynamic Server must abort. 20:30:48 PANIC:
>>>Attempting to bring system down
>
>
>
>
--
Atte,
Jesus Antonio Santos Giraldo
-----------------------------------
jeansagi@myrealbox.com
jeansagi@netscape.net
sending to informix-list
I agree with you 100% Jean.
However, they have IBM involved now (and, without a support contract in
place I trust they are being charged royally for the privilege!), so we've
sort of been elbowed aside now.
They seem to hope that an up-to-date, corrupt database is better than a
3-day-old self-consistent one.
My experience is that they may well eventually end up re-building the old
one.
--
Neil Truby t:01932 724027
Director m:07798 811708
Ardenta Limited e:neil.truby@ardenta.com
"Jean Sagi" <jeansagi@myrealbox.com> wrote in message
news:bu0ub6$edn$1@terabinaries.xmission.com...
>
>
> Wouldn't it be better to recover a consistent system even if two days of
> work are lost?
>
> I think that this is prefereible, in fat some years agi the same (almos
> exactly happen to us: logs to /dev/null, weekend (end of month), no
> recovery plan at all... The only different was that we certainly have a
> support contract but all that was told to us is that we must recover
> from a restore tape.
>
> In the long term some people, includoing us, had to work more everything
> was well.
>
> Chucho!
>
> Neil Truby wrote:
> > Art
> >
> > I've explained the point about the logical recovery to them. They
believe
> > that, because they know which 30 tables were running in the single
> > transaction that was running at the time of failure, it would be
preferable
> > to rebuild those 30 to remove the logical inconsistency, than rebuild
all
> > 300 from 2-day-old ASCII unloads.
> >
> > Reference the RAID: I didn't ask but I did have
/var/adm/syslog/syslog.log
> > checked and there were no SCSI errors ....
> >
> > cheers
> > Neil
> >
> > "Art S. Kagel" <kagel@bloomberg.net> wrote in message
> > news:pan.2004.01.12.17.04.34.187881.12806@bloomberg.net...
> >
> >>On Mon, 12 Jan 2004 16:48:11 -0500, Neil Truby wrote:
> >>
> >>No help for it Neil. Support could truncate the logical logs so the
> >
> > engine
> >
> >>can come online, but the damage is to the data pages not the log pages
so
> >
> > the
> >
> >>data's still hosed. The log is complaining that it could not rollback a
> >>delete (well ... many deletes) because of bad page header/trailer
matching
> >>caused by disk corruption. Sounds like besides a complete lack of
> >>understanding of how archives work, the site is also using RAID5 and has
> >
> > some
> >
> >>very old disks that have now gone flaky on them (I'd put some money on
> >
> > it).
> >
> >>Art S. Kagel
> >>
> >>
> >>
> >>>7.30 UC3 on HP-UX 10.20
> >>>
> >>>A site, not yet a customer, has today had a large single-transaction
> >
> > update
> >
> >>>crash, and now his database server is inaccessible (see log below).
> >>>
> >>>The only have one tape, and they constantly use it to archive at night,
> >
> > than
> >
> >>>back up the logs in the morning. So the tape with the most recent
> >
> > archive,
> >
> >>>Friday night's, has been over-written by Saturday's logs(!)
> >>>
> >>>They unloaded all 300 tables on Friday night, so have the option to
> >
> > re-build
> >
> >>>the database from that. The only alternative I can think of is to get
> >
> > Tech
> >
> >>>Support (and they have no current support contract) and patch the
> >
> > relevant
> >
> >>>area to prevent logical recovery. Of course, their data will be
> >
> > logically
> >
> >>>inconsistent, but they seem quite exceoted by this option.
> >>>
> >>>Any other flashes of genius?
> >>>
> >>>cheers
> >>>Neil
> >>>
> >>>Mon Jan 12 20:25:44 2004
> >>>
> >>>20:25:44 Event alarms enabled. ALARMPROG => >
> > '/usr/informix/etc/no_log.sh'
> >
> >>>20:25:51 DR: DRAUTO is 0 (Off)
> >>>20:25:51 Informix Dynamic Server Version 7.30.UC3A Software Serial> >
> > Number
> >
> >>>AAD#J234612
> >>>20:25:51 Informix Dynamic Server Initialized -- Shared Memory> >
> > Initialized.
> >
> >>>20:25:52 Physical Recovery Started.
> >>>20:25:52 Physical Recovery Complete: 1174 Pages Restored. 20:25:52> >
> > Logical
> >
> >>>Recovery Started.
> >>>20:26:17 Assert Failed: Page Check Error in bfput 20:26:17 Informix
> >>>Dynamic Server Version 7.30.UC3A 20:26:17 Who: Session(11,> >
> > root@sta154, 0,
> >
> >>>0)
> >>> Thread(37, fast_rec, 0, 4)
> >>> File: rsdebug.c Line: 998
> >>>20:26:17 Results: Possible inconsistencies in 'bha:"".' 20:26:17> >
> > Action:
> >
> >>>Run 'oncheck -cDI bha:"".' 20:27:44 See Also: /tmp/af.2502e8,
> >>>shmem.2502e8.0 20:27:44 Error writing '/tmp/shmem.2502e8.0' errno = 28
> >>>20:27:44 Assert Failed: Rollback error 172 20:27:44 Informix Dynamic
> >>>Server Version 7.30.UC3A 20:27:44 Who: Session(11, root@sta154, 0, 0)
> >>> Thread(37, fast_rec, 0, 4)
> >>> File: rstrans.c Line: 2115
> >>>20:27:44 Results: Log record (OLDRSAM:DELITEM) in log 47852, offset> >>>0xd4018 was not rolled back
> >>>20:27:44 Action: Use 'onlog' to view the transaction and repair> >
> > manually.
> >
> >>>20:29:15 See Also: /tmp/af.2502e8, shmem.2502e8.1 20:29:15 Error> >
> > writing
> >
> >>>'/tmp/shmem.2502e8.1' errno = 28 20:29:15 Assert Failed: Rollback
error
> >
> > 172
> >
> >>>20:29:15 Informix Dynamic Server Version 7.30.UC3A 20:29:15 Who:
> >>>Session(11, root@sta154, 0, 0)
> >>> Thread(37, fast_rec, 0, 4)
> >>> File: rsextlog.c Line: 1412
> >>>20:29:15 Results: Log record (OLDRSAM:DELITEM) in log 47852, offset> >>>0xd4018 was not rolled back
> >>>20:29:15 Action: Use 'onlog' to view the transaction and repair> >
> > manually.
> >
> >>>20:30:48 See Also: /tmp/af.2502e8, shmem.2502e8.2 20:30:48 Error> >
> > writing
> >
> >>>'/tmp/shmem.2502e8.2' errno = 28 20:30:48 Assert Failed: Dynamic
Server
> >>>must abort 20:30:48 Informix Dynamic Server Version 7.30.UC3A 20:30:48
> >
> > Who:
> >
> >>>Session(11, root@sta154, 0, 0)
> >>> Thread(37, fast_rec, 0, 4)
> >>> File: rslog.c Line: 3152
> >>>20:30:48 Results: Dynamic Server must abort 20:30:48 Action:> >>>Reinitialize shared memory 20:30:48 stack trace for pid 10183 written
> >
> > to
> >
> >>>/tmp/af.2502e8 20:30:48 See Also: /tmp/af.2502e8, shmem.2502e8.3
> >
> > 20:30:48
> >
> >>>Error writing '/tmp/shmem.2502e8.3' errno = 28 20:30:48 rslog.c, line
> >
> > 3152,
> >
> >>>thread 37, proc id 10183, Dynamic Server must abort. 20:30:48 PANIC:
> >>>Attempting to bring system down
> >
> >
> >
> >
>
> --
>
> Atte,
>
> Jesus Antonio Santos Giraldo
> -----------------------------------
> jeansagi@myrealbox.com
> jeansagi@netscape.net
>
> sending to informix-list