cold restore failed
Posted in 2003
Topics: Backup & Restore, Storage & Space Management, Error Codes & Troubleshooting, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
HP-UX 11.11
IDS 7.31 FD5W6
We experienced problems after doing a cold restore on a different
server
We did a level-0 archive ( whole backup ) with onbar ( omniback II - now
called dataprotector) on system1 and
went to restore this on system2. Disk structure is the same , but I use links
to point to the raw-devices
so should not make a difference. All of the proper symlinks
were there, and pointed at raw devices that were big enough.
The physical restore went great, then the logical restore went
great, but when the engine started to come up it began yelling
about how the temp dbspaces (all 2 that we have) failed sanity
checks.
06:34:09 Physical Restore of rootdbs, dbs_llog, dbs_plog, dbs_assistente,
dbs_conger, dbs_dbssa, dbs_export, dbs_gecex, dbs_logix,
dbs_replic, dbs_sapes, dbs_sca, dbs_sci, dbs_spef Completed.
06:34:09 Checkpoint Completed: duration was 0 seconds.
06:34:09 Checkpoint loguniq 1859, logpos 0xeb6018
06:34:09 Logical Recovery Started.
06:34:09 Checkpoint Completed: duration was 0 seconds.
06:34:09 Checkpoint loguniq 1859, logpos 0xeb6018
06:34:09 Start Logical Recovery - Start Log 1859, End Log ?
06:34:09 Starting Log Position - 1859 0xeb6018
06:34:09 Clearing the physical and logical logs has started
06:35:59 Cleared 8480 MB of the physical and logical logs in 111 seconds
06:36:01 Logical Recovery ABORTED.
Aborted by client.
06:36:03 Assert Failed: Logical Recovery ABORTED.
Dynamic Server must abort
06:36:03 Informix Dynamic Server Version 7.31.FD5W6
06:36:03 Who: Session(21, informix@cp52, 17054, 1670826432)
Thread(34, ontape, c0000000639389f0, 3)
File: rslgr.c Line: 1103
06:36:03 stack trace for pid 17539 written to /backup/af/af.40acf83
06:36:14 See Also: /backup/af/af.40acf83
06:36:14 rslgr.c, line 1103, thread 34, proc id 17539, Logical Recovery
ABORTED.
Dynamic Server must abort.
06:36:14 PANIC: Attempting to bring system down
10:32:19 Segment locked: addr=0x463f3000, size=490192896
Sat Feb 8 10:32:20 2003
10:32:20 Event alarms enabled. ALARMPROG = '/informix/etc/log_full_ADM.sh'
10:32:25 DR: DRAUTO is 0 (Off)
10:32:25 Dynamically allocated new message shared memory segment (size 1128KB)
10:32:26 Informix Dynamic Server Version 7.31.FD5W6 Software Serial Number
AAC#J867926
10:32:26 HPUX Version B.11.11 -> Using flag/select style KAIO
10:32:26 HP KAIO concurrent requests changed from 1000 to 2300
10:32:28 Assert Failed: chunk failed sanity check
10:32:28 Informix Dynamic Server Version 7.31.FD5W6
10:32:28 Who: Session(1, informix@cp52, 0, 1669492776)
Thread(7, main_loop(), c0000000637ef028, 4)
File: rspartn.c Line: 7362
10:32:28 Results: Chunk 20 is being taken OFFLINE.
10:32:28 Action: Restore chunk from archive. If this is a temporary dbspace
chunk, drop and add the dbspace to enable it.
10:32:28 stack trace for pid 12398 written to /backup/af/af.3ef06ec
10:32:29 See Also: /backup/af/af.3ef06ec
10:32:29 I/O error, Primary Chunk '/informix/db_adm/ck_temp_adm_01' -- Offline
(sanity)
10:32:29 Assert Failed: chunk failed sanity check
10:32:29 Informix Dynamic Server Version 7.31.FD5W6
10:32:29 Who: Session(1, informix@cp52, 0, 1669492776)
Thread(7, main_loop(), c0000000637ef028, 4)
File: rspartn.c Line: 7362
10:32:29 Results: Chunk 21 is being taken OFFLINE.
10:32:29 Action: Restore chunk from archive. If this is a temporary dbspace
chunk, drop and add the dbspace to enable it.
10:32:29 stack trace for pid 12398 written to /backup/af/af.3ef06ec
10:32:30 See Also: /backup/af/af.3ef06ec
10:32:30 I/O error, Primary Chunk '/informix/db_adm/ck_temp_adm_02' -- Offline
(sanity)
10:32:30 Informix Dynamic Server Initialized -- Shared Memory Initialized.
10:32:30 Warning: Invalid dbspace 'dbs_temp' listed in DBSPACETEMP.
10:32:30 Physical Recovery Started.
10:32:31 Physical Recovery Complete: 0 Pages Restored.
10:32:31 Logical Recovery Started.
10:32:33 Logical Recovery Complete.
0 Committed, 0 Rolled Back, 0 Open, 0 Bad Locks
10:32:34 Informix Dynamic Server Stopped.
10:32:34 mt_shm_remove: WARNING: may not have removed all/correct segments
10:33:41 Segment locked: addr=0x463f3000, size=490192896
Ricardo Manzo wrote:
> HP-UX 11.11 IDS 7.31 FD5W6
>
> We experienced problems after doing a cold restore on a different
> server
>
> We did a level-0 archive ( whole backup ) with onbar ( omniback II - now
called dataprotector) on system1 and
> went to restore this on system2. Disk structure is the same , but I use
links to point to the raw-devices
> so should not make a difference. All of the proper symlinks
> were there, and pointed at raw devices that were big enough.
>
> The physical restore went great, then the logical restore went
> great, but when the engine started to come up it began yelling
> about how the temp dbspaces (all 2 that we have) failed sanity
> checks.
>
> 06:34:09 Physical Restore of rootdbs, dbs_llog, dbs_plog, dbs_assistente,
dbs_conger, dbs_dbssa, dbs_export, dbs_gecex, dbs_logix,
> dbs_replic, dbs_sapes, dbs_sca, dbs_sci, dbs_spef Completed.
> 06:34:09 Checkpoint Completed: duration was 0 seconds.
> 06:34:09 Checkpoint loguniq 1859, logpos 0xeb6018
>
> 06:34:09 Logical Recovery Started.
> 06:34:09 Checkpoint Completed: duration was 0 seconds.
> 06:34:09 Checkpoint loguniq 1859, logpos 0xeb6018
>
> 06:34:09 Start Logical Recovery - Start Log 1859, End Log ?
> 06:34:09 Starting Log Position - 1859 0xeb6018
> 06:34:09 Clearing the physical and logical logs has started
> 06:35:59 Cleared 8480 MB of the physical and logical logs in 111 seconds
> 06:36:01 Logical Recovery ABORTED.
> Aborted by client.
It looks more like the restore never finished as it was aborted. Then
when you try to bring the server up it complains about reading chunks
that have not been restored yet. The temp dbspaces are not actually
restored but marked as consistent at the end of the restore process.
I would find out what aborted the restore, then run the restore again.
Cheers,
--
Mark.
+----------------------------------------------------------+-----------+
| Mark D. Stock mailto:mdstock@MydasSolutions.com |//////// /|
| Mydas Solutions Ltd http://MydasSolutions.com |///// / //|
| +-----------------------------------+//// / ///|
| |We value your comments, which have |/// / ////|
| |been recorded and automatically |// / /////|
| |emailed back to us for our records.|/ ////////|
+----------------------+-----------------------------------+-----------+
What
version of omniback ?
onbar cold restore failed with OMNIBACK 4.10 on HPUX11.0. When Logical
restore stage, if logs existed in two tape drives then restore failed.
We tested on development and opened case with HP on this issue.
I would suggest to try the following.
When restore failed, I repeated restored again with the following and
restored all logs.
onbar -RESTART
thanks
Ravi Yarlagadda
-----Original Message-----
From: Mark D. Stock [mailto:mdstock@mydassolutions.com]
Sent: Saturday, February 08, 2003 9:34 AM
To: ids@iiug.org
Subject: Re: cold restore failed [305]
Ricardo Manzo wrote:
> HP-UX 11.11 IDS 7.31 FD5W6
>
> We experienced problems after doing a cold restore on a different
> server
>
> We did a level-0 archive ( whole backup ) with onbar ( omniback II - now
called dataprotector) on system1 and
> went to restore this on system2. Disk structure is the same , but I use
links to point to the raw-devices
> so should not make a difference. All of the proper symlinks
> were there, and pointed at raw devices that were big enough.
>
> The physical restore went great, then the logical restore went
> great, but when the engine started to come up it began yelling
> about how the temp dbspaces (all 2 that we have) failed sanity
> checks.
>
> 06:34:09 Physical Restore of rootdbs, dbs_llog, dbs_plog, dbs_assistente,
dbs_conger, dbs_dbssa, dbs_export, dbs_gecex, dbs_logix,
> dbs_replic, dbs_sapes, dbs_sca, dbs_sci, dbs_spef Completed.
> 06:34:09 Checkpoint Completed: duration was 0 seconds.
> 06:34:09 Checkpoint loguniq 1859, logpos 0xeb6018
>
> 06:34:09 Logical Recovery Started.
> 06:34:09 Checkpoint Completed: duration was 0 seconds.
> 06:34:09 Checkpoint loguniq 1859, logpos 0xeb6018
>
> 06:34:09 Start Logical Recovery - Start Log 1859, End Log ?
> 06:34:09 Starting Log Position - 1859 0xeb6018
> 06:34:09 Clearing the physical and logical logs has started
> 06:35:59 Cleared 8480 MB of the physical and logical logs in 111 seconds
> 06:36:01 Logical Recovery ABORTED.
> Aborted by client.
It looks more like the restore never finished as it was aborted. Then
when you try to bring the server up it complains about reading chunks
that have not been restored yet. The temp dbspaces are not actually
restored but marked as consistent at the end of the restore process.
I would find out what aborted the restore, then run the restore again.
Cheers,
--
Mark.
+----------------------------------------------------------+-----------+
| Mark D. Stock mailto:mdstock@MydasSolutions.com |//////// /|
| Mydas Solutions Ltd http://MydasSolutions.com |///// / //|
| +-----------------------------------+//// / ///|
| |We value your comments, which have |/// / ////|
| |been recorded and automatically |// / /////|
| |emailed back to us for our records.|/ ////////|
+----------------------+-----------------------------------+-----------+