Using ISM for a DR test; fails to find a saveset (that is definitely there!)
Posted in 2012
Topics: Backup & Restore, Storage & Space Management, Platform-Specific Issues
IDS9.30HC5, HP-UX 11i
We're running DR testing and have a backup of type "file" we're trying to load onto the DR box.
I've done an ism_catalog -recreate_from and it's reported that all savesets have been correctly recovered; the device is correctly mounted and shows in the ism_show -devices list correctly.
When I run onbar -r -w -p to start the restore, I get the following in bar_act.log:
2012-05-21 12:46:55 5928 5925 /opt/informix/bin/onbar_d -r -w -p
2012-05-21 12:46:55 5928 5925 Successfully connected to Storage Manager.
2012-05-21 12:46:55 5928 5925 Begin cold level 0 restore rootdbs (Storage Mana
ger copy ID: 131678 0).
2012-05-21 12:47:54 5928 5925 Completed cold level 0 restore rootdbs.
2012-05-21 12:47:55 5928 5925 Begin cold level 0 restore loglog (Storage Manag
er copy ID: 131784 0).
2012-05-21 12:47:56 5928 5925 Completed cold level 0 restore loglog.
2012-05-21 12:47:56 5928 5925 Begin cold level 0 restore physlog (Storage Mana
ger copy ID: 131783 0).
2012-05-21 12:47:57 5928 5925 Completed cold level 0 restore physlog.
2012-05-21 12:47:58 5928 5925 Begin cold level 0 restore act_qblobs1 (Storage
Manager copy ID: 131679 0).
2012-05-21 12:47:59 5928 5925 Completed cold level 0 restore act_qblobs1.
2012-05-21 12:47:59 5928 5925 XBSA Error: (BSAGetObject) A system error occurr
ed. Aborting XBSA session.
So the critical dbspaces have recovered ok and the first of our regular dbspaces ditto.
I had a look at the xbsa log file ${INFORMIXDIR}/ism/applogs/xbsa.messages and it says:
XBSA-1.0.1 ISM.2.20.HC1.114 5928 Mon May 21 12:47:59 2012 _nwbsa_open_re
store_session: BSA_RC_ABORT_SYSTEM_ERROR System detected error due to cannot fin
d continued save set for ssid 131685. Operation aborted
...but in the output from the catalog recovery it shows:
scanner: ssid 131684: continued on ssid 131685
scanner: ssid 131684: 1611 MB (of 2000 MB, 1 file(s)) written
scanner: ssid 131685: continued on ssid 131686
scanner: ssid 131685: 2000 MB, 1 file(s)
scanner: ssid 131686: continued on ssid 131687
scanner: ssid 131686: 1417 MB (of 2000 MB, 1 file(s)) written
...so 131685 is continued on 131686, and no error is reported. In the backup path, ls -l shows:
ls -l 13168[456].0
-rw-r--r-- 1 root sys 2058059776 May 4 20:35 131684.0
-rw-r--r-- 1 root sys 2058059776 May 4 20:36 131685.0
-rw-r--r-- 1 root sys 2058059776 May 4 20:38 131686.0
...so the file is there ok.
Corruption? Some other problem?
Any help appreciated as ever...
Hi, ISM can have problems when the files in the directory are not in the right order, that is the order they were initially written in i.e. save set id order. This can happen when you copy the directory, locally or from host to host, and using rsync can do it too. The problem shows up in scanning. So if you haven't copied or changed the files since you ran scanner, then its most likely not this problem. Otherwise, create a new directory, copy the files in alpha/numerical order, then switch the two directories. Try to restore that dbspace again and magically ISM will find the file. HTH, Jason
Jason, thanks - I'll bear that in mind but what then happened is a) the path to the Log logs was unmounted so I mounted it up and b) I ran "ism_catalog -recover" pointing at the bootstrap saveset, reran onbar and it worked...most likely it was the recover that fixed it; the DR boxes have been decommissioned now but we're repeating the exercise in a couple of weeks so I'll have some useful ideas to go on (including the re-ordering you suggested).
More anon, and thanks!