ontape -r Problem
Posted in 2012
A user running IDS 7.31 on SLES11 couldn't complete 'ontape -r' of one nightly tape backup on a test machine: the restore wrote a few rootdbs pages then appeared to hang with no error, on multiple tapes/drives; other databases restored fine, and the suspect DB had fragmented tables. Suggestions: check the source message log and archecker for bad pages/assert failures during the archive, verify variable tape block size, try pipe/netcat-based backup, or use an external (cold/blocked) archive; John Miller noted v7 clears physical/logical logs very slowly at restore end, so it can look hung for hours (watch onstat -g iof). It was also confirmed v7 archives can't be restored into 11.70, so dbexport/dbimport or similar is the upgrade path. No confirmation of the actual cause or fix is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Backup & Restore, Installation, Setup & Upgrades, Storage & Space Management, Platform-Specific Issues
Hello Informix-Experts.
OS: SLES11-SP1 64-Bit
Kernel: Linux testserver 2.6.32.59-0.3-default #1 SMP 2012-04-27 11:14:44
+0200 x86_64 x86_64 x86_64 GNU/Linux
IDS: IBM Informix Dynamic Server Version 7.31.UD10X1
We have the following Problem:
In preparation for the Informix Upgrade to 11.70 we want to restore one of our
productive databases, which runs under 7.31.UD10X1, to test the Ifx-Upgrade.
We had build an testsystem, which has the same structure and configuration.
Here we run into a heavy problem:
Our "ontape -r" on the testsystem, which based on our nightly Backup from the
productive System, fails!
The Restore begin to write some pages in the rootdbs and then the restore is
not continuing.
We`ve received no error message, the page-writing stopps abruptly.
We have tested this with different tape-stations, different tapes and
different Hardware. Every time the same behaviour.
This problem occurs only on one DB-Restore.
We have some other databases, where the restore works without problems.
The difference between the Backups which are working and this problem-backup
is, that we`d made a table fragmentation on this DB, which could not be
restored now.
During the Backup we could not see any problems. Backup says "100 percent
done".
But when we do an "ontape -r" the restore writes some pages and that`s it. The
restore do not work further.
Have somebody experiences with restoring a 7.31er-DB with this strange
behaviour?
Could it be, that the table-fragmentation has corrupt the rootdbs?
We dropped the fragmented tables and test the restore, but it fails, too.
To all appearances the restore hangs in an infinite loop and we had at the
moment no choice to restore this DB.
Have somebody a hint or a workaround to handle this strange behaviour?
Every response is welcome.
Thanks for your efforts.
Sascha
The first thing I would do is to perform an external archive during some
quiet time. At least that way you will have an image that you know you can
restore. I don't remember if 7.31 had built-in support for external
archives (onmode -c block / onmode -c unblock). If not, then you will have
to shutdown the instance in order to take a save chunk archive using OS
tools or some archive utilility.
Once that's done, you could restore the external archive to your test
system and a) test the upgrade, and b) figure out what's wrong with the
production server. I would run onchecks to see if there is any damage to
the reserved pages or other root structures then I would call IBM and see
if they will take a look at this for you. Are you still under support?
Finally, I would plan on upgrading to 11.70 by building a clean empty
instance and import the data from the 7.31 instance either by
dbexport/dbimport (or similar tools like my myexport/myimport utility or
hploader) or by using ER since this server is an ER participant already.
This would get rid of any internal problems in the current server's
structures rather than carrying them forward.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Fri, Jul 13, 2012 at 3:27 AM, SASCHA KURATIS <
sascha.kuratis@westfleisch.de> wrote:
> Hello Informix-Experts.
>
> OS: SLES11-SP1 64-Bit
> Kernel: Linux testserver 2.6.32.59-0.3-default #1 SMP 2012-04-27 11:14:44
> +0200 x86_64 x86_64 x86_64 GNU/Linux
> IDS: IBM Informix Dynamic Server Version 7.31.UD10X1
>
> We have the following Problem:
>
> In preparation for the Informix Upgrade to 11.70 we want to restore one of
> our
> productive databases, which runs under 7.31.UD10X1, to test the
> Ifx-Upgrade.
> We had build an testsystem, which has the same structure and configuration.
> Here we run into a heavy problem:
> Our "ontape -r" on the testsystem, which based on our nightly Backup from
> the
> productive System, fails!
>
> The Restore begin to write some pages in the rootdbs and then the restore
> is
> not continuing.
> We`ve received no error message, the page-writing stopps abruptly.
> We have tested this with different tape-stations, different tapes and
> different Hardware. Every time the same behaviour.
>
> This problem occurs only on one DB-Restore.
> We have some other databases, where the restore works without problems.
>
> The difference between the Backups which are working and this
> problem-backup
> is, that we`d made a table fragmentation on this DB, which could not be
> restored now.
> During the Backup we could not see any problems. Backup says "100 percent
> done".
> But when we do an "ontape -r" the restore writes some pages and that`s it.
> The
> restore do not work further.
>
> Have somebody experiences with restoring a 7.31er-DB with this strange
> behaviour?
> Could it be, that the table-fragmentation has corrupt the rootdbs?
>
> We dropped the fragmented tables and test the restore, but it fails, too.
> To all appearances the restore hangs in an infinite loop and we had at the
> moment no choice to restore this DB.
>
> Have somebody a hint or a workaround to handle this strange behaviour?
>
> Every response is welcome.
>
> Thanks for your efforts.
>
> Sascha
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--14dae9340e1b9acf1a04c4b45ecc
I would look in the message log of the source where the level 0 was created
and see if any warnings/assert failures were written while the archive image
was being created. (Even though the message 'archive complete' is returned
there could have been some bad pages reported that you may have missed). If
you done this and no errors reported then try using archecker to check the
validity of the archive.
other things to check is the tape drive block size (should be variable but
check it anyway).
u can also try archiving and restoring using a fifo pipe (mismasch of rsh, dd
ontape -r, there are some docs out there that describe how to do , not as easyas on Family 11 but doable). But its probably much easier to go the external
route as mentioned in a previous post.
good luck.. and btw wait until friday the 13th is over.
Mark has some good suggestions below. The one think I would add is in=
version 7 informix had a very slow mechanism to clear the physical and
logical logs. This must be done at the end of a restore to ensure
pre-exiting or garbage data does not accidentally look like recover dat=
a.
Many customers aborted the restore because they thought the system was
hung. During this time it looks like no activity is running on the serv=
er.
Many things show no activity such as, tapes not moving, onstat -p does =
not
change. If you have a large physical log and logical logs this can ta=
ke
hours. If memory servers me, the best thing to look at would be an ons=
tat
-g iof. This is the only mechanism which would show activity.
Hope this helps,
John F. Miller III
STSM, Embedability Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
ids-bounces@iiug.org wrote on 07/13/2012 09:16:27 AM:
> From: "MARK JALKIEWICZ" <mark.jalkiewicz@verizon.net>
> To: ids@iiug.org
> Date: 07/13/2012 09:18 AM
> Subject: Re: ontape -r Problem [27597]
> Sent by: ids-bounces@iiug.org
>
> I would look in the message log of the source where the level 0 was
created
> and see if any warnings/assert failures were written while the archiv=
e
image
> was being created. (Even though the message 'archive complete' is
returned
> there could have been some bad pages reported that you may have misse=
d).
If
> you done this and no errors reported then try using archecker to chec=
k
the
> validity of the archive.
>
> other things to check is the tape drive block size (should be variabl=
e
but
> check it anyway).
>
> u can also try archiving and restoring using a fifo pipe (mismasch of=
rsh, dd
> ontape -r, there are some docs out there that describe how to do ,> not as easy
> as on Family 11 but doable). But its probably much easier to go the
external
> route as mentioned in a previous post.
>
> good luck.. and btw wait until friday the 13th is over.
>
>
>
***********************************************************************=
********
> Forum Note: Use "Reply" to post a response in the discussion forum.=
>=
On 13/07/2012 17:16, MARK JALKIEWICZ wrote:
>
> u can also try archiving and restoring using a fifo pipe (mismasch of rsh, dd
> ontape -r, there are some docs out there that describe how to do , not aseasy
> as on Family 11 but doable).
It's very easy - netcat / socat are you friends for this
--
Clive
Vouching for myself, an ifx non-expert, could it be that 7.3 backup's are not compatible with 11.70 restore's? If they're not compatible, would an ETL be a viable alternative?
You are correct, backup must be restore with the same version.
You might want to look at the tool dbexport/dbimport which comes
with both version 7 and version 11 and maybe used to transfer a
database from one version and OS to a different version and/or
OS.
John F. Miller III
STSM, Embedability Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
ids-bounces@iiug.org wrote on 07/13/2012 04:42:10 PM:
> From: "FRANK J. COMPUTER" <frank_in_pr@hotmail.com>
> To: ids@iiug.org
> Date: 07/13/2012 04:43 PM
> Subject: Re: ontape -r Problem [27602]
> Sent by: ids-bounces@iiug.org
>
> Vouching for myself, an ifx non-expert, could it be that 7.3 backup's are
not
> compatible with 11.70 restore's? If they're not compatible, would anETL
be a
> viable alternative?
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Wow, I was just using my common sense on this one, based on the big leap from 7.3 to 11.70. I just recently started tinkering with 11.70.TC4DE (Win), I've been stuck in the SE-4.10.DC1(DOS) world too long :-O