IDS v11.50 Ontape Issue
Posted in 2009
A 1.4 TB IDS 11.50.FC3 instance on Solaris 10 had its weekly ontape level-0 backup (TAPEDEV=STDIO, writing to SAN disk on another server) suddenly fail with "Archive ... ABORTED / Aborted by client" and no further detail in the online log. Suggestions included checking backup space, writing to a file instead of STDIO for a better error message, raising NUMCPUVPS, dropping the quiescent step or using ON-Bar, and capturing ontape output to a log. The actual cause was an inconsistent file system on the SAN; running fsck and remounting fixed it, and fsck checks were added to weekly maintenance.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Backup & Restore, Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Server Administration, Security, Permissions & Auditing
System Information
IDS 11.50.FC3
SQL Dev. 32Bit 7.32.UC3
ESQL/C Client SDK 64Bit
Solaris 10 Intel platform
Database size 1.4 Terabyte
We have been on our new system since March 15, 2009 and have been running
level 0's once a week with no issues. The ontape process is executed via a
shell script that sets environment variables and paths, takes the system to
quiescent mode and executes the ontape command for the Level 0.
The L0 writes to standard out, the onconfig TAPEDEV is set to STDIO.
The backup is written to a separate server with SAN disks and is sized for 2
Terabyte. The level 0 normally takes approximately 5.5 hours to run and starts
at 15:00:00.
The following is the error in the message log.
20:27:20 Archive on rootdbs, dbimageblob, llog, plog, dbaudit, dbcasefile,
dbemploy, dbcasedocs, dbimages, dblogfile, dbh
istory, dbdochist, dbmanage, dbmaintain, dbqueue, dbstatic, dbdisp ABORTED.
20:27:20 Aborted by client.
20:27:21 Started 1 B-tree scanners.
20:27:21 B-tree scanner threshold set at 5000.
20:27:21 B-tree scanner range scan size set to -1.
20:27:21 B-tree scanner ALICE mode set to 6.
20:27:21 B-tree scanner index compression level set to med.
20:27:21 SCHAPI: Started dbScheduler thread.
20:27:21 SCHAPI: Started 2 dbWorker threads.
20:27:22 On-Line Mode
Can someone please suggest possible reasons for this type of error?
Based on your email,
1. your database size is 1.4 Terabyte and your backup SAN is 2 Terabyte. do
you think you may have a backup space issue? check that out too.
2. you don't need to bring your instance to quiescent mode to run ontape
backup; however you want to do it, maybe something is wrong with the backup
script; it looks like something terminates the ontape process and immediately
bring the instance from quiescent to online. It is understood that ontape
script had run fine since March which doesn't it will run with errorfree
forever.
________________________________
From: KATHY DUCHENE <kathy.duchene@state.mn.us>
To: ids@iiug.org
Sent: Monday, June 15, 2009 1:33:10 PM
Subject: IDS v11.50 Ontape Issue [16031]
System Information
IDS 11.50.FC3
SQL Dev. 32Bit 7.32.UC3
ESQL/C Client SDK 64Bit
Solaris 10 Intel platform
Database size 1.4 Terabyte
We have been on our new system since March 15, 2009 and have been running
level 0's once a week with no issues. The ontape process is executed via a
shell script that sets environment variables and paths, takes the system to
quiescent mode and executes the ontape command for the Level 0.
The L0 writes to standard out, the onconfig TAPEDEV is set to STDIO.
The backup is written to a separate server with SAN disks and is sized for 2
Terabyte. The level 0 normally takes approximately 5.5 hours to run and starts
at 15:00:00.
The following is the error in the message log.
20:27:20 Archive on rootdbs, dbimageblob, llog, plog, dbaudit, dbcasefile,
dbemploy, dbcasedocs, dbimages, dblogfile, dbh
istory, dbdochist, dbmanage, dbmaintain, dbqueue, dbstatic, dbdisp ABORTED.
20:27:20 Aborted by client.
20:27:21 Started 1 B-tree scanners.
20:27:21 B-tree scanner threshold set at 5000.
20:27:21 B-tree scanner range scan size set to -1.
20:27:21 B-tree scanner ALICE mode set to 6.
20:27:21 B-tree scanner index compression level set to med.
20:27:21 SCHAPI: Started dbScheduler thread.
20:27:21 SCHAPI: Started 2 dbWorker threads.
20:27:22 On-Line Mode
Can someone please suggest possible reasons for this type of error?
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
The back up script will bring the instance online whether the level 0 fails or completes successfully. That is how the system was brought to online mode.
Ontape for 1.5 terabytes? Why not onbar? And why go quiescent?
I had to debug something like this the other day and never got time to write
it up. So I forget what the problem was - oh but I do remember how I fixed it.
1 - change the STDIO and go to a file - see if that completes, at very least
you'll get a better error message.
2 - I got a disk full error - which was nonsensical, there was plenty of space
in TEMP and on the receiving device.
3 - just for grins I changed the NUMCPUVPS from 2 (on a 4-way) to 4. Suddenly
it worked, we have not been able to break it since.
j.
On Jun 15, 2009, KATHY DUCHENE <kathy.duchene@state.mn.us> wrote:
System Information
IDS 11.50.FC3
SQL Dev. 32Bit 7.32.UC3
ESQL/C Client SDK 64Bit
Solaris 10 Intel platform
Database size 1.4 Terabyte
We have been on our new system since March 15, 2009 and have been running
level 0's once a week with no issues. The ontape process is executed via a
shell script that sets environment variables and paths, takes the system to
quiescent mode and executes the ontape command for the Level 0.
The L0 writes to standard out, the onconfig TAPEDEV is set to STDIO.
The backup is written to a separate server with SAN disks and is sized for 2
Terabyte. The level 0 normally takes approximately 5.5 hours to run and starts
at 15:00:00.
The following is the error in the message log.
20:27:20 Archive on rootdbs, dbimageblob, llog, plog, dbaudit, dbcasefile,
dbemploy, dbcasedocs, dbimages, dblogfile, dbh
istory, dbdochist, dbmanage, dbmaintain, dbqueue, dbstatic, dbdisp ABORTED.
20:27:20 Aborted by client.
20:27:21 Started 1 B-tree scanners.
20:27:21 B-tree scanner threshold set at 5000.
20:27:21 B-tree scanner range scan size set to -1.
20:27:21 B-tree scanner ALICE mode set to 6.
20:27:21 B-tree scanner index compression level set to med.
20:27:21 SCHAPI: Started dbScheduler thread.
20:27:21 SCHAPI: Started 2 dbWorker threads.
20:27:22 On-Line Mode
Can someone please suggest possible reasons for this type of error?
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
It looks like someone killed the shell script running the ontape or the
ontape itself. If this is being run through a terminal emulator, perhaps
the connection to the server timed out or there was a network error. There
is nothing in the online log to indicate why it failed. I would run it
again manually and if it doesn't fail, modify the script to capture all
output from ontape to a file so you can look at it later.
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
On Mon, Jun 15, 2009 at 1:33 PM, KATHY DUCHENE
<kathy.duchene@state.mn.us>wrote:
> System Information
> IDS 11.50.FC3
> SQL Dev. 32Bit 7.32.UC3
> ESQL/C Client SDK 64Bit
> Solaris 10 Intel platform
> Database size 1.4 Terabyte
>
> We have been on our new system since March 15, 2009 and have been running
> level 0's once a week with no issues. The ontape process is executed via a
> shell script that sets environment variables and paths, takes the system to
> quiescent mode and executes the ontape command for the Level 0.
>
> The L0 writes to standard out, the onconfig TAPEDEV is set to STDIO.
>
> The backup is written to a separate server with SAN disks and is sized for
> 2
> Terabyte. The level 0 normally takes approximately 5.5 hours to run and
> starts
> at 15:00:00.
>
> The following is the error in the message log.
> 20:27:20 Archive on rootdbs, dbimageblob, llog, plog, dbaudit, dbcasefile,
> dbemploy, dbcasedocs, dbimages, dblogfile, dbh
> istory, dbdochist, dbmanage, dbmaintain, dbqueue, dbstatic, dbdisp ABORTED.
> 20:27:20 Aborted by client.
> 20:27:21 Started 1 B-tree scanners.
> 20:27:21 B-tree scanner threshold set at 5000.
> 20:27:21 B-tree scanner range scan size set to -1.
> 20:27:21 B-tree scanner ALICE mode set to 6.
> 20:27:21 B-tree scanner index compression level set to med.
> 20:27:21 SCHAPI: Started dbScheduler thread.
> 20:27:21 SCHAPI: Started 2 dbWorker threads.
> 20:27:22 On-Line Mode
>
> Can someone please suggest possible reasons for this type of error?
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--0016e6d27c902f0005046c6aea0c
Thank you for the responses. It turned out to be an inconsistent file system condition with the SAN. We ran fsck on the SAN and remounted. This cleared up the issue. This will be now become part of a weekly maintenance program to ensure there are no inconsistencies with the SAN. This issue is now closed, thank you again.