Backup is Blocked:ARCHIVE when temp space being us
Posted in 2016
Topics: Backup & Restore, Performance & Tuning, Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Versions, Editions & End-of-Life
IDS 9.30HC5 on HP 11i therefore no support :-(
We had a weird hang on our live system last night.
There was a 4gl running which was doing a high volume of writes to a temp
table (created explicitly, so would have been in DBSPACETEMP) and had been
running for several hours, expected to run at least overnight; our instance
backup (ontape -s -L 0, nothing special) started under cron control at 8pm.
NB DBSPACETEMP lists only one dbspace; the relevant database uses Buffered
logging.
We came in this morning and the whole system was hung. The backup had done
nothing (although ps -ef showed 'ontape -s -L 0' was running) but onstats
showed "Blocked:ARCHIVE" in the info line. A simple 4GL running a
single-table query (a high-cost sequential scan, unfortunately, that can't
really be optimised as it mainly filters on fields with only two or three
values in a table of 49,000,000 rows, and also has an OR clause...) under
our job scheduler had hung, as also had the big 4GL.
After killing a couple of processes it all burst back into life, but the
system has been incredibly sluggish all day (we're doing an instance
restart this evening).
I know that a backup switches all transaction activity into DBSPACETEMP for
the duration; would this have been somehow affected by the 4GL which was
already running and using DBSPACETEMP?
--94eb2c1246bac7882f0530d798af
Hi,
pls post if the backup has really started (should be in online.log), find an
entry like "Level 0 Archive started on rootdbs....."
If it has not started, check the output of the ontape (should be in your cron
mail, why it was blocking.
Any other hints in online.log ? Any af dumps ?
There might be a shared memory (I assume ontape is running via shared memory
access) congestion,
which should present random values when calling onstat.
In this case an instance reboot is necessary, you cannot repair the shared
memory in another ways.
Are your log files being backed up (on tape, I suspect) ? check onstat -l
Or is this process blocked, too ?
Marcus Haarmann
----- Ursprüngliche Mail -----
Von: "Malc P" <malcrp@gmail.com>
An: ids@iiug.org
Gesendet: Dienstag, 19. April 2016 16:50:35
Betreff: Backup is Blocked:ARCHIVE when temp space bein.... [36992]
IDS 9.30HC5 on HP 11i therefore no support :-(
We had a weird hang on our live system last night.
There was a 4gl running which was doing a high volume of writes to a temp
table (created explicitly, so would have been in DBSPACETEMP) and had been
running for several hours, expected to run at least overnight; our instance
backup (ontape -s -L 0, nothing special) started under cron control at 8pm.
NB DBSPACETEMP lists only one dbspace; the relevant database uses Buffered
logging.
We came in this morning and the whole system was hung. The backup had done
nothing (although ps -ef showed 'ontape -s -L 0' was running) but onstats
showed "Blocked:ARCHIVE" in the info line. A simple 4GL running a
single-table query (a high-cost sequential scan, unfortunately, that can't
really be optimised as it mainly filters on fields with only two or three
values in a table of 49,000,000 rows, and also has an OR clause...) under
our job scheduler had hung, as also had the big 4GL.
After killing a couple of processes it all burst back into life, but the
system has been incredibly sluggish all day (we're doing an instance
restart this evening).
I know that a backup switches all transaction activity into DBSPACETEMP for
the duration; would this have been somehow affected by the 4GL which was
already running and using DBSPACETEMP?
--94eb2c1246bac7882f0530d798af
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
The backup doesn't shift transaction activity to DBSPACETEMP. It just saves
the before image of all the changed pages (first time each page is changed)
in DBSPACETEMP.
But it does this after the backup effectively starts and it seems that your
system got hang at the checkpoint which is triggered by the backup start.
A possible reason for that would be that some task got stuck into an
exclusive section (X flag in onstat -u). When you kiiled some things it may
have released the checkpoint and the backup probably started.
It's impossible to tell without further data... If the backup has not free
space in DBSPACETEMP to save the before image of the pages being changed it
should exit with an error.
Furthermore.... there are good reasons to be on supported versions.
It's not clear if you mean you're paying support but can't have it because
the version is not supported (that would be even worse) or if you're not
paying support.
Regards.
On Tue, Apr 19, 2016 at 3:50 PM, Malc P <malcrp@gmail.com> wrote:
> IDS 9.30HC5 on HP 11i therefore no support :-(
>
> We had a weird hang on our live system last night.
> There was a 4gl running which was doing a high volume of writes to a temp
> table (created explicitly, so would have been in DBSPACETEMP) and had been
> running for several hours, expected to run at least overnight; our instance
> backup (ontape -s -L 0, nothing special) started under cron control at 8pm.
> NB DBSPACETEMP lists only one dbspace; the relevant database uses Buffered
> logging.
>
> We came in this morning and the whole system was hung. The backup had done
> nothing (although ps -ef showed 'ontape -s -L 0' was running) but onstats
> showed "Blocked:ARCHIVE" in the info line. A simple 4GL running a
> single-table query (a high-cost sequential scan, unfortunately, that can't
> really be optimised as it mainly filters on fields with only two or three
> values in a table of 49,000,000 rows, and also has an OR clause...) under
> our job scheduler had hung, as also had the big 4GL.
>
> After killing a couple of processes it all burst back into life, but the
> system has been incredibly sluggish all day (we're doing an instance
> restart this evening).
>
> I know that a backup switches all transaction activity into DBSPACETEMP for
> the duration; would this have been somehow affected by the 4GL which was
> already running and using DBSPACETEMP?
>
> --94eb2c1246bac7882f0530d798af
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--001a113f2d2af830bd0530d856b3
No, it hadn't even got to the "Archive on rootdbs...etc started" message.
Nothing in onstat -m (and no "Fuzzy Checkpoint completed" messages, and no
af's
On Tue, Apr 19, 2016 at 4:39 PM, Marcus Haarmann <marcus.haarmann@midoco.de>
wrote:
> Hi,
>
> pls post if the backup has really started (should be in online.log), find
> an
> entry like "Level 0 Archive started on rootdbs....."
> If it has not started, check the output of the ontape (should be in your
> cron
> mail, why it was blocking.
> Any other hints in online.log ? Any af dumps ?
>
> There might be a shared memory (I assume ontape is running via shared
> memory
> access) congestion,
> which should present random values when calling onstat.
> In this case an instance reboot is necessary, you cannot repair the shared
> memory in another ways.
>
> Are your log files being backed up (on tape, I suspect) ? check onstat -l
> Or is this process blocked, too ?
>
> Marcus Haarmann
>
> ----- Ursprüngliche Mail -----
>
> Von: "Malc P" <malcrp@gmail.com>
> An: ids@iiug.org
> Gesendet: Dienstag, 19. April 2016 16:50:35
> Betreff: Backup is Blocked:ARCHIVE when temp space bein.... [36992]
>
> IDS 9.30HC5 on HP 11i therefore no support :-(
>
> We had a weird hang on our live system last night.
> There was a 4gl running which was doing a high volume of writes to a temp
> table (created explicitly, so would have been in DBSPACETEMP) and had been
> running for several hours, expected to run at least overnight; our instance
> backup (ontape -s -L 0, nothing special) started under cron control at 8pm.
> NB DBSPACETEMP lists only one dbspace; the relevant database uses Buffered
> logging.
>
> We came in this morning and the whole system was hung. The backup had done
> nothing (although ps -ef showed 'ontape -s -L 0' was running) but onstats
> showed "Blocked:ARCHIVE" in the info line. A simple 4GL running a
> single-table query (a high-cost sequential scan, unfortunately, that can't
> really be optimised as it mainly filters on fields with only two or three
> values in a table of 49,000,000 rows, and also has an OR clause...) under
> our job scheduler had hung, as also had the big 4GL.
>
> After killing a couple of processes it all burst back into life, but the
> system has been incredibly sluggish all day (we're doing an instance
> restart this evening).
>
> I know that a backup switches all transaction activity into DBSPACETEMP for
> the duration; would this have been somehow affected by the 4GL which was
> already running and using DBSPACETEMP?
>
> --94eb2c1246bac7882f0530d798af
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--001a11430efeb98a860530d89568
> > > Furthermore.... there are good reasons to be on supported versions. > It's not clear if you mean you're paying support but can't have it because > the version is not supported (that would be even worse) or if you're not > paying support. > > Ah thereby hangs a tale (I've covered it in previous posts) - let's just say it was a management decision that I strongly disagreed with back in 2001... --94eb2c122d28ad7a970530db338c