performance of ontape slowed down by factor 4
Posted in 1999
Ontape backup performance on IDS 7.13.UC1 degraded from 9 hours to 43 hours without known cause. Respondents suggested an internal Informix algorithm triggered by database volume that causes additional processing during archives, and noted ontape had bugs unfixed until version 7.14. Recommendation: upgrade to 7.30 or later.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Backup & Restore, Performance & Tuning, Storage & Space Management, Versions, Editions & End-of-Life
We have a very big database under IDS 7.13.UC1 and Reliant UNIX 5.43.
The database has around 27 GByte pictures in BLOBS (in BLOBSPACE),
each picture is around 14 KByte in size, together with some
additional information in the table like picture number etc.
The time for a complete backup with "ontape -s -L 0" to a SCSI tape
supporting a transfer of 5 MByte per second enlarged dramatically
from 9h to 43h without any known reason.
I checked the system calls done by the ontape and there are two
semops (to the server) and one write(2) call to the tape with
65536 bytes per second. A test with dd(1) reading with ibs=65536
from disk and writing with obs=65536 to the same physical tape
produces *many* read/write per second.
So it seems that the IDS (which is idle at this moment) does not
return the data fast enough. My questions are:
How is ontape fetching the data from ISD? In structured form or
simply page by page if the page is marked used?
What could be the reason for IDS slowing down so dramatically?
Thanks in advance.
matthias
--
firm: matthias.apitz@sisis.de [voc:+49 89 61308 351, fax: +49 89 61308 188]
priv: guru@thias.muc.de
WWW: http://www.sisis.de/~guru/
Until just recently I worked for a company that ran 7.13 on Seimens-Pyramid
DCOSx (which along with SINIX became Reliant UNIX) I can give you a little info
that may help figure out your problem, we never did. One if you are backing
the db up while it is online and active this can greatly slow things down
depending on the amount of activity since Informix has to go through the
additional steps of checking the times stamps on any pages the engine changes.
If necessary ontape will store the unchanged copy of the page in temp space
which can result in the amount of temp space used increasing dramatically.
More notably, and possibly more helpful to you, we noticed the following very
strange behaviour. For reasons I forget we went about 3 or 4 weeks without
doing an archive of any kind (we never did anything but level 0 archives) on
one of our machines (200+GB db - no blobs though). When we finally started
this archive it was running extremely slowly. After about 6 hours the archive
died (again reason forgotten but probably some kind of tape error). We
restarted the archive and it "flew" through the spaces it had archived in the 6
hr aborted run but when it got to the spaces it had not archived in 3 weeks or
so it slowed to a crawl. Needless to say this caught our attention...Why would
the time between level 0 archives have an impact on the time it takes to do a
level 0? We did some testing and verified this behaviour., We also verified
on NCR's MP-RAS UNIX running 7.13 Having worked on all four (SINIX, RELIANT,
DCOSx &MP-RAS) I can tell you they all pretty muhc straight Bell Labs SVR4
ports so this did not really surprise me. I don't recall if we determined if
the key factor was actual time passed or a given amount of work/db changes
being performed between archives.....I just called an old aquaintance at the
company I left and he said that it turned out to be a designed in "feature".
There is an internal algorythym that dictates/triggers this behaviour and it is
driven more (maybe completely) by volume than time. Apparently it triggers
the engine to do more than just dump the pages to tape during an archive.
Informix should be able to give a much clearer explanation than I have.
Hope this helps
Chris Carter
Systems Engineer - Unix Platforms.
Matthias Apitz wrote:
> We have a very big database under IDS 7.13.UC1 and Reliant UNIX 5.43.
> The database has around 27 GByte pictures in BLOBS (in BLOBSPACE),
> each picture is around 14 KByte in size, together with some
> additional information in the table like picture number etc.
>
> The time for a complete backup with "ontape -s -L 0" to a SCSI tape
> supporting a transfer of 5 MByte per second enlarged dramatically
> from 9h to 43h without any known reason.
>
> I checked the system calls done by the ontape and there are two
> semops (to the server) and one write(2) call to the tape with
> 65536 bytes per second. A test with dd(1) reading with ibs=65536
> from disk and writing with obs=65536 to the same physical tape
> produces *many* read/write per second.
>
> So it seems that the IDS (which is idle at this moment) does not
> return the data fast enough. My questions are:
>
> How is ontape fetching the data from ISD? In structured form or
> simply page by page if the page is marked used?
>
> What could be the reason for IDS slowing down so dramatically?
>
> Thanks in advance.
>
> matthias
> --
> firm: matthias.apitz@sisis.de [voc:+49 89 61308 351, fax: +49 89 61308 188]
> priv: guru@thias.muc.de
> WWW: http://www.sisis.de/~guru/
In article <guru.916308096@almare.sisis.de>, Matthias Apitz
<guru@sisis.de> writes
>
>We have a very big database under IDS 7.13.UC1 and Reliant UNIX 5.43.
>The database has around 27 GByte pictures in BLOBS (in BLOBSPACE),
>each picture is around 14 KByte in size, together with some
>additional information in the table like picture number etc.
>
>The time for a complete backup with "ontape -s -L 0" to a SCSI tape
>supporting a transfer of 5 MByte per second enlarged dramatically
>from 9h to 43h without any known reason.
>
>I checked the system calls done by the ontape and there are two
>semops (to the server) and one write(2) call to the tape with
>65536 bytes per second. A test with dd(1) reading with ibs=65536
>from disk and writing with obs=65536 to the same physical tape
>produces *many* read/write per second.
>
>So it seems that the IDS (which is idle at this moment) does not
>return the data fast enough. My questions are:
>
>How is ontape fetching the data from ISD? In structured form or
>simply page by page if the page is marked used?
Page by page.
>
>What could be the reason for IDS slowing down so dramatically?
>
?? Don't know, are you sure nothing else is running at the same time?
I would probably upgrade to 7.30 since 7.13 is a very old release and
early 7.x did have some problems....
>Thanks in advance.
>
> matthias
>--
>firm: matthias.apitz@sisis.de [voc:+49 89 61308 351, fax: +49 89 61308 188]
>priv: guru@thias.muc.de
> WWW: http://www.sisis.de/~guru/
--
David Williams
David Williams wrote:
> In article <guru.916308096@almare.sisis.de>, Matthias Apitz
> <guru@sisis.de> writes
> >We have a very big database under IDS 7.13.UC1 and Reliant UNIX 5.43.
> >The database has around 27 GByte pictures in BLOBS (in BLOBSPACE),
> >each picture is around 14 KByte in size, together with some
> >additional information in the table like picture number etc.
> >The time for a complete backup with "ontape -s -L 0" to a SCSI tape
> >supporting a transfer of 5 MByte per second enlarged dramatically
> >from 9h to 43h without any known reason.
[SNIP]
> I would probably upgrade to 7.30 since 7.13 is a very old release and
> early 7.x did have some problems....
Also note that the ontape bugs were not fixed until 7.14 and so online
archives of a busy server may miss archiving all pages. Definitely
upgrade.
Art S. Kagel
( by the by, we have run into the missed archive datapages thingy )
I have a couple of questions.
I have read several times over that ontape has changed in it's behavior
between 7.2 ports and 7.3 ( i.e. the "new" way that it archives datapages )
Could someone fill me in on those changes?
PLATFORM INFO: ( if you need it )
NCR 3555XP running MPRAS 3.02
2 gig of ram
16 pentium 133's
ODS 7.24uc5 ( will soon upgrade to 7.3uc6 that's why I ask this question )
214 gig of database space ( mixture of 4 and 2 gig disks running raid 5 using
informix mirroring )
2 - 8 meg of raid controller cache, 2 meg of raid processor cache
Carlos
"Art S. Kagel" wrote:
> David Williams wrote:
>
> > In article <guru.916308096@almare.sisis.de>, Matthias Apitz
> > <guru@sisis.de> writes
>
> > >We have a very big database under IDS 7.13.UC1 and Reliant UNIX 5.43.
> > >The database has around 27 GByte pictures in BLOBS (in BLOBSPACE),
> > >each picture is around 14 KByte in size, together with some
> > >additional information in the table like picture number etc.
>
> > >The time for a complete backup with "ontape -s -L 0" to a SCSI tape
> > >supporting a transfer of 5 MByte per second enlarged dramatically
> > >from 9h to 43h without any known reason.
> [SNIP]
>
> > I would probably upgrade to 7.30 since 7.13 is a very old release and
> > early 7.x did have some problems....
>
> Also note that the ontape bugs were not fixed until 7.14 and so online
> archives of a busy server may miss archiving all pages. Definitely
> upgrade.
>
> Art S. Kagel
Chris Carter wrote: > More notably, and possibly more helpful to you, we noticed the following very > strange behaviour. For reasons I forget we went about 3 or 4 weeks without > doing an archive of any kind (we never did anything but level 0 archives) on > one of our machines (200+GB db - no blobs though). When we finally started > this archive it was running extremely slowly. After about 6 hours the archive > died (again reason forgotten but probably some kind of tape error). We > restarted the archive and it "flew" through the spaces it had archived in the 6 > hr aborted run but when it got to the spaces it had not archived in 3 weeks or > so it slowed to a crawl. Needless to say this caught our attention...Why would > the time between level 0 archives have an impact on the time it takes to do a > level 0? Well, I am probably about to give you more information (and some of it kind of vague, since it's a long time since I looked into this at all) than you really wanted, but since I'm so good at that, I never let it stop me. Let's start with how the archive manages to get all the data at a certain point in time. (Some of you will already know this, but bear with me.) That is, if you start an archive at 8pm, and it takes 10 hours to complete, what happens to all the pages that are modified in the system between 8pm and 6am? We need to make sure we get all the pages as they were at 8pm, even if it is 6am by the time the archive process gets around to reading that chunk. So we compare timestamps. We know what the timestamp was when the archive started, and we compare the timestamp of the page with the timestamp of the archive: if the page timestamp is earlier than the archive timestamp, then we can use it; if it's later, then the page has been modified and we don't want it. (This also means that whenever checkpoints occur, we have to check the physical log to make sure there aren't any old pages we need that are about to get flushed. I'm sure this has been covered before and may even be in the FAQ, so I'm not going to go into it any further.) Here's the key question: how do you know if the timestamp is earlier or later? Timestamps are not actual dates and times, they are merely 4-byte counters, and occasionally (more or less frequently depending on how heavily your system is used) they roll around. So if my archive timestamp is 0x12345678, how do I know if the page stamped 0x87654321 is earlier or later? So if enough time has elapsed (and I don't know how the archive determines this), the first thing the archive process does is read through ALL your pages and re-timestamp them. THEN it can start the actual process of archiving your pages, secure in the knowledge that it isn't going to miss any of your pages because it thought it was timestamped later but it was really 5 weeks earlier. So what you're probably seeing, when your archive is slow after you've gone a long time without an archive, is the archive plodding through all your pages and re-stamping them. When it "flew" through the pages it had archived in the aborted run, that was probably because it recognized that those timestamps were close enough to the present not to be a problem. Hope that's clear, and helpful. June -- june_t@hotmail.com Grounded in Palo Alto, living on M&M's (plain)