ONTAPE Level 1 resource usage
Posted in 2012
A new Informix admin ran a nightly level 0 ontape archive plus hourly level 1 backups on a small VM/SAN instance, and users saw heavy slowdowns during the short level 1 runs. Replies explained that incremental archives must read every page to find changed ones, so they are very read-intensive regardless of output size, and that VM/SAN I/O makes this worse; hourly incrementals are unnecessary when logical logs can be archived instead. The fix adopted: a cron job running 'onmode -l' to switch logs followed by 'ontape -a -d', giving fast (~1 second, ~1MB) scheduled log backups, with restores prompting for logs after the level 0.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Backup & Restore
Hi guys, first post! I have inherited an Informix database running our
companies ERP system and while I'm learning fast I am far from even competent.
The system is running on a RHEL 5.5 server and the usage is not what I would
call very 'intensive'. The problem I have is that we implemented a new backup
schedule where we do a full level 0 backup each night and a level 1 backup
each hour while the system is live.
Now the level 1 backup only takes 5 minutes but the users are complaining of
massive slowdown during that time (we originally had it at 20 minutes).
My question is, is this normal? I would not have expected a level 1 backup to
be so intensive on the system, the resulting file is only between 70-250MB
depending on the time of day.
I've also trialled a level 2 backup and the same slowdown occurs.
I guess a couple of more quick questions:
1. Is this a normal procedure, running level 1 and 2 backups while the system
is in use?
2. Is there a way to limit the usage of ontape to lower the priority
The machine is a VM hosted on a SAN we have assigned 4 processors and 8GB of
RAM to the machine. During normal operation the system usage is really, really
low.
Any other information I need to provide?
Anyway, thanks in advance guys!
Answers in line...
On Thu, Apr 19, 2012 at 12:27 AM, MATTHEW HOUSTON <
matthewh@wedderburn.com.au> wrote:
> Hi guys, first post! I have inherited an Informix database running our
> companies ERP system and while I'm learning fast I am far from even
> competent.
>
> Congratulations. You've entered a wonderful and usually peaceful world of
Informix Administration :)
> The system is running on a RHEL 5.5 server and the usage is not what I
> would
> call very 'intensive'. The problem I have is that we implemented a new
> backup
> schedule where we do a full level 0 backup each night and a level 1 backup
> each hour while the system is live.
>
Each hour?! Why would you want to do that? If your database uses logging,
and you save your logical logs, you'd be able to restore the system by
using a L0 and the logs.
> Now the level 1 backup only takes 5 minutes but the users are complaining
> of
> massive slowdown during that time (we originally had it at 20 minutes).
>
5 minutes is very low considering what it has to do. I assume the instance
is very small.
> My question is, is this normal? I would not have expected a level 1 backup
> to
> be so intensive on the system, the resulting file is only between 70-250MB
> depending on the time of day.
>
L1 and L2 will only save the pages changed since L0 and L1 respectively.
But unfortunately we don't have a "map" of those pages, so every page must
be read to check if it changed or not. So it's very read intensive, even
that the generated file is small (meaning not many changes since the last
backup)
> I've also trialled a level 2 backup and the same slowdown occurs.
>
For exactly the same reason.
> I guess a couple of more quick questions:
>
> 1. Is this a normal procedure, running level 1 and 2 backups while the
> system
> is in use?
>
Yes, but never with that frequency.
> 2. Is there a way to limit the usage of ontape to lower the priority
>
Not any "official" one. But if somehow you can delay the writing, it will
have less impact. The ontape command uses some buffers to communicate with
the database server. On a backup these buffers are filled by the engine and
emptied by the ontape (it's a bit more complex, but let's assume this). If
the command takes too much time to write to the output file/tape, then the
engine will freeze the reads when it fills the buffers.
> The machine is a VM hosted on a SAN we have assigned 4 processors and 8GB
> of
> RAM to the machine. During normal operation the system usage is really,
> really
> low.
>
I've just recently visited a customer where the I/O of a VMWare image
running with virtualized disks of a SAN is causing terrible effects. I
honestly can't explain it, but it's definitively related to the number of
IO operations triggered by Informix. Many operations with small amounts of
data kill it. Few operations with large amounts of data are ok. I currently
suspect of some block sizing issue, but I have no concrete evidences of
that.
Any other information I need to provide?
>
Version of Informix is always good to have.
If your database is logged or not, and why the hourly L1
> Anyway, thanks in advance guys!
>
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--002354790ff82802b404bdfca012
Several comments:
- There is nothing inherently wrong with taking archives during
production usage of the server. Informix invented the concept and is very
good at it. I have a client who does a full level 0 three times a day at
6AM, noon, and 6PM.
- There is actually no good reason to take archives all day long since
you can be archiving your logical logs which is far less overhead and
sufficient along with the nightly level 0 to restore your data to nearer
the point of failure than the hourly archives will permit. (Yes I've told
my other client the same thing, but as long as they keep paying the bills
they can ignore my advice all they want.) Unless of course your databases
are not logged so there are no logical logs. But that would be another
bigger problem.
- IO under a VM is a major performance problem and is probably the
primary reason that your users are experiencing a slowdown during the
archive runs. The VM's IO mapping can't keep up with the combined IO
load. In my testing IO throughput under VMWare was about 30% of running on
raw iron. I know that VMWare and the SAN vendors have been working to
improve things, and the latest VM kernels are a bit better, but the problem
persists and the improvements have not been 3.3x but incremental. I'm
giving a talk partially on this subject at the IIUG Conference in San Diego
next week.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Wed, Apr 18, 2012 at 7:27 PM, MATTHEW HOUSTON <matthewh@wedderburn.com.au
> wrote:
> Hi guys, first post! I have inherited an Informix database running our
> companies ERP system and while I'm learning fast I am far from even
> competent.
>
> The system is running on a RHEL 5.5 server and the usage is not what I
> would
> call very 'intensive'. The problem I have is that we implemented a new
> backup
> schedule where we do a full level 0 backup each night and a level 1 backup
> each hour while the system is live.
>
> Now the level 1 backup only takes 5 minutes but the users are complaining
> of
> massive slowdown during that time (we originally had it at 20 minutes).
>
> My question is, is this normal? I would not have expected a level 1 backup
> to
> be so intensive on the system, the resulting file is only between 70-250MB
> depending on the time of day.
>
> I've also trialled a level 2 backup and the same slowdown occurs.
>
> I guess a couple of more quick questions:
>
> 1. Is this a normal procedure, running level 1 and 2 backups while the
> system
> is in use?
> 2. Is there a way to limit the usage of ontape to lower the priority
>
> The machine is a VM hosted on a SAN we have assigned 4 processors and 8GB
> of
> RAM to the machine. During normal operation the system usage is really,
> really
> low.
>
> Any other information I need to provide?
>
> Anyway, thanks in advance guys!
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--14dae93403c37c6e5804bdfcfa41
Thanks guys for your information, this has got to be the fastest and most
relevant forum I have ever joined!
I had looked into the logical log files about 6 months ago and even put
forward a proposal to implement a logical log backup schedule but the powers
that be considered it to be too 'messy' and at the time I didn't have enough
knowledge to argue the point, I guess I still don't!
So I'm frantically re-reading all the information I can on the logical logs to
re-invent this backup schedule.
Thanks Art for the information on VM performance, I suspected that it was due
to disk IO but as our instance is really small I thought we may have gotten
away with it.
And Fernando, I probably should have read more on the actual level 0-2 backup
processes to understand what they do during the backup, I didn't realise it
was so intensive on the disk.
After my reading this morning it looks like I will be just modifying the
current script to run a 'ontape -a' and backup the resulting logical files.
I haven't got that far in the reading but maybe a quick heads up from you guys
could speed things up:
The reason we settled on the level 1 backups is that we could script the
command and run it on a schedule and back the files up to the network
location. If a restore was needed , the guys could just rename the file they
needed to the level 1 filename and then just run a full restore, level 0 and
level 1.
Is the restore of logical files as straight forward?
Matt:
You will find that the Informix user community is a very tight knit
community. Help is never far away. Were you planning to attend the IIUG
Conference in san Diego next week? You could still come if you haven't
registered already.
Restoring logical logs is as easy as restoring a level 1 or level 2
archive. During the restore process, ontape prompts for any level 1
archives when it finishes restoring the level 0. If you restore a level 1
it prompts for any level 2 archives once the level 1 restore completes.
Once all levels of archive have been restored ontape asks if you want to
restore logical logs. If you say yes, it will tell you which log was
current at the time the last archive it restored was started and asks you
to mount the tape/file that contains that log. If the tape/file contains
multiple logs it restores all that it finds that follow the first one it
prompted for then prompts for any more logs you want to restore until you
say 'no'. Easy peasy.
What version of Informix are you running? The best way to archive the
logical logs is to use the ALARMPROGRAM parameter in the ONCONFIG file to
capture event level 23 which is "Logical log completed." then you run
ontape -a and archive that log as soon as it completes. If you are runninga recent version of Informix (11.50+) this is trivial because you can set
LTAPEDEV to a directory/filesystem and ontape will manage the filenames for
you within that directory. If you are using an earlier version your
ALARMPROGRAM script will have to rename the LTAPEDEV file each time ontape
completes to a name that includes the number of the first log that it
contains (and for good measure the machine and server names so you can't
confuse logs from different servers) and touch a new empty LTAPEDEV file
for the next run.
If you are running 11.50+ the alarmprogram.sh in $INFORMIXDIR/etc/ can be
used by changing a few variables in it (it defaults to using onbar instead
of ontape and log archives are not enabled by default).
Note that the company I work for, Advanced DataTools Corp., gives Informix
training classes that you can attend remotely. Check out our web site (URL
below) for details. We have an advanced DBA performance tuning class
coming up in June and a basic DBA class in September. We can get you
up-to-speed fast.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Wed, Apr 18, 2012 at 8:48 PM, MATTHEW HOUSTON <matthewh@wedderburn.com.au
> wrote:
> Thanks guys for your information, this has got to be the fastest and most
> relevant forum I have ever joined!
>
> I had looked into the logical log files about 6 months ago and even put
> forward a proposal to implement a logical log backup schedule but the
> powers
> that be considered it to be too 'messy' and at the time I didn't have
> enough
> knowledge to argue the point, I guess I still don't!
>
> So I'm frantically re-reading all the information I can on the logical
> logs to
> re-invent this backup schedule.
>
> Thanks Art for the information on VM performance, I suspected that it was
> due
> to disk IO but as our instance is really small I thought we may have gotten
> away with it.
>
> And Fernando, I probably should have read more on the actual level 0-2
> backup
> processes to understand what they do during the backup, I didn't realise it
> was so intensive on the disk.
>
> After my reading this morning it looks like I will be just modifying the
> current script to run a 'ontape -a' and backup the resulting logical files.
>
> I haven't got that far in the reading but maybe a quick heads up from you
> guys
> could speed things up:
>
> The reason we settled on the level 1 backups is that we could script the
> command and run it on a schedule and back the files up to the network
> location. If a restore was needed , the guys could just rename the file
> they
> needed to the level 1 filename and then just run a full restore, level 0
> and
> level 1.
>
> Is the restore of logical files as straight forward?
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--14dae9340d597dbebe04bdfdcd0f
Art:
I would love to attend ANYTHING to get up to speed, unfortunately my current
company is a bit short sighted when it comes to that. I only inherited this
project as it runs on a Linux box and I'm the only one who knows anything
about Linux, I have absolutely zero knowledge in DB admin let alone Informix.
I'm in Sydney, AUS at the moment so it might be a bit of a stretch to get the
company to spring for that sort of flight!
The process for logical logs seems to be just as simple, and easy enough to
teach to the non-Linux guys so I thank you again for your help!
I'm still slogging through the Informix documentation but what I was after was
a timed backup (e.g every 10 minutes), is it possible to have the logs rotate
on a schedule rather then wait until they are full?
Can I get your insight on the two possible scenarios I'm looking at:
1. Increase the log number and decrease the size so that the logs are cycled
more regularly, triggering the ALARMPROGRAM setting. Downside is that no
guarantee when this will run - for instance at the moment the log files are
running at whatever the original guys set it up to and doesn't rotate at all
during the day, the only log I can see generated is at 7:10PM, that's how
little transactions are being put through. Some hours will get a lot of
activity, some not much at all.
2. Rotate the log manually at intervals using a cron job, allows us to specify
exact times for the backup regardless of activity in the DB.
I'm leaning towards the second option as the management is very set on having
specifically timed backups.
I cant believe I didn't include this before, such bad manners, but output of
onstat -version:
$ onstat -version
Program Name: onstat
Build Version: 11.50.FC5WE
Build Number: N104
Build Host: vidar
Build OS: Linux 2.6.9-34.ELsmp
Build Date: Tue Jul 14 19:22:49 CDT 2009
GLS Version: glslib-4.50.FC6
See comments below:
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Wed, Apr 18, 2012 at 9:37 PM, MATTHEW HOUSTON <matthewh@wedderburn.com.au
> wrote:
> Art:
>
> I would love to attend ANYTHING to get up to speed, unfortunately my
> current
> company is a bit short sighted when it comes to that. I only inherited this
> project as it runs on a Linux box and I'm the only one who knows anything
> about Linux, I have absolutely zero knowledge in DB admin let alone
> Informix.
> I'm in Sydney, AUS at the moment so it might be a bit of a stretch to get
> the
> company to spring for that sort of flight!
>
We'll miss you at the conference. The travel cost is why we do training
over Webex, so keep the courses in mind.
>
> The process for logical logs seems to be just as simple, and easy enough to
> teach to the non-Linux guys so I thank you again for your help!
>
> I'm still slogging through the Informix documentation but what I was after
> was
> a timed backup (e.g every 10 minutes), is it possible to have the logs
> rotate
> on a schedule rather then wait until they are full?
>
Sure. You can run onmode -l to change to the next logical log then you can
archive any previous logs that have not be archived yet.
>
> Can I get your insight on the two possible scenarios I'm looking at:
>
> 1. Increase the log number and decrease the size so that the logs are
> cycled
> more regularly, triggering the ALARMPROGRAM setting. Downside is that no
> guarantee when this will run - for instance at the moment the log files are
> running at whatever the original guys set it up to and doesn't rotate at
> all
> during the day, the only log I can see generated is at 7:10PM, that's how
> little transactions are being put through. Some hours will get a lot of
> activity, some not much at all.
>
Are the databases logged? See the is_logging flag column in the
sysmaster:sysdatabases table.
Anyway, you could do that, there's always a trade off between wanting
smaller logs so that they are archived more frequently to improve recovery
and minimize data loss in the event of a catastrophic server failure and
wanting larger logs to minimize the overhead of changing logs and archiving
them too frequently. It looks like your predecessors erred on the side of
larger logs. You can fix that and combine with using onmode periodically
so you don't have to make them tiny and fight the extra overhead during
peak periods.
>
> 2. Rotate the log manually at intervals using a cron job, allows us to
> specify
> exact times for the backup regardless of activity in the DB.
>
> I'm leaning towards the second option as the management is very set on
> having
> specifically timed backups.
>
Again, you can combine things. If you have a cron run onmode -l
periodically, say every 15 minutes or once an hour, that will kick off the
ALARMPROGRAM event code 23 and trigger the archive for you. One thing you
can do is have the alarmprogram touch a file when the archive kicks off and
have the cron job check the age of the file. If the file is new enough,
don't switch logs. That will prevent archiving an empty log during peak
load when it will be changing by itself anyway.
>
> I cant believe I didn't include this before, such bad manners, but output
> of
> onstat -version:>
> $ onstat -version>
> Program Name: onstat
> Build Version: 11.50.FC5WE
> Build Number: N104
> Build Host: vidar
> Build OS: Linux 2.6.9-34.ELsmp
> Build Date: Tue Jul 14 19:22:49 CDT 2009
> GLS Version: glslib-4.50.FC6
>
TMI <smile> onstat - and uname -p -r -s -m output give plenty of detail.
;-)
$ uname -p -s -r -m
Linux 2.6.32-40-generic x86_64 unknown
$ onstat -
IBM Informix Dynamic Server Version 11.70.FC3 -- On-Line -- Up 3 days
10:35:23 -- 508080 Kbytes
$
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--e89a8f3ba75fa5091c04bdff7e69
Just wanted to stop back in and say thanks Art, you really saved my bacon.
I've finally gone with a cron script to run:
onmode -l
then:
ontape -a -d
This allows us the flexibility of knowing when the backups will occur etc. The
log files themselves are only around 1MB and take about a second to complete.
There really is jack all transactions taking place but each one is important.
At least I learned a bit about the backup / restore procedure during all this.
You are welcome.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Mon, Apr 23, 2012 at 6:03 PM, MATTHEW HOUSTON <matthewh@wedderburn.com.au
> wrote:
> Just wanted to stop back in and say thanks Art, you really saved my bacon.
>
> I've finally gone with a cron script to run:
>
> onmode -l>
> then:
>
> ontape -a -d>
> This allows us the flexibility of knowing when the backups will occur etc.
> The
> log files themselves are only around 1MB and take about a second to
> complete.
> There really is jack all transactions taking place but each one is
> important.
>
> At least I learned a bit about the backup / restore procedure during all
> this.
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--bcaec5299a3184de8804be62683a
Related threads
- TABLE AN INDEX REORG
- ROLE FOR USERS
- HDR in Windows
- eliminate duplicate rows
- Fragmentation Elimination Problem