Re: tuning questions
Posted in 1998
Hi Allen,
Allen Jantzen wrote:
>
> I recently doubled the size of my database, and have been tuning for a
> few days now. IIUG CDI Glimpse has been a tremendous help, especially
> the posts of Art Kagel, David Williams, et al. My thanks to them.
>
> My environment:
> HP 9000 J200 , 2 processors, 256 megs, HP-UX 10.01, database dedicated
> Informix ODS 7.22.UC1
> Database around 16 gigs, 10 of which is used
> 6 Dbspaces, all mirrored, 4 cooked 2 raw
> KAIO just enabled
> Basically 1/2 DSS and 1/2 OLTP
How many disks do you have and where do the dbspaces reside ?
> MULTIPROCESSOR 1
> NUMCPUVPS 1
> SINGLE_CPU_VP 1
> LOCKS 100000
> BUFFERS 15000
> NUMAIOVPS 8
> PHYSBUFF 96
> LOGBUFF 48
> CLEANERS 16
> CKPTINTVL 300
> LRUS 16
> LRU_MAX_DIRTY 10
> LRU_MIN_DIRTY 5
> RA_PAGES 24
> RA_THRESHOLD 12> PDQPRIORITY 50
>
> A few questions:
>
> I cannot seem to get my onstat -F LRU write / Chunk Write ratio
> any higher than 1:1 or 1.5:1. During the day my checkpoints average
> only 3-4 seconds:
>
> 08:04:01 Checkpoint Completed: duration was 4 seconds.
> 08:09:04 Checkpoint Completed: duration was 1 seconds.
> 08:14:09 Checkpoint Completed: duration was 1 seconds.
> 08:19:16 Checkpoint Completed: duration was 3 seconds.
> 08:34:32 Checkpoint Completed: duration was 1 seconds.
You might think, this is fine. But more interesting is the
amount of time the users are waiting for the completion
of the checkpoint. That's the basis on which you should make
your decision whether 4 seconds are good or not. If I were
a customer and would try to insert a single row, I would
say I have a bad performance if this action would last 4 seconds.
> but at night when a lot of batch processing is occurring (beginning at
> 3:30 AM for instance) they go up to 30-40 seconds, or even higher:
> 03:15:39 Checkpoint Completed: duration was 1 seconds.
> 03:31:10 Checkpoint Completed: duration was 19 seconds.
> 03:36:44 Checkpoint Completed: duration was 30 seconds.
> 03:42:34 Checkpoint Completed: duration was 47 seconds.
> 03:47:33 Logical Log 5154 Complete.
> 03:47:35 Process exited with return code 133: /bin/sh /bin/sh -c
> /opt/informix/etc/log_full.sh 2 23 "Logical Log 515> 4 Complete." "Logical Log 5154 Complete."
> 03:47:36 Logical Log 5154 - Backup Started
> 03:48:10 Checkpoint Completed: duration was 33 seconds.
> 03:48:10 Logical Log 5154 - Backup Completed
> 03:53:47 Checkpoint Completed: duration was 34 seconds.
The same point of view. If there's just a single process running,
you can ignore the checkpoint duration, because there's no
other session waiting for the completion of the checkpoint.
Checkpoint writes are "sorted and buffered" writes which will
perform faster than LRU-writes. Obviously it takes some time
to write your data. I would say, it's much too long.
Your cache buffer has a size of about 15000 BUFFERS, let's say
about 30MB. If you spread your data across at least 2 disks,
the write activity during the checkpoint should perform faster.
On my very old SCO platform ( since 1995 ), the troughput
during a checkpoint is about 12MB/sec with 3 disks. To flush
a completely filled cache buffer this would take less than 3 seconds.
On your platform you have more than 40seconds. I ran the test
several times on different platforms. On a HP-UX workstation
I measured a throughput of 1.5MB/sec on a single disk for
writes. If you would have 2 disks, this would result in an overall
throughput of 3MB/sec. Since your cache buffer has a size of
30MB, the I/O must be done in about 10seconds. ( I assume the
worst case, that is, your batch jobs load the data very fast,
faster than the CLEANERS can flush the data to the disks between
the checkpoints. That's not really the case because you do not
have Foreground Writes )
There are a few circumstances which reduce the disk I/O.
+ If the data is spread across lots of extents ( disk positioning )
+ If there are several processes accessing the same disk concurrently
( disk positioning )
You told us you have 6 dbspaces but I'm sure you have more
than 6 chunks ( 16gigs data ! ). Now it is interesting, how
you created your dbspaces. Possibly 2 or more of the 8 AIO-vp's
are trying to flush the data to one and the same disk ???
During the checkpoint each CLEANER is responsible to flush
the data of a specific chunk. If the 1st chunk and the 2nd chunk
reside on the same disk, you will get the described behaviour.
The parallel writing to the disk will take about 10 times longer than
the normal sequential writes.
As long as I do not know, where your cooked files and raw devices
reside on the disk, and in which order you created the chunks
it is nearly impossible to say, what you should do at the moment
to improve your I/O performance.
> (As an aside, why would a checkpoint not occur for 16 minutes between
> 3:15 and 3:31 when my CKPTINTVL is 300???)
>
> There are no FG writes. I have decreased LRU_MAX_DIRTY and
> LRU_MIN_DIRTY and increased LRUS and CLEANERS, but no help.
> My Bufwaits are around 3% of Bufreads and Bufwrites. My read and write
> cache% are 99% and 97%.
These high cache rates make sure that you do not see the I/O problems.
> How can I get my LRU write/Chunk write ratio to increase (and
> therefore my checkpoint duration to drop)??
> ----------------------------------------------------------------------------------------------------------
> My onstat -l physical logging pages/io is always higher than 40 while
> my physical log bufsize is 48 - this seems great. However........
> my onstat -l logical logging pages/io is always lower than 5 while my
> logical log bufsize is 24. The logical logging pages/io is way too
> low. How can I increase this?
Obviously your databases are using NO LOGGING or UNBUFFERD LOGGING.
It's not a good idea to change this to BUFFERED LOGGING. It's a
security point.
> -----------------------------------------------------------------------------------------------------------
> My ontape -s lev 0 archives are taking a *LONG* time, like 5 or 6
> hours. What could be causing this?
You must read and write up to 10GB. Once again it's interesting
how fast your tape devices can write 10MB per second.
Use the "dd" command to write 10MB to the tape and measure
the throughput with the "time" command.
>
> Should I make other changes as a result of enabling KAIO?
I don't know whether KAIO will make anything better ? In this
case, each CLEANER would act like an AIO-vp.
> Should NUMAIOVPS go down to 1, even w/ cooked dbspaces??
>
> Would another CPU vp be a good idea? (2 physcial processors)
If you want you can try it. If you want to make sure that
an additional CPU-vp will end in a better performance, then
you must measure the system bottlenecks. There are some reasons
why 2 CPU-vp's are better, and there are some reasons why
you should not set 2 CPU-vp's. Normally you would add an
additional CPU-vp if you have enough CPU power ( sar -u )
but the IDS is not using the power.
Bye
--
Stefan Weideneder
@@NL