Re: Slow system ?! - using Informix 7.x on Sun Sparc 1000
Posted in 1997
In article <5ceths$m9u@cssun.mathcs.emory.edu>, Mario Estrada
<marioestrada@guate.net> writes
>Hi George!
>
>Your problem is with the CHECKPOINT DURATION, and the factors affecting =
>this period of time are: 1) How many dirty pages must be written to disk =
>from the buffer cache and 2) How fast the dirty pages can be written
>
>1) How many dirty pages must be written to disk from the buffer cache.
>
>A) Increase the size of your physical log and Logical Logs Buffers to =
>minimize the amount of physical I/O required for writing to the =
>physycal and Logical Logs, but at the same time
>make them just as large as necessary because they ocuppy shared memory
>space. You can use onstat -l to check if the current size of your =
>physical
>log and Logical Logs are optimal for your Online System. Check the =
>Physical
>Loggin and Logical Loggin in the output section (Fields bufsize and =
>pages i/o
>) if pages i/o is 75% of bufsize value or greater, the buffers are being =
>used=20
>in an efficient manner.
>
>NOTE: You asked that you wanted to increase the physical log size ,
> the size of the physical log only affects the "CHECKPOINT
> INTERVAL" not the "CHECKPOINT DURATION".
Correct - checkpoints occur when physical log is 75% full. Checkpoints
are clearly occuring often enough if they take 45 seconds.
>
>B) LRU parameters, LRU_MIN_DIRTY and LRU_MAX_DIRTY, you can
> monitor the LRU queues with onstat -R. LRU_MAX_DIRTY means
> how much percentage from the LRU Modified Queues will be filled
> up before a page cleaner starts cleaning them until the =
>LRU_MIN_DIRTY> percentage is reached.In your case decreasing this values will give =
> =20
> you LRU WRITE activity "Check it with onstat -F, you can start=20
> with a value of LRU_MAX_DIRTY=3D10 and LRU_MIN_DIRTY =3D 5 and
> then monitor them with onstat -F, if you don't see in the output of
> onstat -F the field LRU Writes increased then reduce these values> until you increase the number LRU Writes and reduce the number
> of Chunk Writes.
>
Correct again.
>2) How fast the dirty pages can be written
> There are other factors, but in your case, I will mention the =
>followings:
>
>A) Number of PAGE CLEANERS, these are threads which will perform
> the cleaning from the LRU Modified Queues and from the Physical Log =
> Buffer , to DISK. Suppose that a checkpoint is requested and =
>your data=20
> must be written to 4 different disks, the performance will be =
>affected if
> you only have configured 1 Page Cleaner, this only one page cleaner
> will be in charge to write to all disks (Written to disks involves =
>another
> subject in performing tunning) so as a GuideLine this parameter =
>should
> be configured initially equal to the number of ACTIVE DISKS,
> =20
> NOTE: You never mentioned the number of page cleaners configured
> in your system.
>
Correct BUT NOTE CHUNKS ARE GIVEN TO PAGE CLEANERS FOR CLEANING IN
CHUNK ID ORDER. i.e. when you create chunks generally create one
chunks per disk. If you have >1 chunk per disk then may sure the
chunks are created round robin across the disks. i.e. for three disks
(and three page cleaners)
create first chunk on disk one
second chunk on disk two
third chnnk on disk three
fourth chunk on disk one
fifth chunk on disk two
sixth chnnk on disk three
seventh chunk on disk one etc.
That way at the start
page cleaner one writes to disk one
page cleaner two writes to disk tweo
page cleaner three writes to disk three
whichever one finishes first starts cleaning chunk four (hopefully
the first one!), the next one to finish cleans chunk four, the next
one to finish cleans chunk siz. Now chunks 4,5,6 and disks 1,2,3 are
in use i.e. minimum disk contention whilst page cleaning.
You have spread the main tables which get updated across chunks so
that tbstat -d shows that disk writes are evenly balanced across the
disks haven't you? If you have OnLine 7.X you can do
select tabname,iswrites from sysmaster:sysptprof
order by 2 desc
to find out which tables are being written to most often.
>B) Number of AIO vps, If your OS supports Kernel AIO configured it=20
> to 2 or 1 for smaller systems (Because AIO VPS will only be used
> to write to control files such as message log, Assert Failure =
>files,
Someone a while back said that on Sun machines you could get
up to a 40% speead increase but using KAIO - something to consider.
cd $INFORMIXDIR/release and look at the files whose names
start with ONLINE for more info.
I could think of more (are you running low on memory and so
paging/swapping alot etc) but it getting late...good night...).
> Dump Files and so on), also Consider this parameter if you are
> using cooked files. If your OS does not support Kernel AIO then=20
> as a GuideLine make this parameter ( NUMAIOVPS ) equal to
> the number of active disks containing database tables. You may
> monitor AIO VPS using ontat -g ioq , look at the field "len" this=20
> is the queue length, when this is large for most AIO queues,=20
> you can increase performance by adding additional AIO vps. In this
> case run onstat -z and after a period of time, monitor the field =20
> maxlen to see if the maximun queue length has decreased. If you
> notice that only the first aio vp has more activity this will be =
>because
> the requests are given to the first idle vp found.
>
>
>So, in your case, physical log does not affect checkpoint duration, so=20
>increase it after monitoring the onstat -l output, you will have to =
>decrease
>the LRU_MAX_DIRTY and LRU_MIN_DIRTY to increase the LRU WRITES
>between checkpoint intervals, and check the number of PAGE CLEANERS
>you have configured so that a single thread cleaner can be used more
>efficiently, and check what kind of AIO you are using, to increase the
>Number of AIO VPS after monitoring them with onstat -g ioq. After this
>I think you will get a better performance whithin your checkpoints, =
>sometimes
>If you don't look LRU WRITES after decreasing LRU_MAX_DIRTY and=20
>LRU_MIN_DIRTY, increase the number of Checkpoint Interval.
>
>PS: You may take "INFORMIX-OnLine Dynamic Server Performance Tuning"
> and "Server System Administration", Yes , the person you send to=20
> this CLASS will be more efficient if he/she has already deal with =
>the
> ENGINE, if you want he/she to take them both and the INFORMIX
> Catalog Class Sched. is not convinient for you, you can contact =
>me.
>
>I hope this helps,
>
>-------------------------------------------------------------------------=
>-----------------
>Mario Estrada / SISTECO,S.A.
>INFORMIX, / Phone (502) 3340214
>Support Department / Fax (502) 3348447
>-------------------------------------------------------------------------=
>