Re: Informix LRU adjustment advice wanted
Posted in 1999
Rob Gale wrote:
>
> Oh wise gurus...your advice please...
First assume a proper lotus position to place your Karma in a
transcendental alignment with ours..... ;-)
Specifics below:
> Given a dedicated database server for heavy-duty OLTP system as follows:
> IDS_7.30UC3 (planning to go to UC7)
> 8 CPU Sun E4000 server
> 4 Gb memory.
> EMC disk subsystem of ~ 200GB user space (apparently striped and mirrored
> under EMC control).
> NUMCPUVPS is 7> Maximum number of BUFFERS configured ~ 750000.
> LRU MIN and MAX are at 1 and 3.
>
> We are not getting the throughput we think we should achieve on this system.
> There are some io hotspots but nothing looks swamped.
> CPU usage is around 20~25%
Check the controllers, you may be swamping them.
> Now, at the moment we have only 7 LRU queues, I believe that this is
Low, see below.
> somewhat low, and intend to raise it.
OK.
> The rational for this is that bufwaits is ever increasing and also that
> onstat -g ioq always shows the kio queues with a maxlen of 16, even shortly> after a stats reset.
Actually your BUFWAITS ratio is in my acceptable range. Anything under
10% is OK but under 7% is best. You are running right around 6%. The
increases are, however, at a higher rate, closer to 7% which is still
acceptable. I suspect that you run OK when you are not running the load
jobs and poorly when they are running. The BUFWAITS is not likely your
real problem. If ou want to empirically test there are two things you
can try. Increasing LRUS and CLEANERS will help somewhat, but there are
few processes loading. Are there many other apps processing smaller
updates/inserts/deletes at the same time? If not it may just be that the
jobs are all hashing to the same LRUS and any change in the LRUS value
will change the hash and may solve the problem, try 32 many sites get good
performance with 32 LRUS and CLEANERS (Kagel's rule: CLEANERS >= LRUS)
with similar load. Another thing you can try is to add the undocumented
ONCONFIG parameter LRUPOLICY to force the apps to each pick one LRU and
stick to it which will eliminate contention for LRUS and buffers (I have
the values if you decide to give this a try).
> Will increasing the number of LRU queues help ? What should I do about
> CLEANERS?
> There are effectively about 10 processes constatntly performing
> simultaeneous batch updates to a common set of 5 tables, although they will
> rarely be going after the same records. These tables may have several
> million records.
>
> Also, I am planning to go from Unbuffered to Buffered logging - any idea how
> much perfomrnace difference this will make ?
>
> I have included some output from onstat -p and onstat -R and -F below as
> well as the ONCONFIG file.
>
> If I have omitted something useful please let me know and I will try to
> supply it.
[SNIP]
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 64800092 101527218 3401510871 98.09 6868834 9821955 32234899 78.69
>
[SNIP]
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 7884672 0 2730403448 0 0 1257 672142 70099
BUFWAITS ratio: (7884672 / (101527218 + 32234899)) * 100 = 5.98454%
That's OK.
[SNIP]
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 64800393 101527530 3401750070 98.10 6868844 9822051 32236046 78.69
>
[SNIP]
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 7884770 0 2730489989 0 0 1257 672174 70103
BUFWAITS ratio: (7884770 / (101527530 + 32236046)) * 100 = 5.89455%
So is that and the ratio of the differences is at about 6.72% which is
still OK.
[SNIP]
> $ onstat -R>
> Informix Dynamic Server Version 7.30.UC3 -- On-Line (CKPT REQ) -- Up 5
> days 21
> :08:31 -- 2736128 Kbytes
> Blocked:CKPT
>
> 7 buffer LRU queue pairs priority levels
> # f/m pair total % of length LOW MED_LOW MED_HIGH HIGH
> 0 f 109696 98.8% 108410 0 74100 32686 1624
> 1 m 1.2% 1286 0 1286 0 0
> 2 f 109695 98.8% 108397 0 74400 32306 1691
> 3 m 1.2% 1298 0 1298 0 0
> 4 F 109695 98.8% 108432 0 74693 32134 1605
> 5 m 1.2% 1263 0 1263 0 0
> 6 f 109696 98.9% 108452 0 74319 32453 1679
> 7 m 1.1% 1244 0 1244 0 0
> 8 f 109695 98.9% 108442 0 74245 32548 1649
> 9 m 1.1% 1253 0 1253 0 0
> 10 f 109695 98.8% 108396 0 74289 32324 1783
> 11 m 1.2% 1299 0 1299 0 0
> 12 f 109695 98.8% 108432 0 74844 31875 1713
> 13 m 1.2% 1263 0 1263 0 0
> 8906 dirty, 767867 queued, 768000 total, 1048576 hash buckets, 2048 buffer
> size
> start clean at 3% (of pair total) dirty, or 3291 buffs dirty, stop at 1%
With 30% of your buffers in MED_HIGH priority you may be hitting the bug
in 7.3x that misidentifies index nodes and leaves and makes them all
MED_HIGH. Watch for the ratio of MED_HIGH buffers to increase over time.
> $ onstat -F | head>
> Informix Dynamic Server Version 7.30.UC3 -- On-Line -- Up 5 days
> 21:08:39 -- 2
> 736128 Kbytes
>
> Fg Writes LRU Writes Chunk Writes
> 3 2454820 4032951
62% of your writes are CHUNK writes. This leads to long checkpoints (look
at the online log, for an OLTP system you want 1-5 second checkpoints
ideally). Adjust LRU_MIN/MAX_DIRTY down to 1,0 to minimize chunk writes.
>
> address flusher state data
> 860424cc 0 I 0 = 0X0
> 86042980 1 I 0 = 0X0
[SNIP]
> ONCONFig follows--------
>
> ROOTNAME rootdbs # Root dbspace name> ROOTPATH /dev/infx_vx/rootdbs_plog # Path for device of root
Does this link name signify that the disk is being used for both the
rootdbs dbspace and the plog dbspace? That would be bad.
> ROOTOFFSET 0 # Offset of root dbspace into device
> (Kbytes)
> ROOTSIZE 20000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH # Path for device containing mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS plog # Location (dbspace) of physical log
> #PHYSFILE 50000 # Physical log file size (Kbytes)
> PHYSFILE 99884 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> #LOGFILES 30 # Number of logical log files
> #LOGSIZE 100000 # Logical log size (Kbytes)
> LOGFILES 3