Bufwaits and Foreground Writes
Posted in 1999
Topics: Performance & Tuning, Installation, Setup & Upgrades, Storage & Space Management, Server Administration, Logging & Checkpoints, Versions, Editions & End-of-Life
Despite our tuning efforts, we are encountering a significant number of
bufwaits and foreground writes
on our production SAP system. This system runs IDS 7.30.UC7XK on a 4-way
RS6000 R50 (AIX).
I attended Informix's Performance Tuning class last week in Menlo Park.
Their suggestions
included:.
1) Use different ONCONFIG at different times of the day.
Our system must be available to online users 24 hours a day, 6 days a
week.
Therefore, this suggestion is not realistic for us.
2) Reduce the number of LRUs to 4 * number of CPU VPs.
Will this really help? Or will it make matters worse?
3) Increase the LRU_MIN_DIRTY and LRU_MAX_DIRTY values to improve write
cache %.
Below are statistics from onstat and SMI queries. This represents less than
50% of our
normal workload due to the holidays. Of concern are:
1) "ovbuff" value in the "onstat -p" and foreground writes in the "onstat
-F".
2) Poor write cache %.
3) Large number of seqscans.
Should I be concerned about these statistics? Or are they possibly caused
by product defects
108509 and 115741? The "seqscans" numbers from the "sysmaster:sysptprof"
table do not seem
realistic. If that many sequential scans were occurring for some of the
large tables, the number
of pages read would be much larger because the entire table could not held
in memory.
Yet we are also seeing a large number of read-aheads which give some
validity to the seqscans
numbers. Does anyone have system-level tuning recommendations? With an SAP
system
upgrade planned for later this year, it is unlikely that I can drum up
support for application-level tuning.
The classroom material from the Performance Tuning class recommend use of
the 'onstat -g spi'
to monitor LRU usage and ensure that adequate queues are allocated.
However, it does not
tell how to interpret the output from this command. The class instructor
suggested that I
monitor the lines containing "vproc" for excessive "Avg Loop/Wait" values.
Is this right?
If so, what values are considered excessive?
Assistance from Informix experts, who participate in this group, would be
greatly appreciated!
=>> 'onstat -p' command output follows:
Informix Dynamic Server Version 7.30.UC7XK -- On-Line -- Up 3 days 19:05:31
-- 1457264 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
34027290 17882518 1257830526 97.29 4359252 3897016 16744531 73.97
isamtot open start read write rewrite delete commit
rollbk
933381413 1817089 75970820 561674437 2554533 1518679 1560434 607349 13
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 2309 150491.52 67514.29 191 382
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
7848325 788 560896897 0 0 247 235256 12909197
ixda-RA idx-RA da-RA RA-pgsused lchwaits
23508356 451550 100557 23854808 466298
=>> Tables with the most seqscans per sysmaster:systprof
table npused bufreads bufwrites seqscans pagreads
pagwrites
knvp 12214 54583788 1177 3751498 17110
523
knvv 4597 25010688 67 3736198 13672
56
vbfa 390751 93265427 238630 2058429 1570756
47803
vbpa 259503 22307877 95060 1657259 425133
3824
vbap 263172 36592659 33988 1072444 835062
9195
knvh 603 13625603 41 263849 11420
24
bsad 198499 1138932 16165 175169 34217
3749
dbsta 22 88887 4 25380 178
4
s066 2108 366205 5406 18109 43414
3754
vapma 62725 8912986 24680 17659 67894
8076
dd04t 13486 834373 0 15397 11890
0
dd04l 2214 1111433 0 15397 8752
0
qals 18760 1944429 4313 14243 8270
1844
dd07t 2084 1790276 0 10436 806
0
dd07l 330 1994037 0 10254 588
0
qpct 101 2076623 0 10216 687
0
hrp10 25 42136 0 8516 44
0
mkol 2 12341 0 6151 15
0
=>> 'onstat -F' command output follows:
Fg Writes LRU Writes Chunk Writes
2309 1121507 374479
=>> 'onstat -P' command output follows:
Informix Dynamic Server Version 7.30.UC7XK -- On-Line -- Up 3 days 19:05:32
-- 1457264 Kbytes
partnum total btree data other resident dirty
Totals: 150000 20661 128688 651 0 2436
Percentages:
Data 85.79
Btree 13.77
Other 0.43
=>> 'onstat -c' command output follows:
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace nameROOTPATH /informix/PRD/sapdata/physdev0/data02
# Path for device containing root dbspace
ROOTOFFSET 16 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 200000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH /informix/PRD/sapdata/physdev21/data212
# Path for device containing mirrored root
MIRROROFFSET 16 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS physdbs # Location (dbspace) of physical log
PHYSFILE 229148 # Physical log file size (Kbytes)092098rb
# Logical Log Configuration
LOGFILES 98 # Number of logical log files
LOGSIZE 500 # initial (!) Logical log size (Kbytes)
# Diagnostics
MSGPATH /informix/PRD/online.wrax39210.prd.log
# System message log file path
# CONSOLE /informix/PRD/console.wrax39210.prd.log
CONSOLE /dev/null
# System console message path
#ALARMPROGRAM /informix/PRD/etc/log_full.sh # Alarm program path120497rb
ALARMPROGRAM /scripts/adsm/log_full.sh # Alarm program path
120497rb
# System Archive Tape Device
#TAPEBLK 256 # Tape block size (Kbytes)
TAPESIZE 1950000 # Maximum amount of data to put on tape
(Kbytes)
TAPEDEV /dev/rmt1 # Tape device path
TAPEBLK 1024 # Tape block size (Kbytes)
# Log Archive Tape Device
LTAPEDEV /tmp/tapedev # Log tape device path
LTAPEBLK 256 # Log tape block size (Kbytes)
LTAPESIZE 1950000 # Max amount of data to put on log tape
(Kbytes)
# Optical
STAGEBLOB # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 20 # Unique id c
Also sprach "Bernstein, Rick" <rbernste@alarismed.com> :
:
:Despite our tuning efforts, we are encountering a significant number of
:bufwaits and foreground writes
:on our production SAP system. This system runs IDS 7.30.UC7XK on a 4-way
:RS6000 R50 (AIX).
:I attended Informix's Performance Tuning class last week in Menlo Park.
:Their suggestions
:included:.
:2) Reduce the number of LRUs to 4 * number of CPU VPs.
: Will this really help? Or will it make matters worse?
I doubt it would help, and would probably make things a bit worse. You might
want to consider increasing your LRUS. As it stands you have around 2450
buffers per queue, which is a bit high.
:3) Increase the LRU_MIN_DIRTY and LRU_MAX_DIRTY values to improve write
:cache %.
This will, in fact, increase your write cache. It will also hurt your
performance. The key thing to understand here is that low writes cache rates do
not equate to poorer performance. In fact, write cache rate is really
meaningless. ALL writes with informix are cached, i.e. all database updates
occur in the buffer pool, and the changed pages are always flushed later. You
could have more flushing done at checkpoint time, which decreases the liklihood
that a given page may be flushed many times during its life in the buffer pool.
But checkpoints block user activity when they occur, so doing more writes at
checkpoint time means longer checkpoints and more holdup of user activity. By
keeping low LRU_MAX_DIRTY/MIN_DIRTY values, you are causing buffers to be
flushed more in between checkpoints, INCREASING the liklihood that a given page
may be flushed many times during its time in the buffer pool. But since
discrete buffer flushes are asynchronous to user activity, and have minimal
impact on it, who cares? Users will see better performance. You will see lower
write cache rates, but as I already said, they are meaningless as a measure of
application performance. As long as your i/o is not causing a bottleneck, you
should be fine.
:
:Below are statistics from onstat and SMI queries. This represents less than
:50% of our
:normal workload due to the holidays. Of concern are:
:1) "ovbuff" value in the "onstat -p" and foreground writes in the "onstat
:-F".
Your ratio of fg writes is very low. Bear in mind that some activities related
to system maintenance seem to always do fg writes. I've never been able to find
out why this is so, or exactly what activities cause them, but I really wouldn't
be too concerned with what you are seeing.
:2) Poor write cache %.
As per above, forget about write cache rates.
:3) Large number of seqscans.
This is curious. Your RA params are not all that high. You are using about 99%
of what you are reading ahead, and most of them are indexed-read access, so
there is decent use of indexes going on. Keep in mind that seq scans may be
being done on very small tables that are frequently accessed. In that case, it
does not equate with disk access, as small tables can be effectively cached in
shared memory, so even seq scans of them are fast. If you were really doing seq
scans of big tables all the time, you would see a much higher value in da-RA.
:
:Should I be concerned about these statistics? Or are they possibly caused
:by product defects
:108509 and 115741? The "seqscans" numbers from the "sysmaster:sysptprof"
:table do not seem
:realistic. If that many sequential scans were occurring for some of the
:large tables, the number
:of pages read would be much larger because the entire table could not held
:in memory.
:Yet we are also seeing a large number of read-aheads which give some
:validity to the seqscans
:numbers. Does anyone have system-level tuning recommendations? With an SAP
:system
:upgrade planned for later this year, it is unlikely that I can drum up
:support for application-level tuning.
:
:The classroom material from the Performance Tuning class recommend use of
:the 'onstat -g spi'
:to monitor LRU usage and ensure that adequate queues are allocated.
:However, it does not
:tell how to interpret the output from this command. The class instructor
:suggested that I
:monitor the lines containing "vproc" for excessive "Avg Loop/Wait" values.
:Is this right?
:If so, what values are considered excessive?
:
:Assistance from Informix experts, who participate in this group, would be
:greatly appreciated!
:
:=>> 'onstat -p' command output follows:
:
:Informix Dynamic Server Version 7.30.UC7XK -- On-Line -- Up 3 days 19:05:31
:-- 1457264 Kbytes
:
:Profile
:dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
:34027290 17882518 1257830526 97.29 4359252 3897016 16744531 73.97
:
:isamtot open start read write rewrite delete commit
:rollbk
:933381413 1817089 75970820 561674437 2554533 1518679 1560434 607349 13
:
:gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
:0 0 0 0 0 0 0
:
:ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
:0 0 2309 150491.52 67514.29 191 382
:
:bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
:7848325 788 560896897 0 0 247 235256 12909197
:
:ixda-RA idx-RA da-RA RA-pgsused lchwaits
:23508356 451550 100557 23854808 466298
:=>> 'onstat -F' command output follows:
:
:Fg Writes LRU Writes Chunk Writes
:2309 1121507 374479
:
Dave
In article <84gb3g$sho$1@news.xmission.com>,
"Bernstein, Rick" <rbernste@alarismed.com> wrote:
>
> Despite our tuning efforts, we are encountering a significant number
of
> bufwaits and foreground writes
> on our production SAP system. This system runs IDS 7.30.UC7XK on a
4-way
> RS6000 R50 (AIX).
I just want to share little bit from my recent experience.
We had 0 foreground writes when we was on 7.30UC7XB
Foreground writes appeared just after upgrade to 7.30UC7XK.
Nothing changed on system - only upgrade.
System is 4CPU, Solaris, SAP 4.5B.
That number is always increasing. Sometimes during checkpoint,
sometimes when engine flushing when % of dirty buffers reaches
LRU_MAX_DIRTY. I tried (without success) to catch is it really
foreground write or it's just wrong statistics. I believe
engine is not performing foreground write at that time, but
there isn't proof of that.
I contacted SAP and got answer something like:
Some of our customers using 7.30.UC7XK also found similar
situation with increasing number of foreground writes ...
not affecting performance ...
you can ignore that value in onstat -F output.
I'm writing from home and can not provide exact answer from
SAP but meaning is above.
Regards
Vardan
from home
vaar@earthlink.net
<SNIP>
> Thank you for any assistance or insight which you can provide.
>
> RIck Bernstein
>
>
--
Vardan Aroustamian
vaar@geocities.com
Sent via Deja.com http://www.deja.com/
Before you buy.
Got to toss in my $0.02US. Everyone is giving you basically good
advice, except those Informix guys (maybe I should give a Performance
Tuning class for Informix consultants and trainers!). I just wanted to
add the weight of my opinion so you take this good advice. Read on.
"Bernstein, Rick" wrote:
>
> Despite our tuning efforts, we are encountering a significant number of
> bufwaits and foreground writes
> on our production SAP system. This system runs IDS 7.30.UC7XK on a 4-way
> RS6000 R50 (AIX).
> I attended Informix's Performance Tuning class last week in Menlo Park.
> Their suggestions
> included:.
> 1) Use different ONCONFIG at different times of the day.
> Our system must be available to online users 24 hours a day, 6 days a
> week.
> Therefore, this suggestion is not realistic for us.
Nor for anyone else running a REAL system. Ask them if they could
bounce the PTS servers to change ONCONFIG files every 12 hours. That
is usually how I get their attention when they suggest I try something
stupid on a production system. Amazingly they ALWAYS give me the same
answer that had them puzzled when I gave it to them a minute earlier:
"Can't do that on a production system without testing it out first."
Of course the conversation usually begins because we or they cannot
duplicate the problem EXCEPT under production conditions.
> 2) Reduce the number of LRUs to 4 * number of CPU VPs.
> Will this really help? Or will it make matters worse?
Ha ha ha ha ha. AAAhhhh ha ha ha ha. Wait my sides hurt. Let me
catch my breath. Let's review, your bufwaits ratio is around 20% when
an ideal ratio is 7% and anything over 10% is death, and they want you
to REDUCE the number of LRUS from 60 to 16? You're killing me! OK let
me explain to you so you can expain to those jokers, bufwaits indicate
contention for access to buffers. This can ONLY happen on one of two
situations: 1) multiple threads want access to the same buffer or 2)
multiple threads need to read a page from disk into an empty or LRU'd
buffer or to create a new page in an empty or LRU'd buffer. For the
first, true buffer contention, the buffer is accessed directly and
there is not much that you can do beyond modifying the general work
flow to reduce contention at the user level. Hoever, the latter, LRU
contention, which is MUCH more common, the problem is that to write
into a buffer a thread must acquire one from the LRU end of some LRU
queue's clean list, more it to MRU end (or to the dirty list if
modifying it), and write the disk page or new data into the buffer.
So the point of contention is the LRU queues themselves either there
are not enough LRUS to prevent contention or all threads are hashing to
the same queue. The latter happens for some reason at LRUS == 64 or 96
only. The solution to the former problem is obvious, add MORE LRUS not
fewer. I say go for broke and up LRUS to 128 (if you get some weird
but harmless message at startup about your logical and physical logs
being zero length change to 127 this is fixed in 7.31 but I don't know
about 7.30).
> 3) Increase the LRU_MIN_DIRTY and LRU_MAX_DIRTY values to improve write
> cache %.
DON'T DO IT! PERIOD! Keep those LRU_MIN/MAX_DIRTY parameters right
where they are. You want better write cache% add more buffers and
find out what indexes to add to reduce sequential scans.
> Below are statistics from onstat and SMI queries. This represents less than
> 50% of our
> normal workload due to the holidays. Of concern are:
> 1) "ovbuff" value in the "onstat -p" and foreground writes in the "onstat
> -F".
Increasing BUFFERS and LRUS will reduce the FGWrites which are caused
when a thread needs a buffer to write to but cannot find one that is
not already dirty after trying several LRU queues.
> 2) Poor write cache %.
More buffers, fewer seqential scans.
> 3) Large number of seqscans.
Add indexes where needed. Update stats using the recommended suite of
UPDATE STATISTICS commands as stated in the Performance Guide andimplemented in my dostats utility (and others).
> Should I be concerned about these statistics? Or are they possibly caused
> by product defects
> 108509 and 115741? The "seqscans" numbers from the "sysmaster:sysptprof"
> table do not seem
> realistic. If that many sequential scans were occurring for some of the
> large tables, the number
> of pages read would be much larger because the entire table could not held
> in memory.
> Yet we are also seeing a large number of read-aheads which give some
> validity to the seqscans
> numbers. Does anyone have system-level tuning recommendations? With an SAP
> system
> upgrade planned for later this year, it is unlikely that I can drum up
> support for application-level tuning.
>
> The classroom material from the Performance Tuning class recommend use of
> the 'onstat -g spi'
> to monitor LRU usage and ensure that adequate queues are allocated.
> However, it does not
> tell how to interpret the output from this command. The class instructor
> suggested that I
> monitor the lines containing "vproc" for excessive "Avg Loop/Wait" values.
Filter the output for the string "LRU" and look at the Loops/Wait
column and number of waits. Compare the latter to the number of isam
start operations to get a sence of the % of requests that wait for an
LRU and look at Loops/Wait to get a sence of how long the waits are.
The former percentage could exceed 100% indicating that a large number
of requests are timing out while waiting and are having to rehash at
least once to another LRU queue to get served.
> Is this right?
> If so, what values are considered excessive?
>
> Assistance from Informix experts, who participate in this group, would be
> greatly appreciated!
>
> =>> 'onstat -p' command output follows:
>
> Informix Dynamic Server Version 7.30.UC7XK -- On-Line -- Up 3 days 19:05:31
> -- 1457264 Kbytes
>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 34027290 17882518 1257830526 97.29 4359252 3897016 16744531 73.97
>
> isamtot open start read write rewrite delete commit
> rollbk
> 933381413 1817089 75970820 561674437 2554533 1518679 1560434 607349 13
If these stats are for the entire 3 days you buffer turnover rate is
about every 45 minutes which is not great but not terrible either. I'd
still recommend increasing BUFFERS by AT LEAST 50% maybe doubling it.
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 0 0 0 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 2309 150491.52 67514.29 191 382
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 7848325 788 560896897 0 0 247 235256
Your bufwaits ratio is about 20% (quick calc) which is killing
performance. Go for 128 LRUS and CLEANERS (CLEANERS should be >= LRUS).
12909197
>
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 23508356 451550
I think your first priority has got to be Application tuning - reducing the sequential scans on your large tables (high npused, high seqscans). For example, every sequential scan of table vbfa (npused = 390751) is going to attempt to clean out your BUFFERS twice over and some (assuming detached indexes). A table of this size (390751 * 4 = over 1.5 Gb ) should *never* be sequentially scanned as part of regular operations. Figure out why these scans are occurring and do something to stop them (update statistics, modify application to use existing indexes, modify existing indexes, add indexes). I would doing the following (a) Modify my instance for obvious discrepancies (I can't see anything really pressing) (b) Pick the top five tables in descending order of "sequential scans * npused" and work on reducing sequential scans. (c) Evaluate impact of such tuning (talk to users, examine instance statistics, examine SAP statistics) (d) Fine tune instance, if required (e) Go back to (b). You may well find dramatic improvements after the first iteration. Rudy "Bernstein, Rick" wrote: > Despite our tuning efforts, we are encountering a significant number of > bufwaits and foreground writeson our production SAP system. This system runs > IDS 7.30.UC7XK on a 4-way RS6000 R50 (AIX). > > Informix Dynamic Server Version 7.30.UC7XK -- On-Line -- Up 3 days 19:05:31-- > 1457264 Kbytes > > Profile > dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached > 34027290 17882518 1257830526 97.29 4359252 3897016 16744531 73.97 > ... > bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans > 7848325 788 560896897 0 0 247 235256 12909197 > > ixda-RA idx-RA da-RA RA-pgsused lchwaits > 23508356 451550 100557 23854808 466298 > > =>> Tables with the most seqscans per sysmaster:systprof > table npused bufreads bufwrites seqscans pagreads pagwrites > knvp 12214 54583788 1177 3751498 17110 523 > knvv 4597 25010688 67 3736198 13672 56 > vbfa 390751 93265427 238630 2058429 1570756 47803 > vbpa 259503 22307877 95060 1657259 425133 3824 > vbap 263172 36592659 33988 1072444 835062 9195 > ...