Some performance questions
Posted in 1999
A DBA saw bufwaits running over 10% of dskreads+bufwrits with LRUS=64 and asked what a healthy ratio is and what the trade-offs of raising LRUS are. Art Kagel explained that LRUS=64 (and 96) hits a known bug causing bufwait storms, and advised avoiding powers of two (32, 127, 128 are safe); more LRUs reduce contention but too many with too few buffers cause LRU rehashes and spinwaits. He suggested aiming for a 3-7% bufwaits ratio, with over 15% being serious, and noted 7.31's LRUPOLICY. The poster switched to 81 but results weren't yet in. Art also explained that btree pages being MED-HIGH priority in 7.3x explains the 98% btree buffers, with no tuning knob available. No confirmed outcome is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning
Currently my bufwaits are in excess of 10% of dskreads+bufwrits. I believe this to be too high (my LRUs are 64 -- and being changed tonight), but don't know what a 'good' value should be. Any suggestions?? What is the tradeoff when I'm increasing LRUs, while holding all other values constant??
Doug Agnew wrote: > > Currently my bufwaits are in excess of 10% of dskreads+bufwrits. I believe > this to be too high (my LRUs are 64 -- and being changed tonight), but don't > know what a 'good' value should be. Any suggestions?? > > What is the tradeoff when I'm increasing LRUs, while holding all other > values constant?? Hi Doug, LRUS==64 is a 'magic' number. There seems to be a bug that sends bufwaits through the roof for that value and a few others. To be on the safe side try a number that is not a power of two (are we DP types creatures of habit or what?!?). Known values that do not trigger the bug are 32, 127, & 128 though there are many others that are OK. The trade off is that with more LRUS you get less contention between sessions needing a clean buffer to dirty or overwrite so that bufwaits tends to go down with higher values. If LRUS is too high and you do not have a large enough number of buffers then there will be no clean buffers in the selected LRU too often causing many LRU rehashes which slows updates and queries down. You can see this situation in increased spinwaits since an app will spin waiting for an LRU latch to clear. More rehashes will probably result in more clashes. If you change LRUS slightly, to say 63 or 65 and things improve then your only problem was the 'magic' number 64 and you do not need more LRUS just a different number than 64. If bufwaits do not improve by changing it slightly try a larger number (not 96, another node in the bug). Art S. Kagel
Art S. Kagel wrote in message <37681A38.C3F3620B@bloomberg.net>...
SNIP
>Hi Doug, LRUS==64 is a 'magic' number. There seems to be a bug that
>sends bufwaits through the roof for that value and a few others. To
>be on the safe side try a number that is not a power of two (are we
>DP types creatures of habit or what?!?). Known values that do not
>trigger the bug are 32, 127, & 128 though there are many others that
>are OK. The trade off is that with more LRUS you get less contention
>between sessions needing a clean buffer to dirty or overwrite so that
>bufwaits tends to go down with higher values. If LRUS is too high and
>you do not have a large enough number of buffers then there will be no
>clean buffers in the selected LRU too often causing many LRU rehashes
>which slows updates and queries down. You can see this situation in
>increased spinwaits since an app will spin waiting for an LRU latch to
>clear. More rehashes will probably result in more clashes.
>
>If you change LRUS slightly, to say 63 or 65 and things improve then
>your only problem was the 'magic' number 64 and you do not need more
>LRUS just a different number than 64. If bufwaits do not improve by
>changing it slightly try a larger number (not 96, another node in the
>bug).
>
>Art S. Kagel
Thanks, Art.
I had seen the bug messages (and your tech notes article), which were the
motivation for changing LRUs. Changed last night to 81, but can't tell yet
if there has been a significant move away from the 10% value. Anyway, what
I'm also interested in is 'what is an acceptable range?'
Also, noted in passing (and poking around) that onstat -P showed that 98+%
of buffers were Btree. That seems awful high, but I recognize that I have
no real control over it. Guess I would just have to keep adding buffers
until I began to see more data buffers.
Doug
In article <37681A38.C3F3620B@bloomberg.net>, Art S. Kagel <kagel@bloomberg.net> writes >Hi Doug, LRUS==64 is a 'magic' number. There seems to be a bug that >sends bufwaits through the roof for that value and a few others. To >be on the safe side try a number that is not a power of two (are we >DP types creatures of habit or what?!?). Known values that do not Us comp.sci types would use prime numbers (too many pure maths courses!!). Has this been confirmed as a bug yet or shall I play with the server at work for a while to try and reproduce it??? >trigger the bug are 32, 127, & 128 though there are many others that >are OK. The trade off is that with more LRUS you get less contention >changing it slightly try a larger number (not 96, another node in the >bug). > >Art S. Kagel -- David Williams
Doug Agnew wrote:
>
> Art S. Kagel wrote in message <37681A38.C3F3620B@bloomberg.net>...
> SNIP
> >Hi Doug, LRUS==64 is a 'magic' number. There seems to be a bug that
> >sends bufwaits through the roof for that value and a few others. To
> >be on the safe side try a number that is not a power of two (are we
> >DP types creatures of habit or what?!?). Known values that do not
> >trigger the bug are 32, 127, & 128 though there are many others that
> >are OK. The trade off is that with more LRUS you get less contention
> >between sessions needing a clean buffer to dirty or overwrite so that
> >bufwaits tends to go down with higher values. If LRUS is too high and
> >you do not have a large enough number of buffers then there will be no
> >clean buffers in the selected LRU too often causing many LRU rehashes
> >which slows updates and queries down. You can see this situation in
> >increased spinwaits since an app will spin waiting for an LRU latch to
> >clear. More rehashes will probably result in more clashes.
> >
> >If you change LRUS slightly, to say 63 or 65 and things improve then
> >your only problem was the 'magic' number 64 and you do not need more
> >LRUS just a different number than 64. If bufwaits do not improve by
> >changing it slightly try a larger number (not 96, another node in the
> >bug).
> >
> >Art S. Kagel
>
> Thanks, Art.
> I had seen the bug messages (and your tech notes article), which were the
> motivation for changing LRUs. Changed last night to 81, but can't tell yet
> if there has been a significant move away from the 10% value. Anyway, what
> I'm also interested in is 'what is an acceptable range?'
You want to see the bufwaits ratio, as you calc'd it in your post, at
from 3-7%. More than that and you may need more LRUS or you may, if
you have 7.31, consider playing with the LRUPOLICY parameter values.
In a full blown bufwaits storm I have seen the bufwaits % as high as
40% though anything much over 15% is death.
> Also, noted in passing (and poking around) that onstat -P showed that 98+%
> of buffers were Btree. That seems awful high, but I recognize that I have
> no real control over it. Guess I would just have to keep adding buffers
> until I began to see more data buffers.
Yes I know. The 7.3x buffer priority stuff does that. Btree buffers
are MED-HIGH, if you have no resident tables they are the highest
priority buffers! I told Cem and the rest of the kernel team that the
buffer priority scheme has a fatal flaw. There is no way to set limits
on the % of buffers that can become HIGH or MED-HIGH and no way for a
buffer assigned to a class to ever be used by a lower priority class
page. I further suggested a system of ONCONFIG parameters to let us
tune the allocations by specifying a maximum % for each of HIGH and
MED-HIGH class buffers. In this way we can fine tune the buffer cache
using these two parameters and BUFFERS.
Their answer, as always, was that they were changing things already in
IDS 7.31 by adding Priority aging which would allow buffers that have
not been access for a while to migrate to a lower priority class thus
making them available to that class again.
The problem, as you have seen, is that btree pages are ALWAYS actively
accessed so what quickly happens is that 98% of your buffers become
MED-HIGH Btree pages and you have very few left over for data caching.
Personally I like my solution. It is simple to implement, simple to
understand and use, elegant in operation, gives us control of the
server (yeahhhh), and best of all from Menlo's standpoint moves the
onus of buffer cache tuning from them to us.
Art S. Kagel