Buffer waits
Posted in 2007
Topics: Performance & Tuning, Platform-Specific Issues, Internationalization & Character Sets
HI, Folks,
Just curious, we have a quite number of buffer waits(onstat -p) and we know
they were caused primarily by a sequential scan of a large table.
As I observed, the percentage (onstat -R) of the free buffers was still
high (> 90%) when the session claimed a buffer wait(onstat -u).
Any idea?
Thanks,
Frank
Build Version: 10.00.UC5
Build Number: N202
Build Host: ibm6c1b
Build OS: AIX 5.2
Build Date: Thu May 18 00:09:08 CDT 2006
GLS Version: glslib-4.00.UC8
Onstat -R shows clean versus dirty buffers, not used versus unused. Onstat -P
on the first line which is partnum zero (0) the 'other' column is the number of
unused buffers, usually zero for any server that's been online for a reasonable
time.
Bufwaits have two sources, one is sessions waiting to modify a particular
buffer that's already being modified by another session. The other is a
session waiting to latch an LRU queue in order to move a clean page to the
dirty queue to modify it or from the LRU end to the MRU end of the clean queue
in order to read a different disk page into it. This latter is the most common
cause of bufwaits and amounts to contention for the LRU queues. You can tell
of this contention is significant using my BR or Bufwaits Ratio metric:
BR = (bufwaits/(pagreads + bufwrits)) * 100
If this value is less than 7 you probably fine. Between 7 and 10 many of your
users are probably experiencing noticable slowness in response times and if
it's greater than 10 likely everyone is complaining about the system being dog
slow.
The solution is:
1) increase BUFFERS if the unused buffers (as above) is zero and/or viewing
onstat -P over time you see a small number of partnums trading largepercentages of the buffer cache back and forth.
2) Increase LRUS and CLEANERS (CLEANERS should always be >= LRUS for best LRU
flush performance) significantly (avoid 32 & 64 there used to be a bug that
cause VERY poor LRU contention at those values that Informix never found.
Don't know if it still exists).
Art S. Kagel
----- Original Message -----
From: Frank <ids@iiug.org>
At: 8/13 10:20:08
HI, Folks,
Just curious, we have a quite number of buffer waits(onstat -p) and we know
they were caused primarily by a sequential scan of a large table.
As I observed, the percentage (onstat -R) of the free buffers was still
high (> 90%) when the session claimed a buffer wait(onstat -u).
Any idea?
Thanks,
Frank
Build Version: 10.00.UC5
Build Number: N202
Build Host: ibm6c1b
Build OS: AIX 5.2
Build Date: Thu May 18 00:09:08 CDT 2006
GLS Version: glslib-4.00.UC8
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Thanks very much, Art!
My BR value is around 1, so, no problem for that.
As we have LRUS=64 and CLEANERS=16 ( see below), the configuration was
based on a paper( or presentation) suggestion on our new disk(Not exactly
know the disk architecture layout at this moment).
But our 64 LRUS may be a concern?
BUFFERPOOL size=4K,buffers=400000,lrus=64,lru_min_dirty=2.000000
,lru_max_dirty=5.000000
CLEANERS 16 # Number of buffer cleaner processes
Thanks again,
Frank
On 8/13/07, ART KAGEL, BLOOMBERG/ 731 LEXIN <kagel@bloomberg.net> wrote:
>
> Onstat -R shows clean versus dirty buffers, not used versus unused. Onstat
> -P
> on the first line which is partnum zero (0) the 'other' column is the
> number
> of
> unused buffers, usually zero for any server that's been online for a
> reasonable
> time.
>
> Bufwaits have two sources, one is sessions waiting to modify a particular
> buffer that's already being modified by another session. The other is a
> session waiting to latch an LRU queue in order to move a clean page to the
> dirty queue to modify it or from the LRU end to the MRU end of the clean
> queue
> in order to read a different disk page into it. This latter is the most
> common
> cause of bufwaits and amounts to contention for the LRU queues. You can
> tell
> of this contention is significant using my BR or Bufwaits Ratio metric:
>
> BR = (bufwaits/(pagreads + bufwrits)) * 100
>
> If this value is less than 7 you probably fine. Between 7 and 10 many of
> your
> users are probably experiencing noticable slowness in response times and
> if
> it's greater than 10 likely everyone is complaining about the system being
> dog
> slow.
>
> The solution is:
> 1) increase BUFFERS if the unused buffers (as above) is zero and/or
> viewing
> onstat -P over time you see a small number of partnums trading large> percentages of the buffer cache back and forth.
>
> 2) Increase LRUS and CLEANERS (CLEANERS should always be >= LRUS for best
> LRU
> flush performance) significantly (avoid 32 & 64 there used to be a bug
> that
> cause VERY poor LRU contention at those values that Informix never found.
> Don't know if it still exists).
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Frank <ids@iiug.org>
> At: 8/13 10:20:08
>
> HI, Folks,
>
> Just curious, we have a quite number of buffer waits(onstat -p) and we
> know
> they were caused primarily by a sequential scan of a large table.
>
> As I observed, the percentage (onstat -R) of the free buffers was still
> high (> 90%) when the session claimed a buffer wait(onstat -u).
>
> Any idea?
>
> Thanks,
> Frank
>
> Build Version: 10.00.UC5
> Build Number: N202
> Build Host: ibm6c1b
> Build OS: AIX 5.2
> Build Date: Thu May 18 00:09:08 CDT 2006
> GLS Version: glslib-4.00.UC8
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
If LRUS=64 is not giving you trouble (and with BR==1 it's not) then it looks
like the bug went away or was fixed (it looked like a quirk in the code that
selects a different LRU queue if it has to wait too long - it seemed that all
waiters were rehashing to the same next LRU so most were just waiting again).
Until you upgrade to 11.10 - where LRU flushing and how fast it completes will
become less important because of the new non-blocking checkpoint code - keep
CLEANERS >= LRUS so set CLEANERS at 64 (assuming you don't have other buffer
pools).
Art S. Kagel
----- Original Message -----
From: Frank <ids@iiug.org>
At: 8/13 11:55:39
Thanks very much, Art!
My BR value is around 1, so, no problem for that.
As we have LRUS=64 and CLEANERS=16 ( see below), the configuration was
based on a paper( or presentation) suggestion on our new disk(Not exactly
know the disk architecture layout at this moment).
But our 64 LRUS may be a concern?
BUFFERPOOL size=4K,buffers=400000,lrus=64,lru_min_dirty=2.000000
,lru_max_dirty=5.000000
CLEANERS 16 # Number of buffer cleaner processes
Thanks again,
Frank
On 8/13/07, ART KAGEL, BLOOMBERG/ 731 LEXIN <kagel@bloomberg.net> wrote:
>
> Onstat -R shows clean versus dirty buffers, not used versus unused. Onstat
> -P
> on the first line which is partnum zero (0) the 'other' column is the
> number
> of
> unused buffers, usually zero for any server that's been online for a
> reasonable
> time.
>
> Bufwaits have two sources, one is sessions waiting to modify a particular
> buffer that's already being modified by another session. The other is a
> session waiting to latch an LRU queue in order to move a clean page to the
> dirty queue to modify it or from the LRU end to the MRU end of the clean
> queue
> in order to read a different disk page into it. This latter is the most
> common
> cause of bufwaits and amounts to contention for the LRU queues. You can
> tell
> of this contention is significant using my BR or Bufwaits Ratio metric:
>
> BR = (bufwaits/(pagreads + bufwrits)) * 100
>
> If this value is less than 7 you probably fine. Between 7 and 10 many of
> your
> users are probably experiencing noticable slowness in response times and
> if
> it's greater than 10 likely everyone is complaining about the system being
> dog
> slow.
>
> The solution is:
> 1) increase BUFFERS if the unused buffers (as above) is zero and/or
> viewing
> onstat -P over time you see a small number of partnums trading large> percentages of the buffer cache back and forth.
>
> 2) Increase LRUS and CLEANERS (CLEANERS should always be >= LRUS for best
> LRU
> flush performance) significantly (avoid 32 & 64 there used to be a bug
> that
> cause VERY poor LRU contention at those values that Informix never found.
> Don't know if it still exists).
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Frank <ids@iiug.org>
> At: 8/13 10:20:08
>
> HI, Folks,
>
> Just curious, we have a quite number of buffer waits(onstat -p) and we
> know
> they were caused primarily by a sequential scan of a large table.
>
> As I observed, the percentage (onstat -R) of the free buffers was still
> high (> 90%) when the session claimed a buffer wait(onstat -u).
>
> Any idea?
>
> Thanks,
> Frank
>
> Build Version: 10.00.UC5
> Build Number: N202
> Build Host: ibm6c1b
> Build OS: AIX 5.2
> Build Date: Thu May 18 00:09:08 CDT 2006
> GLS Version: glslib-4.00.UC8
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.