Forerground writes?
Posted in 2010
After upgrading from IDS 10 to 11.50 on Solaris 10, Mike saw foreground writes in onstat -F (and thousands when LRU min/max was raised to 50/60), despite 128 cleaners, 128 LRUs, 750k buffers and an undersized 2GB physical log. Replies offered diagnoses rather than a confirmed fix: an IBM support engineer noted FG writes can be bumped by btree merges forced to flush dirty pages before a checkpoint; Lester blamed AUTO_LRU_TUNING (turning it off cured it at another site); Art pointed to the shift from LRU writes to chunk writes and slow I/O; Dave Griffen suggested big index builds dirty buffers faster than cleaners can keep up, advising onstat -R monitoring and raising LRUs to 512. No confirmed resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Installation, Setup & Upgrades, Logging & Checkpoints
11.50.UC6W2X4
Solaris 10
All -
Since our upgrade to version 11 (from 10) we are now seeing foreground writes
showing up in our onstat -F output. Not a huge amount but enough to make me
say hmmmm. Admittedly, we do have a poorly configured physical log size right
now - onstat -g ckp suggests that it be increased from the 2 gig it is now to
~50 gig - but excluding that from the scene (if possible or appropriate) what
would cause these new foregroung writes - assuming our processing has remained
the same since v10? We have cleaners maxed out at 128, 750,000 buffers, 128
lrus, and the max/min is back to 2 and 1, (when I increased this to 60 and 50
per the doc. re. new ckpt methods I saw thousands of FG writes. The instance
supports a great deal of drop big table - create big table - load big table -
index big table, but we ran these jobs with the current config under v10 w/o
seeeing ang FG writes - in fact the buffer pool was smaller.
I will say that I am not 100% or even 32% confident in our underlying I/O
system (giant shared SAN) nor the CPU setup on this server (Suns Virtualized
thingee where we have 128 virtual cpu's via cores or whatever). All this said
I am more interested in the mechanics behind why the fg's suddenly started
occuring with v11 - thinking it has something to do with the new non blocking
ckpts. if the answer is 'your physical log is too small' then cool, we will
have the eunuchs admins carve us out bigger luns.
MM
Mike wrote:
11.50.UC6W2X4
Solaris 10
All -
Since our upgrade to version 11 (from 10) we are now seeing foreground writes
showing up in our onstat -F output. Not a huge amount but enough to make me
say hmmmm. Admittedly, we do have a poorly configured physical log size right
now - onstat -g ckp suggests that it be increased from the 2 gig it is now to
~50 gig - but excluding that from the scene (if possible or appropriate) what
would cause these new foregroung writes - assuming our processing has remained
the same since v10? We have cleaners maxed out at 128, 750,000 buffers, 128
lrus, and the max/min is back to 2 and 1, (when I increased this to 60 and 50
per the doc. re. new ckpt methods I saw thousands of FG writes. The instance
supports a great deal of drop big table - create big table - load big table -
index big table, but we ran these jobs with the current config under v10 w/o
seeeing ang FG writes - in fact the buffer pool was smaller.
I will say that I am not 100% or even 32% confident in our underlying I/O
system (giant shared SAN) nor the CPU setup on this server (Suns Virtualized
thingee where we have 128 virtual cpu's via cores or whatever). All this said
I am more interested in the mechanics behind why the fg's suddenly started
occuring with v11 - thinking it has something to do with the new non blocking
ckpts. if the answer is 'your physical log is too small' then cool, we will
have the eunuchs admins carve us out bigger luns.
MM
Response:
Well, most people think of foreground writes as when the engine tries to get a
free buffer for a page, but there are no non-dirty buffers so 1 has to be
written before a free non-dirty buffer is available. However, if you have lru
min/max set to 2 and 1 then there should be pretty much 0% chance this is the
type of foreground write you are seeing. I then was looking into the code to
see when the foreground write profile counter is increased and compared the
two versions. It appears as though nothing has changed in that sense. But 1 of
the ways we can increase the foreground write counter is when doing the btree
merge. If during a bttree merge a checkpoint pops up before the merge pops out
of it's critical section it looks for certain dirty pages that are part of the
merge and it forces those buffers to be flushed, prior to allowing the
checkpoint to happen, and it counts those pages as foreground writes. Of the
places we do increase the foreground write counter, that seems the most likely
place of the 4 or 5 places it gets bumped.
Jacques Renaut
IBM Informix advanced support
APD Team
Mike,
I saw this at another site after the upgrade, AUTO_LRU_TUNING had been turned
on and it did not kick in fast enough when the system would hit a spike in
load so they got fg writes... Turned it off and went back to manual LRU
settings and fg writes disappeared. Checkpoint times also went up during load
spikes because of this.
Regards - Lester
On 9/13/10 11:40 AM, MIKE MAGIE wrote:
> 11.50.UC6W2X4
> Solaris 10
>
> All -
>
> Since our upgrade to version 11 (from 10) we are now seeing foreground writes
> showing up in our onstat -F output. Not a huge amount but enough to make me
> say hmmmm. Admittedly, we do have a poorly configured physical log size right
> now - onstat -g ckp suggests that it be increased from the 2 gig it is now to
> ~50 gig - but excluding that from the scene (if possible or appropriate) what
> would cause these new foregroung writes - assuming our processing has
remained
> the same since v10? We have cleaners maxed out at 128, 750,000 buffers, 128
> lrus, and the max/min is back to 2 and 1, (when I increased this to 60 and 50
> per the doc. re. new ckpt methods I saw thousands of FG writes. The instance
> supports a great deal of drop big table - create big table - load big table -
> index big table, but we ran these jobs with the current config under v10 w/o
> seeeing ang FG writes - in fact the buffer pool was smaller.
>
> I will say that I am not 100% or even 32% confident in our underlying I/O
> system (giant shared SAN) nor the CPU setup on this server (Suns Virtualized
> thingee where we have 128 virtual cpu's via cores or whatever). All this said
> I am more interested in the mechanics behind why the fg's suddenly started
> occuring with v11 - thinking it has something to do with the new non blocking
> ckpts. if the answer is 'your physical log is too small' then cool, we will
> have the eunuchs admins carve us out bigger luns.
>
> MM
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
--
______________________________________________________________________
Lester Knutsen lester@advancedatatools.com
Advanced DataTools Corporation Voice: 703-256-0267 x102
Visit our Web page: http://www.advancedatatools.com
______________________________________________________________________
Thanks for the response guys - MM
I would think that it is the switch from doing mostly LRU writes under 10.00
to doing more Chunk writes under 11.50 caused by increasing
LRU_MIN/MAX_DIRTY. The fact that you saw MASSIVE FG writes with the
recommended settings of 50/60 kind of solidifies that guess. How to fix?
Faster IO, faster KAIO response from the OS (assuming KAIO is enabled). I
don't think that the PLOG is the problem, though a bigger PLOG will permit
fewer buffer flushes at checkpoint time, so it MAY alleviate the situation
slightly.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Mon, Sep 13, 2010 at 11:40 AM, MIKE MAGIE <jmmagie@yahoo.com> wrote:
> 11.50.UC6W2X4
> Solaris 10
>
> All -
>
> Since our upgrade to version 11 (from 10) we are now seeing foreground
> writes
> showing up in our onstat -F output. Not a huge amount but enough to make me
> say hmmmm. Admittedly, we do have a poorly configured physical log size
> right
> now - onstat -g ckp suggests that it be increased from the 2 gig it is now
> to
> ~50 gig - but excluding that from the scene (if possible or appropriate)
> what
> would cause these new foregroung writes - assuming our processing has
> remained
> the same since v10? We have cleaners maxed out at 128, 750,000 buffers, 128
> lrus, and the max/min is back to 2 and 1, (when I increased this to 60 and
> 50
> per the doc. re. new ckpt methods I saw thousands of FG writes. The
> instance
> supports a great deal of drop big table - create big table - load big table
> -
> index big table, but we ran these jobs with the current config under v10
> w/o
> seeeing ang FG writes - in fact the buffer pool was smaller.
>
> I will say that I am not 100% or even 32% confident in our underlying I/O
> system (giant shared SAN) nor the CPU setup on this server (Suns
> Virtualized
> thingee where we have 128 virtual cpu's via cores or whatever). All this
> said
> I am more interested in the mechanics behind why the fg's suddenly started
> occuring with v11 - thinking it has something to do with the new non
> blocking
> ckpts. if the answer is 'your physical log is too small' then cool, we will
> have the eunuchs admins carve us out bigger luns.
>
> MM
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--0016361e88321c580604902722bd
Jacques Renaut Wrote:
Well, most people think of foreground writes as when the engine tries to get a
free buffer for a page, but there are no non-dirty buffers so 1 has to be
written before a free non-dirty buffer is available. However, if you have lru
min/max set to 2 and 1 then there should be pretty much 0% chance this is the
type of foreground write you are seeing. I then was looking into the code to
see when the foreground write profile counter is increased and compared the
two versions. It appears as though nothing has changed in that sense. But 1 of
the ways we can increase the foreground write counter is when doing the btree
merge. If during a bttree merge a checkpoint pops up before the merge pops out
of it's critical section it looks for certain dirty pages that are part of the
merge and it forces those buffers to be flushed, prior to allowing the
checkpoint to happen, and it counts those pages as foreground writes. Of the
places we do increase the foreground write counter, that seems the most likely
place of the 4 or 5 places it gets bumped.
Response:
Do not discard the possibility of 100% dirty buffers even with very low lru
min/max. I found that scenario in one of my instances a few years ago. During
large index builds onstat -R confirmed times when dirty buffers exceeded 99%,
even with lru min/max set to .5/3. If an activity can dirty buffers faster
than they are cleaned, when you start or stop the cleaning has limited impact.
Dave Griffen
MIKE MAGIE wrote:
11.50.UC6W2X4
Solaris 10
All -
Since our upgrade to version 11 (from 10) we are now seeing foreground writes
showing up in our onstat -F output. Not a huge amount but enough to make me
say hmmmm. Admittedly, we do have a poorly configured physical log size right
now - onstat -g ckp suggests that it be increased from the 2 gig it is now to
~50 gig - but excluding that from the scene (if possible or appropriate) what
would cause these new foregroung writes - assuming our processing has remained
the same since v10? We have cleaners maxed out at 128, 750,000 buffers, 128
lrus, and the max/min is back to 2 and 1, (when I increased this to 60 and 50
per the doc. re. new ckpt methods I saw thousands of FG writes. The instance
supports a great deal of drop big table - create big table - load big table -
index big table, but we ran these jobs with the current config under v10 w/o
seeeing ang FG writes - in fact the buffer pool was smaller.
I will say that I am not 100% or even 32% confident in our underlying I/O
system (giant shared SAN) nor the CPU setup on this server (Suns Virtualized
thingee where we have 128 virtual cpu's via cores or whatever). All this said
I am more interested in the mechanics behind why the fg's suddenly started
occuring with v11 - thinking it has something to do with the new non blocking
ckpts. if the answer is 'your physical log is too small' then cool, we will
have the eunuchs admins carve us out bigger luns.
MM
Response:
A few years ago, I noticed a problem where large index builds were dirtying
buffers faster than the LRU system was cleaning. During large index builds I
had long checkpoints and a handful of FG writes. Repeated onstat -R confirmed
times when my buffer pool exceeded 99% dirty, even though my LRU min/max was
set to .5/3. The config settings you noted are in range of what I was using at
the time and you also mention creating indexes on big tables.
Try to capture a repeated onstat -R for the duration of a large index build.
Check it to see if your dirty buffers% comes close to the 100% mark or if it
substantially exceeds max dirty% at any point. If you find the same situation
I had, my first suggestion would be to increase LRU's to 512.
Dave Griffen