Re: Informix FAQ
Posted in 2010
Topics: Performance & Tuning, Server Administration, Versions, Editions & End-of-Life
Reinhard, I don't remember if I answered you last week, so here goes.
That's an interesting observation. I've been looking for a way to make the
BTR more meaningful or develop another similar metric that better estimates
cache churning. I'll have to consider your figures for a while and work it
out.
The idea is to find the stats that can be interpreted as indicating that a
portion of the cache was replaced. Pagwrites don't do it because just
because I wrote out a page to disk doesn't mean that it was replaced in the
cache. Indeed, in a server like yours with a high write cache percentage
any pages being written to are likely at the most recent end of each LRU
queue, so pagwrites is even less meaningful than in general.
Here's my reasons for using pagreads and bufwrits. In order to write to a
buffer, either the page being written to is already in cache, is a new page,
or the page has to be loaded from disk. In the first case, there is no
change to the cache contents after the write. In the other two cases,
however, there is a definite change as existing data in the cache had to be
dropped to make room for the page that was to be written. In a server with
massively high write cache percentages, it's clear that most of the writes
(99.82%) are to the same pages over and over again, so the value of
bufwrites is minimal for this calculation. There you are correct.
The other side of the BTR are the pagreads. This represents pages that were
read from disk into the cache. This has a 100% valid effect on the cache
contents. Every page that is read from disk MUST replace another page in
the cache. If the cache is insufficiently sized many of these pages are
re-reads of pages forced out by previous reads and this WILL affect
performance!
Back to your server. Whether your users are noticing or not, your pagreads
ALONE indicate that your cache is not big enough. The adjusted turnover
calculation - only taking pagreads into account - is:
(338619630 / 2100000) / 9.95hrs = 16.2 turns per hour
So your buffer cache is being churned every 3.7 minutes JUST ON pagreads and
that's if every page in the cache is available for pagreads. They are not.
Your pagwrits indicate that you wrote 2,620,000 buffers per hour to disk.
Most of those were buffers being rewritten many times, however, it does show
that a significant portion of the buffer cache was too active to ever be
replaced by new pages meaning that actually a smaller percentage of the
cache was being churned even more frequently then every 3.7 minutes!
Here's the question you have to ask: What is the valid residence duration
for data to remain useful in my cache? If the answer is more than 3.7
minutes then you ARE suffering from performance degradation due to a too
small, albeit objectively massive, buffer cache! If the answer is 7 minutes
then you MAY need to increase your cache to something over 3,000,000 pages.
If you have a more direct way to determine what your valid working set of
data is and your valid residence duration then you can empirically size your
cache. If not, and most of us cannot do so with any accuracy, then you can
only guess at a new cache size, try it, and calculate the BTR and adjusted
BTR again.
You are correct, the BTR as it stands does not help you look at your server,
but this adjusted BTR does and I'll have to start incorporating it into my
own calculations. For that, thanks!
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
See you at the 2010 IIUG Informix Conference
April 25-28, 2010
Overland Park (Kansas City), KS
www.iiug.org/conf
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Thu, Jan 14, 2010 at 4:23 AM, Habichtsberg, Reinhard <
RHabichtsberg@arz-emmendingen.de> wrote:
> Art,
>
> newratios.ksh reports on one of the instances:
>
> Metric Ratio Report For 2K Cache
>
> Bufwaits Ratio: 0.140000%
> Buffer Turnover Rate: 363.58/hr
>
>
> Reset Time: 2010-01-14 00:04:59
> Time elapsed: 9 Std. 56 Min. 03 Sek. (Modified the script a bit, so you can
> see these values).
>
> We have only 2K Buffers.
>
> The BTR would imply an turnaround of the buffers all 10 seconds.
>
> onstat -p shows:
> IBM Informix Dynamic Server Version 11.50.FC5W4 -- On-Line -- Up 26 days
> 21:46:18 -- 6269952 Kbytes>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits
> %cached
> 179093619 338619630 20158565492 99.11 11559700 26071392 6594149951
> 99.82
>
> You see, bufwrits are 6594149951 and pagewrits only 6594149951. The ratios
> pagwrits/bufwrits is 0.004 or 0.4 %.
>
> Is the following conclusion true: If the ratio pagwrits/bufwrits is (very)
> low, then the significance of BTR isn't (very) high.
>
> BTW: We have 2,100,000 Buffers and we hve not too strong complaints.
>
> TIA, Reinhard.
>
>
>
> -----Original Message-----
> From: informix-list-bounces@iiug.org
> [mailto:informix-list-bounces@iiug.org]On Behalf Of Art Kagel
> Sent: Thursday, January 14, 2010 3:12 AM
> To: fandelau
> Cc: informix-list@iiug.org
> Subject: Re: Informix FAQ
>
>
> Hi Ramon,
>
> OK, I'll take a look at the script again over the weekend and let you know
> what's what. Thanks for trying it out and for the feedback.
>
> Art
>
> Art S. Kagel
> Advanced DataTools ( www.advancedatatools.com
> <http://www.advancedatatools.com> )
> IIUG Board of Directors ( art@iiug.org <mailto:art@iiug.org> )
>
> See you at the 2010 IIUG Informix Conference
> April 25-28, 2010
> Overland Park (Kansas City), KS
> www.iiug.org/conf <http://www.iiug.org/conf>
>
> Disclaimer: Please keep in mind that my own opinions are my own opinions
> and
> do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
> organization with which I am associated either explicitly, implicitly, or
> by
> inference. Neither do those opinions reflect those of other individuals
> affiliated with any entity with which I am affiliated nor those of the
> entities themselves.
>
>
>
>
> On Wed, Jan 13, 2010 at 8:45 PM, fandelau < fandelau@gmail.com
> <mailto:fandelau@gmail.com> > wrote:
>
>
> On Jan 7, 7:43 am, Art Kagel < art.ka...@gmail.com
> <mailto:art.ka...@gmail.com> > wrote:
> >
> <SNIP>
>
> > However, it is still a VERY useful metric. I have determined by
> observation
> > of many servers that if the BTR is ~10 or under the server is likely
> happy
> > and users are not complaining about performance. If the BTR is in higher
> > double digits some users will be experiencing slow performance or
> > intermittent responsiveness problems. If BTR is in the upper double or
> > triple digits then your phone is
Art Kagel <art.kagel@gmail.com> wrote:
> The idea is to find the stats that can be interpreted as indicating that a
> portion of the cache was replaced.
what about "onstat -B" + "oncheck -pe" and creating bufferpool characteristics
for each object (table, index)? if you collect a lot of onstat -B outputs you
will be able to calculate per object data:
- % size of bufferpool and % size of object; example: table 'foo' takes 10% of
buffer pool and is cached in 50%
- number of paged-in pages; egample: from last ostat -B output, 100000 new
pages for table 'foo' appeared in the cache
- number of paged-out pages
you can also create chart for pagein/pageout just to see when
you have some pikes and correlate this with sql trace data from
that moment. warning: sql trace data still have some bugs (<=xc6)
that make it useless in some situations so don't rely on it
after that create a charts for objects that have significant % share in
a bufferpool and simply look at them and try to select objects that
should be all the time in cache (objects cached in >90% of their size almost
whole the time, objects that reappear in the cache just after they were paged out
or the characteristics have a chainsaw shape that shows the "fight" for a cache).
calculate minimum size for those objects that you think should stay in cache,
add some space for other objects (you have to know the application that runs
on this instance), change the bufferpool size and measure all again :-)
--
butthead
Bart (hope you don't mind "Bart" I just can't start a post with "Butthead"):
If you want to track the specific contents of the cache by table (or better
any partition) you can do that more efficiently using the onstat -P report
rather than parsing all of the buffers using the onstat -B output or using
the equivalent SMI table. I do do that myself when I have determined that
there is a problem. Running onstat -P to a file several times over a period
of time and diffing the files is sometimes sufficient to make an educated
guess as to how much additional cache space will suffice to relieve the
contention.
However, that does not replace the purpose of the three basic metrics: BR,
BTR, and RAU that the newratios.ksh script calculate. I developed the
metrics so that I could quickly determine if any of the 46 servers I was
managing needed attention when I got into the office in the morning. There
was no time to run any analysis that would require more than a few seconds
on all 46 servers and still have time to correct the problem. The servers
had to be checked out historically based on the previous day's and overnight
activity. Since the trading day wouldn't be starting for another hour or
two, more detailed differential analysis like what you are hinting at
wouldn't help. There was no significant live activity until shortly before
the trading day began and by then it would be too late. I didn't need
300,000 users calling into support to complain that security price history
functions were running slow.
The discussion with Reinhardt goes to whether there are server load types
where the BTR calculation is less useful than than it was at a more typical
OLTP/DSS mix environment like the one I used to manage and where I developed
the metrics and evaluation levels for them. At Reinhardt's site almost all
of their write activity is rewriting or updating pages that are already in
cache, so including bufwrites in the BTR calculation is producing
artificially high values that do not correlate to performance problems.
Unfortunately, my analysis of the data he sent me indicates that while the
traditional BTR is invalid for him there is still performance degradation
that he simply hasn't recognized due to massive read activity. The analysis
you suggest would help him to determine what tables are affected and may
help him to determine how much additional buffer space he needs to reach
peak read performance.
Now, thanks to Reinhardt, I may have developed a BTR-like metric that is
more valid for DSS/DW installations and those OLTP systems like his which
pull in massive quantities of data but update only a small working set.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
See you at the 2010 IIUG Informix Conference
April 25-28, 2010
Overland Park (Kansas City), KS
www.iiug.org/conf
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Fri, Jan 29, 2010 at 2:46 AM, Bartlomiej Lidke
<ohggurnq@rcs.cy.rot13.invalid> wrote:
> Art Kagel <art.kagel@gmail.com> wrote:
> > The idea is to find the stats that can be interpreted as indicating that
> a
> > portion of the cache was replaced.
>
> what about "onstat -B" + "oncheck -pe" and creating bufferpool
> characteristics
> for each object (table, index)? if you collect a lot of onstat -B outputs
> you
> will be able to calculate per object data:
> - % size of bufferpool and % size of object; example: table 'foo' takes 10%
> of
> buffer pool and is cached in 50%
> - number of paged-in pages; egample: from last ostat -B output, 100000 new
> pages for table 'foo' appeared in the cache
> - number of paged-out pages
>
> you can also create chart for pagein/pageout just to see when
> you have some pikes and correlate this with sql trace data from
> that moment. warning: sql trace data still have some bugs (<=xc6)
> that make it useless in some situations so don't rely on it
>
> after that create a charts for objects that have significant % share in
> a bufferpool and simply look at them and try to select objects that
> should be all the time in cache (objects cached in >90% of their size
> almost
> whole the time, objects that reappear in the cache just after they were
> paged out
> or the characteristics have a chainsaw shape that shows the "fight" for a
> cache).
> calculate minimum size for those objects that you think should stay in
> cache,
> add some space for other objects (you have to know the application that
> runs
> on this instance), change the bufferpool size and measure all again :-)
>
> --
> butthead
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>