7.31 upgrade problems
Posted in 1999
After upgrading from 7.24.UC7 to 7.31, a user saw database-wide slowness and the buffer write cache rate drop from ~98% to ~50%. Respondents pointed to 7.31's new buffer prioritisation: index pages flagged HIGH/MED-HIGH (check onstat -R, -P) were starving data pages out of the cache, because shrinking index nodes could leave the leaf flag incorrectly set (a problem present in both 7.2 and 7.3). Dropping and recreating all indexes helps but only temporarily; John Miller gave SQL against syspaghdr/sysshmhdr to identify affected tables and recommended requesting a fix for bug 115327. No fix date was given, and the poster also noted unrelated I/O imbalance and a low NUMAIOVPS setting, so the thread ends without a confirmed resolution.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Installation, Setup & Upgrades, Storage & Space Management
We recently upgraded from 7.24.UC7 to 7.31 and I am noticing some severe
performance problems. The entire database is running very slowly, and I
am not sure what is causing it.
Here is a snippet from onstat -p:
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
7489 8743 1267093 99.41 1519 2421 3016 49.64
On 7.24, the write was hovering around 98%. Now it sucks. The disk
where the table chunk resides has nothing else on it (indexes are
elsewhere as well).
Any ideas?
Steve Cawley
Check the output of onstat -P. See what the Btree percentages is. If the
number of Btree pages is high, then it is possible that the 'memory resident
table' code in 7.31 is causing some incorrect priorities of buffer
management. If this appears to be the case, please open a case with tech
support.
Stephen F. Cawley wrote:
> We recently upgraded from 7.24.UC7 to 7.31 and I am noticing some severe
> performance problems. The entire database is running very slowly, and I
> am not sure what is causing it.
>
> Here is a snippet from onstat -p:
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 7489 8743 1267093 99.41 1519 2421 3016 49.64
>
> On 7.24, the write was hovering around 98%. Now it sucks. The disk
> where the table chunk resides has nothing else on it (indexes are
> elsewhere as well).
>
> Any ideas?
>
> Steve Cawley
> Here is a snippet from onstat -p:
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 7489 8743 1267093 99.41 1519 2421 3016 49.64
>
> On 7.24, the write was hovering around 98%. Now it sucks. The disk
> where the table chunk resides has nothing else on it (indexes are
> elsewhere as well).
>
> Any ideas?
>
> Steve Cawley
>
>
But the number of reads is much higher, then number of writes... For
detailed analysing send us your ONCONFIG and onstat -a.
--
With best regards, Yuri Dovgart,
SAP R/3, Informix consultant,
"Telecominvest" company.
E-mail y_dovgart@tci.ukrtel.net
Sent via Deja.com http://www.deja.com/
Share what you know. Learn what you don't.
"Stephen F. Cawley" wrote:
>
> We recently upgraded from 7.24.UC7 to 7.31 and I am noticing some severe
> performance problems. The entire database is running very slowly, and I
> am not sure what is causing it.
>
> Here is a snippet from onstat -p:
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 7489 8743 1267093 99.41 1519 2421 3016 49.64
>
> On 7.24, the write was hovering around 98%. Now it sucks. The disk
> where the table chunk resides has nothing else on it (indexes are
> elsewhere as well).
Stephen,
There was a thread over the last 3 weeks about similar experience.
Run onstat -R if most buffers are HIGH or MED-HIGH then index page
caching priorities are starving data pages out of the cache. This
actually seems to be caused by a bug in 7.24 that did not properly set
the index node flags that 7.31 uses for prioritizing index pages in the
buffer cache. If you are seeing this symptom drop and recreate all of
the indexes created under 7.24 and the problem should disappear.
Art S. Kagel
I have just tracked down this issue and dropping and re-creating the indexes
will help,
but only for a period of time. To find out if you have this problem the
following select
will identify the tables.
SELECT trim(dbsname)||":"||trim(tabname)
FROM syspaghdr P, systabnames T
WHERE P.pg_flags = 208
AND P.pg_partnum > 1048570
AND P.pg_partnum = T.partnum
This select is slow, because it must read every single page in your online
system. The
line "P.pg_partnum > 1048570" does not look important, but it is! This will
help
reduce the number of pages read by a significant amount. Without this line
you will
read free, unused pages, physical log, logical logs and all pages which are
not associated with tables.
The problem stems from the fact that shrinking index nodes can leave the leaf
flag set. To
see if you generally do this type of shrinking run the following select
statement and see
if it returns a row:
select * from sysshmhdr where value <> 0 and number = 95;
This problem does exist in both 7.2 and 7.3. The best thing to do is to call
in and ask for bug 115327 to be fixed. This should address the problem from
two sides. First it will ensure
that only node pages and only node pages will be put in the MED_HIGH queue and
second it will stop the creation of pages being flagged as node and leaf
pages. In addition you will
not need to rebuild the indexes.
Hope this helps,
---jmiller
John Miller
Art S. Kagel wrote:
> "Stephen F. Cawley" wrote:
> >
> > We recently upgraded from 7.24.UC7 to 7.31 and I am noticing some severe
> > performance problems. The entire database is running very slowly, and I
> > am not sure what is causing it.
> >
> > Here is a snippet from onstat -p:
> > Profile
> > dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> > 7489 8743 1267093 99.41 1519 2421 3016 49.64
> >
> > On 7.24, the write was hovering around 98%. Now it sucks. The disk
> > where the table chunk resides has nothing else on it (indexes are
> > elsewhere as well).
>
> Stephen,
> There was a thread over the last 3 weeks about similar experience.
> Run onstat -R if most buffers are HIGH or MED-HIGH then index page
> caching priorities are starving data pages out of the cache. This
> actually seems to be caused by a bug in 7.24 that did not properly set
> the index node flags that 7.31 uses for prioritizing index pages in the
> buffer cache. If you are seeing this symptom drop and recreate all of
> the indexes created under 7.24 and the problem should disappear.
>
> Art S. Kagel
Thanks for all of the tips. The `onstat -g iof` showed that the vast
majority of I/O was occurring on the dbspace that contained all of the
tables (~33 IO/sec). The dbspace that came in second was the index dbspace
(5 IO/sec). The NUMAIOVPS was also set to 2, so that was not helping.
Steve
> We recently upgraded from 7.24.UC7 to 7.31 and I am noticing some severe
> performance problems. The entire database is running very slowly, and I
> am not sure what is causing it.
>
> Here is a snippet from onstat -p:
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 7489 8743 1267093 99.41 1519 2421 3016 49.64
>
> On 7.24, the write was hovering around 98%. Now it sucks. The disk
> where the table chunk resides has nothing else on it (indexes are
> elsewhere as well).
<STUFF DELETED FOR BREVITY> >This problem does exist in both 7.2 and 7.3. The best thing to do is to call >in and ask for bug 115327 to be fixed. This should address the problem from >two sides. Isn't tech. info wonderful? I cannot find a way to lookup defect reports by number, only by key words which I do not know. I'm glad I came across this thread since I was just about to recommend an upgrade to 7.31 to my clients. I had a quick chat with a support engineer about this, and he was not sure if/when Informix would fix the problem. I told him that dropping and recreating all indexes was not a real world solution. Most mission critical databases I know have a number of VERY large tables with quite a few indexes. Such a database will have MANY indexes, a large proportion of which maintain referential integrity. Taking the production database out of service for a number of days to address this problem is not on the table. So, anyone from Informix care to tell us when this problem will be addressed? I can't wait to find some issue with 7.30 that a support engineer tells us is fixed in the next release and that we should upgrade...