Re: NOLRUPRIO / LRUPRIORITY IDS 7.31_UC2
Posted in 2000
From: mattb5@ozemail.com.au
>
>I am after some more information and advise regarding the LRU priority
>bug mentioned within this newsgroup.
>
>We are running IDS 7.31_UC2 on a 10way HP T520, 1.75 Gb RAM running
>HPUX 10.20 with KAIO enabled and NO_SUBQF=1 set.
>
>I am able to keep bufwaits > 10% but only by increasing buffers to >
"> 10%"? < 10%, surely?
>35% of total memory. Currently we are forced to bounce the engine every
>Sunday due to degrading performance over time.
You definitely have the LRU Priority scheduling bug, I reckon.
>onstat -P | tail -5 shows the following on Monday morning:>
>Percentages:
>Data 42.78
>Btree 56.66
>Other 0.56
That's pretty bad.
>But by Thursday Btree > 80% and > 90% by Friday.
That's really bad.
>onstat -B | grep " d0 " | wc -l shows 2000 entries on Monday,>increasing to around 7000 by Friday.
>
>I tried to add export NOLRUPRIO=1 to the environment last Sunday but
>the engine fails left right and center with this set, either assert
>fails with numerous different errors on startup or hangs blocked on
>checkpoint required if I manage to get the engine up.
Yup. Nice, innit? Great fix.
>One thing I am confused about is that in some posting LRUPRIORITY is
>mentioned and NOLRUPRIO in others. Are these variables one and the same
>thing, or totally different? How should I be setting these parameters
>in order to work around the problems that we are experiencing?
NOLRUPRIO is the one I know of. There is another which you can use to change
the scheduling policy, called LRUPOLICY.
Here's a post from The Mighty Art Kagel:
The values and descriptions are:
LRUPOLICY 0x0 - Default. Always hash to choose an LRU, rehash if
waiting too long.
LRUPOLICY 0x1 - Use session ID to generate a consistent LRU selection
per session for any 'get'.( I think a 'get' is
obtaining the LRU's LRUed buffer before writing a
clean page to it.
LRUPOLICY 0x2 - Wait on initial LRU for a 'get' operation. Do not
rehash.
LRUPOLICY 0x4 - Wait on initial LRU for a 'clean put', ie replacing a
buffer at the head of the clean LRU after reading a
page from disk into an empty or recycled buffer. Do
not rehash.
LRUPOLICY 0x8 - Wait on initial LRU for a 'dirty put', ie replacing a
buffer at the head of the dirty LRU after updating
its contents. Do not rehash.
LRUPOLICY 0x10 - Use session ID to generate a consistent LRU selection
per session for any 'put' operation (ie returning a
buffer to the LRU queues).
These values can be mathematically OR'd to create values in the range 0-31
which can control all or part of the LRU selection and wait policies. A
value of 0x11 would cause each session to initially always select the same
LRU queue which MAY eliminate LRU contention when the number of sessions is
not significantly larger than LRUS. If sessions < LRUS it will in effect
make each LRU private for a session or two. The even values (2,4,8) will
case a session to remain in wait on the selected LRU and not rehash to try
to find a less hotly contended one.
---End of Clip from Art ----
>Understand that recreating the indexes is the recommended fix to this
>problem, but want to be sure that the symptoms match the bug before
>recreating 100's of indexes on a medium sized database.
I'm not sure if this will work, but have you tried:
SET INDEXES DISABLED;
SET INDEXES ENABLED;
UPDATE STATISTICS LOW DROP DISTRIBUTIONS;
>Is there any way to determine which indexes were created with prior
>versions of Informix, or which indexes I should focus on recreating
>first. The reason being is we have been running 7.31_UC2 for one year
>now and I know "we" have recreated many indexes but "we" don't have
>records of this.
>
>Would an Informix upgrade be of any value in this situation, if so what
>version would you recommended for us?
UC5 or UC7
Or IDS.2000 9.21.
>Informix Dynamic Server Version 7.31.UC2 -- On-Line -- Up 1 days
>01:23:13 -- 1265160 Kbytes
>
>Profile
>dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
>13206162 16437681 778610313 98.30 3020313 4369914 8978046 66.36
>
>isamtot open start read write rewrite delete commit
>rollbk
>512009540 4186475 55106373 303101804 1908916 1466246 1006371
>148896 304
>
>gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
>0 0 0 0 0 0 0
>
>ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
>0 0 0 101442.55 68857.57 52 106
>
>bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
>2210679 3226 409305478 4 0 240 57282 183962
>
>ixda-RA idx-RA da-RA RA-pgsused lchwaits
>4144333 637883 1249500 5671303 897780
>
>
>Also please feel free to pick the hell out of our onconfig
>
>ROOTNAME mprootdbs
>ROOTPATH /opt/informix/dev/mprootdbs.ln
>ROOTOFFSET 0
>ROOTSIZE 20000>
>MIRROR 1
>MIRRORPATH /opt/informix/dev/mprootdbsm.ln
>MIRROROFFSET 0>
>PHYSDBS mpplogdbs
>PHYSFILE 40000
>
>LOGFILES 30
>LOGSIZE 50000>
>MSGPATH /opt/informix/logs/tpramimsp.mlog
>CONSOLE /opt/informix/logs/tpramimsp.clog
>ALARMPROGRAM /opt/informix/current/etc/no_log.sh
>SYSALARMPROGRAM /opt/informix/current/etc/evidence.sh
>TBLSPACE_STATS 0
>WSTATS 0>
>TAPEDEV /dev/rmt/c1t5d0BEST
>TAPEBLK 1024
>TAPESIZE 70000000>
>LTAPEDEV /dev/rmt/c2t5d0BEST
>LTAPEBLK 1024
>LTAPESIZE 70000000>
># System Configuration
>
>SERVERNUM 1
>DBSERVERNAME tpramimsp_shm
>DBSERVERALIASES tpramimsp
>NETTYPE soctcp,9,100,NET
>NETTYPE ipcshm,9,100,CPU>
>(Note I am sure the above NETTYPE is wrong, but I am having trouble
>fully understanding this parameter, other DBA in company believe these
>are correct, 3-tier OLTP application, this instance is on a dedicated
>database server, 700 concurrent users, via a TP)
You're probably over configured on ipcshms, but I don't think it's a major
performance bummer.
>DEADLOCK_TIMEOUT 60
>RESIDENT -1
>MULTIPROCESSOR 1
>NUMCPUVPS 9
>SINGLE_CPU_VP 0
>NOAGE 1
>AFF_SPROC 1
>AFF_NPROCS 9>
># Shared Memory Parameters
>
>LOCKS 200000
I'd allocate more locks (say *10)
>BUFFERS 293750
If you have 1.75GB and you're not using the box for anything else, why not
allocate more buffers?
>NUMAIOVPS 2
>PHYSBUFF 8
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAHHHH!!! 32, at *least*!! Please post onstat
-l output.
>