7.23 loses pages off the buffer hash chain?
Posted in 1999
Informix On-Line Dynamic Server 7.23.UC9
HP K420 4-way, HP-UX 10.10, 2 GB RAM, one dbspace disk.
We're trying to keep a 780 MB database fully cached, and have allocated
550000 buffers (1100 MB) for it. We're also trying to keep it clean so
checkpoints are short, so we've got 3 full-time page cleaner threads
running. LRUS is set to 16 (though we've tried 17 to avoid another
possible bug in the hashing). Only one CPU VP is configured to run.
We pre-load the cache, then start the application running. Application
does a lot of singleton selects, then a tablespace BLOB update and
counter update (this is the hot row) on each user invocation.
onstat -h (apart from the bug in calculating the number of hashedbuffers) shows a steady loss of pages from the hash chains (I calculate
this by taking the number of 2-page chains, multiplying by 2 and adding
the number of 1-page chains to figure out how many are actually hashed).
Cached reads falls gradually, and cached writes and bufwaits start
increasing (but we know we're I/O bound on the one disk; we're working
that as a separate issue).
INFORMIX-OnLine Version 7.23.UC9 -- On-Line -- Up 01:40:41 -- 1257824 Kbytes
buffer hash chain length histogram
# of chains of len
800981 0
125505 1
122090 2
1048576 total chains
247595 hashed buffs
550000 total buffs
calculated 369685
INFORMIX-OnLine Version 7.23.UC9 -- On-Line -- Up 01:41:11 -- 1257824 Kbytes
buffer hash chain length histogram
# of chains of len
800982 0
125562 1
122032 2
1048576 total chains
247594 hashed buffs
550000 total buffs
calculated 369626 ( lost 59 )
INFORMIX-OnLine Version 7.23.UC9 -- On-Line -- Up 01:41:43 -- 1257824 Kbytes
buffer hash chain length histogram
# of chains of len
800976 0
125648 1
121952 2
1048576 total chains
247600 hashed buffs
550000 total buffs
calculated 369552 ( lost 74 )
The profile:
INFORMIX-OnLine Version 7.23.UC9 -- On-Line -- Up 01:41:11 -- 1257824 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
71848 71848 10724659 99.33 481812 560171 503075 4.23
isamtot open start read write rewrite delete commit rollbk
17146594 1013977 1413979 1256115 45 197190 0 98622 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 1972.43 461.92 162 324
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
75276 46427 7711793 0 0 162 20608 14
ixda-RA idx-RA da-RA RA-pgsused lchwaits
26805 0 0 26811 43606
The suspected dropping of pages from the hash gets better with one cleaner,
but then our checkpoint times get too long.
Has anyone seen this before? Can it be explained? (The case with
Informix is pending ...) Can it be configured around?
Regards,
Jim
--
W. Jim Jordan, Nortel Networks, Stop 313 Qualicum,
PO Box 3511 Stn C, Ottawa, ON K1Y 4H7 Canada wjjordan@nortelnetworks.com
I do not speak for Nortel Networks.
I do not do business with spammers.