Re: KAIO vps, and bufwaits
Posted in 1998
In article <6bq06q$8tp$1@news.xmission.com>, Art S. Kagel <kagel@ns1.bloomberg.com> writes > >On Mon, 9 Feb 1998, Tony Tyrwhitt-Drake wrote: > >> Art, >> >> I was intrigued by your mention of a bug in the LRU queue selection. But >> mostly by the number of queues you mentioned. Personally, I have never set >> LRUS to a number anything like as big as 32, let alone 128, prefering to >> stick to a low number such as the number of physical CPUs in the machine >> say 4 ), without getting many bufwaits. >> Can you explain the rationale behind these large ( to me ) numbers of LRU >> queues? >> >> -----Original Message----- >> From: Art S. Kagel <kagel@ns1.bloomberg.com> >> Newsgroups: comp.databases.informix >> To: Tony Tyrwhitt-Drake <tonytd@ttyrwhit.demon.co.uk> >> Date: 09 February 1998 15:34 >> Subject: Re: KAIO vps, and bufwaits >> >> >> >>lots of excellent stuff clipped >> >> >Anyway, if you have LRUS set to either 64 or 96 you may be hitting an >> >as yet unverified bug I have been tracking ........... > >Sure. If you have many users attached and performing updates more LRU >queues prevent spins and BUFWAITS. Hmmm, seems that I was so focused >on this annoying bug that I failed to suggest the obvious, that being >that there may not be ENOUGH LRU queues to begin with which will show >similar symptoms. Let me explain in detail for the benefit of anyone >who is not familiar with the LRU hashing algorithm. > >When a thread needs a buffer to update and the page is not already in >the buffer pool the thread uses a random hash to select an LRU to get >the buffer from. If this LRU is not locked by another process doing >the same thing the thread acquires the lock on that LRU queue takes >the Least Recently Used clean buffer, moves it to the LRU queue's >dirty list, locks the buffer page and releases the lock on the LRU. >If the LRU is already locked it spins (BUFWAIT) for a while trying a >few more times to get a lock on that LRU if not it rehashes to another Are you sure it tries several times, I though the manual said it only tried once and then hashed onto the next queue. PS I wonder what the random hash function is? It can't be that random since it must be only a few instructions long.. >queue. If there are many updaters and few LRUs it is likely that that >LRU will also be locked already and so our poor thread spins again and >so on. Eventually it gets a lock on sone LRU and continues. If there >are more LRUs then the probablility of finding a free one improves >greatly and BUFWAITS go down and throughput goes up. > >If the page to be updated is already in the cache but has not yet been >modified the number of LRUs is still important as the size of the buffer >cache increases. More cache means it is more likely that a page that is >about to be modified is already in the clean cache (or the dirty cache for >that matter but that is not affected by the number of LRUs). With more >LRUs it is less likely that some other process is in the process of moving >some other already cached clean page to the dirty list from the same LRU >that contains the page that your thread needs. Again this reduces >BUFWAITS and improves throughput. > >The bug I suspect is that for certain numbers of LRUs the rehash >algorithm causes many threads to rehash to the same sequence of LRUs >so that the acquisition of LRUs is a race condition and single >threaded. Here one thread at a time gets an LRU and drops out of the >race the symptoms are that updates/inserts/deletes do not scale (ie >adding another updater does not improve throughput significantly) and >BUFWAITS skyrocket. We were seeing both of these symptoms. > >With 128 LRUs (the maximum) we can have 100 update server tasks running >constantly serving update requests from clients and still get near linear >gains from adding up to 40 data sync or cleanup tasks concurrent with >normal production load. > >An additional benefit of many LRUs is that it helps to keep more buffers >flushed to disk by the LRU_MIN_DIRTY/LRU_MAX_DIRTY parameters. Because of >the hashing used to select LRUs the LRUs fill unevenly. More LRUs fill >less evenly than fewer LRUs meaning fewer dirty buffers at CHECKPOINT time >and faster checkpoints. Also more LRUs can take advantage of more CLEANERS >for faster checkpoints. > >Hope this explains my love of LRUs. > >Art S. Kagel, kagel@bloomberg.com > -- David Williams Maintainer of the Informix FAQ Primary site (Beta Version) http://www.smooth1.demon.co.uk Official site http://www.iiug.org/techinfo/faq/faq_top.html I see you standin', Standin' on your own, It's such a lonely place for you, For you to be If you need a shoulder, Or if you need a friend, I'll be here standing, Until the bitter end... So don't chastise me Or think I, I mean you harm... All I ever wanted Was for you To know that I care