Re: Checkpoint Stall problem?
Posted in 2005
This is quite incorrect and very misleading. There is a problem only for some huge systems and it is not as dramatic as this article sounds. The tunig tips are also wrong. As far as I have heard it will be fixed in 9.40.UC7 and 10.0.UC2. There is a problem at checkpoint time for online systems with huge buffer pools, something like > 8GB of buffers. Systems with "more than 500 dirty buffers at checkpoint time" are NOT affected at all if they don't have a huge buffer pool. In those systems that are affected a poll or sql listener thread on cpu 1 cannot run while the checkpoint code is collecting all dirty pages. This lasts for maybe a few seconds with 8GB of buffers (depending on cpu speed) and increases linearly with buffer pool size. If there is no poll thread on cpu vp 1, connected clients are still not affected (They DO migrate to other vps!). If there is a sql listener waiting, new connections must wait for that time. The problem will NOT go away by tuning LRU_MAXDIRTY, etc.It happens even with 0 dirty buffers. This is what happens: At checkpoint time the main_loop thread loops through the whole buffer pool inspecting the DIRTY flag in every page's buffer header and collects a list of dirty pages in memory which is later used by the flushers. This is done in a single thread and without yielding. No other thread can run on cpu 1 for that time. When the list of dirty pages is collected the problem is over. This is BEFORE the first flusher even starts doing anything. This is done twice per checkpoint. There are two flushes. This algorithm has been about like this even in old turbo (version 5) times and it has never been a problem until recently when customers can afford to configure buffer pools in the GB range. What used to take milliseconds even in systems that were huge at the time can now take a couple of seconds. It blocks cpu vp 1 and it increases checkpoint times. The fix will go through the dirty LRU queues to collect all dirty pages if less than 1% of all buffers are dirty. This is much faster. Michael Colin Dawson wrote: > Whilst searching for information on DS_HASHSIZE, I found the following, I > went in search of the results and summary on CDI but couldn't find it. Two > questions arise: 1) Which versions does this apply to? and 2) Does anyone > know where the results are located? A search on CDI archive for Checkpoint > Stall returned 54 pages of results!!!!!!! > > Many thanks > > Colin Dawson > www.ladbrokes.com > > The Checkpoint Stall Problem > We are all aware that updates must wait while a checkpoint completes, and > that SELECTs are not affected. This is not quite true! Read on. > > If you have more than 500 dirty buffers at checkpoint time, the checkpoint > thread will not relinquish control of CPU VP#1 until all buffers have been > handed off to CLEANER threads-as long as there are page cleaner threads > available that are not busy. Since user threads currently executing in CPU > VP#1 are not migrated to other VPs prior to beginning the checkpoint, or you > may only have one CPU VP, this means that even FETCHs and OPENs will hang > for approximately the first half of the checkpoint duration (longer if you > set LRU_MAXDIRTY and LRU_MINDIRTY high). If you have response-time > requirements measured in seconds, this can be deadly. > > If no page cleaner thread is ready, then the checkpoint thread will > relinquish the VP and place itself on a wait queue until a cleaner thread > frees up. With less than 500 dirty buffers, a single cleaner is launched to > handle all dirty buffers and the checkpoint thread relinquishes the CPU VP > to other threads after that launch. This sounds better, but the result is > the same. > > Based on this, I am recommending that sites with large numbers of buffers > and short response time requirements configure with one or only a few page > cleaners so that the checkpoint will have to enter a wait state while the > previous page cleaners complete their last assignments. This should cause > the checkpoint thread to release the VP and permit queries to continue. This > issue is being examined by Informix R&D and, by the time this goes to press, > I should have had an opportunity to test the latest theoretical solution. > Search the C.D.I. archives at http://www.iiug.org/ for my results. > > > sending to informix-list