Re: Checkpoints up to 137 seconds
Posted in 1999
Hi, reading the messages from this list, sometimes I have doubts about the relation between LRU, CLEANERS, AIOVPs, BUFFERS and NUMCPU. I think that LRU and CLEANERS must be identical values. But there exists any predefined relation between the others values ? In my environment, a Sun E-6500 with 14 processors and 9GB RAM, I have 800.000 buffers, 16 LRU, 16 CLEANERS, 100 AIOVP, 13 NUMCPU Thanks in advance fred "Art S. Kagel" wrote: > Bridget Reitsma wrote: > > > > Hi > > > > I'm using informix 730UC5 on a Sequent box, I thought I had everything > > set just right, that was till Wednesday this week, when my checkpoints > > started rocketing. Normally my checkpoints average at about 4 seconds, > > now they are averaging at about 60 seconds. > > > My LRU MAX is set to 2 my LRU_MIN is set to 0. The have 17 LRU queues > > and 8 CLEANERS. The checkpoint interval was recently changed to 600 > > seconds - I have have changed this back to 300 seconds. (I have > > attached a copy of my onconfig file.) > > It might be the increase in the checkpoint interval, but I doubt it. BTW > increasing CLEANERS to 17 or increasing both LRUS and CLEANERS to 31 or 32 > MAY reduce the checkpoint duration but I don't think that any of these > items is to blame for such a substantial increase. More likely it is > a process which is in a critical section causing the checkpoint to wait > for it to release its latch before beginning. The checkpoint blocks new > requests for a critical section latch and then waits for the existing > latches to clear before continuing with the checkpoint. So if there are > a later number of processed in crcitical sections, especially if they are > waiting for each other to free resources, this can cause the checkpoint > to become blocked. > > Just now I had a server checkpoint for 15 seconds with only 53 dirty > buffers and > 3 pages of physical log and 3 pages of logical log buffer to > flush to disk. I run 127 LRUs and 127 CLEANERS and have 49 AIO VPs. This > checkpoint should have taken less than a second. It did not because I > have 21 VERY active middleware processes querying the system. > > > There hasn't been a major increase in users the load is only slightly > > higher than normal. It is the end of the month after all. But does > > anyone have any ideas of what other step I could take to decrease the > > checkpoint interval? > [SNIP] > > Art S. Kagel