Re: Too-frequent checkpoints; which parameters to check?
Posted in 1997
In article <3321A111.3E233C7F@www.weideneder.de>, Stefan Weideneder
<stefan@www.weideneder.de> writes
>David Williams wrote:
>>
>> In article <331FF103.7F5C@lava.net>, Bob Cunningham <bob@lava.net>
>> writes
>> >I've got a situation with On-Line 7.22 (SPARC, Solaris 2.5.1) where
>> >we get flurries of checkpoints at short intervals, during
>> >some update periods that look like:
>> >
>> >23:35:07 Checkpoint Completed: duration was 2 seconds.
>> >23:35:14 Checkpoint Completed: duration was 1 seconds.
>> >23:35:22 Checkpoint Completed: duration was 1 seconds.
>> >23:35:30 Checkpoint Completed: duration was 1 seconds.
>> >23:35:38 Checkpoint Completed: duration was 1 seconds.
>> >23:35:46 Checkpoint Completed: duration was 2 seconds.
>> >23:35:53 Checkpoint Completed: duration was 1 seconds.
>> >...
>> >23:45:06 Checkpoint Completed: duration was 1 seconds.
>> >23:45:14 Checkpoint Completed: duration was 1 seconds.>> >
>> >Needless to say, queries slow to a crawl when there's
>> >a checkpoint every few seconds.
>
>Who is working at midnight ? I think Informix slows down
>because the users are tired :)
:-)
>
>> >Log files aren't a problem, but I'm not sure exactly what is...
>> >
>> >Which tuning parameters should I be looking to change:
>> >PHYSFILE, the LRU* ones, or ???
>>
>> Informix Online Dynamic Server Administrator's Guide, Volume 1
>> Version 7.1 Page 12-52/3 : events that initiate a checkpoint.
>>
>> - The checkpoint interval, specified by the configuration parameter
>> CKPTINTVL, has elapsed and one or more modifications have occured
>> since the last checkpoint.
>>
>> ***CHECK CKPTNTVL***
>>
>> - The phyiscal log on disk becomes 75 percent full.
>>
>> ***CHECK PHYSFILE***
>
>And I thought, Informix prompts a warning when the physical log has
>a size smaller than 30% of the total logical log size.
>
I do remember seeing a message about physical log being too small
- it must be at least 200 K (Online ADmin Guide 7.1 Volume 2 Page 38-
51. Note you will also get an error if the physical log is too small
when building the sysmaster database (possibly that was the error
you are thinking off).
>> - Online detects that the next logical log file to become current
>> contains the most recent checkpoint record.
>>
>> ***USE tbstat -l TO MONITOR LOGICAL LOG USAGE***
> ***USE onstat -l ...
>
>> - The OnLine Administrator initiates a checkpoint from the OnMonitor,
>> Force-Ckpt menu of from the command line using onmode -c.
>>
>> ***CHECK CRON/AT JOBS FOR onmode COMMANDS***
>>
>> - Certain Administrative tasks such as adding a chunk or a dbspace,
>> take place.
>>
>> ***UNLIKELY - MONITOR NUMBER OF CHUNKS WITH onstat -c ***
>
>Look at the number of Checkpoints! How many chunks is he going to add ?
I was just covering everything (another document to add to my Web
Site when I start it...
>
>> --
>> David Williams
>
>It looks as if someone is doing a mass media movement, like an insert
>with thousands of rows, an update or a delete. A good idea is creating
>an index in the late evening or clustering an index.
>Like David told you, look at the cron jobs.
>
>But in your header we can read "some update periods". If you know
>what you are doing and if you want to tune your system for that
>situation: Hm, it's hard to tune an OnLine instance for both,
>OLTP and batch operation. Okay, you can increase the Physical Log
>and this will result in less checkpoints, but the checkpoint duration
>will increase, unless you set your LRU_MAX_DIRTY to the correct value.
>
>BUFFERS * LRU_MAX_DIRTY * avg(disk speed per page) = checkpoint duration
>
>Hope, this will help.
>
>You are using version 7.1. If you migrate to version 7.2 then you
>might set the environment variable LIGHT_SCANS to 1 for the updating
>process. But I'm not quite sure if this setting avoids caching for
>an update process. Any comments ?
>
??? LIGHT_SCANS - I have heard of them in Bersion 7.1 - is there
an environment variable for them in version 7.2???
>
>Bye
>
>
>Stefan
--
David Williams