checkpoint time > 15 sec.
Posted in 1999
A user running Informix 7.3 on an IBM S70 (2GB RAM, 4 CPUs, mirrored SSA disks) asked how to shorten checkpoints taking over 15 seconds, with CKPTINTVL 300, 8 LRUs, LRU_MIN/MAX_DIRTY 2/3 and 4 page cleaners. Replies suggested lowering LRU_MAX/MIN_DIRTY (e.g. 1/0), shrinking the physical log or CKPTINTVL so checkpoints are more frequent, raising LRUS (~32, avoiding 64/96 due to a suspected bug) and matching CLEANERS to LRUS, spreading chunks round-robin across disks, and checking onstat -p/-F/-R/-m/-g iov to see real checkpoint frequency and whether cleaners/AIO VPs keep up. One poster questioned bufwaits as a reliable indicator. No outcome from the original poster is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Logging & Checkpoints
I am currently running informix 7.3 on an IBM S70 with 2GB men and 4
CPU's, and 10 mirrored ssa 4.5 GB disks.
CHKPTINVL 300
LRU 8
LRU_MIN_DIRTY 2
LRU_MAX_DIRTY 3PAGE_CLEANERS 4
any suggestions on improving checkpoint time.
Thanks,
Mike Wallace
wallace@infotechusa.com
Try to lower both lru_min&max_dirty to 1 and 2. It determind how much
data to flush to disk at checkpoint.
jinsy.
michael wallace wrote:
> I am currently running informix 7.3 on an IBM S70 with 2GB men and 4
> CPU's, and 10 mirrored ssa 4.5 GB disks.
> CHKPTINVL 300
> LRU''''''''''''''' 8
> LRU_MIN_DIRTY 2
> LRU_MAX_DIRTY 3> PAGE_CLEANERS 4
>
> any suggestions on improving checkpoint time.
>
> Thanks,
> Mike Wallace
> wallace@infotechusa.com
You haven't really stated what your problem is, but the implication is that
your checkpoints, when they occur, are taking too long. So, try making them
more frequent, either by reduing CHKPTINVL, by reducing the size of your
physical log, or by further reducing LRU_MAX/MIN_DIRTY (although the values
you already have ought to provide for pretty quick checkpoints).
Neil Truby
aracnet Limited
Weybridge, UK
michael wallace wrote in message <36941E5E.1A49C08E@infotechusa.com>...
>I am currently running informix 7.3 on an IBM S70 with 2GB men and 4
>CPU's, and 10 mirrored ssa 4.5 GB disks.
>CHKPTINVL 300
>LRU 8
>LRU_MIN_DIRTY 2
>LRU_MAX_DIRTY 3>PAGE_CLEANERS 4
>
>any suggestions on improving checkpoint time.
>
>Thanks,
>Mike Wallace
>wallace@infotechusa.com
>
Are you using buffered logging ? If so how big are your buffers ? How big is
your physical log ? How much shared memory is configured ( you've said 2GB,
is that all for Informix ? ) ?
You must have a lot of activity on your system if checkpoints are taking >15
secs with these config settings. As Neil says, try reducing the CKPTINVL to
3 minutes and reduce the size of the physical log.
Also, how many chunks do you have ? Try Increasing the number of page
cleaners to the number of chunks that are available to OnLine - don't do it
on the live enironment until you have benchmarked the changes first !!!
covering my own ass here !! )
Hope this helps.
Sean
michael wallace wrote in message <36941E5E.1A49C08E@infotechusa.com>...
>I am currently running informix 7.3 on an IBM S70 with 2GB men and 4
>CPU's, and 10 mirrored ssa 4.5 GB disks.
>CHKPTINVL 300
>LRU 8
>LRU_MIN_DIRTY 2
>LRU_MAX_DIRTY 3>PAGE_CLEANERS 4
>
>any suggestions on improving checkpoint time.
>
>Thanks,
>Mike Wallace
>wallace@infotechusa.com
>
michael wallace wrote:
>
> I am currently running informix 7.3 on an IBM S70 with 2GB men and 4
> CPU's, and 10 mirrored ssa 4.5 GB disks.
> CHKPTINVL 300
> LRU 8
> LRU_MIN_DIRTY 2
> LRU_MAX_DIRTY 3> PAGE_CLEANERS 4
Recommendatations are hard without more info like onstat -p, onstat -F,
onstat -R, onstat -m output and even onstat -g iov output to look at.But here
are some suggestions in a vacuum:
CKPTINTVL OKIs it possible that the physical log is so small that you are
checkpointing more frequently than 5 minutes such that the LRU_MAX/MIN
parameters are not having any effect? Look at onstat -F and see if
most writes are Chunk Writes then look in the log for the actual
checkpoint frequency. I'd bet it's more often than every 5 minutes.
LRUS 32 (if you have fewer than 100,000 buffers or even higher if you
have more but avoid the values 64 and 96 these trigger a suspected bug).
You can generally tell if you need more LRUS by checking the onstat -p
output and compare bufwaits to (pagreads + bufwrits). Bufwaits should
be only a few % of the sum.
Increase CLEANERS to equal LRUS.
If you need more help you will have to post more info.
Art S. Kagel
In article <3694E2AF.1CEC@bloomberg.net>, Art S. Kagel
<kagel@bloomberg.net> writes
>michael wallace wrote:
>>
>> I am currently running informix 7.3 on an IBM S70 with 2GB men and 4
>> CPU's, and 10 mirrored ssa 4.5 GB disks.
>> CHKPTINVL 300
>> LRU 8
>> LRU_MIN_DIRTY 2
>> LRU_MAX_DIRTY 3>> PAGE_CLEANERS 4
>
>Recommendatations are hard without more info like onstat -p, onstat -F,
>onstat -R, onstat -m output and even onstat -g iov output to look at.>But here
>are some suggestions in a vacuum:
>
>CKPTINTVL OK>Is it possible that the physical log is so small that you are
>checkpointing more frequently than 5 minutes such that the LRU_MAX/MIN
>parameters are not having any effect? Look at onstat -F and see if
>most writes are Chunk Writes then look in the log for the actual
>checkpoint frequency. I'd bet it's more often than every 5 minutes.
>
>LRUS 32 (if you have fewer than 100,000 buffers or even higher if you
>have more but avoid the values 64 and 96 these trigger a suspected bug).
>You can generally tell if you need more LRUS by checking the onstat -p
>output and compare bufwaits to (pagreads + bufwrits). Bufwaits should
>be only a few % of the sum.
>
>Increase CLEANERS to equal LRUS.
>
>If you need more help you will have to post more info.
>
>Art S. Kagel
Don't forget to create your chunks round robin across the disks.
When you have N cleaners they start writing to chunks 1-N in parallel,
clearly you want these chunks to be on separate disks to avoid device
contention and the disk heads seeking back and forth between different
chunks on the same disk!!
--
David Williams
Art S. Kagel wrote:
> michael wallace wrote:
> >
> > I am currently running informix 7.3 on an IBM S70 with 2GB men and 4
> > CPU's, and 10 mirrored ssa 4.5 GB disks.
> > CHKPTINVL 300
> > LRU 8
> > LRU_MIN_DIRTY 2
> > LRU_MAX_DIRTY 3> > PAGE_CLEANERS 4
>
> Recommendatations are hard without more info like onstat -p, onstat -F,
> onstat -R, onstat -m output and even onstat -g iov output to look at.> But here
> are some suggestions in a vacuum:
>
> CKPTINTVL OK> Is it possible that the physical log is so small that you are
> checkpointing more frequently than 5 minutes such that the LRU_MAX/MIN
> parameters are not having any effect? Look at onstat -F and see if
> most writes are Chunk Writes then look in the log for the actual
> checkpoint frequency. I'd bet it's more often than every 5 minutes.
Or that the page cleaners are not able to keep up with the write activity,
so that the target values for LRU_MAX/MIN are not being met. (Or maybe that
the AIOVP's are not able to keep up? I'm not sure whether that would affect
your checkpoint duration, though I suspect it would.) Check onstat -R to
see whether your number of dirty pages exceeds LRU_MAX_DIRTY. Also onstat
-g iov to see how your AIOVP's are doing.
> LRUS 32 (if you have fewer than 100,000 buffers or even higher if you
> have more but avoid the values 64 and 96 these trigger a suspected bug).
> You can generally tell if you need more LRUS by checking the onstat -p
> output and compare bufwaits to (pagreads + bufwrits). Bufwaits should
> be only a few % of the sum.
Hmmm. Now this is interesting to me. I have never been able to make any
meaningful deductions out of bufwaits. It seems to me that bufwaits can be
incremented by enough different things that it is difficult to attribute it
to LRU contention or anything else, although what you say here makes sense
to me (sort of). But if there are only 4 page cleaners (cleaning at most 4
LRU's at the same time), and 4 CPU's (with, I suspect, <=4 CPU VPs, although
this is not necessarily true), and 8 LRU's, I'm not convinced that there is
any contention that can be fixed. In fact, I think it is theoretically
possible that lower LRU_MAX/MIN values can increase your bufwaits, because
the page cleaners must lock the page that is being written, which means that
anyone who wants that page (specifically THAT page, because of the data that
is on it, not just any clean page) must wait for it to be cleaned.
June
--
june_t@hotmail.com
Grounded in Palo Alto, living on M&M's (plain)
michael wallace wrote:
>
> I am currently running informix 7.3 on an IBM S70 with 2GB men and 4
> CPU's, and 10 mirrored ssa 4.5 GB disks.
> CHKPTINVL 300
> LRU 8
> LRU_MIN_DIRTY 2
> LRU_MAX_DIRTY 3> PAGE_CLEANERS 4
>
> any suggestions on improving checkpoint time.
>
> Thanks,
> Mike Wallace
> wallace@infotechusa.com
Hi Mike,
how many BUFFERS do you have ?
I would recommend to set LRU_MAX_DIRTY to 1 and LRU_MIN_DIRTY to 0.
( Sure, this would not work on every version, but in 7.3 ).
Depending on your BUFFERS' size I would try to get faster disks.
Best regards,
Stefan Weideneder