Checkpoint durations
Posted in 2008
Neil Truby asked whether 5-8 second checkpoints (roughly 100MB of dirty pages, ~12-20MB/s) were reasonable on a large HP-UX 11.31 / IDS 10 box with 10M buffers, LRU_MAX_DIRTY 0.5% and an EMC Symmetrix SAN, and sought comparable figures from other large sites. Suggestions included checking page-cleaner balance (onstat -F, onstat -u 'F' flags), spreading dbspaces over many chunks since single-chunk layouts serialise flushing, iostat, and dd baselines. His dd tests looked slow (1GB in ~4.5 min; ~2 min read-only), but HP and EMC had passed the hardware and an IBM PMR was still open. No resolution or definitive comparison numbers are recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
IDS 10.0FC8W2 on HP-UX 11.31 I'm tuning IDS on quite a large system (18 cpus, 65g memory, Tier One SAN). I have BUFFERS set to 10,000,000 (ie 20g), and lru_max_dirty to 0.5%.(ie 50,000 pages or 100m). lru_min_dirty is effectively zero. At times I'm seeingcheckpoints of 5 to 8 seconds. According to the checkpoint tracer nearly all of this is waiting on disk. And my question is: on a lareg OLTP system, does this seem reasonable throughput, a maximum of 100m in 5-8 seconds?). Obviously I can reduce lru_max_dirty further, but I'm just interested to know if this 12-20MByte/s rate is comparable to others' experience. Checkpoint statistics from other large users would be most welcome. thx Neil
Neil Truby wrote: > IDS 10.0FC8W2 on HP-UX 11.31 > > I'm tuning IDS on quite a large system (18 cpus, 65g memory, Tier One SAN). > > I have BUFFERS set to 10,000,000 (ie 20g), and lru_max_dirty to 0.5%.(ie > 50,000 pages or 100m). lru_min_dirty is effectively zero. > > At times I'm seeingcheckpoints of 5 to 8 seconds. According to the > checkpoint tracer nearly all of this is waiting on disk. And my > question is: on a lareg OLTP system, does this seem reasonable > throughput, a maximum of 100m in 5-8 seconds?). > > Obviously I can reduce lru_max_dirty further, but I'm just interested to > know if this 12-20MByte/s rate is comparable to others' experience. > > Checkpoint statistics from other large users would be most welcome. > > thx > Neil Neil, Are you sure you're not having more than 100MB to write? Are you basing this number on observations or just because of LRU_MAX_DIRTY? You should also check the cleaners work during checkpoint. Are they all reasonably busy, or are you left with one or two for the last seconds? -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently...
"Fernando Nunes" <domusonline@gmail.com> wrote in message news:ggpu5d$bht$1@news.motzarella.org... > Neil Truby wrote: >> IDS 10.0FC8W2 on HP-UX 11.31 >> >> I'm tuning IDS on quite a large system (18 cpus, 65g memory, Tier One >> SAN). >> >> I have BUFFERS set to 10,000,000 (ie 20g), and lru_max_dirty to 0.5%.(ie >> 50,000 pages or 100m). lru_min_dirty is effectively zero. >> >> At times I'm seeingcheckpoints of 5 to 8 seconds. According to the >> checkpoint tracer nearly all of this is waiting on disk. And my question >> is: on a lareg OLTP system, does this seem reasonable throughput, a >> maximum of 100m in 5-8 seconds?). >> >> Obviously I can reduce lru_max_dirty further, but I'm just interested to >> know if this 12-20MByte/s rate is comparable to others' experience. >> >> Checkpoint statistics from other large users would be most welcome. >> >> thx >> Neil > > Neil, > > Are you sure you're not having more than 100MB to write? > Are you basing this number on observations or just because of > LRU_MAX_DIRTY? > You should also check the cleaners work during checkpoint. Are they all > reasonably busy, or are you left with one or two for the last seconds? Yes, I am sure. I have a script that monitors the number of dirty pages etc, and see it change in real time. Also, checkpoint tracing gives you the number of pages written. Not too sure how to check the cleaners; any advice? Can you answer my original question from experience: do the checkpoint durations I cite above look good, or look long to you? cheers Neil
Neil Truby wrote:
>
> "Fernando Nunes" <domusonline@gmail.com> wrote in message
> news:ggpu5d$bht$1@news.motzarella.org...
>> Neil Truby wrote:
>>> IDS 10.0FC8W2 on HP-UX 11.31
>>>
>>> I'm tuning IDS on quite a large system (18 cpus, 65g memory, Tier One
>>> SAN).
>>>
>>> I have BUFFERS set to 10,000,000 (ie 20g), and lru_max_dirty to
>>> 0.5%.(ie 50,000 pages or 100m). lru_min_dirty is effectively zero.
>>>
>>> At times I'm seeingcheckpoints of 5 to 8 seconds. According to the
>>> checkpoint tracer nearly all of this is waiting on disk. And my
>>> question is: on a lareg OLTP system, does this seem reasonable
>>> throughput, a maximum of 100m in 5-8 seconds?).
>>>
>>> Obviously I can reduce lru_max_dirty further, but I'm just interested
>>> to know if this 12-20MByte/s rate is comparable to others' experience.
>>>
>>> Checkpoint statistics from other large users would be most welcome.
>>>
>>> thx
>>> Neil
>>
>> Neil,
>>
>> Are you sure you're not having more than 100MB to write?
>> Are you basing this number on observations or just because of
>> LRU_MAX_DIRTY?
>> You should also check the cleaners work during checkpoint. Are they
>> all reasonably busy, or are you left with one or two for the last
>> seconds?
>
> Yes, I am sure. I have a script that monitors the number of dirty pages
> etc, and see it change in real time. Also, checkpoint tracing gives you
> the number of pages written. Not too sure how to check the cleaners; any
> advice?
>
> Can you answer my original question from experience: do the checkpoint
> durations I cite above look good, or look long to you?
>
> cheers
> Neil
onstat -F during checkpoint will show cleaner activity.Column "data" is the chunk number... If the I/O load is not balanced you'll
probably see a "burst" of activity on all of them when checkpoint starts, then
possibly most of them will get back to sleep and one or two may be there
longer. This can also happen if by some weird reason the chunk is on "slow
disk" (hard to understand on modern systems...)
I can't give you any numbers at the moment... and they wouldn't be comparable
to your system.
The higher load system I work daily has a much smaller number of buffers, very
old hardware and HDR... And I'm not tracing the checkpoints (I can't even
remember the max_dirty/min_dirty :) ) I could risk something like 15K to 17K
buffers in about 3/4 seconds... that would be 30-34MB in about half the time...
If this numbers are correct maybe your numbers are not very good, but I need to
check these numbers.
At first glance I wouldn't be shocked by your numbers... It has to order the
buffers, and write them. Maybe you can test your disk throughput with dd? It's
a completely different situation. It should be clearly higher than what you get
at checkpoints, but maybe it can give you an idea...
Let's wait for more...
P.S.: I don't need to remind *you* but for others, IDS 11 would be nice for
that system...
Regards,
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
"Fernando Nunes" <domusonline@gmail.com> wrote in message news:ggq4au$uk5$1@news.motzarella.org... > Neil Truby wrote: > At first glance I wouldn't be shocked by your numbers... It has to order > the buffers, and write them. Maybe you can test your disk throughput with > dd? It's a completely different situation. It should be clearly higher > than what you get at checkpoints, but maybe it can give you an idea... Yes, I have tried dd. the results are somewhat variable, and therefore inconclusive, hence the question about checkpoint times on other users' large systems.
Neil Truby wrote: > Yes, I have tried dd. the results are somewhat variable, and therefore > inconclusive, hence the question about checkpoint times on other users' > large systems. Can you please post them up so that people can understand what you mean by "somewhat variable"? -- Cheers, Obnoxio The Clown http://obotheclown.blogspot.com
On 28 Nov, 19:45, "Neil Truby" <neil.tr...@ardenta.com> wrote:
> IDS 10.0FC8W2 on HP-UX 11.31
>
> I'm tuning IDS on quite a large system (18 cpus, 65g memory, Tier One SAN).
>
> I have BUFFERS set to 10,000,000 (ie 20g), and lru_max_dirty to 0.5%.(ie
> 50,000 pages or 100m). lru_min_dirty is effectively zero.
>
> At times I'm seeingcheckpoints of 5 to 8 seconds. According to the
> checkpoint tracer nearly all of this is waiting on disk. And my question
> is: on a lareg OLTP system, does this seem reasonable throughput, a maximum
> of 100m in 5-8 seconds?).
>
> Obviously I can reduce lru_max_dirty further, but I'm just interested to
> know if this 12-20MByte/s rate is comparable to others' experience.
>
> Checkpoint statistics from other large users would be most welcome.
>
> thx
> Neil
Well a cetain customer with an IBM p690 with 32 cpus was treating 5-11
second checkpoints as normal but that we IDS 9.40.
Another Informix customer had a HP-UX Superdome with 32cpus and a
fully loaded XP12000 array capable of 500,000 io/s. They
created a 60GB chunk then tried to load 100 million rows via lots of
parallel loads. This was IDS 9.40 with the checkpoint interval set to
15 minutes. As this was all to one chunk the checkpoints were single
threaded and taking >15 minutes per checkpoint!! Once they partitioned
across lots of 2GB chunks checkpoints started using lots of cpus and
dropped to 30sec-1 minute.
It depends on how much to flush and also how many chunks the dirty
pages are spread across and disk performance.
Run onstat -u and look for entries with the flags column ending in F,
see how the writes are balanced across page flushers.
Try running iostat -x 4 or whatever the HP-UX equivalent is to see
disk i/o times per device. Possible HP-UX has something equivalent to
vxstat that can see i/o performence per volume.
What array is behind the SAN, something like an XP12000 has lots of
stats/settings that can be monitoring/tuned by the HP guys.
<david@smooth1.co.uk> wrote in message news:f09d1744-ab19-4e08-9a9d-4e6f8cdf14a4@y18g2000yqn.googlegroups.com... On 28 Nov, 19:45, "Neil Truby" <neil.tr...@ardenta.com> wrote: > IDS 10.0FC8W2 on HP-UX 11.31 What I really need, David, is someone who works for a huge company with lots of big tin - er, exactly like you as it happens! ;-) - to tell me what throughout *you* get on checkpoints etc, so I can know for sure if this customer's is bad. >> It depends on how much to flush and also how many chunks the dirty pages are spread across and disk performance. In this test case there is a root, physlog, llog, tempdbs and appdbs, each with one chunk. >> Try running iostat -x 4 or whatever the HP-UX equivalent is to see disk i/o times per device. Possible HP-UX has something equivalent to vxstat that can see i/o performence per volume. Service time on the "disks" seem brilliant, a millisecond or two at least >> What array is behind the SAN, something like an XP12000 has lots of stats/settings that can be monitoring/tuned by the HP guys. An EMC Symmetrix. The Daddy of them All! OK, to address Obnoxio's question, a dd of 1g in 2k pages gives: # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=50000 # 500000+0 records in 500000+0 records out real 4:35.5 user 0.4 sys 21.1 This is much worse that a small VM at Ardenta attached to an EMC Clariion, and worse even than a tiny Itanium at Ardenta copying within its own single internal disk. But the customer has had HP and EMC check performance thoroughly, and has been given a clean bill of health. I have an IBM PMR open; they suspect a disk problem. So I'm seeking some comparative data from others on checkpoint times, or even dd rates to try to verify that there is actually no problem.
Neil Truby wrote: > > <david@smooth1.co.uk> wrote in message > news:f09d1744-ab19-4e08-9a9d-4e6f8cdf14a4@y18g2000yqn.googlegroups.com... > On 28 Nov, 19:45, "Neil Truby" <neil.tr...@ardenta.com> wrote: >> IDS 10.0FC8W2 on HP-UX 11.31 > > What I really need, David, is someone who works for a huge company with > lots of big tin - er, exactly like you as it happens! ;-) - to tell me > what throughout *you* get on checkpoints etc, so I can know for sure if > this customer's is bad. > >>> It depends on how much to flush and also how many chunks the dirty > pages are spread across and disk performance. > > In this test case there is a root, physlog, llog, tempdbs and appdbs, > each with one chunk. > >>> Try running iostat -x 4 or whatever the HP-UX equivalent is to see > disk i/o times per device. Possible HP-UX has something equivalent to > vxstat that can see i/o performence per volume. > > Service time on the "disks" seem brilliant, a millisecond or two at least > >>> What array is behind the SAN, something like an XP12000 has lots of > stats/settings that can be monitoring/tuned by the HP guys. > > An EMC Symmetrix. The Daddy of them All! > > OK, to address Obnoxio's question, a dd of 1g in 2k pages gives: > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=50000 > # 500000+0 records in > 500000+0 records out > > real 4:35.5 > user 0.4 > sys 21.1 > > This is much worse that a small VM at Ardenta attached to an EMC > Clariion, and worse even than a tiny Itanium at Ardenta copying within > its own single internal disk. > > But the customer has had HP and EMC check performance thoroughly, and > has been given a clean bill of health. I have an IBM PMR open; they > suspect a disk problem. > > So I'm seeking some comparative data from others on checkpoint times, or > even dd rates to try to verify that there is actually no problem. Maybe you should repeat that writing to /dev/null in order to check what proportion of this is read time. -- Ian Hotmail is for spammers. Real mail address is igoddard at nildram co uk
"Ian Goddard" <goddai01@hotmail.co.uk> wrote in message news:gLOdnTSQGug9JqzUnZ2dnUVZ8h2dnZ2d@pipex.net... > Maybe you should repeat that writing to /dev/null in order to check what > proportion of this is read time. Yes, maybe. What I'd really like is someone to answer my original questions about their expereince of checkpoint throughputs and dd times ... ;-)
"Ian Goddard" <goddai01@hotmail.co.uk> wrote in message news:gLOdnTSQGug9JqzUnZ2dnUVZ8h2dnZ2d@pipex.net... > Neil Truby wrote: > Maybe you should repeat that writing to /dev/null in order to check what > proportion of this is read time. What does this tell me then? # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=500000 500000+0 records in 500000+0 records out real 4:24.8 user 0.4 sys 21.8 # time dd if=/dev/vg00/lvol7 of=/dev/null bs=2k count=500000 500000+0 records in 500000+0 records out real 2:07.9 user 0.4 sys 12.3