Re: Checkpoint durations
Answered: amber (solid confidence) — Continuation of the 'Checkpoint durations' conversation (see also threads 12222/12224). Richard Kofler provides concrete comparative throughput figures and a credible root-cause theory (EMC write-cache exhaustion/upstaging plus SRDF replication overhead); Neil Truby thanks him but the underlying customer issue is never confirmed fixed.
Advisory only.
Posted in 2008
Thread asks how to judge whether checkpoint times are bad on IDS 10.0FC8W2 under HP-UX 11.31 with an EMC Symmetrix SAN; a dd of 1GB in 2K blocks took 4:35, yet HP and EMC declared the hardware healthy and an IBM PMR was open. Advice: compare against the customer's acceptable checkpoint time rather than other sites, spread data/indexes over more dbspaces/chunks, check CPU/paging and per-volume IOPS, and question EMC's guaranteed IOPS. Suggestions to move to IDS 11's non-blocking checkpoints were rejected (app not certified). One poster shared reference figures (~185MB/s, 23K IOPS over 60 LUNs) and suspected write-cache saturation; planned SRDF replication would slow writes further. No resolution recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
On 29 Nov, 20:38, "Neil Truby" <neil.tr...@ardenta.com> wrote: > <da...@smooth1.co.uk> wrote in message > > news:f09d1744-ab19-4e08-9a9d-4e6f8cdf14a4@y18g2000yqn.googlegroups.com... > On 28 Nov, 19:45, "Neil Truby" <neil.tr...@ardenta.com> wrote: > > > IDS 10.0FC8W2 on HP-UX 11.31 > > What I really need, David, is someone who works for a huge company with lots > of big tin - er, exactly like you as it happens! ;-) - to tell me what > throughout *you* get on checkpoints etc, so I can know for sure if this > customer's is bad. > No you do not, I run different applications producing a different load using different machines with potentially different SAN switches and using different arrays with a different disk layout with different performance requirements. What you need to know is what is the maximum checkpoint time your customers is willing to accept? > >> It depends on how much to flush and also how many chunks the dirty > > pages are spread across and disk performance. > > In this test case there is a root, physlog, llog, tempdbs and appdbs, each > with one chunk. How are the writes spread across the dbspaces in the test case? Can you create more dbspaces to spread the load across different chunks e.g. would seperating data and indexes into different dbspaces help? Or different tables into different dbspaces? One dbspace for an app is not enough if you want good checkpoint times. > > >> Try running iostat -x 4 or whatever the HP-UX equivalent is to see > > disk i/o times per device. Possible HP-UX has something equivalent to > vxstat that can see i/o performence per volume. > > Service time on the "disks" seem brilliant, a millisecond or two at least OK so running top whilst the checkpoint is running how many cpus are busy out of the 18? Is something else using cpu time? Is the box swapping/paging heavily? How much memory is free on the box? > > >> What array is behind the SAN, something like an XP12000 has lots of > > stats/settings that can be monitoring/tuned by the HP guys. > > An EMC Symmetrix. The Daddy of them All! How many i/os per second are you getting in the last 2-3 seconds of the checkpoint and across how many different volumes? > > OK, to address Obnoxio's question, a dd of 1g in 2k pages gives: > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=50000 > # 500000+0 records in > 500000+0 records out > > real 4:35.5 > user 0.4 > sys 21.1 > > This is much worse that a small VM at Ardenta attached to an EMC Clariion, > and worse even than a tiny Itanium at Ardenta copying within its own single > internal disk. > > But the customer has had HP and EMC check performance thoroughly, and has > been given a clean bill of health. I have an IBM PMR open; they suspect a > disk problem. > > So I'm seeking some comparative data from others on checkpoint times, or > even dd rates to try to verify that there is actually no problem. EMC with Synmetrix used to guaranttee so many i/os per second do they still do that? What do they say for this setup? Can they tune the array to be able to go faster given your test case? Your two dd's above are the same aren't they? Ask EMC why the times vary so much. Are other machines accessing the array at the same time?
Hate to say the obvious, but... you could try upgrading to IDS 11 and not worry about checkpoints at all since it has nonblocking checkpoint technology. It also includes automatically configuring LRU flushing and lots of nice performance features. Just a suggestion :) On Nov 29, 9:34 pm, "da...@smooth1.co.uk" <da...@smooth1.co.uk> wrote: > On 29 Nov, 20:38, "Neil Truby" <neil.tr...@ardenta.com> wrote:> <da...@smooth1.co.uk> wrote in message > > >news:f09d1744-ab19-4e08-9a9d-4e6f8cdf14a4@y18g2000yqn.googlegroups.com... > > On 28 Nov, 19:45, "Neil Truby" <neil.tr...@ardenta.com> wrote: > > > > IDS 10.0FC8W2 on HP-UX 11.31 > > > What I really need, David, is someone who works for a huge company with lots > > of big tin - er, exactly like you as it happens! ;-) - to tell me what > > throughout *you* get on checkpoints etc, so I can know for sure if this > > customer's is bad. > > No you do not, I run different applications producing a different > load using different machines with > potentially different SAN switches and using different arrays with > a different disk layout with different performance requirements. > > What you need to know is what is the maximum checkpoint time your > customers is willing to accept? > > > >> It depends on how much to flush and also how many chunks the dirty > > > pages are spread across and disk performance. > > > In this test case there is a root, physlog, llog, tempdbs and appdbs, each > > with one chunk. > > How are the writes spread across the dbspaces in the test case? > Can you create more dbspaces to spread the load across different > chunks e.g. would seperating data and indexes into different > dbspaces help? Or different tables into different dbspaces? One > dbspace for an app is not enough if you want good checkpoint times. > > > > > >> Try running iostat -x 4 or whatever the HP-UX equivalent is to see > > > disk i/o times per device. Possible HP-UX has something equivalent to > > vxstat that can see i/o performence per volume. > > > Service time on the "disks" seem brilliant, a millisecond or two at least > > OK so running top whilst the checkpoint is running how many cpus > are busy out of the 18? > > Is something else using cpu time? Is the box swapping/paging > heavily? How much memory is free on the box? > > > > > >> What array is behind the SAN, something like an XP12000 has lots of > > > stats/settings that can be monitoring/tuned by the HP guys. > > > An EMC Symmetrix. The Daddy of them All! > > How many i/os per second are you getting in the last 2-3 seconds of > the checkpoint and across how many different volumes? > > > > > > > OK, to address Obnoxio's question, a dd of 1g in 2k pages gives: > > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=50000 > > # 500000+0 records in > > 500000+0 records out > > > real 4:35.5 > > user 0.4 > > sys 21.1 > > > This is much worse that a small VM at Ardenta attached to an EMC Clariion, > > and worse even than a tiny Itanium at Ardenta copying within its own single > > internal disk. > > > But the customer has had HP and EMC check performance thoroughly, and has > > been given a clean bill of health. I have an IBM PMR open; they suspect a > > disk problem. > > > So I'm seeking some comparative data from others on checkpoint times, or > > even dd rates to try to verify that there is actually no problem. > > EMC with Synmetrix used to guaranttee so many i/os per second do they > still do that? What do they say for this setup? > > Can they tune the array to be able to go faster given your test case? > > Your two dd's above are the same aren't they? Ask EMC why the times > vary so much. > Are other machines accessing the array at the same time?
<david@smooth1.co.uk> wrote in message news:a9b22443-9ba1-44c3-b2f3-830da196ade2@v4g2000yqa.googlegroups.com... > On 29 Nov, 20:38, "Neil Truby" <neil.tr...@ardenta.com> wrote: >> <da...@smooth1.co.uk> wrote in message >> >> news:f09d1744-ab19-4e08-9a9d-4e6f8cdf14a4@y18g2000yqn.googlegroups.com... >> On 28 Nov, 19:45, "Neil Truby" <neil.tr...@ardenta.com> wrote: >> >> > IDS 10.0FC8W2 on HP-UX 11.31 >> >> What I really need, David, is someone who works for a huge company with >> lots >> of big tin - er, exactly like you as it happens! ;-) - to tell me what >> throughout *you* get on checkpoints etc, so I can know for sure if this >> customer's is bad. >> > No you do not, I run different applications producing a different > load using different machines with > potentially different SAN switches and using different arrays with > a different disk layout with different performance requirements. > > What you need to know is what is the maximum checkpoint time your > customers is willing to accept? Well, zero would probably be acceptable ;-) And of course I can tune Informix to do virtually everythign as LRU writes to achieve that. What I would like to assess is whether, by doing so, I'm addressing the problem rather than the cause. In this respect then, even though you are using different hardware, SAN and application profile, if you or anyone else could provide some checkpoint write rates/dd run times run on other Big Tin, it really *would* help me to assess whether the rates we are seeing here are reasonable.
>> <pokeyman76@yahoo.com> wrote in message >> news:b1d117a3-e7ec-4443-944e-c19b8cf758be@z1g2000yqn.googlegroups.com... Hate to say the obvious, but... you could try upgrading to IDS 11 and not worry about checkpoints at all since it has nonblocking checkpoint technology. It also includes automatically configuring LRU flushing and lots of nice performance features. It isn't an option here as the underlying application is not certified for 11.
Neil Truby schrieb: > <david@smooth1.co.uk> wrote in message > news:a9b22443-9ba1-44c3-b2f3-830da196ade2@v4g2000yqa.googlegroups.com... >> On 29 Nov, 20:38, "Neil Truby" <neil.tr...@ardenta.com> wrote: >>> <da...@smooth1.co.uk> wrote in message >>> >>> news:f09d1744-ab19-4e08-9a9d-4e6f8cdf14a4@y18g2000yqn.googlegroups.com... >>> >>> On 28 Nov, 19:45, "Neil Truby" <neil.tr...@ardenta.com> wrote: >>> >>> > IDS 10.0FC8W2 on HP-UX 11.31 >>> >>> What I really need, David, is someone who works for a huge company >>> with lots >>> of big tin - er, exactly like you as it happens! ;-) - to tell me what >>> throughout *you* get on checkpoints etc, so I can know for sure if this >>> customer's is bad. >>> >> No you do not, I run different applications producing a different >> load using different machines with >> potentially different SAN switches and using different arrays with >> a different disk layout with different performance requirements. >> >> What you need to know is what is the maximum checkpoint time your >> customers is willing to accept? > > Well, zero would probably be acceptable ;-) > And of course I can tune Informix to do virtually everythign as LRU > writes to achieve that. > > What I would like to assess is whether, by doing so, I'm addressing the > problem rather than the cause. In this respect then, even though you > are using different hardware, SAN and application profile, if you or > anyone else could provide some checkpoint write rates/dd run times run > on other Big Tin, it really *would* help me to assess whether the rates > we are seeing here are reasonable. Hi Neil, here are some outdated figures from a former cutomer of mine, now running Oracle. Version was 9.40. There were 32 CPUs and a multipathed 1Gbit, full duplex FC network (bonding 2 paths actually). 2 I/O Subsystems were used exclusively by the database server. Each I/O subsystem was quad connected (4x 1Gbit FC). The host used 8 FC boards. We saw 24 flushers beeing active almost to the end of checkpoint. Thruput was around 185 MB/sec, a bit more than 23K IOPS doing 8KB writes we were monitoring as normal behavior, evenly spread over both I/O subsystems. To reach this we had to ensure: - that long blocks are written (8KB per start-I/O) - spread the load evenly on both I/O-subsystems - spread the load over 60 LUNs as even as possible This was neither easy nor a goal to be reached very fast. When we had like 40% of this thruput everyone and her sister tried to convince me, that everything is OK...... But in fact - it was not - that made ne say since then: Best practice does not always mean 'good' ;) And another well hated lil signature of mine: Why spend +50% money to get 50% less performance out of the boxes? The figures you present from dd makes me believe, that if u are on a DMX3, your test was eaviliy hampered by upstaging, i.e. had to wait for physical disks, because the write cache in the EMC was full at all times. This happens when large block sequenial writes are there at the same time as your short block writes (8KB during checkpoint). On the other hand I do know from other customers of mine, that on a DMX3 a write thruput of upto 300MB/sec is possible, though not too long (less than 1 minute) because if the write cache in the subsystem is full .... (see above) One more thing to pay attention to: Is the I/O- subsystem doing replication (like SRDF)?. If so, all depends on the interlink speed and latency and a typical value is that your writes are slower by 30%, even on 4Gbit FC attachment system when the interlink is 1 GBIT only. HTH dic_k -- Richard Kofler SOLID STATE EDV Dienstleistungen GmbH Vienna/Austria/Europe
"Richard Kofler" <richard.kofler@chello.at> wrote in message news:d71c5$4932c7e3$506c1b69$20540@news.chello.at... > One more thing to pay attention to: Is the I/O- subsystem doing > replication (like SRDF)?. If so, all depends on the interlink > speed and latency and a typical value is that your writes are > slower by 30%, even on 4Gbit FC attachment system when the interlink > is 1 GBIT only. Thanks for that. There is no SRDF at present, but the intention is to start SRDF sync. replication, which will make the disk times worse of course.