Re: Checkpoint durations
Answered: amber (solid confidence) — Art Kagel gives a thorough, credible root-cause analysis (I/O bound checkpoint, RAID5 vs RAID10, chunk/AIO layout); the thread continues into a long comparative-data/tuning discussion and ends with a real 20x throughput fix reported by a third party (LUN segment size), but the original poster (Neil Truby) never confirms his own checkpoint problem was solved.
Advisory only.
Posted in 2008
Thread about long IDS checkpoints (5–8s) on HP-UX with an EMC SAN, where the DBA used simple dd tests (2K blocks) to gauge disk speed and HP/EMC insisted their kit was fine. Replies argued I/O is the bottleneck and suggested: avoid RAID5 in favour of RAID10, check SAN port/channel load and multipathing (powermt), use raw chunks with KAIO, enough AIO VPs, and more/smaller chunks for parallel chunk writes. Others noted dd with 2K is misleading (Informix coalesces writes, big buffers), advised parallel dd runs with cache busting, and one poster reported going from 5MB/s to 100MB/s by changing SAN segment size from 4K to 32K. A side argument on kernel tuning/RESIDENT ensued. No confirmed fix for the original poster is recorded; an Oracle ORION-style tool was suggested but isn't available on HP-UX.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Logging & Checkpoints, Networking & sqlhosts Configuration
It's telling you that it takes about 3.2 seconds to write 500,000 Informix pages to disk on your SAN Taking half of the dev/null time as reading time and half as writing time. If your checkpoints are 5-8 seconds then the IO speed is the biggest part of the bottleneck which is what IDS is telling you - no matter what EMC says. May be the network connections to the SAN, the SAN itself, the disk array configuration (can you say NO RAID5????). Another possibility is the dbspace setup. At chunk write time (ie during a checkpoint) IDS writes out each chunk's dirty pages with a single IO thread. If you are pushing all of that changed data to a small number of very large chunks you are not taking very good advantage of IDS's parallelism. Similarly, are you using RAW chunks or COOKED? If COOKED two points 1) do you have enough AIO VPS configured (you should have at least 1-1.5 AIO VPs per chunk), 2) know that on HPUX enabling and properly tuning KAIO and using RAW chunks will make a significant difference in IO throughput (>25%). Finally make sure that the SA's and EMC have not configured the SAN with a RAID5 array configuration with a very large block size. Databases, especially but not only Informix, do MUCH better with the performance of RAID10 than RAID5 (not to mention that RAID5 is NOT SAFE AT ANY SPEED -see the papers at www.baarf.com) and with small block sizes (16K - 64K at most). EMC is in love with 256K blocks which work well for filesystems but are poor for database style IO. Art On Sat, Nov 29, 2008 at 6:47 PM, Neil Truby <neil.truby@ardenta.com> wrote: > "Ian Goddard" <goddai01@hotmail.co.uk> wrote in message > news:gLOdnTSQGug9JqzUnZ2dnUVZ8h2dnZ2d@pipex.net... > > Neil Truby wrote: > > > Maybe you should repeat that writing to /dev/null in order to check what > > proportion of this is read time. > > What does this tell me then? > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=500000 > 500000+0 records in > 500000+0 records out > > real 4:24.8 > user 0.4 > sys 21.8 > > # time dd if=/dev/vg00/lvol7 of=/dev/null bs=2k count=500000 > 500000+0 records in > 500000+0 records out > > real 2:07.9 > user 0.4 > sys 12.3 > > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list > -- Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves.
Art Kagel wrote: > It's telling you that it takes about 3.2 seconds to write 500,000 > Informix pages to disk on your SAN Taking half of the dev/null time as > reading time and half as writing time. If your checkpoints are 5-8 > seconds then the IO speed is the biggest part of the bottleneck which is > what IDS is telling you - no matter what EMC says. May be the network > connections to the SAN, the SAN itself, the disk array configuration > (can you say NO RAID5????). Another possibility is the dbspace setup. > At chunk write time (ie during a checkpoint) IDS writes out each chunk's > dirty pages with a single IO thread. If you are pushing all of that > changed data to a small number of very large chunks you are not taking > very good advantage of IDS's parallelism. Similarly, are you using RAW > chunks or COOKED? If COOKED two points 1) do you have enough AIO VPS > configured (you should have at least 1-1.5 AIO VPs per chunk), 2) know > that on HPUX enabling and properly tuning KAIO and using RAW chunks will > make a significant difference in IO throughput (>25%). Finally make > sure that the SA's and EMC have not configured the SAN with a RAID5 > array configuration with a very large block size. Databases, especially > but not only Informix, do MUCH better with the performance of RAID10 > than RAID5 (not to mention that RAID5 is NOT SAFE AT ANY SPEED -see the > papers at www.baarf.com <http://www.baarf.com>) and with small block > sizes (16K - 64K at most). EMC is in love with 256K blocks which work > well for filesystems but are poor for database style IO. > > Art In addition to what Art said... Check with your SAN Admin to see how many servers and disks are on the same ports that these drives are on. On my last project we found 200, yes 200 servers on one port. The previous SAN Admin had EMC training and still allowed this to happen. If too many systems are on one fibre channel no amount of tuning of your software is going to matter. A good SAN Admin is going to distribute the load across all the available ports and channels to make sure your system gets the right throughput. Hint: You can get this information with powermt commands. Ask your system administrator or SAN Admin to run reports on I/Os and get a powermt display of your ports. You should see multiple ports--at least two--for each disk. If you don't then you should ask why. This alone has a major impact on performance. EMC can get you those I/O reports as well, all their people are trained to do this. Regarding EMC RAID5, EMC loves their RAID5, so it's most likely already set up this way. I'd be very surprised if it isn't. EMC loves to sell you on their RAID5 being 'superior' to conventional RAID5. The reason the customer uses it is most likely because they cry when they figure out how expensive RAID1-0 is going to be on Symmetrix, so they fall back to RAID5 and move on. Flash drives are coming, hide your wallet! -ID- > > On Sat, Nov 29, 2008 at 6:47 PM, Neil Truby <neil.truby@ardenta.com > <mailto:neil.truby@ardenta.com>> wrote: > > "Ian Goddard" <goddai01@hotmail.co.uk > <mailto:goddai01@hotmail.co.uk>> wrote in message > news:gLOdnTSQGug9JqzUnZ2dnUVZ8h2dnZ2d@pipex.net... > > Neil Truby wrote: > > > Maybe you should repeat that writing to /dev/null in order to > check what > > proportion of this is read time. > > What does this tell me then? > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=500000 > 500000+0 records in > 500000+0 records out > > real 4:24.8 > user 0.4 > sys 21.8 > > # time dd if=/dev/vg00/lvol7 of=/dev/null bs=2k count=500000 > 500000+0 records in > 500000+0 records out > > real 2:07.9 > user 0.4 > sys 12.3 > > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org <mailto:Informix-list@iiug.org> > http://www.iiug.org/mailman/listinfo/informix-list > > > > > -- > Art S. Kagel > Oninit (www.oninit.com <http://www.oninit.com>) > IIUG Board of Directors (art@iiug.org <mailto:art@iiug.org>) > > Disclaimer: Please keep in mind that my own opinions are my own opinions > and do not reflect on my employer, Oninit, the IIUG, nor any other > organization with which I am associated either explicitly or implicitly. > Neither do those opinions reflect those of other individuals affiliated > with any entity with which I am affiliated nor those of the entities > themselves. >
>> "Art Kagel" <art.kagel@gmail.com> wrote in message >> news:mailman.400.1228011609.874.informix-list@iiug.org... It's telling you that it takes about 3.2 seconds to write 500,000 Informix pages to disk on your SAN Taking half of the dev/null time as reading time and half as writing time. If your checkpoints are 5-8 seconds then the IO speed is the biggest part of the bottleneck which is what IDS is telling you - no matter what EMC says. May be the network connections to the SAN, the SAN itself, the disk array configuration (can you say NO RAID5????). Another possibility is the dbspace setup. At chunk write time (ie during a checkpoint) IDS writes out each chunk's dirty pages with a single IO thread. If you are pushing all of that changed data to a small number of very large chunks you are not taking very good advantage of IDS's parallelism. Similarly, are you using RAW chunks or COOKED? If COOKED two points 1) do you have enough AIO VPS configured (you should have at least 1-1.5 AIO VPs per chunk), 2) know that on HPUX enabling and properly tuning KAIO and using RAW chunks will make a significant difference in IO throughput (>25%). Finally make sure that the SA's and EMC have not configured the SAN with a RAID5 array configuration with a very large block size. Databases, especially but not only Informix, do MUCH better with the performance of RAID10 than RAID5 (not to mention that RAID5 is NOT SAFE AT ANY SPEED -see the papers at www.baarf.com) and with small block sizes (16K - 64K at most). EMC is in love with 256K blocks which work well for filesystems but are poor for database style IO. Thanks Art. Not sure where you get 3.2 seconds. The figures from dd were 264s to write 500,000 2k pages to disk, and 127s to "write" the same to /dev/null. Usually we do the SAN administration for customers so I have a view overv all this stuff. But here, the customer pays EMC to do it, so I'm in the dark really. But we informed that the mirroring is RAID-10. I take all your points about the layout of the SANs and dbs. I know all this stuff but maybe I can't see the wood for the trees. What I would really like to see, if anyone could provide it, is some checkpoint write rates/dd run times on other Big Tin to assess whether the rates we are seeing here are reasonable.
"InDeep" <indeep@indeep.com> wrote in message news:49321236$0$14699$6c36adad@news.usenetserver.com... > Art Kagel wrote: >> It's telling you that it takes about 3.2 seconds to write 500,000 >> Informix pages to disk on your SAN Taking half of the dev/null time as >> reading time and half as writing time. If your checkpoints are 5-8 >> seconds then the IO speed is the biggest part of the bottleneck which is >> what IDS is telling you - no matter what EMC says. May be the network >> connections to the SAN, the SAN itself, the disk array configuration (can >> you say NO RAID5????). Another possibility is the dbspace setup. At >> chunk write time (ie during a checkpoint) IDS writes out each chunk's >> dirty pages with a single IO thread. If you are pushing all of that >> changed data to a small number of very large chunks you are not taking >> very good advantage of IDS's parallelism. Similarly, are you using RAW >> chunks or COOKED? If COOKED two points 1) do you have enough AIO VPS >> configured (you should have at least 1-1.5 AIO VPs per chunk), 2) know >> that on HPUX enabling and properly tuning KAIO and using RAW chunks will >> make a significant difference in IO throughput (>25%). Finally make sure >> that the SA's and EMC have not configured the SAN with a RAID5 array >> configuration with a very large block size. Databases, especially but >> not only Informix, do MUCH better with the performance of RAID10 than >> RAID5 (not to mention that RAID5 is NOT SAFE AT ANY SPEED -see the papers >> at www.baarf.com <http://www.baarf.com>) and with small block sizes >> (16K - 64K at most). EMC is in love with 256K blocks which work well for >> filesystems but are poor for database style IO. HP and EMC have told the customer that their parts of the system are performing well. What I would really like to see, if anyone could provide it, is some checkpoint write rates/dd run times on other Big Tin to assess whether the rates we are seeing here are reasonable. If they aren't I could perhaps press the issue. Otherwise I can't. Sending me tuning tips is great, and wuld be the next valuable step. But firstly I would really value some statistics from others so that i can assess if I even have a problem or not. thanks Neil
Neil Truby wrote: > What I would really like to see, if anyone could provide > it, is some checkpoint write rates/dd run times on other Big Tin to assess > whether the rates we are seeing here are reasonable. But if your "crappy virtual machine" has faster dd times, what more do you need? -- Cheers, Obnoxio The Clown http://obotheclown.blogspot.com
Obnoxio The Clown wrote: > Neil Truby wrote: >> What I would really like to see, if anyone could provide it, is some >> checkpoint write rates/dd run times on other Big Tin to assess whether >> the rates we are seeing here are reasonable. > > But if your "crappy virtual machine" has faster dd times, what more do > you need? > Good point. There should be two metrics in play. Your disk speed is indicated by the disk drives' specs. The difference is what you're getting from the SAN and the software. You can't go faster than the drive itself, so everything from that point forward is your performance penalty. Like I said before, EMC can give you I/O reports, their people know how to get that for you, so you know what your raw I/O is. The speed of the database, and the OS, to read/write are next. You the DBA can get your stats, and really, it doesn't matter what you see from other sites, the difference from what the disk drive is capable of and what you are getting is all you need to worry about. I'd be willing to bet you're probably going to be faster than others simply because you care, most of the shops I work in don't think about it, they assume the SAN is set up with the right multi-channel architecture, and that the SAN people, yes, even EMCs' own people, know what they are doing. Most of them I've seen I can't give the highest marks, but they will do what you need if you know what to ask--and they can ask their own people if they don't know what to do. EMC likes to sell you the goods, but they expect you to go to their training, and learn how to use it. If you buy it and don't know what to ask, in most cases they do not offer much unless you pay for their professional services. -ID-
"Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message news:mailman.0.1228065491.1132.informix-list@iiug.org... > Neil Truby wrote: >> What I would really like to see, if anyone could provide it, is some >> checkpoint write rates/dd run times on other Big Tin to assess whether >> the rates we are seeing here are reasonable. > > But if your "crappy virtual machine" has faster dd times, what more do you > need? Well my crappy virtual machine is attached to a Clariion which is likely doing nothing else, whereas (as someone else suggested) the customer's DMX may be busy servicing many other calls. But more generally the message back was that a 2k dd was a simplistic and unrealistic test.
(In response to a suggestion from Art): Well yes, 8k is ***far*** superior to 2k. But, what does this tell me, and how can it more representative of what Informix does ... which is write 2k pages, isn't it? # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=500000 500000+0 records in 500000+0 records out real 4m31.09s user 0m0.46s sys 0m21.65s # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=8k count=125000 125000+0 records in 125000+0 records out real 1m19.58s user 0m0.11s sys 0m5.21s # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=8k count=125000 125000+0 records in 125000+0 records out real 1m15.62s user 0m0.12s sys 0m5.91s cheers N
Neil Truby wrote: > "Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message > news:mailman.0.1228065491.1132.informix-list@iiug.org... >> Neil Truby wrote: >>> What I would really like to see, if anyone could provide it, is some >>> checkpoint write rates/dd run times on other Big Tin to assess whether >>> the rates we are seeing here are reasonable. >> But if your "crappy virtual machine" has faster dd times, what more do you >> need? > > Well my crappy virtual machine is attached to a Clariion which is likely > doing nothing else, whereas (as someone else suggested) the customer's DMX > may be busy servicing many other calls. > But more generally the message back was that a 2k dd was a simplistic and > unrealistic test. Well, they would say that, wouldn't they? However, it's been a pretty standard test that has served me and others well for two decades now. -- Cheers, Obnoxio The Clown http://obotheclown.blogspot.com
Obnoxio The Clown wrote: >> But more generally the message back was that a 2k dd was a simplistic >> and unrealistic test. > > Well, they would say that, wouldn't they? However, it's been a pretty > standard test that has served me and others well for two decades now. FYI - over in O-land we used to hear the "unrealistic test" refrain a lot. So to overcome that we have the Orion tool - see http://www.oracle.com/technology/software/tech/orion/index.html Not going to be applicable here, as it models Oracle I/O, but something similar should be reasonably easy to do for Informix, and is a very useful tool for customers wishing to make informed choices around storage.
Neil Truby schrieb: > (In response to a suggestion from Art): > > Well yes, 8k is ***far*** superior to 2k. But, what does this tell me, > and how can it more representative of what Informix does ... which is > write 2k pages, isn't it? > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=500000 > 500000+0 records in > 500000+0 records out > > real 4m31.09s > user 0m0.46s > sys 0m21.65s > > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=8k count=125000 > 125000+0 records in > 125000+0 records out > > real 1m19.58s > user 0m0.11s > sys 0m5.21s > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=8k count=125000 > 125000+0 records in > 125000+0 records out > > real 1m15.62s > user 0m0.12s > sys 0m5.91s > > cheers > N > Hi Neil, testing using dd is not an easy thing. You must consider: If your host has a lot of memory and is running LINUX or SOLARIS you will see that if the size of your sequetial file fits into memory, every repetition of your test will be I/O-less, i.e. very fast. You must destroy your cache inside the host memory, and you can do this by shrinking memory by allocating it to another process, which eats memory. I use an IDS instance to do this and I set RESIDENT to 1. Or you can read another huge file big enough to be sure, that it destroys your cache from your dd. Also you should use more than 1 dd in parallel. Best of course, if you already know how many cleaners are running in parallel during checkpoint. During ckeckpoint you will never see all cleaners staring to run in parallel ending at the same time, but generally you will have many cleaners active in the beginning 75% of you waiting time, and only 1 will run at the very end. It is impossible to map this behavior to dd but you can simulate the beginning, or the 'avarage' behavior. Note that starting dd in parallel will flood - your input side of dd (disk reads) - your output side of dd (writes to the I/O-subsystem) - cache(s) in your I/O-subsystem - the interconnecting network So if you read and write from the same I/O-subsystem (which is typical) you must add input and output MB/sec and IOPS. Many FC switches are able to monitor and display their thruput channelwise, and there you can see the network load. EMC can monitor caherates and I/O rates per host. You should ask for these figures during your test. A good test on large iron will use 15-20GB filesize and start 4,8 and 12 dds in parallel. This will give you an impression when you hit the ceiling, and what resource is limiting your speed. dic_k -- Richard Kofler SOLID STATE EDV Dienstleistungen GmbH Vienna/Austria/Europe
Mark Townsend wrote: > Obnoxio The Clown wrote: > > >>> But more generally the message back was that a 2k dd was a simplistic >>> and unrealistic test. >> >> Well, they would say that, wouldn't they? However, it's been a pretty >> standard test that has served me and others well for two decades now. > > > FYI - over in O-land we used to hear the "unrealistic test" refrain a > lot. So to overcome that we have the Orion tool - see > http://www.oracle.com/technology/software/tech/orion/index.html > > Not going to be applicable here, as it models Oracle I/O, but something > similar should be reasonably easy to do for Informix, and is a very > useful tool for customers wishing to make informed choices around storage. To me, this would seem more applicable than dd tests. At least it will (hopefully) show up what would be "bad" for Oracle.
On Dec 1, 3:00 am, Richard Kofler <richard.kof...@chello.at> wrote: > You must destroy your cache inside the host memory, and you can do > this by shrinking memory by allocating it to another process, which > eats memory. I use an IDS instance to do this and I set RESIDENT to 1. > Or you can read another huge file big enough to be sure, that it > destroys your cache from your dd. Slightly off topic, the whole "memory resident" thing isn't a good idea. The short reason... Your Unix Kernel usually tunes itself prior to the start up of IDS during the boot process. So you're tying up memory that the system anticipates as being available. The other part of the argument is that if IDS is being utilized it will be in memory so you don't need to set the memory resident flag. (When was the last time you set your kernel's parameters manually?) -G
Ian Michael Gumby wrote: > On Dec 1, 3:00 am, Richard Kofler <richard.kof...@chello.at> wrote: > >> You must destroy your cache inside the host memory, and you can do >> this by shrinking memory by allocating it to another process, which >> eats memory. I use an IDS instance to do this and I set RESIDENT to 1. >> Or you can read another huge file big enough to be sure, that it >> destroys your cache from your dd. > > Slightly off topic, the whole "memory resident" thing isn't a good > idea. > > The short reason... Your Unix Kernel usually tunes itself prior to the > start up of IDS during the boot process. So you're tying up memory > that the system anticipates as being available. The other part of the > argument is that if IDS is being utilized it will be in memory so you > don't need to set the memory resident flag. > > (When was the last time you set your kernel's parameters manually?) > > -G Gumpy, 1. Type this command on your linux command-line: /sbin/sysctl -a 2. Pull foot out of mouth
On Dec 1, 10:40 pm, InDeep <ind...@indeep.com> wrote: > Ian Michael Gumby wrote: > > On Dec 1, 3:00 am, Richard Kofler <richard.kof...@chello.at> wrote: > > >> You must destroy your cache inside the host memory, and you can do > >> this by shrinking memory by allocating it to another process, which > >> eats memory. I use an IDS instance to do this and I set RESIDENT to 1. > >> Or you can read another huge file big enough to be sure, that it > >> destroys your cache from your dd. > > > Slightly off topic, the whole "memory resident" thing isn't a good > > idea. > > > The short reason... Your Unix Kernel usually tunes itself prior to the > > start up of IDS during the boot process. So you're tying up memory > > that the system anticipates as being available. The other part of the > > argument is that if IDS is being utilized it will be in memory so you > > don't need to set the memory resident flag. > > > (When was the last time you set your kernel's parameters manually?) > > > -G > > Gumpy, > > 1. Type this command on your linux command-line: > > /sbin/sysctl -a > > 2. Pull foot out of mouth Again, how often do you tune your kernel? It used to be that you had to. Then it was limited because some parameters were functions of other parameters. Now most tend to do this automatically. Of course I'm going back to SCO days, HP-UX and AIX many moons ago. Back when you were doing that porn thing on Sybase. ;-)
On Nov 30, 9:13 pm, "Neil Truby" <neil.tr...@ardenta.com> wrote:
> (In response to a suggestion from Art):
>
> Well yes, 8k is ***far*** superior to 2k. But, what does this tell me, and
> how can it more representative of what Informix does ... which is write 2k
> pages, isn't it?
Neil,
I don't think it necessarily does it, at least not if you use KAIO.
This is output from "onstat -g iob" on my main instance:
IBM Informix Dynamic Server Version 10.00.FC8W2 -- On-Line (Prim) --
Up 5 days 10:07:46 -- 58418704 Kbytes
AIO big buffer usage summary:
class reads writes
pages ops pgs/op holes hl-ops hls/op pages ops
pgs/op
kio 96831278 46041964 2.10 5094718 851714 5.98 6224761
1581086 3.94
So, almost 4 pages per average IO write request. And this is on an
average day, not peak load.
But back to your original question, some numbers finally :)
Had the very similar situation as you until few months ago, even
worse, we flushed approx 5 MB/sec on our SAN. Did all possible tests:
dd (with parallel threads), pfread, table (un)loading, index building
and modeling the main business processes/application requests.
Always got 5MB/sec max, no matter what. Until we reconfigured the SAN,
pretty much along the lines of Art's suggestions. Now we can flush 80
to 100 MB per second - a factor 20 increase!
Since you're already tracing the checkpoints you have "a proof" it's
the dskflush() is taking that checkpoint time, and not wait4critex().
What we found out is it's not really the throughput problem as MBs/
sec, it's the number of IO requests outstanding on physical disks in
SAN. The reason for this was LUN's were built using 4k "segment size"
because we also thought Informix only writes page by page (4k in our
case, AIX). With such small segment/stripe sizes you get a lot of IOs
on a busy day.
We changed that to 8 pages so 32k (we've gathered AIX block size stats
and decided 32k is a good starting point and trade-off between
optimising for average day and peak time load; also, Informix big
buffers should be 8 pages if that's relevant at all). So we slashed
the number of IO requests by factor of 8 (theoretically, but close to
truth during the peak load). And voila, using the same tests as before
we went from 5 MB/s to 100 MB/s.
HTH
Davorin
On 30 Nov, 21:13, "Neil Truby" <neil.tr...@ardenta.com> wrote: > (In response to a suggestion from Art): > > Well yes, 8k is ***far*** superior to 2k. But, what does this tell me, and > how can it more representative of what Informix does ... which is write 2k > pages, isn't it? > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=2k count=500000 > 500000+0 records in > 500000+0 records out > > real 4m31.09s > user 0m0.46s > sys 0m21.65s > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=8k count=125000 > 125000+0 records in > 125000+0 records out > > real 1m19.58s > user 0m0.11s > sys 0m5.21s > > # time dd if=/dev/vg00/lvol7 of=/dev/vg01/rneil_1 bs=8k count=125000 > 125000+0 records in > 125000+0 records out > > real 1m15.62s > user 0m0.12s > sys 0m5.91s > > cheers > N Informix does not always use 2k a) it can use 16k "bug buffers" for i/o b) it can coalecse writes together at checkpoint time
Ian Michael Gumby wrote: > On Dec 1, 10:40 pm, InDeep <ind...@indeep.com> wrote: >> Ian Michael Gumby wrote: >>> On Dec 1, 3:00 am, Richard Kofler <richard.kof...@chello.at> wrote: >>>> You must destroy your cache inside the host memory, and you can do >>>> this by shrinking memory by allocating it to another process, which >>>> eats memory. I use an IDS instance to do this and I set RESIDENT to 1. >>>> Or you can read another huge file big enough to be sure, that it >>>> destroys your cache from your dd. >>> Slightly off topic, the whole "memory resident" thing isn't a good >>> idea. >>> The short reason... Your Unix Kernel usually tunes itself prior to the >>> start up of IDS during the boot process. So you're tying up memory >>> that the system anticipates as being available. The other part of the >>> argument is that if IDS is being utilized it will be in memory so you >>> don't need to set the memory resident flag. >>> (When was the last time you set your kernel's parameters manually?) >>> -G >> Gumpy, >> >> 1. Type this command on your linux command-line: >> >> /sbin/sysctl -a >> >> 2. Pull foot out of mouth > > Again, how often do you tune your kernel? > Every single installation. Every single kernel compile. > It used to be that you had to. Then it was limited because some > parameters were functions of other parameters. Now most tend to do > this automatically. > > Of course I'm going back to SCO days, HP-UX and AIX many moons ago. > Back when you were doing that porn thing on Sybase. ;-) I'm surprised you don't understand kernel tuning better, having more kernel experience than a lot of people. Oracle installations require certain kernel params, same for Informix. -ID-
I do have more experience tuning kernels. Thats why I don't try to tune them unless I have to. HP-UX starting around version 9 started using formulas when it came to tunable parameters. If you replace the formula with a hard value, you could screw things up with unintended consequences. If you modify the formula with a different formula, you could screw up other variables and again, screw things up with unintended consequences. Also there became this separation of DBA and Sys Admin roles. So you have Logical DBAs, Physical DBAs, Sys Admins, and network Admins who all get involved on the back end. If its a web based app, toss in your webmaster/app server master too. Since most machines served dual or multiple purposes, rather than just a dedicated database server, you erred on the side of caution by not tuning the kernel. In today's machines, when the machines boot, they do a self discovery and they tend to tune themselves. In addition, it takes a bit of time to make the changes and test them. This costs money. Any performance boost you may get is minimal and you end up facing the statement ... "Just add more CPU/Memory/Disk. Its cheaper and lasts longer than the cost of a good consultant.". Also, all of the changes you just made get tossed out of whack when they decide to add an additional application on the server. The point is that you don't muck with the kernel. You're better off tuning the engine to what you have than mucking up the kernel. There are other areas that you can change that will have a greater impact on performance. To give you an example... I have an Browning A-Bolt Rem 7mm sitting in storage since I don't get to hunt much anymore. I spent a lot of time and money in ammo to get my rifle sighted in to 200yrds and harmonicly balanced to a specific brand of ammo. (Hornaddy 139gr molly coated SSTs) I can get sub MOA shots under good conditions from a bench. I use this ammo primarily for coyotes and ferral dogs, but also effective on deer. If I shift to a different ammunition and a heavier bullet, my grouping expands and the point of impact will shift slightly. (You want a heavier slug for larger game and different bullets have different characteristics upon impact). We're talking 1-3" depending on bullet shape, weight, quality of cartridge, and weather conditions. (And of course, I'm estimating the distance too which has the largest impact on performance). The point is that I could spend the time and money in ammunition to sight my rifle for each different load. But why? In the field, I'm still able to put the target down because the rifle performs well enough. There are other factors which will have a greater impact on performance than trying to tune the barrel harmonics. (Oh and if I switch to the muzzle break, all bets are off and you have to tune it again, although you'll have a good starting point.) Your trying to mod the kernel is akin to tuning the barrel harmonics. You can do it, but you'd be better off spending your money elsewhere. -G > Date: Tue, 2 Dec 2008 19:35:26 -0800 > From: indeep@indeep.com > Subject: Re: Checkpoint durations > To: informix-list@iiug.org > > Ian Michael Gumby wrote: > > On Dec 1, 10:40 pm, InDeep <ind...@indeep.com> wrote: > >> Ian Michael Gumby wrote: > >>> On Dec 1, 3:00 am, Richard Kofler <richard.kof...@chello.at> wrote: > >>>> You must destroy your cache inside the host memory, and you can do > >>>> this by shrinking memory by allocating it to another process, which > >>>> eats memory. I use an IDS instance to do this and I set RESIDENT to 1. > >>>> Or you can read another huge file big enough to be sure, that it > >>>> destroys your cache from your dd. > >>> Slightly off topic, the whole "memory resident" thing isn't a good > >>> idea. > >>> The short reason... Your Unix Kernel usually tunes itself prior to the > >>> start up of IDS during the boot process. So you're tying up memory > >>> that the system anticipates as being available. The other part of the > >>> argument is that if IDS is being utilized it will be in memory so you > >>> don't need to set the memory resident flag. > >>> (When was the last time you set your kernel's parameters manually?) > >>> -G > >> Gumpy, > >> > >> 1. Type this command on your linux command-line: > >> > >> /sbin/sysctl -a > >> > >> 2. Pull foot out of mouth > > > > Again, how often do you tune your kernel? > > > > Every single installation. Every single kernel compile. > > > It used to be that you had to. Then it was limited because some > > parameters were functions of other parameters. Now most tend to do > > this automatically. > > > > Of course I'm going back to SCO days, HP-UX and AIX many moons ago. > > Back when you were doing that porn thing on Sybase. ;-) > > I'm surprised you don't understand kernel tuning better, having more > kernel experience than a lot of people. Oracle installations require > certain kernel params, same for Informix. > > > -ID- > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list _________________________________________________________________ Send e-mail faster without improving your typing skills. http://windowslive.com/Explore/hotmail?ocid=TXT_TAGLM_WL_hotmail_acq_speed_122008
"Mark Townsend" <markbtownsend@sbcglobal.net> wrote in message news:4933879C.60903@sbcglobal.net... > Obnoxio The Clown wrote: > > >>> But more generally the message back was that a 2k dd was a simplistic >>> and unrealistic test. >> >> Well, they would say that, wouldn't they? However, it's been a pretty >> standard test that has served me and others well for two decades now. > > > FYI - over in O-land we used to hear the "unrealistic test" refrain a lot. > So to overcome that we have the Orion tool - see > http://www.oracle.com/technology/software/tech/orion/index.html > > Not going to be applicable here, as it models Oracle I/O, but something > similar should be reasonably easy to do for Informix, and is a very useful > tool for customers wishing to make informed choices around storage. Have had a look and think it *might* be very useful in diagnosing disk problems. Unfortunately for this case it is not available for HP-UX. rgds Neil
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g