Re: High lngspins - IDS 9.30 UC2
Posted in 2003
Topics: Performance & Tuning, Versions, Editions & End-of-Life
Mark D. Stock wrote: > > I would like to see at least three on any system, especially one > making heavy use of temp tables. Preferably on different disks. > Although that may be difficult with your striping. And anyway, striping is the enemy of performance on standard Intel systems in these days of multi-multi-GB disks. It overloads the disk with too many active access threads, which reduces performance to the minimum the disk can offer instead of the maximum it can offer. That's typically a factor of about 4:1.
On Sun, 19 Oct 2003 19:18:16 -0400, Andrew Hamm wrote: > Mark D. Stock wrote: >> >> I would like to see at least three on any system, especially one making heavy >> use of temp tables. Preferably on different disks. Although that may be >> difficult with your striping. > > And anyway, striping is the enemy of performance on standard Intel systems in > these days of multi-multi-GB disks. It overloads the disk with too many active > access threads, which reduces performance to the minimum the disk can offer > instead of the maximum it can offer. That's typically a factor of about 4:1. That requires further elucidation, Mark. Please!?! Art S. Kagel
On Mon, 20 Oct 2003 09:19:09 -0400, Art S. Kagel wrote: Oops, should have addressed my question to Andrew, not Mark. Sorry. Art S. Kagel > On Sun, 19 Oct 2003 19:18:16 -0400, Andrew Hamm wrote: > >> Mark D. Stock wrote: >>> >>> I would like to see at least three on any system, especially one making >>> heavy use of temp tables. Preferably on different disks. Although that may >>> be difficult with your striping. >> >> And anyway, striping is the enemy of performance on standard Intel systems in >> these days of multi-multi-GB disks. It overloads the disk with too many >> active access threads, which reduces performance to the minimum the disk can >> offer instead of the maximum it can offer. That's typically a factor of about >> 4:1. > > > That requires further elucidation, Mark. Please!?! > > Art S. Kagel
Andrew Hamm wrote: > > Mark D. Stock wrote: > > > > I would like to see at least three on any system, especially one > > making heavy use of temp tables. Preferably on different disks. > > Although that may be difficult with your striping. > > And anyway, striping is the enemy of performance on standard Intel systems > in these days of multi-multi-GB disks. It overloads the disk with too many > active access threads, which reduces performance to the minimum the disk can > offer instead of the maximum it can offer. That's typically a factor of > about 4:1. I now for a while testing on an Intel box with two 3-ware HW RAID controllers on 2 PCI busses 64bit/66MHz 1 RAID-5 with 4 parallel 160 GB IDE disks 1 RAID 10 with 4 parallel 160 GB IDE disks Running 6 bonnie++ tests in parallel shows me better read perf from RAID5. Ratio RAID 10 : RAID 5 is like 2:3 (reads from 3 instead of 2 disks at any given time, so it makes sense) Write performance of the RAID5 just suxx (20 MB/sec). bonnie++ random write performance of the RAID10 is over 105MB/sec HPL performance on this system beats the heavy $$$$ SANs by a magnitude! (if NOT loading onto the RAID5): I can load 80.000 2 KB pages per sec reading from RAID5 loading into 8 chunks which sit on the RAID10. (Version 9.40.UC2 on SuSE 8.2) I con unload >120.000 pg/sec from the RAID5 and 81.000 from RAID10 Parallism of more than 8 HPL threads tends to increase the overall loading duration, whereas 7 parallel HPL threads is not as fast as 8. 8-way parallel on this 4 disk RAID10 is the sweet spot for writes, so it seems. dic_k -- Richard Kofler SOLID STATE EDV Dienstleistungen GmbH Vienna/Austria/Europe
Art S. Kagel wrote: > Andrew Hamm wrote: >> >> And anyway, striping is the enemy of performance on standard Intel >> systems in these days of multi-multi-GB disks. It overloads the disk >> with too many active access threads, which reduces performance to >> the minimum the disk can offer instead of the maximum it can offer. >> That's typically a factor of about 4:1. > > That requires further elucidation, Mark^H^H^H^HAndrew. Please!?! welllllll what follows, I have found does NOT apply to large funky highly-resourced disk arrays like Clariions, however no "ordinary" hardware (ie using SCSI and raid controllers) on Intel escapes the following principles. I know also that typical HPs are affected. I think I tested a Sun box two years ago and saw the same results. After that, I can't say, however I will make an educated guess that any box directly interfacing to SCSI controllers, or raid boxes based around SCSI controllers, will fit the following model. The Clariion arrays escape because they have massive non-volatile cache which completely hides the performance characteristics of the disks; for any controller system to escape, it must also have smarts that are at least equivalent to Clariion arrays. on with the story... Starting at the disk, you can make multiple requests on disks which are queued and processed by the disks on-board processor. The multiple requests can come from multiple processes which require service. IDS is a beautiful example of this, because it literally has multiple processes using one file handle per chunk to perform I/O. If you are using KAIO then you just have to map through the indirection that KAIO offers, but it still applies... Each file handle used by IDS - consider it one "thread" of access. One thread of access might return performance of 400. Two threads 600, three 970, upto 1500 for 8 threads, then suddenly the performance drops to 800, 500, 400, 400, 400, ..... for 9, 10, 11,.... threads. These numbers are from an actual measurement of an HP 3 years ago - the only sample files I could find :-) The numbers measure the KB/second when performing random transfers of 2K pages (somewhat like a busy IDS engine running an OLTP system:-) The numbers are similar proportionally but of course with higher throughput when the activity is sequential accesses. So, if you put too many ACTIVE threads on a disk, it will ultimately limit performance to the baseline (ie 400 in this case) just as surely as having too few ACTIVE threads. Now, combine striping with IDS's 2GB/chunk limit (prior to 9.40) and a multi-GB disk set, and you will overload the disks. For example, if you have 4 x 40Gb disks, and you allocate your chunks using striping: 2Gb striped over 4 way will mean that each disk will contribute 0.5G to each chunk. That leaves 39.5Gb on each disk, so you'll probably allocate a few more striped chunks. Allocate 8 striped chunks, and you will have consumed 4Gb only on each disk. At this point, assuming each chunk is generally busy, you will have fully loaded the disks with all the threads it can handle. Your performance might be the 1500 times 4 due to the striping. Now add just two more striped chunks and make them busy. Suddenly each disk is overloaded, their performance slumps to 400, so total performance is only 400 x 4. In this case, almost 4 times slower than maximum, and you still haven't consumed more than 1/8th the available space. Bad news. Clariions manage to avoid this problem. I guess they use their massive cache to absorb data as quickly as possible, and then the on-board processing doles out the disk activity in a scientifically optimised manner to make the best use of the disks. In other words, a Clariion must map take your N threads to 8 threads per disk no matter what your N is. For any box to offer this performance, it must be as smart. I have no measurements of SANs or other manufacturers smart disk arrays, but it's very easy to test this for yourself. If you are using standard SCSI hardware (and this includes most commodity boxes with Intel, HP, Sun, etc etc etc inside) then you can measure your disk system and find out how lucky you are. Don't count on it. So, although striping was phat news in the early and mid 90's when disks were a few Gb if you were lucky, these days you simply can't afford to stripe across a disk with dozens and dozens of Gb. What might help you avoid this situation? IDS 9.40 now supports >2Gb, correct? PHEW!!! Allocate chunks >2Gb which directly cover your needs. Decision to use striping depends on your application: I keep capitalising ACTIVE chunks. If you have a data warehouse, then presumably you CAN profit from striping. You just need to be absolutely sure about typical access patterns. Just don't overload the disk. One trick that can help with DW performance - fragment a large table onto the same disk! consider your typical fragment elimination etc, and if you can get upto 8 busy fragments from the same disk (and perhaps another 8 on another disk, etc) then you can count on a maximum throughput of 1500 per disk for a single-table scan. Jack Parker has pushed this trick in the past. Very clever. People running OLTP cannot afford the luxury of the fantasy that only one thing will be happening at once. You need to spread your data evenly across all available disks so that each disk keeps equally busy. Striping multiplies the thread load on a disk, so it is now an enemy of performance on todays standard multi-gig disks. A busy OLTP system will be simultaneously perform random reads of indexes and tables, sorts, scans, updates, LRU writes, Checkpoint writes. A lot of people have issues with checkpoint durations; applying these principles can cause your checkpoints to equally load all of your disks to the maximum of their capability, and that can result in a startling performance benefit. For example, if a system is badly setup so that mostly one chunk gets all the new write activity, then it's checkpoint duration could be 20 seconds (an example of an appalling checkpoint time which typically causes people to get as much LRU writes as possible). Now, spread out the load to max on that disk, and immediately your checkpoint might drop to 8 seconds. Finally, spread the load also to 2, 3, 4 disks, and your checkpoint times will drop to 4, 3, 2 seconds. That's a FANTASTIC improvement of checkpoint times. Even if you believe that LRU writes are King, guess what? Spread the load out properly, and you'll realise the same performance benefits for LRU writes. You might even be able to ease off on the LRU percentages, which means there's more performance available for queries, which is a handy bonus on top of the fact that query performance will ALSO benefit equally from spreading the load around. Last few years, when setting up OLTP, I've had to try to get creative when allocating spaces. You have to find quiet but large tables that can be put into spaces that aren't busy, and therefore will not affect normal performance. Other tricks might be to find tables that do partake in large sequential scans, and spread them around. Even on an OLTP system, some of the big ugly tables do tend