Informix IDS and RAID 5
Posted in 1999
A user asked why RAID 5 is commonly discouraged for Informix IDS. Respondents explained it's mainly a performance issue: RAID 5's rotating parity forces read-modify-write cycles on updates, roughly halving write throughput versus RAID 0, and striping across all disks removes the ability to isolate tables on separate spindles. Art Kagel added further drawbacks: undetected partial media failures corrupting parity, total data loss if a second drive fails, and heavy degradation during rebuilds, recommending RAID 10 (or RAID 3/4) instead. No problem to fix; the thread ends with the advice accepted.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: General Discussion
I have seen several recommendations to avoid RAID 5. Can someone tell me exactly what the problems with RAID 5 are?
On Mon, 29 Nov 1999 17:18:51 -0500, Barry Lloyd <bdl62@42_dirm_po.dmh.state.sc.us> wrote: >I have seen several recommendations to avoid RAID 5. Can someone tell >me exactly what the problems with RAID 5 are? > We use Raid 5 with IDS (raw chunks less than 2 GB) since a year. No problems yet. Reinhard
Barry Lloyd wrote: > I have seen several recommendations to avoid RAID 5. Can someone tell > me exactly what the problems with RAID 5 are? There are no functionality problems. The problem is performance as RAID 5 is designed giving data security with minimum disk space. RAID 5 has to distribute data and parity information across several disks and such multiplying IO's, especially in write. Whereas RAID 1/0 is very similar to conventional mirroring (fast, disks can be written in parallel, but needs double disk space). Some RAID systems (e.g. EMC) supply RAID S which needs less disk space than RAID 1/0 but is considerabely more performant than RAID 5. Regards -- Helmut Leininger Bull AG / Vienna Open Systems Support Email: h.leininger@bull.at helmut.leininger@bull.net This opinion is mine and not necessarily that of my employer. No guarantees whatsoever.
The primary issue is performance. Basically, RAID 5 uses "rotating parity," which means that the byte at position X of one of the disks can be recalculated by using the bytes at position X of the other disks in the group. This means that one disk worth of data in a RAID 5 grouping is sacrificed for this rotating parity. Example: If you've got a 5-disk RAID 5 rank, then you've really only got four disks worth of data, because the fifth part is for parity. Since the parity rotates, basically one fifth of each disk is reserved for this. What's bad about this? It means that every time you read or write to the disk, you can't just read or write; you also have to do the parity calculation. If your RAID 5 is done in hardware, then the performance hit is usually pretty small. If it's software-based, however, you're going to take a huge performance hit. The advantage to RAID 5 is that you can lose any one disk in each rank without losing any data. Any two disks go bad, though, and the entire rank is lost. With mirroring, exactly the right two disks have to go bad, not just any two. Also, if you do lose a disk in a RAID 5 rank, your accesses will become even slower until that disk is replaced and reconstructed. The other downside to RAID 5 is actually the same as any type of striping. That is, it's tempting to just stripe across all the available disks, giving yourself one big rank; this means that you can't isolate tables from one another because you don't have independently addressable spindles. I address this issue in my forthcoming article, "Optimal Disk Layouts for Informix Dynamic Server Systems," tentatively scheduled for publication in the 2Q2000 edition of Informix Tech Notes. Hope this helps, - Tom Girsch "Barry Lloyd" <bdl62@42_dirm_po.dmh.state.sc.us> wrote in message news:3842FBCB.43C016BE@42_dirm_po.dmh.state.sc.us... > I have seen several recommendations to avoid RAID 5. Can someone tell > me exactly what the problems with RAID 5 are? >
Reinhard Habichtsberg wrote: > > On Mon, 29 Nov 1999 17:18:51 -0500, Barry Lloyd > <bdl62@42_dirm_po.dmh.state.sc.us> wrote: > > >I have seen several recommendations to avoid RAID 5. Can someone tell > >me exactly what the problems with RAID 5 are? Hi Barry, Most of those disparaging postings are from me. There are two problems with RAID5. The first is performance which is the one most people notice and if you can live with write throughput which is 50% of the equivalent RAID0 stripe set then that is fine. The performance hit is caused because RAID5 ONLY reads the one drive containing the requested sector leaving the other drives free to return other sectors from different stripe blocks. This is the reason that RAID5 is preferred to RAID3 or RAID4 for filesystems, this feature improved small random read performance. However, since the parity and the balance of the stripe block were not read, if you rewrite the block (which databases do far more frequently than filesytems) the other drives must all be read and a new parity calculated and then both the modified block and the parity block must be written back to disk. This READ-WRITE-READ-WRITE for each modified block is the reason RAID5 is so poor in terms of write throughput. Large RAID controller caches and on controller firmware level RAID implementations alleviate the problem somewhat but not completely and write performance still hovers at around half what a pure stripe (RADI0) would get. The second problem, despite what others have said IS a FUNDAMENTAL problem with the design of RAID5 which various implementors have tried to correct with varying levels of success. The problem is that if a drive fails slowly over time, known as partial media failure, where periodically a sector or two goes bad, this is NOT detected by RAID5's parity and so is propagated to the parity when that sector is rewritten which means that if another drive fails catastrophically its data will be rebuilt utilizing damaged parity resulting in two sectors with garbage. Now this may not even be noticed for a long time as modern SCSI drives automatically remap bad sectors to a set of sectors set asside for the purpose but the corrected error is NOT reported to the OS or the administrators. Over time if the drive is going it will run out of remap sectors and will have to begin returning data reconstructed from the drive's own ECC codes. Eventually the damage will exceed the ECC's ability to rebuild a single bit error per byte and will return garbage. RAID3 and RAID4 are superior in both areas. In both all drives are read for any block which improves sequential read performance (Informix Read Ahead depends on sequential read performance) over RAID5 and parity can be (and in most implementations IS) checked at read time so that partial media failure problems can be detected. Write performance is approximately the same as RAID0 for large writes or smaller stripe block sizes. One problem with early implementations of RAID3/4 was slow parity checking since it has to be calculated for every read and every write. Modern controller based RAID systems use the on-board processor on the SCSI controller to perform the parity checks without impacting system performance by tying up the a system CPU to check and produce parity. These RAID levels require the exact same number of drives as RAID5. RAID10 provides the best protection and performance with read performance exceeding any other RAID level (since both drives of a mirrored pair can be reading different sectors on parallel) and write performance is closest to pure striping. Indeed in a hardware/firmware implemented RAID10 array with on-board cache apparent write throughput can exceed RAID0 for brief periods due to the two drives of each pair being written to independently though the gain is not sustainable over time. A third problem with ALL RAID3/4/5 from which RAID10 does not suffer is multiple drive failure. (Ever get a batch or 200 bad drives? We have!) If one drive in a RAID3/4/5 array fails catastrophically you are at risk for complete data loss if ANY of the remaining 4 (or more) drives should fail before the original failed drive can be replaced and rebuilt. With RAID10, since it is made up as a stripe set of N mirrored pairs, when a drive fails you are only at risk for complete data loss if that one drives particular mirror partner should fail. Make each mirrored pair from drives selected from different manufacturer's lots and the probability of this happening become vanishingly small. Fourth problem. During drive rebuild RAID3/4/5 (and RAID01 mirrored stripe sets) performance of the array during the rebuild can degrade by as much as 80%! Some RAID systems let you tune the relative priority of rebuild versus production to reduce the performance hit to as low as about 40% degradation but this will increase the recovery time increasing the number of production requests that are degraded and increasing the risk of the previous problem with a second drive failure. RAID10, since only one drive is involved in mirror recovery, the array's performance (for a 4 drive array) is degraded only a maximum of 80% for reads and writes against the failed pair and only slightly (due to controller traffic) for accesses to the other drives, on average, since the one pair comprises only 20% of accesses, performance is affected no more than 16% during recovery and the risk of catastrophic data loss is reduced. Well this concludes my quarterly RAID5 rant for anyone who has had trouble finding my earlier ones. Art S. Kagel
Art S. Kagel <kagel@bloomberg.net> schrieb in im Newsbeitrag: 384534F4.D896A91A@bloomberg.net... > Well this concludes my quarterly RAID5 rant for anyone who has had trouble > finding my earlier ones. > > Art S. Kagel Thank you for this detailed explanations. Reinhard
Thomas J. Girsch <tgirsch@iname.com> schrieb in im Newsbeitrag: 3844139c_1@news2.one.net... > > I address this issue in my forthcoming article, "Optimal Disk Layouts for > Informix Dynamic Server Systems," tentatively scheduled for publication in > the 2Q2000 edition of Informix Tech Notes. > Is this article also available in german? Regards Reinhard