Re: To RAID or not to RAID? -that is the question...
Posted in 1999
Dennis, Here is a discussion of RAID5 and other RAID taken from the Informix list. Thought you might be interested. > ---- Original Msg from: Dibb <jdibb@clariion.com> At: 7/ 9 13:45 > > Hello all, > Art has some good info here, but maybe I can add some as well. > I'll cut the parts where I agree. The more information the better, thanks. > Art S. Kagel <kagel@bloomberg.net> wrote in message > news:3784CA6A.95F0BC39@bloomberg.net... > > OK time for some ACCURATE information! I know Obnoxio has been waiting > > for me to join in. > > > > Tony Platt wrote: > > > > > > Tony C wrote in message <931333286.482062@cacheraq001>... > > > >Not even close :) > > > > RAID5 is the same as RAID4 EXCEPT that instead of a dedicated parity > > drive containing all of the parity blocks the parity block is rotated > > round robin over all of the drives in the stripe-set. So block 1's > > parity may be on drive 5 (in a 5 drive RAID5 set) then block 2's parity > > is on drive 1, block 3's parity on drive 2 etc. round and round. RAID > > 5 always calculates, for a particular logical block, which drive > > contains the data and which drive has the parity and reads ONLY the > > corresponding block from those two drives. Therefore the other N-2 > > drives in the RAID set are available for similar 2-at-a-time reads and > > so RAID 5 performs better on small random reads than it's cousin RAID 4 > > which does better at sequential read since it always reads all N drives > > every time. > > > When you're talking about 'reads' of two drives, I assume this in the > context of a write operation (kind of lost context from quoted material). > When you are doing host read operations, you only need to read 1 drive, > leaving the other N-1 for other commands. Most RAID5 descriptions I have seen have the single data drive and the parity block read together, though I have never understood why that is when the parity is not used to verify the read anyway. If this is not true great! It frees another drive for interleaved reads. This detail does not change my point that RAID5 allows for interleaved concurrent reads of multiple partial stripe blocks while RAID4 always reads one entire stripe block in parallel which is why RAID5 does better at small random reads while RAID4 tends to do better with larger and sequential reads. > Most RAID5 systems use a large RAID block size to improve > > sequential read performance (64K+). Unlike RAID3 and RAID4, RAID5 does > > not EVER check or verify parity for data reads. That means that if a > > drive becomes flaky (that's a technical term) and the media begins to > > return garbage (yes it can even happen on modern SCSI drives with > > automatic sector remapping, lost remap lists are finite you know!) then > > not only will the RAID5 set return garbage but when you write back the > > parity will be recalculated with the garbage an become trash itself. > > > > > If you're going to talk about drives returning data other than what was > written, without letting you know, then that is a big problem that will > affect raid 1/ raid 1/0 as well as Raid 5. Most Raid 1 systems, usually > being tuned for best performance, will allow the two drives to service data > requests independantly. The data is not checked for consistency on each > command, therefore a 'flaky' drive will affect them the same way. OK, I missed this aspect. My thinking, and that of others before me, is that since RAID1 writes independently to both drives it is less likely to propagate errors to the redundant drive. But I see your point, if garbage is read, which is what I was discussing, then garbage will be written to both drive, true. What is still true of RAID5 but not RAID1 is that more data is damaged, potentially, if a drive fails. If one RAID10 drive, part of a RAID1 pair, fails only those sectors trashed on that one pair, are damaged beyond recovery. Once a flaky sector is written back on a RAID5 set if another drive fails the damaged parity will write garbage to the corresponding sectors of the recovered drive as well propagating the damage to another drive. This RAID10 cannot do. > > > Raid 5 offers > > > Best protection > > > > RAID 5 offers NO PROTECTION against multiple drive failure or against > > partial media failure. ONLY RAID1 and it's derivative RAID10 can be > > > called BEST PROTECTION. If a RAID1 drive becomes flaky the mirror, > > which is written independently, will be fine and can be used to build > > a replacement for the flaky drive. > > Only if you know that the drive is 'flaky'. Unless you're always reading > both (and sacrificing half the performance) you won't know any more than on > a raid 5. I know that I am forgetting something that supports my original position but logic says you are correct. I'll concede this point. > > > Increased throughput > > > > Over what? Over singleton drives? Yes. Over RAID0? No. Over RAID1? > > Yes. Over RAID10? No! No! NO! > > > > > > Most cost effective > > > > Actually exactly as cost effective as RAID3 and RAID4 which are better > > for databases than RAID5. Many RAID systems now calculate RAID3&4 > > parity in hardware on the controller so the main objection to RAID3&4 > > is no longer valid. And RAID3&4 do NOT suffer from reduced write > > performance as RAID5 does since all of the data blocks in a stripe > > block are read together (drives are spindle locked) there is not > > additional read needed to calculate parity before writing. > > This is not necessarily true. Raid4 (block level striping, dedicated > parity) may require read operations before writing, if writes are done in > non-stripe aligned accesses. TRUTH. That is why for all striped RAID arrays, RAID0, 3, 4, 5, & 10; it is very important to size the stripe block to match how the application will use it. Since for general filesystem use this is nearly impossible, except for filesystems dedicated to one or a few applications, RAID5 will always be best overall for filesystem performance. But for databases, which is essentially a single application use with predetermined I/O behavior (Informix ALWAYS writes and reads either 2K or 8K per I/O and it's readahead feature will gang I/Os together to pre-fetch data but the number of pages fetched is user configurable so this is controllable also) one can adjust the stripe block size for optimal I/O so that multiple block I/O does not happen. It is also why I always recommend much smaller RAID stripe block sizes for ANY RAID level for Informix use, 16K seems best, than is normally recommended for filesystems (64K or 128K). > Since > > parity and all data blocks are read concurrently for ALL reads RAID3&4 > > systems can, and most do, check parity on read which can trap and > > correct most partial media failure problems at read time which is when > > they are still correctable. > > Because of this behavior, Raid3 & 4 will always be able to process LESS > operations per second than a Raid 5. Except during sequential read operations and write operations. Because of the nature of Informix's buffer cache, read ahead, "lite scans", and LRU flush operations most Inform