RE: General RAID5 question - dangers?
Posted in 2000
Topics: Performance & Tuning
I will admit to having some fear of RAID 5 but for different reasons. I like being able to lose 3 or 4 disks and still have at least a chance of not having to restore my data. This is mainly because I am a lazy bum who does not like the idea of doing work which I do not have to. I am interested in the theory that due to the fact that parity is calculated it is more susceptible to error. Would you mind adding a little more meat to that argument, or is it based on more complex=more problems? Or maybe he fact that it reads all the other disks to get the information to change the parity so there is more oppurtunity for a problem? I have never been priveleged enough to witness a site using RAID 5 have serious issues, and am trying to understand where mirroring would help in such a situation Thanks, Will >===== Original Message From "Obnoxio The Clown" <obnoxio@hotmail.com> ===== >From: psyklops@my-deja.com >> >>Wouldn't parity verification on RAID-5 seem redundant since >>the "reverse" of the parity should be the actual user data written. If >>the parity is incorrect, then the data would also be incorrect and if >>so, would be of greater concern than the parity data itself? > >But the parity gets "calculated", and is therefore more susceptible to >faults. And the other issue is not RAID parity, but the low-level disk >parity that can fail all by itself... :-) > >>RAID-1 whether hardware or software RAID needs to verify that data is >>identical on both volumes (hence mirrored). If the data is not >>identical on both volumes and a disk fails, data integrity is >>compromised and implementing a RAID-1 solution would not be useful. >> >>As for the XP256, the array is cache centric. Therefore all reads and >>writes hit the cache first and then serviced to the host/disk. As RAID- >>5 is much slower in writes, it is generally not preferred in database >>environments where a large number of writes is done. However, the >>XP256 will read the entire RAID-5 stripe to cache memory regardless of >>the size of the read/write. All read's and writes are then serviced >>from cache memory until the array decides to de-stage the stripe to >>disk. The XP256 can accomodate RAID-1 and RAID-5 within the same array >>for environments that require it. For maximum tuning, certain heavily >>hit volumes "hot spots" can be permanently locked in cache using a >>special software package for the XP256 - all writes/reads to this area >>(possibly redo logs) can be serviced directly from cache memory... > >So it has a bit more than 64MB cache? >________________________________________________________________________ >Get Your Private, Free E-mail from MSN Hotmail at http://www.hotmail.com ------------------------------------------------------------ This e-mail has been sent to you courtesy of OperaMail, as a free service from Opera Software, makers of the award-winning Web Browser, Opera. Visit us at http://www.opera.com/ or our portal at: http://www.myopera.com/ Your free e-mail account is waiting at: http://www.operamail.com/ ------------------------------------------------------------
William Rice wrote: > > I will admit to having some fear of RAID 5 but for different reasons. I like > being > able to lose 3 or 4 disks and still have at least a chance of not having to > restore my data. This is mainly because I am a lazy bum who does > not like the idea of doing work which I do not have to. I am interested > in the theory that due to the fact that parity is calculated it is more > susceptible > to error. Would you mind adding a little more meat to that argument, or > is it based on more complex=more problems? Or maybe he fact that it > reads all the other disks to get the information to change the parity so > there is more oppurtunity for a problem? OK, I will deal with why RAID10 (and by extention RAID1) is not as much at risk. Remember that the problem is not the partial media failure trashing the data on the one drive that is in the process of progressively failing over time but the problem that good data will be read from another stripe member (call it drive #1) for the stripe block containing damaged data on the failing drive (call it drive #2) and modified causing a read of the remaining stripe members and the calculation of new parity after which the modified block will be written back to drive #1 and the parity back to that stripe's parity drive (perhaps drive#3). Now since the data from drive #2 was not needed by the application (here IDS) the fact that it was damaged was not detected and so the new parity was calculated using garbage resulting in a parity that can ONLY be used accurately to recreate the trashed block on drive #2. Now suppose that drive #4 suffers from a catastrophic failure and has to be replaced. The damaged drive has continued to fail and now returns a different pattern of bits than was used to calculate the trashed parity block on drive #3. Now when the missing block data on the drive #4 replacement is calculated it too will become garbage and two disk blocks are now unusable, the damage has propagated. Now that dealt with what happens if the bad block was NOT read directly and detected by IDS. If it was IDS will mark the chunk OFFLINE and refuse to use it until you repair the damage. The only way you can do that is if you restore from backup or remove the partially damaged drive and try to rebuild it from the parity as if it had completely failed. HOWEVER, if all, or even several, of the drives in that array are from the same manufacturing lot or are even of similar age, there is a good chance that the previous problem has already trashed the parity of other blocks so that you will possibly be reconstructing a new drive that will have more bad data blocks than the one it replaces. With RAID10, each drive in each mirrored pair is written independently. If a block on drive #1a is trashed the data on drive #1b (its mirror) is fine. If the bad data is read from the drive that is failing (say 1a) the engine will recognize it and mark the chunk down. All you have to do is remove drive 1a and mark the chunk back online rebuilding the mirror online. No problems and less chance that there are other damaged blocks on the one remaining mirror than on any of the 5 or more drives in a RAID5 stripe. If the data is NOT read from 1a but from 1b and modified it will be rewritten to BOTH drives improving the chances that it will be correctly readable if read from the failing 1a next time just because the flux changes will have been renewed, if the platter is too far gone we are just back to the possibility that the bad block will be read and flagged by IDS later. In no case can the data on 2a/2b, 3a/3b, 4a/4b, etc be damaged. Yes, if we were talking about ANY old data file on RAID10 the damage might propagate but since IDS has its own methods for detecting bad data reads this probability is vanishingly small (the damage to the block would have to NOT alter the contents of the first 28 bytes or the last 4 to 1020 bytes thus not damaging the page header, page trailer, or the slot table to not be detected by IDS). > I have never been priveleged enough to witness a site using RAID 5 have > serious issues, and am trying to understand where mirroring would help > in such a situation Lucky man! I have and let me assure you it is no privilege! Art S. Kagel
You are correct, RAID-5 does need calculation for parity and therefore with this extra "step" is therefore susceptible to error. Additionaly, the XP256 is capable of delivering both a RAID 0/1 or RAID 5 solution in the single array wihtout comprimising user data. In terms of performance, the XP256's RAID-5 (although not as quick as it's RAID-1) can still provide the throughput that most DBA's need. The XP256 can accomodate up to 16GB of cache In article <8i30k5$beo$1@news.xmission.com>, William Rice <ricew@operamail.com> wrote: > >I will admit to having some fear of RAID 5 but for different reasons. I like >being >able to lose 3 or 4 disks and still have at least a chance of not having to >restore my data. This is mainly because I am a lazy bum who does >not like the idea of doing work which I do not have to. I am interested >in the theory that due to the fact that parity is calculated it is more >susceptible >to error. Would you mind adding a little more meat to that argument, or >is it based on more complex=more problems? Or maybe he fact that it >reads all the other disks to get the information to change the parity so >there is more oppurtunity for a problem? > >I have never been priveleged enough to witness a site using RAID 5 have >serious issues, and am trying to understand where mirroring would help >in such a situation > >Thanks, >Will > >>===== Original Message From "Obnoxio The Clown" <obnoxio@hotmail.com> ===== >>From: psyklops@my-deja.com >>> >>>Wouldn't parity verification on RAID-5 seem redundant since >>>the "reverse" of the parity should be the actual user data written. If >>>the parity is incorrect, then the data would also be incorrect and if >>>so, would be of greater concern than the parity data itself? >> >>But the parity gets "calculated", and is therefore more susceptible to >>faults. And the other issue is not RAID parity, but the low- level disk >>parity that can fail all by itself... :-) >> >>>RAID-1 whether hardware or software RAID needs to verify that data is >>>identical on both volumes (hence mirrored). If the data is not >>>identical on both volumes and a disk fails, data integrity is >>>compromised and implementing a RAID-1 solution would not be useful. >>> >>>As for the XP256, the array is cache centric. Therefore all reads and >>>writes hit the cache first and then serviced to the host/disk. As RAID- >>>5 is much slower in writes, it is generally not preferred in database >>>environments where a large number of writes is done. However, the >>>XP256 will read the entire RAID-5 stripe to cache memory regardless of >>>the size of the read/write. All read's and writes are then serviced >>>from cache memory until the array decides to de-stage the stripe to >>>disk. The XP256 can accomodate RAID-1 and RAID-5 within the same array >>>for environments that require it. For maximum tuning, certain heavily >>>hit volumes "hot spots" can be permanently locked in cache using a >>>special software package for the XP256 - all writes/reads to this area >>>(possibly redo logs) can be serviced directly from cache memory... >> >>So it has a bit more than 64MB cache? >>_______________________________________________________________ _________ >>Get Your Private, Free E-mail from MSN Hotmail at http://www.hotmail.com > >------------------------------------------------------------ >This e-mail has been sent to you courtesy of OperaMail, as a free service from >Opera Software, makers of the award-winning Web Browser, Opera. Visit us at >http://www.opera.com/ or our portal at: http://www.myopera.com/ Your free e-mail >account is waiting at: http://www.operamail.com/ >------------------------------------------------------------ > > > * Sent from RemarQ http://www.remarq.com The Internet's Discussion Network * The fastest and easiest way to search and participate in Usenet - Free!
You are correct, RAID-5 does need calculation for parity and therefore with this extra "step" is therefore susceptible to error. Additionaly, the XP256 is capable of delivering both a RAID 0/1 or RAID 5 solution in the single array wihtout comprimising user data. In terms of performance, the XP256's RAID-5 (although not as quick as it's RAID-1) can still provide the throughput that most DBA's need. The XP256 can accomodate up to 16GB of cache In article <8i30k5$beo$1@news.xmission.com>, William Rice <ricew@operamail.com> wrote: > >I will admit to having some fear of RAID 5 but for different reasons. I like >being >able to lose 3 or 4 disks and still have at least a chance of not having to >restore my data. This is mainly because I am a lazy bum who does >not like the idea of doing work which I do not have to. I am interested >in the theory that due to the fact that parity is calculated it is more >susceptible >to error. Would you mind adding a little more meat to that argument, or >is it based on more complex=more problems? Or maybe he fact that it >reads all the other disks to get the information to change the parity so >there is more oppurtunity for a problem? > >I have never been priveleged enough to witness a site using RAID 5 have >serious issues, and am trying to understand where mirroring would help >in such a situation > >Thanks, >Will > >>===== Original Message From "Obnoxio The Clown" <obnoxio@hotmail.com> ===== >>From: psyklops@my-deja.com >>> >>>Wouldn't parity verification on RAID-5 seem redundant since >>>the "reverse" of the parity should be the actual user data written. If >>>the parity is incorrect, then the data would also be incorrect and if >>>so, would be of greater concern than the parity data itself? >> >>But the parity gets "calculated", and is therefore more susceptible to >>faults. And the other issue is not RAID parity, but the low- level disk >>parity that can fail all by itself... :-) >> >>>RAID-1 whether hardware or software RAID needs to verify that data is >>>identical on both volumes (hence mirrored). If the data is not >>>identical on both volumes and a disk fails, data integrity is >>>compromised and implementing a RAID-1 solution would not be useful. >>> >>>As for the XP256, the array is cache centric. Therefore all reads and >>>writes hit the cache first and then serviced to the host/disk. As RAID- >>>5 is much slower in writes, it is generally not preferred in database >>>environments where a large number of writes is done. However, the >>>XP256 will read the entire RAID-5 stripe to cache memory regardless of >>>the size of the read/write. All read's and writes are then serviced >>>from cache memory until the array decides to de-stage the stripe to >>>disk. The XP256 can accomodate RAID-1 and RAID-5 within the same array >>>for environments that require it. For maximum tuning, certain heavily >>>hit volumes "hot spots" can be permanently locked in cache using a >>>special software package for the XP256 - all writes/reads to this area >>>(possibly redo logs) can be serviced directly from cache memory... >> >>So it has a bit more than 64MB cache? >>_______________________________________________________________ _________ >>Get Your Private, Free E-mail from MSN Hotmail at http://www.hotmail.com > >------------------------------------------------------------ >This e-mail has been sent to you courtesy of OperaMail, as a free service from >Opera Software, makers of the award-winning Web Browser, Opera. Visit us at >http://www.opera.com/ or our portal at: http://www.myopera.com/ Your free e-mail >account is waiting at: http://www.operamail.com/ >------------------------------------------------------------ > > > * Sent from RemarQ http://www.remarq.com The Internet's Discussion Network * The fastest and easiest way to search and participate in Usenet - Free!