Re: General RAID5 question - dangers?
Posted in 2000
Topics: Storage & Space Management
--=_095177BE.36573BAC Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: quoted-printable Content-Disposition: inline Art, i'm missing something here... must be a result of the many nights out on = the town. but my question Art is this... if in Raid 10 we are relying = quite a bit on IDS detecting bad data on a drive, why are'nt we transferrin= g the same dependence that IDS will detect bad data onto Raid 5. Bad data = is detected when a drive is accessed... and the probability of accessing a = bad drive would be random in both cases... right? In other words detection = of bad data by IDS would be a game of luck in either case. i think i'm = missing something. -kudzi >>> "Art S. Kagel" <kagel@bloomberg.net> 06/13/00 09:36AM >>> William Rice wrote: >=20 > I will admit to having some fear of RAID 5 but for different reasons. I = like > being > able to lose 3 or 4 disks and still have at least a chance of not having = to > restore my data. This is mainly because I am a lazy bum who does > not like the idea of doing work which I do not have to. I am interested > in the theory that due to the fact that parity is calculated it is more > susceptible > to error. Would you mind adding a little more meat to that argument, or > is it based on more complex=3Dmore problems? Or maybe he fact that it > reads all the other disks to get the information to change the parity so > there is more oppurtunity for a problem? OK, I will deal with why RAID10 (and by extention RAID1) is not as much = at=20 risk. Remember that the problem is not the partial media failure = trashing=20 the data on the one drive that is in the process of progressively = failing=20 over time but the problem that good data will be read from another = stripe=20 member (call it drive #1) for the stripe block containing damaged data = on=20 the failing drive (call it drive #2) and modified causing a read of the=20 remaining stripe members and the calculation of new parity after which = the=20 modified block will be written back to drive #1 and the parity back to = that=20 stripe's parity drive (perhaps drive#3). Now since the data from drive = #2=20 was not needed by the application (here IDS) the fact that it was = damaged=20 was not detected and so the new parity was calculated using garbage=20 resulting in a parity that can ONLY be used accurately to recreate the=20 trashed block on drive #2. Now suppose that drive #4 suffers from a=20 catastrophic failure and has to be replaced. The damaged drive has=20 continued to fail and now returns a different pattern of bits than was = used=20 to calculate the trashed parity block on drive #3. Now when the missing=20= block data on the drive #4 replacement is calculated it too will become=20 garbage and two disk blocks are now unusable, the damage has propagated. = =20 Now that dealt with what happens if the bad block was NOT read directly = and=20 detected by IDS. If it was IDS will mark the chunk OFFLINE and refuse = to=20 use it until you repair the damage. The only way you can do that is if=20 you restore from backup or remove the partially damaged drive and try = to=20 rebuild it from the parity as if it had completely failed. HOWEVER, if = all,=20 or even several, of the drives in that array are from the same manufacturin= g=20 lot or are even of similar age, there is a good chance that the previous=20= problem has already trashed the parity of other blocks so that you will=20 possibly be reconstructing a new drive that will have more bad data = blocks=20 than the one it replaces. With RAID10, each drive in each mirrored pair is written independently. = If=20 a block on drive #1a is trashed the data on drive #1b (its mirror) is = fine.=20 If the bad data is read from the drive that is failing (say 1a) the = engine=20 will recognize it and mark the chunk down. All you have to do is = remove=20 drive 1a and mark the chunk back online rebuilding the mirror online. = No=20 problems and less chance that there are other damaged blocks on the one=20 remaining mirror than on any of the 5 or more drives in a RAID5 stripe. = If=20 the data is NOT read from 1a but from 1b and modified it will be = rewritten=20 to BOTH drives improving the chances that it will be correctly readable = if=20 read from the failing 1a next time just because the flux changes will = have=20 been renewed, if the platter is too far gone we are just back to the=20 possibility that the bad block will be read and flagged by IDS later. = In=20 no case can the data on 2a/2b, 3a/3b, 4a/4b, etc be damaged. Yes, if = we=20 were talking about ANY old data file on RAID10 the damage might = propagate=20 but since IDS has its own methods for detecting bad data reads this=20 probability is vanishingly small (the damage to the block would have to = NOT=20 alter the contents of the first 28 bytes or the last 4 to 1020 bytes = thus=20 not damaging the page header, page trailer, or the slot table to not be=20 detected by IDS). > I have never been priveleged enough to witness a site using RAID 5 have > serious issues, and am trying to understand where mirroring would help > in such a situation Lucky man! I have and let me assure you it is no privilege! Art S. Kagel --=_095177BE.36573BAC Content-Type: text/plain Content-Disposition: attachment; filename="TEXT.htm" <!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN"> <HTML><HEAD> <META content="text/html; charset=iso-8859-1" http-equiv=Content-Type> <META content="MSHTML 5.00.2919.6307" name=GENERATOR></HEAD> <BODY bgColor=#ffffff style="FONT: 10pt Arial; MARGIN-LEFT: 2px; MARGIN-TOP: 2px"> <DIV>Art,</DIV> <DIV><BR>i'm missing something here... must be a result of the many nights out on the town. but my question Art is this... if in Raid 10 we are relying quite a bit on IDS detecting bad data on a drive, why are'nt we transferring the <STRONG>same dependence </STRONG>that IDS will detect bad data onto Raid 5. Bad data is detected when a drive is accessed... and the probability of accessing a bad drive would be random in both cases... right? In other words detection of bad data by IDS would be a game of luck in either case. i think i'm missing something.</DIV> <DIV> </DIV> <DIV>-kudzi</DIV> <DIV><BR>>>> "Art S. Kagel" <kagel@bloomberg.net> 06/13/00 09:36AM >>><BR><BR><BR>William Rice wrote:<BR>> <BR>> I will admit to having some fear of RAID 5 but for different reasons. I like<BR>> being<BR>> able to lose 3 or 4 disks and still have at least a chance of not having to<BR>> restore my data. This is mainly because I am a lazy bum who does<BR>> not like the idea of doing work which I do not have to. I am interested<BR>> in the theory that due to the fact that parity is calculated it is more<BR>> susceptible<BR>> to error. Would you mind adding a little more meat to that argument, or<BR>> is it based on more complex=more problems? Or maybe he fact that it<BR>> reads all the other d
Kudzi Muchaka wrote: > > Art, > > i'm missing something here... must be a result of the many nights out on = > the town. but my question Art is this... if in Raid 10 we are relying = > quite a bit on IDS detecting bad data on a drive, why are'nt we transferrin= > g the same dependence that IDS will detect bad data onto Raid 5. Bad data = > is detected when a drive is accessed... and the probability of accessing a = > bad drive would be random in both cases... right? In other words detection = > of bad data by IDS would be a game of luck in either case. i think i'm = > missing something. True, however, an error on a RAID10 drive will not be propagated to its mirror, since the drives are written separately, and may or may not eventually be detected before the drive fails completely. Only if the mirror fails first will ANY corrupt data be propagated to the mirror and then ONLY the partially damaged sector(s). In RAID5, because bad parity is calculated from garbage that may NEVER be read again, and therefore never caught by the engine (in either RAID5 or RAID10 but cause no problems in RAID10) if ANY of the other drives should fail (and in a five drive array there is a 4X higher probability that another drive will fail than that the damaged drive's mirror will fail in RAID10, the reconstructed drive's block that is a member of the same stripe block as the damaged drive's data block will now ALSO be garbage and you now have two bad sectors with only one bad drive! In RAID5 if the engine detects bad data read from the array which is caused by partial media failure, and that stripe block has been written since the damage occured, there is no way to recover the damaged data since the parity is poultry poop. In RAID10 if a bad block is read from a partially failed drive you can recover the ENTIRE drive completely from its mirror WITHOUT data loss! THAT is the difference! I thought I explained this on Thursday? Art S. Kagel > -kudzi > > >>> "Art S. Kagel" <kagel@bloomberg.net> 06/13/00 09:36AM >>> > > William Rice wrote: > >=20 > > I will admit to having some fear of RAID 5 but for different reasons. I = > like > > being > > able to lose 3 or 4 disks and still have at least a chance of not having = > to > > restore my data. This is mainly because I am a lazy bum who does > > not like the idea of doing work which I do not have to. I am interested > > in the theory that due to the fact that parity is calculated it is more > > susceptible > > to error. Would you mind adding a little more meat to that argument, or > > is it based on more complex=3Dmore problems? Or maybe he fact that it > > reads all the other disks to get the information to change the parity so > > there is more oppurtunity for a problem? > > OK, I will deal with why RAID10 (and by extention RAID1) is not as much = > at=20 > risk. Remember that the problem is not the partial media failure = > trashing=20 > the data on the one drive that is in the process of progressively = > failing=20 > over time but the problem that good data will be read from another = > stripe=20 > member (call it drive #1) for the stripe block containing damaged data = > on=20 > the failing drive (call it drive #2) and modified causing a read of the=20 > remaining stripe members and the calculation of new parity after which = > the=20 > modified block will be written back to drive #1 and the parity back to = > that=20 > stripe's parity drive (perhaps drive#3). Now since the data from drive = > #2=20 > was not needed by the application (here IDS) the fact that it was = > damaged=20 > was not detected and so the new parity was calculated using garbage=20 > resulting in a parity that can ONLY be used accurately to recreate the=20 > trashed block on drive #2. Now suppose that drive #4 suffers from a=20 > catastrophic failure and has to be replaced. The damaged drive has=20 > continued to fail and now returns a different pattern of bits than was = > used=20 > to calculate the trashed parity block on drive #3. Now when the missing=20= > > block data on the drive #4 replacement is calculated it too will become=20 > garbage and two disk blocks are now unusable, the damage has propagated. = > =20 > Now that dealt with what happens if the bad block was NOT read directly = > and=20 > detected by IDS. If it was IDS will mark the chunk OFFLINE and refuse = > to=20 > use it until you repair the damage. The only way you can do that is if=20 > you restore from backup or remove the partially damaged drive and try = > to=20 > rebuild it from the parity as if it had completely failed. HOWEVER, if = > all,=20 > or even several, of the drives in that array are from the same manufacturin= > g=20 > lot or are even of similar age, there is a good chance that the previous=20= > > problem has already trashed the parity of other blocks so that you will=20 > possibly be reconstructing a new drive that will have more bad data = > blocks=20 > than the one it replaces. > > With RAID10, each drive in each mirrored pair is written independently. = > If=20 > a block on drive #1a is trashed the data on drive #1b (its mirror) is = > fine.=20 > If the bad data is read from the drive that is failing (say 1a) the = > engine=20 > will recognize it and mark the chunk down. All you have to do is = > remove=20 > drive 1a and mark the chunk back online rebuilding the mirror online. = > No=20 > problems and less chance that there are other damaged blocks on the one=20 > remaining mirror than on any of the 5 or more drives in a RAID5 stripe. = > If=20 > the data is NOT read from 1a but from 1b and modified it will be = > rewritten=20 > to BOTH drives improving the chances that it will be correctly readable = > if=20 > read from the failing 1a next time just because the flux changes will = > have=20 > been renewed, if the platter is too far gone we are just back to the=20 > possibility that the bad block will be read and flagged by IDS later. = > In=20 > no case can the data on 2a/2b, 3a/3b, 4a/4b, etc be damaged. Yes, if = > we=20 > were talking about ANY old data file on RAID10 the damage might = > propagate=20 > but since IDS has its own methods for detecting bad data reads this=20 > probability is vanishingly small (the damage to the block would have to = > NOT=20 > alter the contents of the first 28 bytes or the last 4 to 1020 bytes = > thus=20 > not damaging the page header, page trailer, or the slot table to not be=20 > detected by IDS). > > > I have never been priveleged enough to witness a site using RAID 5 have > > serious issues, and am trying to understand where mirroring would help > > in such a situation > > Lucky man! I have and let me assure you it is no privilege! > > Art S. Kagel > > --=_095177BE.36573BAC > Content-Type: text/plain > Content-Disposition: attachment; filename="TEXT.htm" > > <!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN"> > <HTML><HEAD> > <META content="text/html; charset=iso-8859-1" http-equiv=Content-Type> > <META content="MSHTML 5.00.2919.6307" name=GENERATOR></HEAD> > <BODY
Related threads
- Conversion to differeent characters sets
- Problem in changing locale via dbexport/dbimport
- RE: openlink error "Unable to load locale categories"