RE: General RAID5 question - dangers?
Posted in 2000
--=_7F270501.482945D3 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: quoted-printable Content-Disposition: inline ding... (sound of light bulb coming on, thanks Mr Rice)... ah huh!!!! can = i then make the assumption that if data on disk 1 is updated the parity = that is rewritten will just not be an update of the changed data, but an = entire rewrite of data stripe from disk 1 and data stripe from disk 2 to = recreate the parity? is this always the case that there is no selective = parity updates but just entire parity rewrites when data on a disk is = updated? TIA =20 -kudzi >>> William Rice <ricew@operamail.com> 06/19/00 07:50AM >>> Even though this is addressed to Art I figured I would try to answer. RAID 5, say you have three disks. 1. A sector goes bad on disk2 . 2. You read read data from disk 1,=20 3. You write data to disk 1 4. The parity of disk 1 and disk 2 is taken to write to the parity sector = on=20 disk 3. =20 5. An hour later disk 2 blows up.=20 6. You have no recovery. Because the read from disk 2 when it calculated parity was invalid. It also never went through Informix to be = flagged as=20 bad. Hope this helps, Will >=3D=3D=3D=3D=3D Original Message From "Kudzi Muchaka" <KMuchaka@HAFresno.o= rg> =3D=3D=3D=3D=3D >Art, > >i'm missing something here... must be a result of the many nights out on = the=20 town. but my question Art is this... if in Raid 10 we are relying quite a = bit=20 on IDS detecting bad data on a drive, why are'nt we transferring the = same=20 dependence that IDS will detect bad data onto Raid 5. Bad data is = detected=20 when a drive is accessed... and the probability of accessing a bad drive = would=20 be random in both cases... right? In other words detection of bad data by = IDS=20 would be a game of luck in either case. i think i'm >missing something. > >-kudzi > >>> "Art S. Kagel" <kagel@bloomberg.net> 06/13/00 09:36AM >>> > > >William Rice wrote: >> >> I will admit to having some fear of RAID 5 but for different reasons. = I=20 like >> being >> able to lose 3 or 4 disks and still have at least a chance of not = having to >> restore my data. This is mainly because I am a lazy bum who does >> not like the idea of doing work which I do not have to. I am interested= >> in the theory that due to the fact that parity is calculated it is more >> susceptible >> to error. Would you mind adding a little more meat to that argument, = or >> is it based on more complex=3Dmore problems? Or maybe he fact that it >> reads all the other disks to get the information to change the parity = so >> there is more oppurtunity for a problem? > >OK, I will deal with why RAID10 (and by extention RAID1) is not as much = at >risk. Remember that the problem is not the partial media failure = trashing >the data on the one drive that is in the process of progressively failing >over time but the problem that good data will be read from another stripe >member (call it drive #1) for the stripe block containing damaged data on >the failing drive (call it drive #2) and modified causing a read of the >remaining stripe members and the calculation of new parity after which = the >modified block will be written back to drive #1 and the parity back to = that >stripe's parity drive (perhaps drive#3). Now since the data from drive = #2 >was not needed by the application (here IDS) the fact that it was damaged >was not detected and so the new parity was calculated using garbage >resulting in a parity that can ONLY be used accurately to recreate the >trashed block on drive #2. Now suppose that drive #4 suffers from a >catastrophic failure and has to be replaced. The damaged drive has >continued to fail and now returns a different pattern of bits than was = used >to calculate the trashed parity block on drive #3. Now when the missing >block data on the drive #4 replacement is calculated it too will become >garbage and two disk blocks are now unusable, the damage has propagated. >Now that dealt with what happens if the bad block was NOT read directly = and >detected by IDS. If it was IDS will mark the chunk OFFLINE and refuse to >use it until you repair the damage. The only way you can do that is if >you restore from backup or remove the partially damaged drive and try to >rebuild it from the parity as if it had completely failed. HOWEVER, if = all, >or even several, of the drives in that array are from the same manufacturi= ng >lot or are even of similar age, there is a good chance that the previous >problem has already trashed the parity of other blocks so that you will >possibly be reconstructing a new drive that will have more bad data = blocks >than the one it replaces. > >With RAID10, each drive in each mirrored pair is written independently. = If >a block on drive #1a is trashed the data on drive #1b (its mirror) is = fine. >If the bad data is read from the drive that is failing (say 1a) the = engine >will recognize it and mark the chunk down. All you have to do is remove >drive 1a and mark the chunk back online rebuilding the mirror online. No >problems and less chance that there are other damaged blocks on the one >remaining mirror than on any of the 5 or more drives in a RAID5 stripe. = If >the data is NOT read from 1a but from 1b and modified it will be = rewritten >to BOTH drives improving the chances that it will be correctly readable = if >read from the failing 1a next time just because the flux changes will = have >been renewed, if the platter is too far gone we are just back to the >possibility that the bad block will be read and flagged by IDS later. In >no case can the data on 2a/2b, 3a/3b, 4a/4b, etc be damaged. Yes, if we >were talking about ANY old data file on RAID10 the damage might propagate >but since IDS has its own methods for detecting bad data reads this >probability is vanishingly small (the damage to the block would have to = NOT >alter the contents of the first 28 bytes or the last 4 to 1020 bytes thus >not damaging the page header, page trailer, or the slot table to not be >detected by IDS). > >> I have never been priveleged enough to witness a site using RAID 5 have >> serious issues, and am trying to understand where mirroring would help >> in such a situation > >Lucky man! I have and let me assure you it is no privilege! > >Art S. Kagel ------------------------------------------------------------ This e-mail has been sent to you courtesy of OperaMail, as a free = service from Opera Software, makers of the award-winning Web Browser, Opera. Visit = us at http://www.opera.com/ or our portal at: http://www.myopera.com/ Your free = e-mail=20 account is waiting at: http://www.operamail.com/ ------------------------------------------------------------ --=_7F270501.482945D3 Content-Type: TEXT/HTML Content-Transfer-Encoding: base64 Content-Disposition: attachment; filename="TEXT.htm" PCFET0NUWVBFIEhUTUwgUFVCTElDICItLy9XM0MvL0RURCBIVE1MIDQuMCBUcmFuc2l0aW9uYWwv L0VOIj4NCjxIVE1MPjxIRUFEPg0KPE1FVEEgY29udG