RE: Cook device vs Raw device
Posted in 1999
Topics: Performance & Tuning, Storage & Space Management, Platform-Specific Issues
This is a response to the messages I have read so far. One of the loads took around 15 minutes. During which time we were completely i/o bound. I will admit, the disk was over 96% usage the whole time. So maybe we could not see the difference in performance due to the fact the disk was pretty much pegged either way. We were going to the raw character device. We do have a good cache rate on our machines. It is difficult for me to justify the time and resources my company will have to use to go to raw disk space if I can not show how making this change will benefit the company. Previously when on raw spaces this company had issues with informix not handling remappable i/o errors as well as the AIX operating system does. It caused the chunk to go offline, and the options offered by informix I believe were let them dial in to fix it or do a restore. When running on filesystems they have not experienced such issues.(I was not here at the point in time this occurred so do not have the specifics of what occurred) Given what I have seen so far, I understand and agree with my current employers recluctance to go to raw disk. This is not a testimonial against raw disk space. It is a statement that for the system we run I have not seen a case which can be build for going to raw disk. I opened this question up because somewhere deep in my gut I feel raw disk has to be a good bit faster because you are bypassing the operating systems buffering(This is probably due to reading lots of Art Kagel postings) But my gut is by no means justification. Just that from what they have seen, for their situation raw disk space does not seem to have much of a performance benefit, and has in the past caused parts of the business to be offline. Will ------------------------------------------------------------ This e-mail has been sent to you courtesy of OperaMail, a free web-based service from Opera Software, makers of the award-winning Web Browser - http://www.operasoftware.com ------------------------------------------------------------
William Rice wrote: [SNIP] > > Previously when on raw spaces this company had issues > with informix not handling remappable i/o errors as well as the AIX operating > system does. It caused the chunk to go offline, and the options offered by > informix I believe were let them dial in to fix it or do a restore. When > running > on filesystems they have not experienced such issues.(I was not here at > the point in time this occurred so do not have the specifics of what occurred) We have the same problem with DG Clariion. DG says that they do not report remappable disk and controller errors and that the data is replaced from mirror or the controller is swapped without reporting errors to the application (ie Informix) unless the application registers for extended error handling and makes certain DGUX and Clariion specific ioctl() calls. Informix insists that all the code does is open()-read/write-read/write repeat seven times when an errno is detected and report the simple errno received. Chunks go down when controllers or power supplies fail or when a RAID5 correction occurs due to a disk failure or even a rebuild begin event. I have taken to marking the chunks and dbspaces up myself (makes tech support crazy) because I cannot bear to wait through a two hour support call to Singapore or Sidney on Sunday night. (No I cannot share my chunk and dbspace cleanup utilities, I promised, sorry). And DG just keeps saying, "It does not happen!" and Informix keeps saying, "It's not OUR problem!" Going from RAID5 to RAID10 at least relieved the more common RAID events (we do not get errors from the RAID10 module if a drive goes down). > Given what I have seen so far, I understand and agree with my current > employers recluctance to go to raw disk. This is not a testimonial against > raw disk space. It is a statement that for the system we run I have not seen > a case which can be build for going to raw disk. I opened this question up > because somewhere deep in my gut I feel raw disk has to be a good bit faster > because you are bypassing the operating systems buffering(This is probably > due to reading lots of Art Kagel postings) But my gut is by no means > justification. Just that from what they have seen, for their situation > raw disk space does not seem to have much of a performance benefit, and > has in the past caused parts of the business to be offline. Only testing yourself is going to produce the justification that you, correctly, seek. I post based on my own tests supported by the reports of others, but I'm just another DBA/programmer. When your firm's business is on the line "not invented here" is not a bad policy! Art S. Kagel
On Fri, 17 Sep 1999 11:13:30 -0400, "Art S. Kagel" <kagel@bloomberg.net> wrote: >William Rice wrote: >[SNIP] > >We have the same problem with DG Clariion. DG says that they do not >report remappable disk and controller errors and that the data is replaced >from mirror or the controller is swapped without reporting errors to the >repeat seven times when an errno is detected and report the simple errno >received. Chunks go down when controllers or power supplies fail or when >a RAID5 correction occurs due to a disk failure or even a rebuild begin >event. I have taken to marking the chunks and dbspaces up myself (makes >tech support crazy) because I cannot bear to wait through a two hour >Art S. Kagel Hi Art, for us -newbies-, can you explain more about raw and raid systems. If a bad spot is detected at the controller level and remapped, will that upset the Informix engine (7.3)? Also if a drive is replaced and then rebuilt from a raid 5, is Informix going to have problems with that also? Thanks - Gary Quiring
Gary Quiring wrote: > > On Fri, 17 Sep 1999 11:13:30 -0400, "Art S. Kagel" <kagel@bloomberg.net> > wrote: > > >William Rice wrote: > >[SNIP] > > > >We have the same problem with DG Clariion. DG says that they do not > >report remappable disk and controller errors and that the data is replaced > >from mirror or the controller is swapped without reporting errors to the > > >repeat seven times when an errno is detected and report the simple errno > >received. Chunks go down when controllers or power supplies fail or when > >a RAID5 correction occurs due to a disk failure or even a rebuild begin > >event. I have taken to marking the chunks and dbspaces up myself (makes > >tech support crazy) because I cannot bear to wait through a two hour > > >Art S. Kagel > > Hi Art, for us -newbies-, can you explain more about raw and raid systems. > If a bad spot is detected at the controller level and remapped, will that > upset the Informix engine (7.3)? Also if a drive is replaced and then > rebuilt from a raid 5, is Informix going to have problems with that also? It's going to depend on your RAID sub-system and how well/poorly it reports/hides such errors. Note that my problem experience is ONLY with DG/UX and DG Clariion RAID Arrays. I cannot report on other systems. I guess it's been a while since I ranted against RAID5 so here goes: First note that RAID5 controllers CANNOT detect ANY bad spots on disk drives unless additional diagnostics are built in beyond the requirements of RAID5. DG Clariion tries to do this by adding a 8 byte checksum to each disk block but I have seen DG fail to detect errors also. Remember RAID5 checksums are NOT checked on read or write they are only used to reconstruct a single completely defunct drive in the array and CANNOT detect or correct what you describe which is called "Partial Media Failure". Normally the drives themselves detect and correct these types of errors and SCSI drives go so far as to remap sectors that return errors repeatedly to sectors which have been held back from the declared drive size for that purpose. However, in a drive that has been in use for more than say two years it is likely that the sector remapping list has been or is nearly exhausted and SCSI devices are not required to, and do not, report when sectors have been remapped nor when a non-remapped sector's data was reconstructed from its ECC codes. So the first time you know that the remapping table is empty is when another sector has gone bad and remains un-remapped until it degrades beyond the 1 or 2 bit error correction capability of the drive's ECC coding. By then your RAID array will be using that BAD data to calculate a checksum block which will be useful only to recreate garbage. If the slowly failing drive croaks you will have lost only the one sector which is already garbage. No worse than a singleton drive. However, if another drive should fail the faulty checksum will now be used to reconstruct and write garbage to a second drive! This is the main reason I refuse to use RAID5 EVER! The write performance problem of RAID5 is well up there but not paramount. Good performance can keep me from getting stuck in the office late but preventing data loss keeps me from losing sleep and my job! So back to your question. There is no RAID controller remapping. What I have experienced is that Informix gets errno 5 or 6 when a controller fails and the RAID sub-system switches to the backup controller, when a drive fully fails and its data is being reconstructed, or when a power supply fails (though the cabinet, controllers, and drives can continue to operate with two out of three power supplies online). This dispite assurances by DG that it does not EVER return an errno under these circumstances so long as the backup systems are functional. You judge for yourself. OH and yes, I did once even get an error when a RAID5 drive rebuild completed successfully. At least that was the best explanation anyone could come up with. Informix got an errno on an I/O and marked several chunks on the RAID array down. At the time there were no events in the UNIX system logs but the ONLY event in the Clariion logs on the array was a rebuild complete at about the same time. Again you judge. In all these cases, BTW, the data was safe. I have lost data to the very RAID problems I described above but these spurious errors were just that completely spurious. I now use RAID10 EXCLUSIVELY on all servers configured in the last three years or more. If you ABSOLUTELY cannot afford to use RAID10, and I find it harder and harder to accept that argument with drive prices plummetting these days, then use RAID3 or RAID4 with synchronized spindles. Performance is superior to RAID5 on writes, comparable on sequential reads (and databases do NOT perform extensive random reads BTW), and safety is excellent as all drives are always read and checksums are always checked. Oh yeah, with RAID10 I also no longer get Informix errors when a drive crashes! Same Clariions, same controllers, same Informix! Hmmmmmmm! Hope that is what you were looking for Gary. Art S. Kagel