RE: Informix and RAID
Posted in 2004
Topics: High Availability & Replication, Storage & Space Management, Logging & Checkpoints
Hi Art, You certainly make a strong argument against RAID 5. I'm wondering what you reaction is to some of the reasons we have chosen to use RAID 5 in the past. Namely: About 4 years ago when we purchased our database servers, some of them were 4-way systems with 12 8G drives and some were 2-way's with 6 drives. On the servers with 6 drives we didn't have much choice. Due to the size of the database, we had to use RAID 5 to get the space we needed (5 drives in a RAID 5 and 1 hot spare). On the servers with 12 drives we have experimented with different options. One of our concerns has been spreading the load evenly over as many spindles as possible to reduce contention for a single disk. With RAID 5, that happens automatically to a certain extent, but I suspect that it can work against you too if you have parellel processes that are both tying up several disks at once. With 12 drives to work with, our current configuration is: a 2 disk RAID 1 mirror for the OS and rootdbs, a 2-disk RAID 1 mirror for physical, log, temp dbspaces, and /var/tmp, a 6 disk RAID 5 array for the database dbspace and a large scratch filesystem which we use for maintenance activities like performing a datbase export/import to disk, and 2 hot spares The newer servers that are available today (i.e. they are "approved" by IT, they are compatible with our system software, and they are within our budget) have lots of space, but not that many internal hard drives. RAID 5 is always available for the array controllers, but not always RAID 10. So figuring out how to make the most effective use of all the spindles is still a challenge. BTW, our group has supported as many as dozen critical OLTP database instances at a time for 20+ years, many of them running 24x7 for most of the year. We're very concerned about not losing data. We use HDR, level 0 archives 1-2 times daily, and we continuously back up the logical logs to tape. Under these conditions, data loss hasn't been a problem, even though we use RAID 5. In fact, I can't remember a case where disk failures have ever been the root cause of a production downtime. John sending to informix-list
"Spurgeon, John P" <john.p.spurgeon@intel.com> schrieb: >BTW, our group has supported as many as dozen critical OLTP database >instances at a time for 20+ years, many of them running 24x7 for most of >the year. We're very concerned about not losing data. We use HDR, level >0 archives 1-2 times daily, and we continuously back up the logical logs >to tape. Under these conditions, data loss hasn't been a problem, even >though we use RAID 5. In fact, I can't remember a case where disk >failures have ever been the root cause of a production downtime. I am a very cautions and diligent driver and have driven safely without an accident for 20+ years. I am very concerned about not losing my life. I am using power steering, power brakes, electronic assist systems for brakes and driving stability. Do you think that I should stop using my seat belt and disable the airbag? Regards, Richard
Spurgeon, John P wrote: > Hi Art, > > You certainly make a strong argument against RAID 5. I'm wondering what > you reaction is to some of the reasons we have chosen to use RAID 5 in > the past. Namely: > > About 4 years ago when we purchased our database servers, some of them > were 4-way systems with 12 8G drives and some were 2-way's with 6 > drives. On the servers with 6 drives we didn't have much choice. Due to > the size of the database, we had to use RAID 5 to get the space we > needed (5 drives in a RAID 5 and 1 hot spare). Actually you did have a choice. It would have cost you an extra $75US to get and external SCSI drive cabinet or two or four with more drives in them. (You weren't using IDS drives I hope, 'cause they're many times more susceptible to partial media failure data corruption then SCSI drives are and that's my BIGGEST argument against RAID5?) > On the servers with 12 drives we have experimented with different > options. One of our concerns has been spreading the load evenly over as > many spindles as possible to reduce contention for a single disk. With And RAID10 does that better than RAID5 putting twice as many spindles to work for the same capacity resulting in 100% better burst read times. No, it's not sustainable, but most high demand read situations are short lived. > RAID 5, that happens automatically to a certain extent, but I suspect > that it can work against you too if you have parellel processes that are > both tying up several disks at once. Again, RAID10 performs much better in multi-thread/multi-process environments because it can read two sectors from the same 'drive' in parallel. > With 12 drives to work with, our > current configuration is: > > a 2 disk RAID 1 mirror for the OS and rootdbs, > > a 2-disk RAID 1 mirror for physical, log, temp dbspaces, and /var/tmp, > > a 6 disk RAID 5 array for the database dbspace and a large scratch > filesystem which we use for maintenance activities like performing a > datbase export/import to disk, and > > 2 hot spares > > The newer servers that are available today (i.e. they are "approved" by > IT, they are compatible with our system software, and they are within > our budget) have lots of space, but not that many internal hard drives. That's what external drive cabinets, SANs, and NASs are for. > RAID 5 is always available for the array controllers, but not always > RAID 10. So figuring out how to make the most effective use of all the > spindles is still a challenge. Then you're purchasing your RAID controllers from the wrong vendors. I can name at least three excellent SCSI controllers that do RAID10. Then there's always a good Volume Manager (and there are several) that can do RAID10 in software. Performance is about 90% what a pure hardware solution will give you but the gain in configuration flexibility is enormous. You can easily build your array from pairs on different controllers in different cabinets to maximize bandwidth and minimize loss potential from crashing controllers and power supplies. > BTW, our group has supported as many as dozen critical OLTP database > instances at a time for 20+ years, many of them running 24x7 for most of > the year. We're very concerned about not losing data. We use HDR, level > 0 archives 1-2 times daily, and we continuously back up the logical logs > to tape. Under these conditions, data loss hasn't been a problem, even > though we use RAID 5. In fact, I can't remember a case where disk > failures have ever been the root cause of a production downtime. And when I was lead DBA here I supported 48 servers with over 100 TB of disk all 24x7x365. The ONLY time I EVER had a server down due to disk or controller problems was when we used RAID5. Once we went to RAID10 I stopped getting calls at 3AM about a server complaining about a down chunk. Before I won the RAID war here, I used to spend one to three nights every month rebuilding indexes overnight. One time two out of three servers in our News system were down with three or four bad indexes each. The tables were so large even then that it took 8-12 hours to rebuild one (Online 5.08) so I was looking at an all nighter plus the next day to rebuild. If there were any important news the next day (like the Lauraine Bobbet story just 3 months earlier) we'd need all three servers online to handle the load and we'd only have one. Needless to say the BIG bosses were WAY beyond nervous: "If we lose that one we're out of business, Art! Are you going to fix this? Am I going to have a News system in the morning?" It took some slight of hand but I got a second server up before dawn. Phew. The index corruption was due to partial media failure on one or two of the RAID5 drives. It took several days to replace half of the drives in a 30 drive RAID5 plaid, one at a time, until the culprits were found (don't ask why the SAs didn't just build a new array and restore, I dunno). I shudder to think about what data records were trashed, what stories had been lost. Fortunately only old news was affected and noone noticed that part, but we needed the indexes to be whole and consistent. That was a year or so before the mass drive failure finally convinced the SAs and managers to switch to RAID10. In the 5 years since I gave up that role, none of the DBAs have gotten even one of those calls due to disk or controller problems and we've only gotten one corrupted index in all that time on all those servers. Have we lost drives, controllers or power supplies in all that time? Sure, many, we have many hundreds of the things after all. It's inevitable. In any given week something somewhere is failing. But the DBAs here don't care. I used to know exactly how many cabinets and controllers and drives were on each and every server. I had to to survive. Today the DBAs haven't a clue. Is it because they don't care as much as I did? No, it because it no longer affects their job performance as it did mine. Do you REALLY wonder why I prefer RAID10 over RAID5 so vehemently? Art S. Kagel > John > > sending to informix-list