chunk transplant
Posted in 2001
Topics: High Availability & Replication, Storage & Space Management, Stored Procedures & SPL, Platform-Specific Issues
We have been investigating ways to keep data current on a local and identical backup server. We are aware of IER and HDR. My boss suggested that we keep a mirrored chunk on a cross mounted drive that resides on the backup server. Then in the event of a failure of the primary server, bring the backup server online using the primary servers chunk. I skeptically attempted this and it worked! The conditions are: The schemas, database spaces, chunk size and number of chunks are the same on both servers. The chunks use cooked files. The databases are created in rootdbs and the tables, indexes are created in dedicated chunks. The chunks from the primary server are copied over the backup server chunk files before the backup instance is brought online. We are using 9.2uc3 on solaris 2.7. Question: Under what conditions would this not work (within the conditions above) ? I either need to demonstrate why not to use the technique or make it work reliably? Thanks, Bob Carts
In article <9422kf$k08$1@knight.vf.lmco.com>, Robert Carts <rcarts@agccs.lmco.com> writes >We have been investigating ways to keep data current on a local and >identical backup server. We are aware of IER and HDR. My boss suggested >that we keep a mirrored chunk on a cross mounted drive that resides on the >backup server. Then in the event of a failure of the primary server, bring >the backup server online using the primary servers chunk. I skeptically >attempted this and it worked! > >The conditions are: The schemas, database spaces, chunk size and number of >chunks are the same on both servers. The chunks use cooked files. The >databases are created in rootdbs and the tables, indexes are created in >dedicated chunks. The chunks from the primary server are copied over the >backup server chunk files before the backup instance is brought online. We ?? If the primary server is down how can you do the copy? I'm not sure I follow, take the root dbspace as an example. What chunks are on each machine. What is the copying which is done? >are using 9.2uc3 on solaris 2.7. > >Question: Under what conditions would this not work (within the conditions >above) ? I either need to demonstrate why not to use the technique or make >it work reliably? > >Thanks, Bob Carts > > -- David Williams
Robert Carts wrote in message <9422kf$k08$1@knight.vf.lmco.com>... >We have been investigating ways to keep data current on a local and >identical backup server. We are aware of IER and HDR. My boss suggested >that we keep a mirrored chunk on a cross mounted drive that resides on the >backup server. Then in the event of a failure of the primary server, bring >the backup server online using the primary servers chunk. I skeptically >attempted this and it worked! > <SNIP> > >Question: Under what conditions would this not work (within the conditions >above) ? I either need to demonstrate why not to use the technique or make >it work reliably? > >Thanks, Bob Carts *) Performance will be much worse just from using cooked files. My experience shows typically 10 times worse. *) Cooked files are intrinsically unreliable because the engine sometimes needs to guarantee that pages are physically written in special sequences, and you just can't get that reliability when the O/S uses buffering and picks it's own sequence of flushes to disk. *) Performance will be even worse than that since one of the mirror pairs will be at the end of a network connection and NFS. *) If the network has a hiccup or the machine with the mirrors goes offline, then the mirrors will be marked down on the first attempt to write them, and it will perform a full update of the mirror chunks once the connection is reestablished. This will cause mass traffic on the network and hit both machines really hard. *) If the original machine goes down, and you attempt to bring up an instance on the mirror machine, then there is a really strong chance that the so-called mirror instance will be in an inconstent state, due to the unpredictability of the sequencing of the network writes and the UNIX writes to cooked files. In a sense, for reliability, a remote cooked file is as bad as a local cooked file only much worse. If the physical data is inconsistent then the results will be anything from complete failure to bring up the mirr or instance, to grubby damaged data (which is probably worse because you will be tempted to imagine that it's clean and reliable data) SUMMARY: don't do it. If you want fallback to another machine, then: *) use the replication offerings, or *) improve the hardware redundancy within the source machine: mirrored disks, multiple disk controllers, duplicate power supplies, hot-swappable disks and even controllers and power supplies - whatever it takes to reach the level of reliability you require for 24*7*365 uptime. *) If you are not concerned with 100% uptime, but just want the convenience of a mirror machine, consider setting up identical machines (including partitioning layouts) and identical config files etc. Then, you could archive the main machine (as you should) and restore to the mirror machine, which would give you the luxury of having a test of the quality of your archives. By a strange coincidence, our development machine blew up it's power supply last night, and one of our chunks got scribbled in the process because a big mass-compile was running at the time. We haven't mirrored the development databases, because we don't see the need. I'm still very happy with that decision after today's experience - we had to restore from an archive tape and now we're back in business. Total time down: 15 minutes to screw in a new power supply, plus the time to restore the archive; typically 130% of the time archives take to write in my experience. This is the bread and butter of maintaining the security of your databases. Cute or interesting ideas will inevitably end in tears.
Comments embedded In article <3a64f31a@news.iprimus.com.au>, "Andrew Hamm" <ahamm@sanderson.net.au> wrote: > Robert Carts wrote in message <9422kf$k08$1@knight.vf.lmco.com>... > >We have been investigating ways to keep data current on a local and > >identical backup server. We are aware of IER and HDR. My boss suggested > >that we keep a mirrored chunk on a cross mounted drive that resides on the > >backup server. Then in the event of a failure of the primary server, bring > >the backup server online using the primary servers chunk. I skeptically > >attempted this and it worked! > > > <SNIP> > > > >Question: Under what conditions would this not work (within the conditions > >above) ? I either need to demonstrate why not to use the technique or make > >it work reliably? > > > >Thanks, Bob Carts > > *) Performance will be much worse just from using cooked files. My > experience shows typically 10 times worse. > I would be surprised at a full order of magnitude difference in performance on i/o. If your cache rates are high enough it shouldn't have much of an effect on the performance of the engine. Of course on some systems, the cache rate isn't high enough to be able to do this. > *) Cooked files are intrinsically unreliable because the engine sometimes > needs to guarantee that pages are physically written in special sequences, > and you just can't get that reliability when the O/S uses buffering and > picks it's own sequence of flushes to disk. If O_SYNC is used on the open, which informix does use, the write is does not complete until it is delivered to the underlying hardware. > > *) Performance will be even worse than that since one of the mirror pairs > will be at the end of a network connection and NFS. Unless he is planning on having a raid array connect to both systems, and just swing the disks over to the other system in the case of a failure. > > *) If the network has a hiccup or the machine with the mirrors goes offline, > then the mirrors will be marked down on the first attempt to write them, and > it will perform a full update of the mirror chunks once the connection is > reestablished. This will cause mass traffic on the network and hit both > machines really hard. > > *) If the original machine goes down, and you attempt to bring up an > instance on the mirror machine, then there is a really strong chance that > the so-called mirror instance will be in an inconstent state, due to the > unpredictability of the sequencing of the network writes and the UNIX writes > to cooked files. In a sense, for reliability, a remote cooked file is as bad > as a local cooked file only much worse. If the physical data is inconsistent > then the results will be anything from complete failure to bring up the mirr > or instance, to grubby damaged data (which is probably worse because you > will be tempted to imagine that it's clean and reliable data) See above argument about O_SYNC. Due to that it is just like bringing the instance up on the primary machine after the crash. <SNIP> If done properly it seems like a viable solution to me. Being very cafeful is important as always. Will Sent via Deja.com http://www.deja.com/
Thanks for blowing away all my advice! I don't think you missed anything!-/ William Rice wrote in message <944iko$4m9$1@nnrp1.deja.com>... > >I would be surprised at a full order of magnitude difference in >performance on i/o. If your cache rates are high enough it shouldn't >have much of an effect on the performance of the engine. Of course on >some systems, the cache rate isn't high enough to be able to do this. > These are just my real-world measurements from a few years ago, and the rule of thumb I therefore go with. Perhaps I've missed some wonderful UNIX optimisations that completely obsolete raw spaces? Your statements about cache rates may apply on a basically read-only engine, but I can't see large engine caches being any use when it's time to write altered pages back to physical disk. Hell - I try to setup engines based on checkpoint writes only, if I can get the times down enough, but I've met plenty of people on this newsgroup who set their LRU percents very low to get continuous LRU writes and 0 checkpoint writes. I'm thinking of trying it myself. What do you think the measurements will be in either of these cases? The odd thing is, some newbies to Informix seem to be afraid of raw spaces; in particular, some Oraclers I've met recently seem obsessed with cooked files, and I tell them "hey, you'll never get used to them if you don't even try them on your test or new development machine. Go raw and get over it." - hmmm - sounds like a slogan for a nudist camp. >If O_SYNC is used on the open, which informix does use, the write is >does not complete until it is delivered to the underlying hardware. > Quality of this depends on the platform I thought. And if O_SYNC is used, then you can guarantee a lot more waits inside the engine... There's no such thing as a free lunch. O_SYNCS still go via the buffers because the page MAY be already in a buffer. Perhaps some UNIX have optimisations on that if O_EXCL is also used? >> *) Performance will be even worse than that since one of the mirror >>pairs will be at the end of a network connection and NFS. > >Unless he is planning on having a raid array connect to both systems, >and just swing the disks over to the other system in the case of a >failure. > I didn't get that implication from his message, but fair enough. However he can't swing the disks over if it's them that's carked it. Unfortunately, the original information is too much on the vague side so we're both guessing. >> *) If the network has a hiccup or the machine with the mirrors goes >> offline, then the mirrors will be marked down on the first attempt to >> write them, and it will perform a full update of the mirror chunks >> once the connection is reestablished. This will cause mass >> traffic on the network and hit both machines really hard. > No arguments one that one? Cool - at least I got one right;-) >> *) If the original machine goes down, and you attempt to bring up an >> instance on the mirror machine, then there is a really strong chance >> that the so-called mirror instance will be in an inconstent state, due >> to the unpredictability of the sequencing of the network writes and >> the UNIX writes to cooked files. In a sense, for reliability, a remote >> cooked file is as bad as a local cooked file only much worse. If the >> physical data is inconsistent then the results will be anything from >> complete failure to bring up the mirror instance, to grubby damaged >> data (which is probably worse because you will be tempted to >> imagine that it's clean and reliable data) > >See above argument about O_SYNC. Due to that it is just like bringing >the instance up on the primary machine after the crash. > Ditto see my refutation! But even more, there is no O_SYNC in NFS and network traffic in general. Packets can get switched, flipped, broken and reassembled at any time. Before claims are made about it not being a problem on a local LAN without even so much as a router in the way, consider this: NFS typically has several nfsd and biod daemons which do the business. I note that my machine has 5 nfsd and 4 biod. Each one sits there waiting to receive requests, and it's basically first-in first-served. If one is busy, then the others capture the work. On a machine with a lot of NFS traffic (ie in the case of a mirror recovery) you can bet your boots that all 5 would be busy, and if you have a sharp-eyed UNIX administrator on the machine, then chances are they might decide to add yet more daemons to handle the load. Now, we are talking about several parallel processes being tossed around by the wind of UNIX process scheduling (or NT - we are both making assumtions here!) and the currents of network activity, collisions and retries. Traditional NFS can't even boast the use of "reliable" TCP - it deliberately uses UDP so as to maintain the illusion of a stateless service. Can you guarantee now which page will be written first? I'm trying to find out in what circumstances modern NFS can use TCP, but I'm not holding my breath for a nice surprise. Perhaps some other reader can enlighten us. >If done properly it seems like a viable solution to me. Being very >cafeful is important as always. > Above all, I just can never see the point of picking a more complicated route to a solution than an easier one. If you are going to get basically duplicate machinery as security, then there are more boring, simpler and more reliable methods to achieve the fallback. Depending on your needs for usage and reliability: 1) Add hardware and space redundancy to one or both machines 2) Replicate manually using archive and restore 3) Setup the Informix replication offerings and setup also your clients to fallback to the alternative box in the event of a failure. Straightforward solutions may not be as exciting, but I can tell you it's extremely unexciting going without sleep for a night if you have to recover from a mess. I think you've been in hard-core support from the style of your messages. Do you really enjoy that kind of excitement? I'd prefer to get my thrills from a particularly clever select statement or something. Maybe the plan can be made to work and be considered reliable enough to work for a while or indefinitely as long as nothing goes wrong, but it's far from the best solution. That's all I'm attempting to pass across to the quissical new-user of informix. -- <insert quote about Occam's Razor>