Re: chunk transplant
Posted in 2001
Topics: Performance & Tuning, Storage & Space Management, Stored Procedures & SPL, Logging & Checkpoints, Networking & sqlhosts Configuration
A few comments embedded. In article <3a66364d@news.iprimus.com.au>, "Andrew Hamm" <ahamm@sanderson.net.au> wrote: > Thanks for blowing away all my advice! I don't think you missed anything!-/ > > William Rice wrote in message <944iko$4m9$1@nnrp1.deja.com>... > > > >I would be surprised at a full order of magnitude difference in > >performance on i/o. If your cache rates are high enough it shouldn't > >have much of an effect on the performance of the engine. Of course on > >some systems, the cache rate isn't high enough to be able to do this. > > > These are just my real-world measurements from a few years ago, and the rule > of thumb I therefore go with. Perhaps I've missed some wonderful UNIX > optimisations that completely obsolete raw spaces? Your statements about > cache rates may apply on a basically read-only engine, but I can't see large > engine caches being any use when it's time to write altered pages back to > physical disk. > > Hell - I try to setup engines based on checkpoint writes only, if I can get > the times down enough, but I've met plenty of people on this newsgroup who > set their LRU percents very low to get continuous LRU writes and 0 > checkpoint writes. I'm thinking of trying it myself. What do you think the > measurements will be in either of these cases? > If you were relying on checkpoint writes I wouldn't have been overly surprised at 40 percent decrease in peformance at some points in time. that is all I was saying. > The odd thing is, some newbies to Informix seem to be afraid of raw spaces; > in particular, some Oraclers I've met recently seem obsessed with cooked > files, and I tell them "hey, you'll never get used to them if you don't even > try them on your test or new development machine. Go raw and get over it." - > hmmm - sounds like a slogan for a nudist camp. > I agree(about the slogan and the fear of raw spaces) > >If O_SYNC is used on the open, which informix does use, the write is > >does not complete until it is delivered to the underlying hardware. > > > Quality of this depends on the platform I thought. And if O_SYNC is used, > then you can guarantee a lot more waits inside the engine... There's no such > thing as a free lunch. O_SYNCS still go via the buffers because the page MAY > be already in a buffer. Perhaps some UNIX have optimisations on that if > O_EXCL is also used? > To be honest I don't know what they do besides O_SYNC, I just know my data is safe :) > >> *) Performance will be even worse than that since one of the mirror > >>pairs will be at the end of a network connection and NFS. > > > >Unless he is planning on having a raid array connect to both systems, > >and just swing the disks over to the other system in the case of a > >failure. > > > I didn't get that implication from his message, but fair enough. However he > can't swing the disks over if it's them that's carked it. Unfortunately, the > original information is too much on the vague side so we're both guessing. Quite the truth. > > >> *) If the network has a hiccup or the machine with the mirrors goes > >> offline, then the mirrors will be marked down on the first attempt to > >> write them, and it will perform a full update of the mirror chunks > >> once the connection is reestablished. This will cause mass > >> traffic on the network and hit both machines really hard. > > > No arguments one that one? Cool - at least I got one right;-) I agree with a decent bit of stuff, I just sometimes seem contradictory, because when I disagree, I post. Or maybe I am contradictory :) If I don't post, I am not giving the other person a chance to tell me why I am wrong... > > >> *) If the original machine goes down, and you attempt to bring up an > >> instance on the mirror machine, then there is a really strong chance > >> that the so-called mirror instance will be in an inconstent state, due > >> to the unpredictability of the sequencing of the network writes and > >> the UNIX writes to cooked files. In a sense, for reliability, a remote > >> cooked file is as bad as a local cooked file only much worse. If the > >> physical data is inconsistent then the results will be anything from > >> complete failure to bring up the mirror instance, to grubby damaged > >> data (which is probably worse because you will be tempted to > >> imagine that it's clean and reliable data) > > > >See above argument about O_SYNC. Due to that it is just like bringing > >the instance up on the primary machine after the crash. > > > Ditto see my refutation! But even more, there is no O_SYNC in NFS and > network traffic in general. Packets can get switched, flipped, broken and > reassembled at any time. Before claims are made about it not being a problem > on a local LAN without even so much as a router in the way, consider this: > > NFS typically has several nfsd and biod daemons which do the business. I > note that my machine has 5 nfsd and 4 biod. Each one sits there waiting to > receive requests, and it's basically first-in first-served. If one is busy, > then the others capture the work. On a machine with a lot of NFS traffic (ie > in the case of a mirror recovery) you can bet your boots that all 5 would be > busy, and if you have a sharp-eyed UNIX administrator on the machine, then > chances are they might decide to add yet more daemons to handle the load. > > Now, we are talking about several parallel processes being tossed around by > the wind of UNIX process scheduling (or NT - we are both making assumtions > here!) and the currents of network activity, collisions and retries. > Traditional NFS can't even boast the use of "reliable" TCP - it deliberately > uses UDP so as to maintain the illusion of a stateless service. Can you > guarantee now which page will be written first? I'm trying to find out in > what circumstances modern NFS can use TCP, but I'm not holding my breath for > a nice surprise. Perhaps some other reader can enlighten us. > Using NFS (Unless we are talking about a SAN solution) is most likely a bad idea. > >If done properly it seems like a viable solution to me. Being very > >cafeful is important as always. > > > Above all, I just can never see the point of picking a more complicated > route to a solution than an easier one. If you are going to get basically > duplicate machinery as security, then there are more boring, simpler and > more reliable methods to achieve the fallback. > > Depending on your needs for usage and reliability: > 1) Add hardware and space redundancy to one or both machines > 2) Replicate manually using archive and restore > 3) Setup the Informix replication offerings and setup also your clients to > fallback to the alternative box in the event of a failure. > > Straightforward solutions may not be as exciting, but I can tell you it's > extremely unexciting going without sleep for a night if you have to recover > from a mess. I think you've been in hard-core s
William Rice wrote in message <945gk6$1c6$1@nnrp1.deja.com>...
>
>......... That is actually why I like the
>idea of swinging the disks being attached to two different systems. As
>long as you build in redundancy into the disk solution(mirror across
>disk arrays) you should be ok. I am pretty sure you dont even have to
>use cooked space, you can use should be able to use raw as easily. My
>system administration skills are to rusty for me to be sure on that,
>but I would be appalled were I incorrect.
>
>This way there are no difficult details of setting up replication, and
>there isn't the pain of having a down system for a long time. The only
>pain is the time it takes to switch the disks over, and whatever needs
>to be done to switch the application over.
>
>Hope this post isn't to long...
Not at all. I'm hoping someone can enlighten us further on NFS and TCP. If
Robert is thinking of using NT boxes, probably similar arguments could be
used against the use of the underlying SMB protocol or whatever the tell NT
uses these days for file sharing.
The swinging or hot-swapping of disks works from all I've heard from other
people. In fact, it's quite fun on a machine you setup with hot-swappable
disks, to suddenly pull out one of a mirror pair, especially after hours and
hours of machine setup time. Everybody's jaws hit the ground, and some
people even start sweating or screaming. You will too if you accidentally
pull a non-mirror pair member...
Then you give 'em a big smile, plug the bugger back in, run the onspaces to
bring the chunk(s) back online, and then sit mesmerised for 30 minutes while
the disk lights flash in a pretty pattern recovering the spaces. Ultimately,
the customers are happy with this vivid experience of hardware redundancy
and they feel all warm and fuzzy inside.
I would perhaps suggest this party trick is not tried by anyone setting up
in a military site in the presence of strapped weaponry!
In article <3a664be5@news.iprimus.com.au>, "Andrew Hamm" <ahamm@sanderson.net.au> wrote: > William Rice wrote in message <945gk6$1c6$1@nnrp1.deja.com>... > > > >......... That is actually why I like the > >idea of swinging the disks being attached to two different systems. As > >long as you build in redundancy into the disk solution(mirror across > >disk arrays) you should be ok. I am pretty sure you dont even have to > >use cooked space, you can use should be able to use raw as easily. My > >system administration skills are to rusty for me to be sure on that, > >but I would be appalled were I incorrect. > > > >This way there are no difficult details of setting up replication, and > >there isn't the pain of having a down system for a long time. The only > >pain is the time it takes to switch the disks over, and whatever needs > >to be done to switch the application over. > > > >Hope this post isn't to long... > > Not at all. I'm hoping someone can enlighten us further on NFS and TCP. If > Robert is thinking of using NT boxes, probably similar arguments could be > used against the use of the underlying SMB protocol or whatever the tell NT > uses these days for file sharing. This is getting close to off topic ... On a solaris 2.6 box. -- Ouput from man nfsd -l Set connection queue length for the NFS TCP over a connection-oriented transport. The default value is 32 entries. given this option, I think it can be said that tcp is possible for NFS on at least one O/S. I have not done enough research on the behaviour of NFS to be completely comfortable with the idea, though I seem to recall Net Appliance uses NFS mounts for you to access their SAN, and I seem to recall Informix supports this, so it must be reliable in some cases... Will P.S. I claim no expertise on this topic, just an uninformed opinion :) <SNIP> Sent via Deja.com http://www.deja.com/
William Rice wrote in message <946t9j$3to$1@nnrp1.deja.com>... > >P.S. I claim no expertise on this topic, just an uninformed opinion :) Heh heh. I'm rapidly approaching the uninformed side too, 'cos I've got the old Nutshell book on NFS and NIS and it goes into detail about using UDP to implement a stateless server which simplifies crash recovery. It specifically points out the problems with using TCP. So I'm stuffed if I know what the reasons for allowing TCP are these days - perhaps just better programming and possibly make it an option when you can almost guarantee crash-free machinery on a local LAN. I dunno. Possibly even TCP doesn't get around the parallelism and peculiar exploitation of transport unreliability in the interests of service reliability.
Wow, I wondered if anyone would respond! First thanks to everyone that responded. -I work with both raw and cooked spaces on different projects. For some reason I haven't seen an appreciable performance difference given our system setup (except on a dec/alpha datawarehouse project). We even tested the difference a few years back and the cooked configuration won. Obviously something else came into play. I have also never seen the dreaded incomplete write/file system buffer problem on solaris 2.51 in 4.5 years. -I think the best reason I read for avoiding the chunk transplant approach is that the chunk may become inconsistent when the primary instance chokes. Then I wouldn't have squat. -I'm not sure I fall into the "quissical new-user of informix" after getting my sys admin certification, 4.5 years of Informix DBA experience and presenting at the annual user conference on IER, (sorry for the resume). Perhaps "quissical old-user of informix" would be better. -I think I will keep the technique limited to a quick way to populate another development instance for the time being and suggest using IER to keep the backup server very fresh. Thanks again, Bob Carts Andrew Hamm <ahamm@sanderson.net.au> wrote in message news:3a677742$1@news.iprimus.com.au... > William Rice wrote in message <946t9j$3to$1@nnrp1.deja.com>... > > > >P.S. I claim no expertise on this topic, just an uninformed opinion :) > > Heh heh. I'm rapidly approaching the uninformed side too, 'cos I've got the > old Nutshell book on NFS and NIS and it goes into detail about using UDP to > implement a stateless server which simplifies crash recovery. It > specifically points out the problems with using TCP. So I'm stuffed if I > know what the reasons for allowing TCP are these days - perhaps just better > programming and possibly make it an option when you can almost guarantee > crash-free machinery on a local LAN. I dunno. Possibly even TCP doesn't get > around the parallelism and peculiar exploitation of transport unreliability > in the interests of service reliability. > > >