HDR performance Sydney - London
Posted in 2007
A user testing HDR between a London primary and Sydney secondary (IDS 10.FC5, Solaris 10) found a load that took under a minute without HDR took 28 minutes with it, despite the same load running fine between two London servers and scp of the file taking only 3 minutes. Ping showed ~377ms round-trip latency. Suggestions included raising LOGBUFF (which halved the time but enlarged checkpoints), using buffered logging, tuning the ASF netbuffer size in sqlhosts, examining sqlexec thread states via onstat -g ath/stk, and cascading through local secondaries at each site. No resolution is recorded; the poster was still questioning why scp was so much faster than HDR.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Performance & Tuning, Server Administration, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
I am currently testing a new HDR environment (IDS 10.FC5 on Solaris 10) for a customer where the primary and secondary are in London and Sydney respectively. As an initial test of performance I have been trying a relatively small "load from ... insert into ..." statement. With HDR turned off this completes in less than a minute but with HDR on it is 28 minutes (DRINTERVAL is set to 30). Running a ping between the two boxes during the load shows no difference in response time compared to when informix is down on each. At the moment there is no other activity either on the servers or across the network as I am more or less the only user. If I make the secondary a standard and run the same load there it also completes within a minute. If I pair the London primary with another instance in London and do the same tests the load completes in less than a minute with or without HDR which inevitable points to the network. However if I scp the load file between the London and Sydney boxes that completes in 3 minutes, obviously significantly slower than informix is able to do it through HDR. The only way I have been able to improve the performance with HDR switched on is to increase LOGBUFF - increasing it from 32 to 256 halves the load time. I would guess that this points to contention with the HDR log buffer as there is only one of these compared to 3 log buffers? Even though this makes the load twice as quick it has the draw back of leading to very much larger checkpoints and increases the possible data loss in a crash. Any help or comments on using HDR over such a distance would be greatly appreciated as previously I have either used it on the same lan or just between sites within the UK. In particular I would be interested to understand: - whether there are any other ONCONFIG settings I could experiment with - whether informix is doing an excessive amount of "handshaking" when sending the log information which may be improved at a network level by adjusting the size of packets sent Please let me know if there is any more information that I can provide. Kind Regards, Andy
andrew.watson@ardenta.com said: > - whether informix is doing an excessive amount of "handshaking" when > sending the log information which may be improved at a network level > by adjusting the size of packets sent That's the only thing I can think of. -- Bye now, Obnoxio "I don't read newspapers anymore except the local rag which I do weekly to cheer myself trying to see if anyone I hate has been stabbed." -- Horribilis XVI -- This message has been scanned for viruses and dangerous content by OpenProtect(http://www.openprotect.com), and is believed to be clean.
> The only way I have been able to improve the performance with HDR
> switched on is to increase LOGBUFF - increasing it from 32 to 256
> halves the load time. I would guess that this points to contention
> with the HDR log buffer as there is only one of these compared to 3
> log buffers? Even though this makes the load twice as quick it has the
> draw back of leading to very much larger checkpoints and increases the
> possible data loss in a crash.
>
Well there are actually 12 HDR log buffers, not just 1. You don't
mention if the database is using buffered or unbuffered logging.
Although when you said changing the LOGBUFF to 256 did speed up the
performance, I'll guess buffered, but weirder things have happened.
If it does happen to be unbuffered logging, you might try it changing
to buffered and see the effect.
>
> - whether informix is doing an excessive amount of "handshaking" when
> sending the log information which may be improved at a network level
> by adjusting the size of packets sent
When we send a HDR buffer to the secondary we are not specifically
don't a lot of handshaking ourselves. However, tcp might. If I
remember correctly we break the HDR buffer down into asf netbuffs,
which defaults to 4k I think. So we are doing 4k tcp sends. The tcp
layer has to ack each of those sends, but at the Informix level, we
only do 1 ack of the send after the entire log buff is sent. So we do
wait for 1 ack after each HDR buffer is sent, but tcp will be doing
several acks since each HDR buffer will be broken into multiple
pieces. I think you can configure the asf netbuffer size with an
option in the sqlhost file (make sure you make the same change to both
servers). I'm not 100% certain that would after the HDR layer, but
you could try it. Also, I'd be curious what the state of the sqlexec
thread is in during those 28 minutes, from an onstat -g ath or do
onstat -g stk <thread id> to collect stack traces to see where thethread is waiting around.
>
> Please let me know if there is any more information that I can
> provide.
>
> Kind Regards,
> Andy
Jacques Renaut
IBM/Informix Resolution Team
andrew.watson@ardenta.com wrote: > I am currently testing a new HDR environment (IDS 10.FC5 on Solaris > 10) for a customer where the primary and secondary are in London and > Sydney respectively. As an initial test of performance I have been > trying a relatively small "load from ... insert into ..." statement. > With HDR turned off this completes in less than a minute but with HDR > on it is 28 minutes (DRINTERVAL is set to 30). Running a ping between > the two boxes during the load shows no difference in response time > compared to when informix is down on each. At the moment there is no > other activity either on the servers or across the network as I am > more or less the only user. If I make the secondary a standard and run > the same load there it also completes within a minute. What is the ping time between the two servers? > > If I pair the London primary with another instance in London and do > the same tests the load completes in less than a minute with or > without HDR which inevitable points to the network. However if I scp > the load file between the London and Sydney boxes that completes in 3 > minutes, obviously significantly slower than informix is able to do it > through HDR. > > The only way I have been able to improve the performance with HDR > switched on is to increase LOGBUFF - increasing it from 32 to 256 > halves the load time. I would guess that this points to contention > with the HDR log buffer as there is only one of these compared to 3 > log buffers? Even though this makes the load twice as quick it has the > draw back of leading to very much larger checkpoints and increases the > possible data loss in a crash. > > Any help or comments on using HDR over such a distance would be > greatly appreciated as previously I have either used it on the same > lan or just between sites within the UK. In particular I would be > interested to understand: > > - whether there are any other ONCONFIG settings I could experiment > with > > - whether informix is doing an excessive amount of "handshaking" when > sending the log information which may be improved at a network level > by adjusting the size of packets sent > > Please let me know if there is any more information that I can > provide. > > Kind Regards, > Andy >
"Madison Pruet" <mpruet@tx.rr.com> wrote in message news:45C222EF.8020406@tx.rr.com... > andrew.watson@ardenta.com wrote: >> I am currently testing a new HDR environment (IDS 10.FC5 on Solaris >> 10) for a customer where the primary and secondary are in London and >> Sydney respectively. As an initial test of performance I have been >> trying a relatively small "load from ... insert into ..." statement. >> With HDR turned off this completes in less than a minute but with HDR >> on it is 28 minutes (DRINTERVAL is set to 30). Running a ping between >> the two boxes during the load shows no difference in response time >> compared to when informix is down on each. At the moment there is no >> other activity either on the servers or across the network as I am >> more or less the only user. If I make the secondary a standard and run >> the same load there it also completes within a minute. > > What is the ping time between the two servers? [informix@londb1:~] /usr/sbin/ping -s syddb1 PING syddb1: 56 data bytes 64 bytes from syddb1 (10.1.2.183): icmp_seq=0. time=378. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=1. time=378. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=2. time=377. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=3. time=377. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=4. time=383. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=5. time=377. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=6. time=377. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=7. time=377. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=8. time=377. ms 64 bytes from syddb1 (10.1.2.183): icmp_seq=9. time=377. ms ^C ----syddb1 PING Statistics---- 10 packets transmitted, 10 packets received, 0% packet loss round-trip (ms) min/avg/max/stddev = 377./377.7/383./2.02
>From: "Neil Truby" <neil.truby@ardenta.com> >"Madison Pruet" <mpruet@tx.rr.com> wrote in message >news:45C222EF.8020406@tx.rr.com... > > andrew.watson@ardenta.com wrote: > >> I am currently testing a new HDR environment (IDS 10.FC5 on Solaris > >> 10) for a customer where the primary and secondary are in London and > >> Sydney respectively. As an initial test of performance I have been > >> trying a relatively small "load from ... insert into ..." statement. > >> With HDR turned off this completes in less than a minute but with HDR > >> on it is 28 minutes (DRINTERVAL is set to 30). Running a ping between > >> the two boxes during the load shows no difference in response time > >> compared to when informix is down on each. At the moment there is no > >> other activity either on the servers or across the network as I am > >> more or less the only user. If I make the secondary a standard and run > >> the same load there it also completes within a minute. > > > > What is the ping time between the two servers? > >[informix@londb1:~] /usr/sbin/ping -s syddb1 >PING syddb1: 56 data bytes >64 bytes from syddb1 (10.1.2.183): icmp_seq=0. time=378. ms Hmmm. So you have what 1/3 of a second latency between packets? That's a killer. (Looks like you've got a nice network if you only have 1/3 a second latency.) I doubt you'll get the performance that you want. One option would be to have two servers at each location. A production box to a local redundant box which is tied to the redundant box at the other location... A1=A2<--->B2=B1 So at either location production isn't effected. (At least in theory. Its possible that the latency between the two locations may impact system performance by holding resources longer.) _________________________________________________________________ Search for grocery stores. Find gratitude. Turn a simple search into something more. http://click4thecause.live.com/search/charity/default.aspx?source=hmemtagline_gratitude&FORM=WLMTAG
FYI The database is buffered. I will try the asf netbuff in sqlhosts and get back to you if I can see any differences in the ath and stk outputs before and after. The main point I struggle to understand is how sending the data through a scp only takes 3 minutes compared to the HDR transfer of 28 minutes - wouldn't any general network latency issues bring these two figures closer together? Thanks for the comments
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g