ER backing up on network
Posted in 2008
Topics: High Availability & Replication, Installation, Setup & Upgrades, Platform-Specific Issues
Version: Informix Dynamic Server 2000 Version 9.21.UC5XF
Yes, we are in the plans with our 3rd party vendor responsible for
upgrades to upgrade to 10, finally.
However.... we are running ER between a server in Denver and a server
in Dallas. Both are rs6000, one runs AIX 4.3, the other runs AIX 5.2.
Using very simple replication (i.e. "select * from xxx"). Initially,
we were load-balanced between the 2 servers. Currently, all processing
is being done on 1 server at this time with ER running between the 2.
We ran into an issue a little over a month ago where the replication
queues stopped moving back
and forth, data quickly backs up, then we get this:
09:52:27 CDR connection to server lost, id 11, name <east_serv>Reason: shutdown server
Ping response between both servers was approx 18ms...
Both servers then reconnect, "some" data gets flushed, then it starts
to build back up.
When initially diagnosing the issue, we assumed that there was
something wrong with ER (and not network related), so we totally
rebuild all replicates. THEN, after working on it for over 24 hours,
one of our network techs decided to work on the "primary" MPLS circuit
in Denver... within seconds of him switching all network traffic to
the "secondary" MPLS circuit... ALL data immediately flushed and
nothing now gets backed up.
We've been running on the secondary since them. We've done some
weekend testing on the primary circuits and traces show alot of
duplicate acks in the traces. (I'm not a network expert so bear with
me).
Tuesday we switched back to the primary. It ran fine until yesterday
when data inexplicably started backing up again. As soon as we
switched back to the secondary, within seconds again all was cleared.
I know we are running an OLD version of IDS, but has anybody
encountered this before and if so, is there anything that can be done
to alleviate it? Or is there anything we can tell our network provider
what to look for, or what to adjust on the MPLS side?
Thanks in advance,
Pete
Pete wrote:
> Version: Informix Dynamic Server 2000 Version 9.21.UC5XF
> Yes, we are in the plans with our 3rd party vendor responsible for
> upgrades to upgrade to 10, finally.
>
> However.... we are running ER between a server in Denver and a server
> in Dallas. Both are rs6000, one runs AIX 4.3, the other runs AIX 5.2.
> Using very simple replication (i.e. "select * from xxx"). Initially,
> we were load-balanced between the 2 servers. Currently, all processing
> is being done on 1 server at this time with ER running between the 2.
>
> We ran into an issue a little over a month ago where the replication
> queues stopped moving back
> and forth, data quickly backs up, then we get this:
>
> 09:52:27 CDR connection to server lost, id 11, name <east_serv>> Reason: shutdown server
>
> Ping response between both servers was approx 18ms...
> Both servers then reconnect, "some" data gets flushed, then it starts
> to build back up.
>
> When initially diagnosing the issue, we assumed that there was
> something wrong with ER (and not network related), so we totally
> rebuild all replicates. THEN, after working on it for over 24 hours,
> one of our network techs decided to work on the "primary" MPLS circuit
> in Denver... within seconds of him switching all network traffic to
> the "secondary" MPLS circuit... ALL data immediately flushed and
> nothing now gets backed up.
>
> We've been running on the secondary since them. We've done some
> weekend testing on the primary circuits and traces show alot of
> duplicate acks in the traces. (I'm not a network expert so bear with
> me).
>
> Tuesday we switched back to the primary. It ran fine until yesterday
> when data inexplicably started backing up again. As soon as we
> switched back to the secondary, within seconds again all was cleared.
>
> I know we are running an OLD version of IDS, but has anybody
> encountered this before and if so, is there anything that can be done
> to alleviate it? Or is there anything we can tell our network provider
> what to look for, or what to adjust on the MPLS side?
>
I would talk to your WAN vendor and get the primary circuit fixed.
Art S. Kagel
Oninit
> Thanks in advance,
> Pete
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
>
>
Pete wrote:
> Version: Informix Dynamic Server 2000 Version 9.21.UC5XF
> Yes, we are in the plans with our 3rd party vendor responsible for
> upgrades to upgrade to 10, finally.
ER can apply in parallel if the source node is at least 9.3. Otherwise,
it can only apply serially.
The reason for this is that we generate a hash value for each row being
replicated and that hash value is then used to determine if two
transactions are attempting to update the same row. Think of this hash
value as a 'pre-lock'. We have to serialize replicated transactions
which are updating a common subset of data.
The hash value logic went into IDS 9.3. While we support mixing 7.31
and 9.21 with more recent servers, we can only do so by supporting the
common subset of functionality. Since you are using a pre-9.3 server as
the source, we can only take advantage of the pre-9.3 functionality.
That means that the increased parallelism is not available because the
hash locks are not being generated.
You can work around this somewhat by creating multiple replicates on the
tables being replicated. This would partition the table and allow data
belonging to different replicates to be applied in parallel. This may
mean, however, that you will need to remove referential constraints on
the target system.
>
> However.... we are running ER between a server in Denver and a server
> in Dallas. Both are rs6000, one runs AIX 4.3, the other runs AIX 5.2.
> Using very simple replication (i.e. "select * from xxx"). Initially,
> we were load-balanced between the 2 servers. Currently, all processing
> is being done on 1 server at this time with ER running between the 2.
>
> We ran into an issue a little over a month ago where the replication
> queues stopped moving back
> and forth, data quickly backs up, then we get this:
>
> 09:52:27 CDR connection to server lost, id 11, name <east_serv>> Reason: shutdown server
>
> Ping response between both servers was approx 18ms...
> Both servers then reconnect, "some" data gets flushed, then it starts
> to build back up.
>
> When initially diagnosing the issue, we assumed that there was
> something wrong with ER (and not network related), so we totally
> rebuild all replicates. THEN, after working on it for over 24 hours,
> one of our network techs decided to work on the "primary" MPLS circuit
> in Denver... within seconds of him switching all network traffic to
> the "secondary" MPLS circuit... ALL data immediately flushed and
> nothing now gets backed up.
>
> We've been running on the secondary since them. We've done some
> weekend testing on the primary circuits and traces show alot of
> duplicate acks in the traces. (I'm not a network expert so bear with
> me).
>
> Tuesday we switched back to the primary. It ran fine until yesterday
> when data inexplicably started backing up again. As soon as we
> switched back to the secondary, within seconds again all was cleared.
>
> I know we are running an OLD version of IDS, but has anybody
> encountered this before and if so, is there anything that can be done
> to alleviate it? Or is there anything we can tell our network provider
> what to look for, or what to adjust on the MPLS side?
>
> Thanks in advance,
> Pete