HDR hardware requirements and capabilities
Posted in 2004
Topics: High Availability & Replication
I have a 7.31.UC6X2 engine running on a HP L3000 with 4 (750), 6G Ram processors. The server is running HDR and replicating across a T1 to the secondary box, a HP K570 with 6 processors (200) 4.5G Ram. We cannot keep HDR running without the secondary dying and then some secondary failure takes out the primary because of a loss of contact with the secondary. Does HDR absolutely REQUIRE the hardware to be the same? The DR box is the very box that ran the database before the L3000 came online. Part of me says that there is no way the primary should be susceptible to changes in state on the secondary. Instead of bullet-proofing your site, you are creating another point of failure. HDR has been the cause of three D's so far and we're ready to just scrap it and say if we truly want uptime we have to do away with HDR, or watch the secondary like a hawk. We are running the patched version because we had two outages regarding HDR 7.31.UC6. Informix support offered us this patch but we had the same issues and now they are recommending going to 7.31.UD7. Has anyone else had this kind of experience? Thanks
T.Paakki wrote: > I have a 7.31.UC6X2 engine running on a HP L3000 with 4 (750), 6G > Ram processors. The server is running HDR and replicating across a > T1 to the secondary box, a HP K570 with 6 processors (200) 4.5G > Ram. We cannot keep HDR running without the secondary dying and > then some secondary failure takes out the primary because of a loss > of contact with the secondary. Does HDR absolutely REQUIRE the > hardware to be the same? The DR box is the very box that ran the > database before the L3000 came online. Not identical - but they must both be running the same IDS binaries, and the disk layouts need to be the same. If the binaries are different, you're on a hiding to nothing. Additionally, the secondary must be able to keep up with the workload on the primary. I'm guessing that your primary has 4x750 MHz CPUs, for 3000 MHz nominal processing capacity, and the secondary has 6x200 MHz CPUs, for 1200 MHz nominal processing capacity. I'm not sure whether your secondary is fast enough to keep up with the primary; probably, but I'm not sure of that. Are you using DRAUTO? How reliable is your T1 line? Is the data able to get across with all the other traffic also on the connection? What is the volume of data being shipped across? Is there any consistency to when the secondary fails? Is there a peak in activity on the primary? Why is the secondary dying? What does the online log file contain in the way of diagnostics? Is there any evidence that the secondary can't keep up with the workload from the primary? > Part of me says that there is no way the primary should be susceptible > to changes in state on the secondary. Instead of bullet-proofing your > site, you are creating another point of failure. HDR has been the > cause of three D's so far and we're ready to just scrap it and say if > we truly want uptime we have to do away with HDR, or watch the > secondary like a hawk. It depends on what you mean by the 'the primary should [not] be susceptible to changes in state on the secondary'. The primary is aware of the secondary; it keeps an eye on whether the secondary is still there at the other end of the network. > We are running the patched version because we had two outages > regarding HDR 7.31.UC6. Informix support offered us this patch but we > had the same issues and now they are recommending going to 7.31.UD7. > > Has anyone else had this kind of experience? Not me, personally - no, I've not had that kind of experience. If Madison chips in with alternative suggestions, take his advice where it conflicts with mine; he's the guru on ER and HDR. Not that I've given much advice - more questions than answers. -- Jonathan Leffler #include <disclaimer.h> Email: jleffler@earthlink.net, jleffler@us.ibm.com Guardian of DBD::Informix v2003.04 -- http://dbi.perl.org/
T.Paakki wrote: > I have a 7.31.UC6X2 engine running on a HP L3000 with 4 (750), 6G Ram > processors. The server is running HDR and replicating across a T1 to > the secondary box, a HP K570 with 6 processors (200) 4.5G Ram. We > cannot keep HDR running without the secondary dying and then some > secondary failure takes out the primary because of a loss of contact > with the secondary. Does HDR absolutely REQUIRE the hardware to be > the same? The DR box is the very box that ran the database before the > L3000 came online. > > Part of me says that there is no way the primary should be susceptible > to changes in state on the secondary. Instead of bullet-proofing your > site, you are creating another point of failure. HDR has been the > cause of three D's so far and we're ready to just scrap it and say if > we truly want uptime we have to do away with HDR, or watch the > secondary like a hawk. > As Jonhatan wrote, the hardware doesn't have to be the same. The binaries and disk layout yes. There are a few problems with HDR which can cause instability. Most of them are already fixed or will be in UD8. Some of them have more or less easy workarounds. I've been having problems for some long time. Currently I'm a bit confident I've trace them. > We are running the patched version because we had two outages > regarding HDR 7.31.UC6. Informix support offered us this patch but we > had the same issues and now they are recommending going to 7.31.UD7. UD7 is a good bet for HDR (at least it has many bugs corrected). But if you're having the same problems I'm facing it won't solve them... Please give us more details. A few messages from the online logs would be a start. Also, indicate how you have DBSPACETEMP configured (in the engine and client environment). And finally indicate us the DR* parameters from the onconfig file and the logging mode of your databases. By the way... most of my problems don't happen on recent 9.30 (and presumably on 9.40). What this means is that the problems have been fixed, but not (yet) backported to 7.31. Regards.
Jonathan Leffler <jleffler@earthlink.net> wrote in message news:<toS3c.14343$%06.11688@newsread2.news.pas.earthlink.net>...
> T.Paakki wrote:
> > I have a 7.31.UC6X2 engine running on a HP L3000 with 4 (750), 6G
> > Ram processors. The server is running HDR and replicating across a
> > T1 to the secondary box, a HP K570 with 6 processors (200) 4.5G
> > Ram. We cannot keep HDR running without the secondary dying and
> > then some secondary failure takes out the primary because of a loss
> > of contact with the secondary. Does HDR absolutely REQUIRE the
> > hardware to be the same? The DR box is the very box that ran the
> > database before the L3000 came online.
>
> Not identical - but they must both be running the same IDS binaries,
> and the disk layouts need to be the same. If the binaries are
> different, you're on a hiding to nothing.
OK, this binaries are exactly the same as is the disk layout.
>
> Additionally, the secondary must be able to keep up with the workload
> on the primary. I'm guessing that your primary has 4x750 MHz CPUs,
> for 3000 MHz nominal processing capacity, and the secondary has 6x200
> MHz CPUs, for 1200 MHz nominal processing capacity. I'm not sure
> whether your secondary is fast enough to keep up with the primary;
> probably, but I'm not sure of that.
>
> Are you using DRAUTO?
No, DRAUTO is set to 0. Our switch to DR would require a DNS change
anyway, so that wouldn't work for us.
>
> How reliable is your T1 line? Is the data able to get across with all
> the other traffic also on the connection? What is the volume of data
> being shipped across? Is there any consistency to when the secondary
> fails? Is there a peak in activity on the primary?
That was my first suspect. I had the comm guys run the graphs on the
routers, and traffic was normal with no outages or spikes.
>
> Why is the secondary dying? What does the online log file contain in
> the way of diagnostics? Is there any evidence that the secondary
> can't keep up with the workload from the primary?
I bug was encountered on the secondary that was, I feel,
environmental, or not Informix's fault. The Secondary tried to
allocate memory and the system was out. My log goes from a long
series of "Checkpoint Completed" to a long series of "DR: Server
State Incompatible". This happened on Saturday at 2:00 PM.
I'm fine with the secondary crashing, it should have, it needed more
RAM to keep up with processing from the primary and couldn't get it.
It's what happened on the primary that I find hard to believe. An
hour after the secondary goes into it's incompatible state my log on
the primary if filled with a series of:
DR: Primary server connected
DR: Receive error
DR: Failure recovery error (2)
DR: Turned off on primary server
Checkpoint Completed: duration was 0 seconds.
Every 5 seconds for the next 32 hours! Then at a 01:30 Monday, I get
an assert fail:
01:30:28 Assert Failed: No Exception Handler
01:30:28 Informix Dynamic Server Version 7.31.UC6X2
01:30:28 Who: Session(72650, informix@, 0, -1015616584)
Thread(68634, dr_prsend, c3745374, 1)
File: mtex.c Line: 446
01:30:28 Results: Exception Caught. Type: MT_EX_OS, Context: mem
01:30:28 Action: Please notify Informix Technical Support.
01:30:28 stack trace for pid 1994 written to /hpkme/tmpi/af.1002e694
01:30:37 See Also: /hpkme/tmpi/af.1002e694, shmem.1002e694.0
01:30:38 Error writing '/hpkme/tmpi/shmem.1002e694.0' errno = 28
01:30:38 mtex.c, line 446, thread 68634, proc id 1994, No Exception
Handler.
01:30:38 PANIC: Attempting to bring system down
Darnit. Informix support analyzed the .af and came up with the
following explanation:
161050 REGRESSION TO FIX FOR BUG 156250 CAN CAUSE SEGV ASSERT FAIL VIA
LOGWAKEUP
134945 RACE CONDITION BETWEEN POLL THREAD AND SQLEXEC THREAD CAUSES
SERVER TO CRASH
The gist of it is that a race condition occurs between a thread trying
to talk to the DR box.
Shouldn't this be handled by a seperate process altogether so that if
you try to connect and fail you could just check for that from the
calling process? It would seem that this method actually causes
problems rather than fixing them.
>
> > Part of me says that there is no way the primary should be susceptible
> > to changes in state on the secondary. Instead of bullet-proofing your
> > site, you are creating another point of failure. HDR has been the
> > cause of three D's so far and we're ready to just scrap it and say if
> > we truly want uptime we have to do away with HDR, or watch the
> > secondary like a hawk.
>
> It depends on what you mean by the 'the primary should [not] be
> susceptible to changes in state on the secondary'. The primary is
> aware of the secondary; it keeps an eye on whether the secondary is
> still there at the other end of the network.
But the condition above, where the secondary barfs and subsequently
the primary dies with it, just dosen't make sense to me.
>
> > We are running the patched version because we had two outages
> > regarding HDR 7.31.UC6. Informix support offered us this patch but we
> > had the same issues and now they are recommending going to 7.31.UD7.
> >
> > Has anyone else had this kind of experience?
>
> Not me, personally - no, I've not had that kind of experience. If
> Madison chips in with alternative suggestions, take his advice where
> it conflicts with mine; he's the guru on ER and HDR. Not that I've
> given much advice - more questions than answers.
Thanks!
[cutting] > > How reliable is your T1 line? Is the data able to get across with all > the other traffic also on the connection? What is the volume of data > being shipped across? Is there any consistency to when the secondary > fails? Is there a peak in activity on the primary? [cutting] We had a Cisco Pix periodically shutting down the informix port, the immediate fix was to reboot the firewall, long term there was firmware upgrade -- Paul Watson # Oninit Ltd # Growing old is mandatory Tel: +44 1436 672201 # Growing up is optional Fax: +44 1436 678693 # Mob: +44 7818 003457 # www.oninit.com #
Can you give us the online log for the secondary?