HDR
Posted in 2009
A DBA planning HDR between a production and DR site (Informix 10 on HP-UX) asked six setup questions: DRAUTO, other DR ONCONFIG parameters, ontape -r vs -p, sync vs async, how to stop HDR while keeping production up, and whether restarting HDR needs a fresh restore. Art Kagel, Madison Pruet and Paul Mosser answered: set DRAUTO=0 and the DR* parameters identically on both servers; use ontape -p (physical restore only, logs come from the primary); DRINTERVAL=-1 gives SYNC (logs shipped before commit completes), otherwise ASYNC; stopping the secondary stops HDR and it restarts automatically; a new level-0 restore is only needed if the required logical logs have been overwritten, otherwise onmode -d primary plus oninit resyncs. A side question about an instance refusing onmode commands drew only the answer of killing the master CPU VP.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Backup & Restore, Server Administration
I'm getting ready to turn on HDR between our production site and our D/R
site located 7 miles away. We are running Jenzabar CX on Informix 10
and HPUX 11i v2. Our 2 servers are connected via a 20 MB pipe. I've
gone through the documentation but I have some questions below. We have
a very busy system so I'm trying to avoid little downtime on our
production site.
1. We don't want our DR server to take over automatically in any
case. If we have a planned shutdown of production, I want to manually
get the D/R site ready. Would I set the DRAUTO parameter in the
ONCONFIG file to 0 - Manual? And do I do this on the primary server
only?
2. Are there any other settings in the ONCONFIG file I need to
change? I've seen other DR parameters in the documentation, i.e.
DRIDXAUTO, DRINTERVAL, etc. And again, just on the primary server?
3. Can I use ontape -r? This is what I've been using for a full
restore on my D/R box. When beginning replication, the documentation
says to use ontape -p? Just curious what the difference is between the
two.
4. From most of the readings, I would think I would be using
asynchronous communication, correct? Is this also defined on just the
primary? What is the difference between the synchronous and
asynchronous communication?
5. Once I have replication going, there may be times where I need
to do maintenance on one or both of the servers. What would be the
process for stopping HDR on both servers but keeping my production
engine up and running? I want to develop a good procedure, if I'm out
of town or not available and for whatever reason, a backup admin needs
to turn HDR off but keep production running. My main concern is there a
need to stop the Informix engine on production?
6. Finally, let's just say I had HDR turned off for several hours
and I'm ready to turn on again. Do I need to go through the steps again
including the full tape restore or can it pickup by starting HDR again?
My boss is pressuring me to get this working since we are in hurricane
season. Right now, my plan is to run a complete restore using HP's
Ignite for the file system and then restore the Informix database from
tape. This works but is a timely process. Any help would be
appreciated. Thanks.
Andy Pistocchi
DBA / Systems Administrator
The University of Tampa
813-257-3196
See responses below:
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
On Thu, Aug 13, 2009 at 9:46 AM, ANDREW PISTOCCHI <APISTOCCHI@ut.edu> wrote:
> I'm getting ready to turn on HDR between our production site and our D/R
> site located 7 miles away. We are running Jenzabar CX on Informix 10
> and HPUX 11i v2. Our 2 servers are connected via a 20 MB pipe. I've
> gone through the documentation but I have some questions below. We have
> a very busy system so I'm trying to avoid little downtime on our
> production site.
>
> 1. We don't want our DR server to take over automatically in any
> case. If we have a planned shutdown of production, I want to manually
> get the D/R site ready. Would I set the DRAUTO parameter in the
> ONCONFIG file to 0 - Manual? And do I do this on the primary server
> only?
No, do this on both servers.
>
>
> 2. Are there any other settings in the ONCONFIG file I need to
> change? I've seen other DR parameters in the documentation, i.e.
> DRIDXAUTO, DRINTERVAL, etc. And again, just on the primary server?
Yes, you should configure these parameters as well unless you are OK with
the defaults.
>
>
> 3. Can I use ontape -r? This is what I've been using for a full
> restore on my D/R box. When beginning replication, the documentation
> says to use ontape -p? Just curious what the difference is between the
> two.
Use ontape -p. The difference is that ontape -r also restores the logical
logs or tells the engine that there are no logical logs to restore. The
ontape -p does not restore local logical logs and leaves the engine waitingfor the primary to start shipping over its logical logs. All of the logs
since the time the archive was take must be on disk (ie not yet overwritten)
on the primary in order for the secondary server initialization to succeed.
>
>
> 4. From most of the readings, I would think I would be using
> asynchronous communication, correct? Is this also defined on just the
> primary? What is the difference between the synchronous and
> asynchronous communication?
Define it on both in case you decide after a primary crash to make the
recovered primary the secondary for a time until you can find a maintenance
window to reverse them again. You use synchronous if you are using the
secondary actively for queries that must return the exact data as the
primary as soon as possible and if it is critical that the secondary be in
exact synch with the primary. In asynch mode there is a possibility that a
transaction committed on the primary will not have been transmitted to the
secondary before the primary crashed and so they will be out-of-synch when
the secondary takes over as primary. That last transaction(s) will be lost
when you have to restore the primary from the secondary because they are not
in synch and a catchup is no longer possible.
>
>
> 5. Once I have replication going, there may be times where I need
> to do maintenance on one or both of the servers. What would be the
> process for stopping HDR on both servers but keeping my production
> engine up and running? I want to develop a good procedure, if I'm out
> of town or not available and for whatever reason, a backup admin needs
> to turn HDR off but keep production running. My main concern is there a
> need to stop the Informix engine on production?
To your last: sometimes. If you have to do maintenance on the machines,
you'll stop replication and restore it afterwards. If you have to do
maintenance on the servers themselves, it will depend on what you are
doing. Some things can be done with HDR online, others cannot.
>
>
> 6. Finally, let's just say I had HDR turned off for several hours
> and I'm ready to turn on again. Do I need to go through the steps again
> including the full tape restore or can it pickup by starting HDR again?
As long as the last logical log that the secondary knows about is still
available on disk on the primary, then you will not have to perform a
restore again, the primary will ship over the logs that the secondary missed
when you reestablish HDR. If that oldest log has been overwritten on the
primary already, then yes you have to restore again. You should have enough
logical logs on an HDR server to allow the server to be offline for any
reasonable maintenance or quickly fixed hardware failure without overwriting
the logs. Say 24 to 96 hours depending on your uptime and recovery
requirements.
>
>
> My boss is pressuring me to get this working since we are in hurricane
> season. Right now, my plan is to run a complete restore using HP's
> Ignite for the file system and then restore the Informix database from
> tape. This works but is a timely process. Any help would be
> appreciated. Thanks.
>
> Andy Pistocchi
>
> DBA / Systems Administrator
>
> The University of Tampa
>
> 813-257-3196
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--001636c5acf8b7e1fb0471069ba2
=
"ANDREW =
PISTOCCHI" =
<APISTOCCHI@UT.ED =
To
U> ids@iiug.org =
Sent by: =
cc
ids-bounces@iiug. =
org Subj=
ect
HDR [16679] =
=
08/13/2009 08:46 =
AM =
=
=
Please respond to =
ids@iiug.org =
=
=
I'm getting ready to turn on HDR between our production site and our D/=
R
site located 7 miles away. We are running Jenzabar CX on Informix 10
and HPUX 11i v2. Our 2 servers are connected via a 20 MB pipe. I've
gone through the documentation but I have some questions below. We have=
a very busy system so I'm trying to avoid little downtime on our
production site.
1. We don't want our DR server to take over automatically in any
case. If we have a planned shutdown of production, I want to manually
get the D/R site ready. Would I set the DRAUTO parameter in the
ONCONFIG file to 0 - Manual? And do I do this on the primary server
only?
Set DRAUTO=3D0 on both nodes.
2. Are there any other settings in the ONCONFIG file I need to
change? I've seen other DR parameters in the documentation, i.e.
DRIDXAUTO, DRINTERVAL, etc. And again, just on the primary server?
No - HDR will check several things to make sure that the onconfig file =
is
the same on the primary and secondary servers. The DR*** parapmeters a=
re
included.
3. Can I use ontape -r? This is what I've been using for a full
restore on my D/R box. When beginning replication, the documentation
says to use ontape -p? Just curious what the difference is between the
two.
ontape -p is 'physical restore' That is what must be used to start HDR=
.The log transfer will
be done from the primary to the secondary once onmode -d secondary is r=
un.
4. From most of the readings, I would think I would be using
asynchronous communication, correct? Is this also defined on just the
primary? What is the difference between the synchronous and
asynchronous communication?
I think that you are confusing communication protocols with HDR SYNC/AS=
YNC
mode.
With HDR SYNC mode, the log transfer to the secondary occurs as part of=
the
log flush.
With HDR Async mode the log buffer transfer occurs asynchronously to th=
e
log flush.
To use SYNC mode set DRINTERVAL to -1. To use ASYNC mode set DRINTERVA=
L to
the max time
that a buffer flushed buffer is allowed to wait before transmitting to =
the
secondary.
5. Once I have replication going, there may be times where I need
to do maintenance on one or both of the servers. What would be the
process for stopping HDR on both servers but keeping my production
engine up and running? I want to develop a good procedure, if I'm out
of town or not available and for whatever reason, a backup admin needs
to turn HDR off but keep production running. My main concern is there a=
need to stop the Informix engine on production?
If you break the connection (i.e. bring the secondary down) then HDR wi=
ll
automatically stop.
When you bring both nodes back up, then it will automatically restart.
6. Finally, let's just say I had HDR turned off for several hours
and I'm ready to turn on again. Do I need to go through the steps again=
including the full tape restore or can it pickup by starting HDR again?=
It might be easier to start from scratch, depending on if the logs have=
wrapped and/or
the size of the database. If the logs have not wrapped, then all you w=
ould
have to do
is to simply bring up the secondary.
My boss is pressuring me to get this working since we are in hurricane
season. Right now, my plan is to run a complete restore using HP's
Ignite for the file system and then restore the Informix database from
tape. This works but is a timely process. Any help would be
appreciated. Thanks.
Andy Pistocchi
DBA / Systems Administrator
The University of Tampa
813-257-3196
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
> > I'm getting ready to turn on HDR between our production site and our
> D/R
> > site located 7 miles away. We are running Jenzabar CX on Informix 10
> > and HPUX 11i v2. Our 2 servers are connected via a 20 MB pipe. I've
> > gone through the documentation but I have some questions below. We
> have
> > a very busy system so I'm trying to avoid little downtime on our
> > production site.
> >
> > 1. We don't want our DR server to take over automatically in any
> > case. If we have a planned shutdown of production, I want to manually
> > get the D/R site ready. Would I set the DRAUTO parameter in the
> > ONCONFIG file to 0 - Manual? And do I do this on the primary server
> > only?
>
> No, do this on both servers.
>
<< snipped >>
All good stuff from Art and Madison, of course. I'll just throw in my .02 as
someone that has used & admin'ed HDR for about 10 years now.
We run our HDR with asynch. Yes, I know that this *can* be risky, technically
speaking. But, even with many planned and unplanned failovers, we have never
lost a transaction (knock, knock), AFAIK. The biggest reason that we use
asynch is that response time is *extremely* critical in this application. In
synch mode, the data must be committed on the secondary before the commit is
complete on the primary (Madison, pls correct me if I'm wrong here). Our
primary and secondary are hundreds of miles apart, with lots of network stuff
between them. If a network blip caused an outage of replication, with synch
mode, it would also hold up txns on the primary, something that the business
cannot afford. That's our situation, tho -- YMMV. Also, keep in mind that
checkpoints are ALWAYS synch'd -- that is where the primary and secondary make
sure that they are in synch, regardless of your DRINTERVAL setting.
To start/stop replication, I always take down the secondary first via 'onmode
-yuck'. Then I switch the primary to standard mode via 'onmode -d standard'.
When I am ready to restart replication, I first look at which logical log the
primary is currently at. If the logs have wrapped around since replication was
stopped, such that some of the logs that the secondary would need are now only
available via log backups, then I usually just run another level-0 archive on
the primary (we use ontape to disk files), do the physical restore (ontape -p)
of that archive onto the secondary, and then restart replication again from
scratch. However, if the logs have *NOT* wrapped, then you can very simply
restart replication by issuing the 'onmode -d primary
<secondary_dbserver_name>', and then simply bring the secondary server back up
via 'oninit -v'. They will find each other and re-synch automagically.
HTH,
Paul M.
Almost right.
With SYNC mode, we don't have to wait until the transaction has been
applied on the secondary, but we do wait until all of the logs of that
transaction have been shipped to the secondary.
=
"mosserp@wellsfar =
go.com" =
<mosserp@wellsfar =
To
go.com> ids@iiug.org =
Sent by: =
cc
ids-bounces@iiug. =
org Subj=
ect
RE: HDR [16690] =
=
08/14/2009 12:57 =
PM =
=
=
Please respond to =
ids@iiug.org =
=
=
> > I'm getting ready to turn on HDR between our production site and ou=
r
> D/R
> > site located 7 miles away. We are running Jenzabar CX on Informix 1=
0
> > and HPUX 11i v2. Our 2 servers are connected via a 20 MB pipe. I've=
> > gone through the documentation but I have some questions below. We
> have
> > a very busy system so I'm trying to avoid little downtime on our
> > production site.
> >
> > 1. We don't want our DR server to take over automatically in any
> > case. If we have a planned shutdown of production, I want to manual=
ly
> > get the D/R site ready. Would I set the DRAUTO parameter in the
> > ONCONFIG file to 0 - Manual? And do I do this on the primary server=
> > only?
>
> No, do this on both servers.
>
<< snipped >>
All good stuff from Art and Madison, of course. I'll just throw in my .=
02
as
someone that has used & admin'ed HDR for about 10 years now.
We run our HDR with asynch. Yes, I know that this *can* be risky,
technically
speaking. But, even with many planned and unplanned failovers, we have
never
lost a transaction (knock, knock), AFAIK. The biggest reason that we us=
e
asynch is that response time is *extremely* critical in this applicatio=
n.
In
synch mode, the data must be committed on the secondary before the comm=
it
is
complete on the primary (Madison, pls correct me if I'm wrong here). Ou=
r
primary and secondary are hundreds of miles apart, with lots of network=
stuff
between them. If a network blip caused an outage of replication, with s=
ynch
mode, it would also hold up txns on the primary, something that the
business
cannot afford. That's our situation, tho -- YMMV. Also, keep in mind th=
at
checkpoints are ALWAYS synch'd -- that is where the primary and seconda=
ry
make
sure that they are in synch, regardless of your DRINTERVAL setting.
To start/stop replication, I always take down the secondary first via
'onmode
-yuck'. Then I switch the primary to standard mode via 'onmode -d
standard'.
When I am ready to restart replication, I first look at which logical l=
og
the
primary is currently at. If the logs have wrapped around since replicat=
ion
was
stopped, such that some of the logs that the secondary would need are n=
ow
only
available via log backups, then I usually just run another level-0 arch=
ive
on
the primary (we use ontape to disk files), do the physical restore (ont=
ape
-p)
of that archive onto the secondary, and then restart replication again =
from
scratch. However, if the logs have *NOT* wrapped, then you can very sim=
ply
restart replication by issuing the 'onmode -d primary
<secondary_dbserver_name>', and then simply bring the secondary server =
back
up
via 'oninit -v'. They will find each other and re-synch automagically.
HTH,
Paul M.
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
Paul, with you exeprience with HDR, have you hit problems where the instance
will not accept an onmode command? If so how did you fix it short of killing
the master daemon?
> Paul, with you exeprience with HDR, have you hit problems where the
> instance
> will not accept an onmode command? If so how did you fix it short of
> killing
> the master daemon?
>
<< snipped >>
Not sure which "Paul" you are asking, but this one will take a stab.
Yes, we have occasionally hit that problem. Very rarely, but we have come
across it. The only solution that I have found is to kill the master daemon,
i.e., the vp for cpu #1 from the 'onstat -g glo' listing.
HTH,
Paul M.
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g