ER Target Server Crashed (unto death)
Posted in 2014
Topics: High Availability & Replication, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
Hi IDS 9.4 on AIX 5.3 as an ER Source to two Targets: IDS 10 on Solaris 5.8 and IDS 11 on RedHat Enterprise. Following a building power issue and some aircon getting turned off the Solaris Server has crashed, been down for two days and is unlikely to be recoverable. As it no longer 'exists' with respect to ER we are unable to cdr delete repl to individual replicates not cdr delete serv the server, both just hang. Call to IBM overnight suggests a Catch22 where we need to delete a failed server but can only do so if it exists to be deleted !! Obviously I am beginning to run out of logical log space and need to drop this server from the ER environment P.D.Q. before it takes down my OLTP system. I would like to manage it without interrupting the other ER (or HDR in which the AIX server is the Primary). Any and all suggestion gratefully received. My best thought so far is that I need to stop both engines, start them without ER, drop the syscdr databases, stop and restart both engines with ER and then redefine the two servers and all replicates. This is possible as I have scripts available, but as this Source is 24 x 7 I would like to avoid if possible. Also I am on hols the other side of the world and will need to do this all very remotely or by guiding 'intelligent' hands many Thanks Keith --001a1135f33a1de7a20500678cce
Hi
All resolved, onmode -yuck, oninit -D, rename syscdr to syscdrold, onmode
-yuck, oninit on both ER servers that are still alive has sorted the issue
of ER. HDR didn't flinch.
Just got to resurrect the dead now :-)
Keith
On 12 August 2014 06:00, Keith Simmons <smiley73@gmail.com> wrote:
> Hi
>
> IDS 9.4 on AIX 5.3 as an ER Source to two Targets:
> IDS 10 on Solaris 5.8 and IDS 11 on RedHat Enterprise.
> Following a building power issue and some aircon getting turned off the
> Solaris Server has crashed, been down for two days and is unlikely to be
> recoverable.
> As it no longer 'exists' with respect to ER we are unable to cdr delete
> repl to individual replicates not cdr delete serv the server, both just
> hang. Call to IBM overnight suggests a Catch22 where we need to delete a
> failed server but can only do so if it exists to be deleted !!
> Obviously I am beginning to run out of logical log space and need to drop
> this server from the ER environment P.D.Q. before it takes down my OLTP
> system. I would like to manage it without interrupting the other ER (or HDR
> in which the AIX server is the Primary). Any and all suggestion gratefully
> received.
>
> My best thought so far is that I need to stop both engines, start them
> without ER, drop the syscdr databases, stop and restart both engines with
> ER and then redefine the two servers and all replicates.
>
> This is possible as I have scripts available, but as this Source is 24 x 7
> I would like to avoid if possible. Also I am on hols the other side of the
> world and will need to do this all very remotely or by guiding
> 'intelligent' hands
>
> many Thanks
>
> Keith
>
> --001a1135f33a1de7a20500678cce
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--001a11c300b81a71ce05006be17b
Any non-leaf node can be used to administer ER. That includes commands =
such
as change replicate and delete server. While it is true that to comple=
tely
remove node B from a three node domain of nodes A, B, and C; if node B =
is
down for an extended time and you want to remove references to it from =
node
A and C, then you should simply execute 'CDR delete server' on node A.=
Once you repair what is preventing node B from starting, then you would=
execute 1) oninit -D followed by 2) cdr delete server --force on node B=
.
I would never suggest that anyone manually try to drop ER by renaming o=
r
dropping the syscdr database because that will not completely remove al=
l of
the objects that ER use, and will cause subsequent issues.
Sent from my iPad
> On Aug 11, 2014, at 11:01 PM, "Keith Simmons" <smiley73@gmail.com> wr=
ote:
>
> Hi
>
> IDS 9.4 on AIX 5.3 as an ER Source to two Targets:
> IDS 10 on Solaris 5.8 and IDS 11 on RedHat Enterprise.
> Following a building power issue and some aircon getting turned off t=
he
> Solaris Server has crashed, been down for two days and is unlikely to=
be
> recoverable.
> As it no longer 'exists' with respect to ER we are unable to cdr dele=
te
> repl to individual replicates not cdr delete serv the server, both ju=
st
> hang. Call to IBM overnight suggests a Catch22 where we need to delet=
e a
> failed server but can only do so if it exists to be deleted !!
> Obviously I am beginning to run out of logical log space and need to =
drop
> this server from the ER environment P.D.Q. before it takes down my OL=
TP
> system. I would like to manage it without interrupting the other ER (=
or
HDR
> in which the AIX server is the Primary). Any and all suggestion
gratefully
> received.
>
> My best thought so far is that I need to stop both engines, start the=
m
> without ER, drop the syscdr databases, stop and restart both engines =
with
> ER and then redefine the two servers and all replicates.
>
> This is possible as I have scripts available, but as this Source is 2=
4 x
7
> I would like to avoid if possible. Also I am on hols the other side o=
f
the
> world and will need to do this all very remotely or by guiding
> 'intelligent' hands
>
> many Thanks
>
> Keith
>
> --001a1135f33a1de7a20500678cce
>
>
>
***********************************************************************=
********
> Forum Note: Use "Reply" to post a response in the discussion forum.=
>=
Madison
Thanks for this, while I agree that using cdr del serv is the correct way,
what do you do if it is still running after some 4 hours, apparently doing
nothing and you are strating to run out of log space ?
This was our situation. We logged a support call and it was they that said
that fact server B was down (and will not be back for some days) was
preventing the command completing. They recommended the database rename of
syscdr (although I suspect a drop would have been better as we seem to have
some structure remaining in the CDR Blob Space (would these be removed if
we now dropped the old syscdr database ?)).
We seem to be in a stable situation at the moment but need to add space to
the CDR Blob Space.
Keith
On 12 August 2014 20:46, Madison Pruet <mpruet@us.ibm.com> wrote:
> Any non-leaf node can be used to administer ER. That includes commands =
> such
> as change replicate and delete server. While it is true that to comple=
> tely
> remove node B from a three node domain of nodes A, B, and C; if node B =
> is
> down for an extended time and you want to remove references to it from =
> node
> A and C, then you should simply execute 'CDR delete server' on node A.=
>
> Once you repair what is preventing node B from starting, then you would=
>
> execute 1) oninit -D followed by 2) cdr delete server --force on node B=
> ..
>
> I would never suggest that anyone manually try to drop ER by renaming o=
> r
> dropping the syscdr database because that will not completely remove al=
> l of
> the objects that ER use, and will cause subsequent issues.
>
> Sent from my iPad
>
> > On Aug 11, 2014, at 11:01 PM, "Keith Simmons" <smiley73@gmail.com> wr=
> ote:
> >
> > Hi
> >
> > IDS 9.4 on AIX 5.3 as an ER Source to two Targets:
> > IDS 10 on Solaris 5.8 and IDS 11 on RedHat Enterprise.
> > Following a building power issue and some aircon getting turned off t=
> he
> > Solaris Server has crashed, been down for two days and is unlikely to=
> be
> > recoverable.
> > As it no longer 'exists' with respect to ER we are unable to cdr dele=
> te
> > repl to individual replicates not cdr delete serv the server, both ju=
> st
> > hang. Call to IBM overnight suggests a Catch22 where we need to delet=
> e a
> > failed server but can only do so if it exists to be deleted !!
> > Obviously I am beginning to run out of logical log space and need to =
> drop
> > this server from the ER environment P.D.Q. before it takes down my OL=
> TP
> > system. I would like to manage it without interrupting the other ER (=
> or
> HDR
> > in which the AIX server is the Primary). Any and all suggestion
> gratefully
> > received.
> >
> > My best thought so far is that I need to stop both engines, start the=
> m
> > without ER, drop the syscdr databases, stop and restart both engines =
> with
> > ER and then redefine the two servers and all replicates.
> >
> > This is possible as I have scripts available, but as this Source is 2=
> 4 x
> 7
> > I would like to avoid if possible. Also I am on hols the other side o=
> f
> the
> > world and will need to do this all very remotely or by guiding
> > 'intelligent' hands
> >
> > many Thanks
> >
> > Keith
> >
> > --001a1135f33a1de7a20500678cce
> >
> >
> >
> ***********************************************************************=
> ********
>
> > Forum Note: Use "Reply" to post a response in the discussion forum.=
>
> >=
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--001a1135f33af4009e0500776dad