ER error after restore -DDR down
Posted in 2005
After restoring a level 0 backup on IDS 9.20 (SCO UnixWare), Enterprise Replication wouldn't start: the log showed "CDR Grouper: replay logical log ID (95142) is not present" and the Grouper/FanOut threads aborted, because the log holding ER's replay position was no longer on disk or in the archive. Suggestions were to restore an older archive and roll forward so the missing log is available, or call support. Madison Pruet explained the cause: the backups were taken in quiescent mode, where ER isn't running, so the engine didn't know the replay position and omitted the logs ER needs; backing up while online avoids it. The poster worked around the immediate failure by redefining all replication.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Backup & Restore, Logging & Checkpoints
Informix Dynamic Server 9.20
Sco Unixware 7.1.x (Intel)
I have the following problem with ER.
After a level 0 backup restore on ER root server DDR cannot start;
the messages in online.log are:
10:38:34 CDR GC: catalog recovery complete
10:38:34 CDR queuer initialization complete
10:38:34 CDR Grouper: replay logical log ID (95142) is not present
10:38:35 CDR Grouper Fan Out thread is aborting.
10:38:35 CDR Grouper FanOut thread is aborting.
Here is $onstat -l :
a49b13c 41 F------ 0 409a81 1000 0 0.00
a49b158 42 F------ 0 409e69 1000 0 0.00
a49b174 43 U-B---- 95143 40a251 1000 1000 100.00
a49b190 44 U-B---- 95144 40a639 1000 1000 100.00
..................................................
a49b740 96 U-B---- 95196 417159 1000 1000 100.00
a49b75c 97 U-B---- 95197 417541 1000 1000 100.00
a49b778 98 U---C-L 95198 417929 1000 789 78.90
a49b794 99 F------ 0 417d11 1000 0 0.00
a49b7b0 100 F------ 0 4180f9 1000 0 0.00
We have about 80 servers and we restored all of them Sunday evening (
we backups the server friday evening and restore sunday evening),
after a series of tests done during the last week-end. After restore,
the error appears on many servers, not only on root server. Another
point
is: there are specific servers (with the same/likewise configuration)
having
NO issue what-so-ever to start the replication (i.e. they are working
just
fine).
What can we do to solve the issue, or if it is not possible, to avoid
the
same type of errors from occurring again?
Regards,
Cristi.
__________________________________________________
Do You Yahoo!?
Tired of spam? Yahoo! Mail has the best spam protection around
http://mail.yahoo.com
your in deep... trouble 10:38:34 CDR Grouper: replay logical log ID (95142) is not present --------onstat -l------------- shows------- a49b158 42 F------ 0 a49b174 43 U-B---- 95143 means you have not got log 95142 on disk so ER can not start from where it is suppose to start. i guess this happened after clearing the logs during the restore. if possible try and restore from an earlier archive and rollforward all the logs so 95142 ends up on disk (try test system not connected to your network please!! ) or as tech support to hack your system... Ask the experts: Madison: comments please; in case i am way off sounds like warm restore enterprise replication feature wanted.... Superboer.
Was ER active when the level 0 backup was taken?
"cristizaharioiu" <cristizaharioiu@gmail.com> wrote in message
news:1119263040.696866.230280@g44g2000cwa.googlegroups.com...
> Informix Dynamic Server 9.20
> Sco Unixware 7.1.x (Intel)
>
> I have the following problem with ER.
> After a level 0 backup restore on ER root server DDR cannot start;
> the messages in online.log are:
>
> 10:38:34 CDR GC: catalog recovery complete
> 10:38:34 CDR queuer initialization complete
> 10:38:34 CDR Grouper: replay logical log ID (95142) is not present
> 10:38:35 CDR Grouper Fan Out thread is aborting.
> 10:38:35 CDR Grouper FanOut thread is aborting.>
>
> Here is $onstat -l :
>
> a49b13c 41 F------ 0 409a81 1000 0 0.00
> a49b158 42 F------ 0 409e69 1000 0 0.00
> a49b174 43 U-B---- 95143 40a251 1000 1000 100.00
> a49b190 44 U-B---- 95144 40a639 1000 1000 100.00
>
> ..................................................
>
> a49b740 96 U-B---- 95196 417159 1000 1000 100.00
> a49b75c 97 U-B---- 95197 417541 1000 1000 100.00
> a49b778 98 U---C-L 95198 417929 1000 789 78.90
> a49b794 99 F------ 0 417d11 1000 0 0.00
> a49b7b0 100 F------ 0 4180f9 1000 0 0.00
>
>
> We have about 80 servers and we restored all of them Sunday evening (
> we backups the server friday evening and restore sunday evening),
> after a series of tests done during the last week-end. After restore,
> the error appears on many servers, not only on root server. Another
> point
> is: there are specific servers (with the same/likewise configuration)
> having
> NO issue what-so-ever to start the replication (i.e. they are working
> just
> fine).
>
> What can we do to solve the issue, or if it is not possible, to avoid
> the
> same type of errors from occurring again?
>
> Regards,
> Cristi.
>
> __________________________________________________
> Do You Yahoo!?
> Tired of spam? Yahoo! Mail has the best spam protection around
> http://mail.yahoo.com
>
Madison,
All level 0 backups are done in quiescent mode. ER was active at the
backup's time.
Yesterday I redefined all replication and now it's ok.
But how can I do to avoid this problem?
Why few servers found replay logical log on archive and almost all
servers didn't ? Can I do something else trying to have replay
position close to current log position?
It would be better to make the archive on-line ?
Madison Pruet wrote:
> Was ER active when the level 0 backup was taken?
>
>
> "cristizaharioiu" <cristizaharioiu@gmail.com> wrote in message
> news:1119263040.696866.230280@g44g2000cwa.googlegroups.com...
> > Informix Dynamic Server 9.20
> > Sco Unixware 7.1.x (Intel)
> >
> > I have the following problem with ER.
> > After a level 0 backup restore on ER root server DDR cannot start;
> > the messages in online.log are:
> >
> > 10:38:34 CDR GC: catalog recovery complete
> > 10:38:34 CDR queuer initialization complete
> > 10:38:34 CDR Grouper: replay logical log ID (95142) is not present
> > 10:38:35 CDR Grouper Fan Out thread is aborting.
> > 10:38:35 CDR Grouper FanOut thread is aborting.>
> >
> > Here is $onstat -l :
> >
> > a49b13c 41 F------ 0 409a81 1000 0 0.00
> > a49b158 42 F------ 0 409e69 1000 0 0.00
> > a49b174 43 U-B---- 95143 40a251 1000 1000 100.00
> > a49b190 44 U-B---- 95144 40a639 1000 1000 100.00
> >
> > ..................................................
> >
> > a49b740 96 U-B---- 95196 417159 1000 1000 100.00
> > a49b75c 97 U-B---- 95197 417541 1000 1000 100.00
> > a49b778 98 U---C-L 95198 417929 1000 789 78.90
> > a49b794 99 F------ 0 417d11 1000 0 0.00
> > a49b7b0 100 F------ 0 4180f9 1000 0 0.00
> >
> >
> > We have about 80 servers and we restored all of them Sunday evening (
> > we backups the server friday evening and restore sunday evening),
> > after a series of tests done during the last week-end. After restore,
> > the error appears on many servers, not only on root server. Another
> > point
> > is: there are specific servers (with the same/likewise configuration)
> > having
> > NO issue what-so-ever to start the replication (i.e. they are working
> > just
> > fine).
> >
> > What can we do to solve the issue, or if it is not possible, to avoid
> > the
> > same type of errors from occurring again?
> >
> > Regards,
> > Cristi.
> >
> > __________________________________________________
> > Do You Yahoo!?
> > Tired of spam? Yahoo! Mail has the best spam protection around
> > http://mail.yahoo.com
> >
cristizaharioiu wrote: > Madison, > > All level 0 backups are done in quiescent mode. ER was active at the > backup's time. > <snip> If the engine is in quiescent mode, then ER is effectively down. i.e. there are no ER threads running. Why do you bring the engine to quiescent mode to do the backups?
I'm pretty sure that in your version that the replay position is not known
by the engine when ER is not running.
In a backup is taken, we also backup the logs that are required to bring the
system to a consistant state as of the backup checkpoint. This is necessary
to bring the engine into an online mode (i.e. roll back any active
transactions.) This will be needed if you chose not to recover logs as part
of the restore.
In a system where ER is running, we also check to see where the replay
positiion is, and ensure that we will backup the logs that ER will need to
recover as well. Since ER was not running because you were in quiescent
mode, those logs were not backed up.
We might be able to tighten this up a bit, but that would not be going into
9.20. So it is best for you not to go into a quiescent mode while making
the backups.
M.P.
"cristizaharioiu" <cristizaharioiu@gmail.com> wrote in message
news:1119425352.850652.298140@o13g2000cwo.googlegroups.com...
> Madison,
>
> All level 0 backups are done in quiescent mode. ER was active at the
> backup's time.
>
> Yesterday I redefined all replication and now it's ok.
> But how can I do to avoid this problem?
>
> Why few servers found replay logical log on archive and almost all
> servers didn't ? Can I do something else trying to have replay
> position close to current log position?
>
> It would be better to make the archive on-line ?
>
> Madison Pruet wrote:
> > Was ER active when the level 0 backup was taken?
> >
> >
> > "cristizaharioiu" <cristizaharioiu@gmail.com> wrote in message
> > news:1119263040.696866.230280@g44g2000cwa.googlegroups.com...
> > > Informix Dynamic Server 9.20
> > > Sco Unixware 7.1.x (Intel)
> > >
> > > I have the following problem with ER.
> > > After a level 0 backup restore on ER root server DDR cannot start;
> > > the messages in online.log are:
> > >
> > > 10:38:34 CDR GC: catalog recovery complete
> > > 10:38:34 CDR queuer initialization complete
> > > 10:38:34 CDR Grouper: replay logical log ID (95142) is not present
> > > 10:38:35 CDR Grouper Fan Out thread is aborting.
> > > 10:38:35 CDR Grouper FanOut thread is aborting.> >
>
> > >
> > > Here is $onstat -l :
> > >
> > > a49b13c 41 F------ 0 409a81 1000 0
0.00
> > > a49b158 42 F------ 0 409e69 1000 0
0.00
> > > a49b174 43 U-B---- 95143 40a251 1000 1000
100.00
> > > a49b190 44 U-B---- 95144 40a639 1000 1000
100.00
> > >
> > > ..................................................
> > >
> > > a49b740 96 U-B---- 95196 417159 1000 1000
100.00
> > > a49b75c 97 U-B---- 95197 417541 1000 1000
100.00
> > > a49b778 98 U---C-L 95198 417929 1000 789
78.90
> > > a49b794 99 F------ 0 417d11 1000 0
0.00
> > > a49b7b0 100 F------ 0 4180f9 1000 0
0.00
> > >
> > >
> > > We have about 80 servers and we restored all of them Sunday evening (
> > > we backups the server friday evening and restore sunday evening),
> > > after a series of tests done during the last week-end. After restore,
> > > the error appears on many servers, not only on root server. Another
> > > point
> > > is: there are specific servers (with the same/likewise configuration)
> > > having
> > > NO issue what-so-ever to start the replication (i.e. they are working
> > > just
> > > fine).
> > >
> > > What can we do to solve the issue, or if it is not possible, to avoid
> > > the
> > > same type of errors from occurring again?
> > >
> > > Regards,
> > > Cristi.
> > >
> > > __________________________________________________
> > > Do You Yahoo!?
> > > Tired of spam? Yahoo! Mail has the best spam protection around
> > > http://mail.yahoo.com
> > >
>