Enterprise Replication - Questions on monitoring Replication
Posted in 1999
Topics: High Availability & Replication, Server Administration, Security, Permissions & Auditing, Platform-Specific Issues, Jobs, Consulting & Announcements
Hello all, We have been using ER now for around 3-4 months and have found the manuals and information supplied rather limiting. We have two sets of two Sun Solaris boxes running replication (one set as test, the second is live). We have managed to set up the servers and replicates and have by trial and (many, many) error(s) managed to stumble along so far, but we now have some questions on monitoring and using replication. Questions: 1. How do we monitor what data is still unsent on a server? (How do we know when both sides are up to date?) (We have been selecting bytesqued from sysmaster:syscdrq but some tables seem to always have a bytesqued count.) 2. How do we monitor the volume of data replicated (in rows / bytes) for each table. 3. We have both Primary / Target and Update anywhere replicates defined and have taken the following approach: - Set conflict resolution to ignore for Primary / Target replicates because there should not be any update conflict. - Set conflict to timestamp and add CRCOLS for Update anywhere replicates. - When using timestamp conflict resolution the times on both servers become critical. 4. We have turned on both ATS and RIS spooling and check the directories for errors, what else should / can we do to check for replication errors. 5. We use cdr suspend on both servers when we want to take one server off line (for DBA or reboots etc...), we presume cdr suspend keeps the data queued (we have tested this and it seems correct, the manual contradicts itself in different sections). What steps should we follow when we want to reboot a server? 6. We limit the use of cdr stop, this is only used when we have exclusive use of both servers (usually when we need to do DBA work on replicated tables). 7. What steps should we follow to remove replication between the two servers? (especially if we want to run our cdr define server scripts again later). 8. How do we know that both servers are in sync (especially for proof for the Auditors) should we add CRCOLS and timestamp conflict resolution just so we can check the tables are identical? If yes, then I guess this means we would have to change all Primary / Target replicates to update anywhere. (When trying to set a P/T replicate to timestamp conflict I get the following error from the cdr define replicate command failed -- conflict mode for replicate not ignore (69)). 9. What other monitoring / error checking should we do? 10. What standard commands are others using? TIA and Regards, Stephen
First of all --- I hope that you are on 7.31 or are considering moving to 7.31.
For ER it is proving to be far superior to prior versions.
Secondly, we are currently working on a web-based monitor tool to replace the
old ER GUI. This should make monitoring much easer than before.
See below for specific comments...
Stephen Manga wrote:
> Hello all,
>
> We have been using ER now for around 3-4 months and have found the manuals
> and information supplied rather limiting.
>
> We have two sets of two Sun Solaris boxes running replication (one set as
> test, the second is live).
>
> We have managed to set up the servers and replicates and have by trial and
> (many, many) error(s) managed to stumble along so far, but we now have some
> questions on monitoring and using replication.
>
> Questions:
>
> 1. How do we monitor what data is still unsent on a server? (How do we
> know when both sides are up to date?)
> (We have been selecting bytesqued from sysmaster:syscdrq but
> some tables seem to always have a bytesqued count.)
The best way is to dump the TRG send queue. In 7.31 this is done by "onstat -g
rqm full" and on prior versions by "onstat -g que SQlock". Each transaction is
identified by the address in the log file of that transactions commit point.
Thus if a transaction's commit record is in unique log file 17 at address 0x422,
ER's identifier for that transaction would be <group_id_number>/17/0x422/0.
Each transaction within the queue has a "need ack" bit map of the servers that
have not yet ACKed the receipt of that transaction. Thus by examining the need
ack bit map and the output of onstat -g cat, you can determine which servers
have not yet ACKed the transaction.
We maintain a "progress table" on each server which contains the most recently
ACKed transaction's ER key (<cdrid
#>/<unique_log_number>/<commit_logpos>/<counter>). There is one row for each
defined replicate for each target server. By examining that you can also
determine "where replication is currently".
However, I think that the easiest way to do this is to actually replicate data.
Create a table as follows:
create database er_db with log;
create table er_status (
serverid int primary key,
servertime datetime ) with CRCOLS lock mode row;
Now defined this to be replicated on all servers - update anywhere.
Then create a cron job which once every 10 minutes or so runs the following
script. This needs to run on each server.
------------------------------------------------
dbaccess er_db <<!
update er_status set servertime = current;!
------------------------------------------------
Since the table er_db is replicated to all servers and the cron job is run on
all servers, by examining the table er_db on any server, you can easily
determine what "time" you've replicated to.
> 2. How do we monitor the volume of data replicated (in rows / bytes)
> for each table.
onstat -g dss
> 3. We have both Primary / Target and Update anywhere replicates defined
> and have taken the following approach:
> - Set conflict resolution to ignore for Primary / Target replicates
> because there should not be any update conflict.
> - Set conflict to timestamp and add CRCOLS for Update anywhere
> replicates.
> - When using timestamp conflict resolution the times on both servers
> become critical.
That's what I would do.
> 4. We have turned on both ATS and RIS spooling and check the
> directories for errors, what else should / can we do to check for
> replication errors.
Monitor the message log file.
> 5. We use cdr suspend on both servers when we want to take one server
> off line (for DBA or reboots etc...), we presume cdr suspend keeps the data
> queued (we have tested this and it seems correct, the manual contradicts
> itself in different sections). What steps should we follow when we want to
> reboot a server?
"cdr suspend" suspends communication to remote servers. The data is queued, but
not sent. "cdr stop" simply stops the cdr threads. When "cdr start" is run,
then the log snooping will resume with the replay point that was current when
the cdr stop was run, IF THAT LOG FILE STILL EXISTS. With a simple reboot, you
should not have to do either.
> 6. We limit the use of cdr stop, this is only used when we have
> exclusive use of both servers (usually when we need to do DBA work on
> replicated tables).
Again, cdr start will resume snooping of the log files with the replay point
that was current when cdr stop was run. However, while the ER threads are not
running, there is nothing to prevent a log wrap. Thus if cdr start is run, it
is possible that you will not be able to start snoopy because the log file is no
longer on the system.
> 7. What steps should we follow to remove replication between the two
> servers? (especially if we want to run our cdr define server scripts again
> later).
cdr delete server
I hate to even mention the next stuff. There is an environmental variable,
CDRBLOCKOUT, which many customers have used to prevent ER threads from starting
up. They have run the server with this set and then manually dropped the
database syscdr. THIS IS A VERY BAD PRACTICE AND IS TOTALLY UNSUPPORTED!!!!!
If this is done, then it is very possible that you will end up with corruption
in sysmaster as well as in your user databases. This is because the practice
will NOT to a complete cleanup of ER in the instance. This can (and often does)
cause later problems with the instance. PLEASE DO NOT USE THIS TECHNIQUE.
> 8. How do we know that both servers are in sync (especially for proof
> for the Auditors) should we add CRCOLS and timestamp conflict resolution
> just so we can check the tables are identical? If yes, then I guess this
> means we would have to change all Primary / Target replicates to update
> anywhere. (When trying to set a P/T replicate to timestamp conflict I get
> the following error from the cdr define replicate command failed -- conflict
> mode for replicate not ignore (69))
We are addressing this issue in 9.3.
> .
> 9. What other monitoring / error checking should we do?
> 10. What standard commands are others using?
>
> TIA and Regards,
>
> Stephen
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g