Enterprise replication error
Posted in 2004
Topics: High Availability & Replication, Storage & Space Management, Logging & Checkpoints, Platform-Specific Issues
I have Informix 9.4 running on two AIX machines (A and B) with
replication defined between them.
They have been working smoothly for several weeks.
A few days ago, one of the systems (B) crashed due to disk failure.
When B came backe up, all the primaary chunks were off line. After I
they were brought on line, things seemed to be OK. I restarted the
replication a few days later (today) and discovered severe error (100)
on system A.
Any help explain what is wrong with machine A would be very much
appreciated.
Thnaks,
chariya
Below are output from cdr start repl, onstat -g nif, and online.log when
I started the server.
1. Bothe server are connected
SERVER ID STATE STATUS QUEUE CONNECTION CHANGED
-----------------------------------------------------------------------
class_ncdc 2 Active Local 0
class_saa 1 Active Connected 0 Mar 10 16:29:18
2. I created a new replication on machine A, but when I ran cdr stat
repl I got the error
informix@saatest3 $ cdr start repl test_replicate_repl
command failed -- fatal server error (100)
However, There is no problem start this replicate on machine B.
3. Here is onstat -g nif from A:
Informix Dynamic Server Version 9.40.UC1 -- On-Line -- Up 00:50:56
-- 60112 Kbytes
NIF anchor Block: 50b6f018
nifGState RUN
RetryTimeout 300
Detailed Site Instance Block: 50b6f220
siteId 2
siteState 256 <RUN>
siteVersion 7
siteCompress 0
siteNote 0
Send Thread 42 <CDRNsT2>
Recv Thread 43 <CDRNr2>
Connection Start (1078954158) 2004/03/10 16:29:18
Last Send (1078956128) 2004/03/10 17:02:08
Idle Timeout 30000
Flowblock Sent 0 Receive 0
NifInfo: 509b3050
Last Updated (1078954158) 2004/03/10 16:29:18
State Connected
Total sent 13 (1837 bytes)
Total recv'd 14 (1613 bytes)
Retry 0 (0 attempts)
Connected 1
Protocol asf
Proto block 507fb7e0
assoc 509f7538
state 0
signal 0
Recv Buf 0
Recv Data 0
Send Count 0
Send Avail 0
---------------------------------------
Here is onstat -g nif from machine 2
informix@dublin $ onstat -g nif 1
IBM Informix Dynamic Server Version 9.40.UC2 -- On-Line -- Up
02:50:05 -- 154704 Kbytes
NIF anchor Block: 50f0e078
nifGState RUN
RetryTimeout 300
Detailed Site Instance Block: 50e8ee90
siteId 1
siteState 257 <RUN>
siteVersion 7
siteCompress 0
siteNote 0
Send Thread 344 <CDRNsA1>
Recv Thread 345 <CDRNrA1>
Connection Start (1078954158) 2004/03/10 16:29:18
Last Send (1078957209) 2004/03/10 17:20:09
Idle Timeout <Forever>
Flowblock Sent 0 Receive 0
NifInfo: 50d75050
Last Updated (1078954158) 2004/03/10 16:29:18
State Connected
Total sent 28621 (4118888 bytes)
Total recv'd 7075 (316856 bytes)
Retry 0 (8 attempts)
Connected 4
Protocol asf
Proto block 50d1ef38
assoc 50908088
state 0
signal 0
Recv Buf 0
Recv Data 0
Send Count 0
Send Avail 0
Below is what's in the online.log
-----------------
16:54:43 Maximum server connections 2
16:59:43 Fuzzy Checkpoint Completed: duration was 0 seconds, 2 buffersnot flushed,
timestamp: 149430959.
16:59:43 Checkpoint loguniq 182, logpos 0x2362058, timestamp: 149430959
16:59:43 Maximum server connections 2
17:00:49 CDR GC: operation grouper start replicate failed (error 1).
17:00:49 CDR GC peer request failed: command: start repl, error 100,CDR server 2
17:01:27 CDR GC peer request failed: command: stop repl, error 30, CDRserver 2
17:01:38 CDR GC: operation grouper start replicate failed (error 1).
17:01:45 CDR GC: operation grouper start replicate failed (error 1).
17:01:48 CDR GC: operation grouper start replicate failed (error 1).
17:02:09 CDR GC: operation grouper start replicate failed (error 1).
17:02:09 CDR GC peer request failed: command: start repl, error 100,CDR server 2
17:04:44 Fuzzy Checkpoint Completed: duration was 0 seconds, 7 buffersnot flushed,
timestamp: 149433132.
17:04:44 Checkpoint loguniq 182, logpos 0x2376100, timestamp: 149433132
17:04:44 Maximum server connections 3
17:09:44 Fuzzy Checkpoint Completed: duration was 0 seconds, 7 buffersnot flushed,
timestamp: 149433143.
17:09:44 Checkpoint loguniq 182, logpos 0x2377094, timestamp: 149433143
---------------
Chariya Peterson <Chariya.Peterson@noaa.gov> wrote in message news:<404F95BB.5E4C522F@noaa.gov>...
> I have Informix 9.4 running on two AIX machines (A and B) with
> replication defined between them.
> They have been working smoothly for several weeks.
>
> A few days ago, one of the systems (B) crashed due to disk failure.
> When B came backe up, all the primaary chunks were off line. After I
> they were brought on line, things seemed to be OK. I restarted the
> replication a few days later (today) and discovered severe error (100)
> on system A.
>
> Any help explain what is wrong with machine A would be very much
> appreciated.
> Thnaks,
> chariya
>
>
> Below are output from cdr start repl, onstat -g nif, and online.log when
> I started the server.
>
>
>
>
> 1. Bothe server are connected
> SERVER ID STATE STATUS QUEUE CONNECTION CHANGED
> -----------------------------------------------------------------------
> class_ncdc 2 Active Local 0
> class_saa 1 Active Connected 0 Mar 10 16:29:18
>
>
> 2. I created a new replication on machine A, but when I ran cdr stat
> repl I got the error
> informix@saatest3 $ cdr start repl test_replicate_repl
> command failed -- fatal server error (100)
>
> However, There is no problem start this replicate on machine B.
>
>
> 3. Here is onstat -g nif from A:
> Informix Dynamic Server Version 9.40.UC1 -- On-Line -- Up 00:50:56
> -- 60112 Kbytes
>
> NIF anchor Block: 50b6f018
> nifGState RUN
> RetryTimeout 300
>
> Detailed Site Instance Block: 50b6f220
> siteId 2
> siteState 256 <RUN>
> siteVersion 7
> siteCompress 0
> siteNote 0
> Send Thread 42 <CDRNsT2>
> Recv Thread 43 <CDRNr2>
> Connection Start (1078954158) 2004/03/10 16:29:18
> Last Send (1078956128) 2004/03/10 17:02:08
> Idle Timeout 30000
> Flowblock Sent 0 Receive 0
> NifInfo: 509b3050
> Last Updated (1078954158) 2004/03/10 16:29:18
> State Connected
> Total sent 13 (1837 bytes)
> Total recv'd 14 (1613 bytes)
> Retry 0 (0 attempts)
> Connected 1
> Protocol asf
> Proto block 507fb7e0
> assoc 509f7538
> state 0
> signal 0
> Recv Buf 0
> Recv Data 0
> Send Count 0
> Send Avail 0
>
>
> ---------------------------------------
>
> Here is onstat -g nif from machine 2
> informix@dublin $ onstat -g nif 1
>
> IBM Informix Dynamic Server Version 9.40.UC2 -- On-Line -- Up
> 02:50:05 -- 154704 Kbytes>
> NIF anchor Block: 50f0e078
> nifGState RUN
> RetryTimeout 300
>
> Detailed Site Instance Block: 50e8ee90
> siteId 1
> siteState 257 <RUN>
> siteVersion 7
> siteCompress 0
> siteNote 0
> Send Thread 344 <CDRNsA1>
> Recv Thread 345 <CDRNrA1>
> Connection Start (1078954158) 2004/03/10 16:29:18
> Last Send (1078957209) 2004/03/10 17:20:09
> Idle Timeout <Forever>
> Flowblock Sent 0 Receive 0
> NifInfo: 50d75050
> Last Updated (1078954158) 2004/03/10 16:29:18
> State Connected
> Total sent 28621 (4118888 bytes)
> Total recv'd 7075 (316856 bytes)
> Retry 0 (8 attempts)
> Connected 4
> Protocol asf
> Proto block 50d1ef38
> assoc 50908088
> state 0
> signal 0
> Recv Buf 0
> Recv Data 0
> Send Count 0
> Send Avail 0
>
>
>
>
> Below is what's in the online.log
> -----------------
>
> 16:54:43 Maximum server connections 2
> 16:59:43 Fuzzy Checkpoint Completed: duration was 0 seconds, 2 buffers> not flushed,
> timestamp: 149430959.
> 16:59:43 Checkpoint loguniq 182, logpos 0x2362058, timestamp: 149430959
>
> 16:59:43 Maximum server connections 2
> 17:00:49 CDR GC: operation grouper start replicate failed (error 1).
> 17:00:49 CDR GC peer request failed: command: start repl, error 100,> CDR server 2
> 17:01:27 CDR GC peer request failed: command: stop repl, error 30, CDR> server 2
> 17:01:38 CDR GC: operation grouper start replicate failed (error 1).
> 17:01:45 CDR GC: operation grouper start replicate failed (error 1).
> 17:01:48 CDR GC: operation grouper start replicate failed (error 1).
> 17:02:09 CDR GC: operation grouper start replicate failed (error 1).
> 17:02:09 CDR GC peer request failed: command: start repl, error 100,> CDR server 2
> 17:04:44 Fuzzy Checkpoint Completed: duration was 0 seconds, 7 buffers> not flushed,
> timestamp: 149433132.
> 17:04:44 Checkpoint loguniq 182, logpos 0x2376100, timestamp: 149433132
>
> 17:04:44 Maximum server connections 3
> 17:09:44 Fuzzy Checkpoint Completed: duration was 0 seconds, 7 buffers> not flushed,
> timestamp: 149433143.
> 17:09:44 Checkpoint loguniq 182, logpos 0x2377094, timestamp: 149433143
> ---------------
I would suggest opening a case with the following diagnostics....
You will need to run the following as user informix...
xtrace size 30000
xtrace heavy -c XTF_CDR_GCC
xtrace on
cdr start replicate ....
<wait for the errors >
xtrace view > some_file
xtrace off
The file "some_file" will contain an internal trace of the start
replicate command and might give us some idea as to what happened.
One of the reason for this failure might be that snoopy(log reader)
and grouper components are down. The reason for this can be that ER
replay/snoopy position might have been overwritten(logwrap situation).
Please check the output of 'onstat -g ddr'. If the output shows
(--Down--) then snoopy(log reader) thread is down.
If the replay position overrun is detected while starting ER then you
will see the following messages in the server message log file.
20:07:35 CDR queuer initialization complete
20:07:35 CDR NIF Initialization complete
20:07:35 CDR Grouper: replay logical log ID (22) is not present
20:07:35 rmiProcessAdjust(200) Entry Count: 6
20:07:36 CDR Grouper Fan Out thread is aborting.
20:07:36 CDR Grouper FanOut thread is aborting.
If the replay position overrun is detected while is ER is in active
state then you will see the the following messages in the server
message log file.
15:44:32 WARNING: The replay position was overrun, data may not bereplicated.
15:44:34 DDR Log Snooping - Catchup phase started, userthreadsblocked
16:06:08 CDR Grouper Fanout interface with the Dynamic Server 2000has failed.
16:06:08 CDR Grouper FanOut thread is aborting.
16:06:09 DDR Log Snooping - Shutdown
Thanks & Regards,
Nagaraju
mpruet@comcast.net (mpruet) wrote in message news:<40676184.0403102140.6ffaec1a@posting.google.com>...
> Chariya Peterson <Chariya.Peterson@noaa.gov> wrote in message news:<404F95BB.5E4C522F@noaa.gov>...
> > I have Informix 9.4 running on two AIX machines (A and B) with
> > replication defined between them.
> > They have been working smoothly for several weeks.
> >
> > A few days ago, one of the systems (B) crashed due to disk failure.
> > When B came backe up, all the primaary chunks were off line. After I
> > they were brought on line, things seemed to be OK. I restarted the
> > replication a few days later (today) and discovered severe error (100)
> > on system A.
> >
> > Any help explain what is wrong with machine A would be very much
> > appreciated.
> > Thnaks,
> > chariya
> >
> >
> > Below are output from cdr start repl, onstat -g nif, and online.log when
> > I started the server.
> >
> >
> >
> >
> > 1. Bothe server are connected
> > SERVER ID STATE STATUS QUEUE CONNECTION CHANGED
> > -----------------------------------------------------------------------
> > class_ncdc 2 Active Local 0
> > class_saa 1 Active Connected 0 Mar 10 16:29:18
> >
> >
> > 2. I created a new replication on machine A, but when I ran cdr stat
> > repl I got the error
> > informix@saatest3 $ cdr start repl test_replicate_repl
> > command failed -- fatal server error (100)
> >
> > However, There is no problem start this replicate on machine B.
> >
> >
> > 3. Here is onstat -g nif from A:
> > Informix Dynamic Server Version 9.40.UC1 -- On-Line -- Up 00:50:56
> > -- 60112 Kbytes
> >
> > NIF anchor Block: 50b6f018
> > nifGState RUN
> > RetryTimeout 300
> >
> > Detailed Site Instance Block: 50b6f220
> > siteId 2
> > siteState 256 <RUN>
> > siteVersion 7
> > siteCompress 0
> > siteNote 0
> > Send Thread 42 <CDRNsT2>
> > Recv Thread 43 <CDRNr2>
> > Connection Start (1078954158) 2004/03/10 16:29:18
> > Last Send (1078956128) 2004/03/10 17:02:08
> > Idle Timeout 30000
> > Flowblock Sent 0 Receive 0
> > NifInfo: 509b3050
> > Last Updated (1078954158) 2004/03/10 16:29:18
> > State Connected
> > Total sent 13 (1837 bytes)
> > Total recv'd 14 (1613 bytes)
> > Retry 0 (0 attempts)
> > Connected 1
> > Protocol asf
> > Proto block 507fb7e0
> > assoc 509f7538
> > state 0
> > signal 0
> > Recv Buf 0
> > Recv Data 0
> > Send Count 0
> > Send Avail 0
> >
> >
> > ---------------------------------------
> >
> > Here is onstat -g nif from machine 2
> > informix@dublin $ onstat -g nif 1
> >
> > IBM Informix Dynamic Server Version 9.40.UC2 -- On-Line -- Up
> > 02:50:05 -- 154704 Kbytes> >
> > NIF anchor Block: 50f0e078
> > nifGState RUN
> > RetryTimeout 300
> >
> > Detailed Site Instance Block: 50e8ee90
> > siteId 1
> > siteState 257 <RUN>
> > siteVersion 7
> > siteCompress 0
> > siteNote 0
> > Send Thread 344 <CDRNsA1>
> > Recv Thread 345 <CDRNrA1>
> > Connection Start (1078954158) 2004/03/10 16:29:18
> > Last Send (1078957209) 2004/03/10 17:20:09
> > Idle Timeout <Forever>
> > Flowblock Sent 0 Receive 0
> > NifInfo: 50d75050
> > Last Updated (1078954158) 2004/03/10 16:29:18
> > State Connected
> > Total sent 28621 (4118888 bytes)
> > Total recv'd 7075 (316856 bytes)
> > Retry 0 (8 attempts)
> > Connected 4
> > Protocol asf
> > Proto block 50d1ef38
> > assoc 50908088
> > state 0
> > signal 0
> > Recv Buf 0
> > Recv Data 0
> > Send Count 0
> > Send Avail 0
> >
> >
> >
> >
> > Below is what's in the online.log
> > -----------------
> >
> > 16:54:43 Maximum server connections 2
> > 16:59:43 Fuzzy Checkpoint Completed: duration was 0 seconds, 2 buffers> > not flushed,
> > timestamp: 149430959.
> > 16:59:43 Checkpoint loguniq 182, logpos 0x2362058, timestamp: 149430959
> >
> > 16:59:43 Maximum server connections 2
> > 17:00:49 CDR GC: operation grouper start replicate failed (error 1).
> > 17:00:49 CDR GC peer request failed: command: start repl, error 100,> > CDR server 2
> > 17:01:27 CDR GC peer request failed: command: stop repl, error 30, CDR> > server 2
> > 17:01:38 CDR GC: operation grouper start replicate failed (error 1).
> > 17:01:45 CDR GC: operation grouper start replicate failed (error 1).
> > 17:01:48 CDR GC: operation grouper start replicate failed (error 1).
> > 17:02:09 CDR GC: operation grouper start replicate failed (error 1).
> > 17:02:09 CDR GC peer request failed: command: start repl, error 100,> > CDR server 2
> > 17:04:44 Fuzzy Checkpoint Completed: duration was 0 seconds, 7 bu
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g