RE: Blocked:DDR
Posted in 2000
Topics: High Availability & Replication, Backup & Restore, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues, Versions, Editions & End-of-Life
That's what I had thought initially, but when I looked at onstat -g act
only the polling threads were active.
It couldn't have been that easy, experiencing a problem that I had just
read about in the news groups.
It's seems to be fixed, I just wish I knew what caused the problem. There
is not a whole lot of activity from the Primary Server to that Target,
maybe a couple of hundred replications a day max.
Andrew
-----Original Message-----
From: Steve Romankiw [SMTP:sromankiw@execrisk.com]
Sent: Wednesday, May 31, 2000 1:34 PM
To: Andrew Ford
Cc: Informix-List (E-mail)
Subject: Re: Blocked:DDR
Andrew:
If you were to look at the active threads running, onstat -g act, you might
notice that CDRGeval# threads are consuming a sh*% load of CPU
time. This was a bug in prior versions of IDS, ie. 7.30 family. It has to
do with the sorting algorithm used by the evaluator thread.
When I saw that you were running v7.31, I was surprised that use were
experiencing this problem. The sorting algorithm was fixed in this
version.
I am not sure if you are out of the woods yet. I am not sure if the next
time you cycle your instance without CDRBLOCKOUT if you will
expierence the same problem. Or will the Send/Recieve queue be cleared out
by then?
Madison, are you out there? Any thoughts?
Steve Romankiw
Andrew Ford wrote:
> Thanks for the advice, I was able to solve the problem by
>
> export CDRBLOCKOUT=1
> bounce engine
> export CDRBLOCKOUT=0
> bounce engine
>
> I didn't have to rebuild syscdr, that was my next plan of attack.
>
> Why does this occur?
>
> Thanks,
>
> Andrew
>
> -----Original Message-----
> From: Steve Romankiw [SMTP:sromankiw@execrisk.com]
> Sent: Wednesday, May 31, 2000 11:55 AM
> To: Andrew Ford
> Cc: Informix-List (E-mail)
> Subject: Re: Blocked:DDR
>
> Andrew:
>
> I had the same problem, here is how I solved the Block:DDR.
>
> ** NOTE: this solution may be too drastic for your situation. This
solutions suggests dropping and rebuilding CDR ***
>
> 1. setenv CDRBLOCKOUT=1
> 2. recycle instance
> 3. unset CDRBLOCKOUT
> 4. DROP DATABASE syscdr
> 5. Run $INFORMIXDIR/etc/buildsmi. Note that sysutils will be rebuilt.
So onbar information will have to be reset.
> 6. Recycle instance.
> 7. Redefine your replication environment.
>
> Madison, any alternate methods?
>
> Steve Romankiw
>
> Andrew Ford wrote:
>
> > Was wondering if anyone could help me with this problem.
> >
> > We are running IDS 7.31.UC4 on HP-UX 11.0.
> > ER is defined to replicate between 4 servers, 1 server is a Primary and
the other 3 are Targets (Read only).
> >
> > One of our target servers has locked all user threads because it is
Blocked:DDR.
> > No querries can be performed on this instance.
> >
> > I tried bouncing the engine but when the server came back online it
went back to the same state, has anyone experienced this before?
> >
> > Thanks,
> >
> > Andrew Ford
> >
> > Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 00:25:47
-- 401824
> > Kbytes
> > Blocked:DDR
> >
> > 11:09:32 CDR GC: catalog recovery begin
> > 11:09:32 CDR GC: server recovery complete
> > 11:09:32 Checkpoint Completed: duration was 1 seconds.
> > 11:09:36 CDR GC: replicate recovery complete
> > 11:09:37 CDR GC: replicate pending state recovery complete
> > 11:09:37 CDR GC: group recovery complete
> > 11:09:37 CDR GC: group pending state recovery complete
> > 11:09:37 CDR GC: delete table recovery complete
> > 11:09:37 CDR GC: delete mapping recovery complete
> > 11:09:37 CDR GC: catalog recovery complete
> > 11:09:37 CDR queuer initialization complete
> > 11:09:37 CDR NIF site: 3 <g_dbsvr2> clocks differ by 37 seconds
> > 11:09:38 DDR Log Snooping - Snooping started in log 100
> > 11:09:38 DDR Log Snooping - Catchup phase started, userthreads blocked
> > 11:09:38 CDR NIF site: 3 <g_pinsvr1> clocks differ by 27 seconds
> > 11:14:23 Checkpoint Completed: duration was 0 seconds.
> > 11:29:13 Logical Log 117 Complete.
> > 11:29:15 Logical Log 117 - Backup Started> >
> >
#**************************************************************************
> > #
> > # INFORMIX SOFTWARE, INC.
> > #
> > # Title: onconfig.std
> > # Description: Informix Dynamic Server Configuration Parameters
> > #
> >
#**************************************************************************
> >
> > # Root Dbspace Configuration
> >
> > ROOTNAME rootdbs # Root dbspace name
> > ROOTPATH /dev/links/vg00rl13 # Path for device containing rootdbspace
> > ROOTOFFSET 0 # Offset of root dbspace into device
(Kbytes)
> > ROOTSIZE 200000 # Size of root dbspace (Kbytes)> >
> > # Disk Mirroring Configuration Parameters
> >
> > MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> > MIRRORPATH # Path for device containing mirroredroot
> > MIRROROFFSET 0 # Offset into mirrored device (Kbytes)> >
> > # Physical Log Configuration
> >
> > PHYSDBS physdbs # Location (dbspace) of physical log
> > PHYSFILE 149894 # Physical log file size (Kbytes)> >
> > # Logical Log Configuration
> >
> > LOGFILES 33 # Number of logical log files
> > LOGSIZE 500 # Logical log size (Kbytes)> >
> > # Diagnostics
> >
> > MSGPATH /usr/informix/online.log # System message log file path
> > CONSOLE /dev/console # System console message path
> > ALARMPROGRAM /usr/informix/etc/log_full.sh # Alarm program path> > SYSALARMPROGRAM /usr/informix/etc/evidence.sh # System Alarm program
path
> > TBLSPACE_STATS 1> >
> > # System Archive Tape Device
> >
> > TAPEDEV /dev/null # Tape device path
> > TAPEBLK 16 # Tape block size (Kbytes)
> > TAPESIZE 10240 # Maximum amount of data to put on tape
(Kbytes)> >
> > # Log Archive Tape Device
> >
> > LTAPEDEV /var/simplified/log/logarchive/logfile # Log tape
device path
> > LTAPEBLK 16 # Log tape block size (Kbytes)
> > LTAPESIZE 2000000 # Max amount of data to put on log tape
(Kbytes)> >
> > # Optical
> >
> > STAGEBLOB # Informix Dynamic Server/Optical
staging area
> >
> > # System Configuration
> >
> > SERVERNUM 3 # Unique id corresponding to a DynamicServer in
> > stance
> > DBSERVERNAME dbsvr1 # Name of default database server
> > DBSERVERALIASES dbsvr1_shm # List of alternate dbservernames
> > DEADLOCK_TIMEOUT 60 # Max time to wait of lock indistributed env.
> > RESIDENT -1 # Forced residency flag (Yes = 1, No =
0)
> >
> > MULTIPROCESSOR 2 # 0 for single-processor, 1 formulti-processor
> > NUMCPUVPS 4 # Number of
If you are currently running with CDRBLOCKOUT=0, then you still have
CDRBLOCKOUT set. We only check to see the the environmental variable exists,
not what its value is. So you are running with ER not running.
In this case the only thing that you can really do is to bump LTXHWM and
LTXEHWM up as high as you dare.
I'm sure that someone has mentioned the suggested configuration for ER sites.
If not you need to ask the case owner for this problem to give it to you. I'm
currently on vacation and only have 2400 baud so, just to write this response
is really slow. I can't reply with the suggested parameters.
For any more information, please contact the case owner.
Andrew Ford wrote:
> That's what I had thought initially, but when I looked at onstat -g act
> only the polling threads were active.
>
> It couldn't have been that easy, experiencing a problem that I had just
> read about in the news groups.
>
> It's seems to be fixed, I just wish I knew what caused the problem. There
> is not a whole lot of activity from the Primary Server to that Target,
> maybe a couple of hundred replications a day max.
>
> Andrew
>
> -----Original Message-----
> From: Steve Romankiw [SMTP:sromankiw@execrisk.com]
> Sent: Wednesday, May 31, 2000 1:34 PM
> To: Andrew Ford
> Cc: Informix-List (E-mail)
> Subject: Re: Blocked:DDR
>
> Andrew:
>
> If you were to look at the active threads running, onstat -g act, you might
> notice that CDRGeval# threads are consuming a sh*% load of CPU
> time. This was a bug in prior versions of IDS, ie. 7.30 family. It has to
> do with the sorting algorithm used by the evaluator thread.
>
> When I saw that you were running v7.31, I was surprised that use were
> experiencing this problem. The sorting algorithm was fixed in this
> version.
>
> I am not sure if you are out of the woods yet. I am not sure if the next
> time you cycle your instance without CDRBLOCKOUT if you will
> expierence the same problem. Or will the Send/Recieve queue be cleared out
> by then?
>
> Madison, are you out there? Any thoughts?
>
> Steve Romankiw
>
> Andrew Ford wrote:
>
> > Thanks for the advice, I was able to solve the problem by
> >
> > export CDRBLOCKOUT=1
> > bounce engine
> > export CDRBLOCKOUT=0
> > bounce engine
> >
> > I didn't have to rebuild syscdr, that was my next plan of attack.
> >
> > Why does this occur?
> >
> > Thanks,
> >
> > Andrew
> >
> > -----Original Message-----
> > From: Steve Romankiw [SMTP:sromankiw@execrisk.com]
> > Sent: Wednesday, May 31, 2000 11:55 AM
> > To: Andrew Ford
> > Cc: Informix-List (E-mail)
> > Subject: Re: Blocked:DDR
> >
> > Andrew:
> >
> > I had the same problem, here is how I solved the Block:DDR.
> >
> > ** NOTE: this solution may be too drastic for your situation. This
> solutions suggests dropping and rebuilding CDR ***
> >
> > 1. setenv CDRBLOCKOUT=1
> > 2. recycle instance
> > 3. unset CDRBLOCKOUT
> > 4. DROP DATABASE syscdr
> > 5. Run $INFORMIXDIR/etc/buildsmi. Note that sysutils will be rebuilt.
> So onbar information will have to be reset.
> > 6. Recycle instance.
> > 7. Redefine your replication environment.
> >
> > Madison, any alternate methods?
> >
> > Steve Romankiw
> >
> > Andrew Ford wrote:
> >
> > > Was wondering if anyone could help me with this problem.
> > >
> > > We are running IDS 7.31.UC4 on HP-UX 11.0.
> > > ER is defined to replicate between 4 servers, 1 server is a Primary and
> the other 3 are Targets (Read only).
> > >
> > > One of our target servers has locked all user threads because it is
> Blocked:DDR.
> > > No querries can be performed on this instance.
> > >
> > > I tried bouncing the engine but when the server came back online it
> went back to the same state, has anyone experienced this before?
> > >
> > > Thanks,
> > >
> > > Andrew Ford
> > >
> > > Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 00:25:47
> -- 401824
> > > Kbytes
> > > Blocked:DDR
> > >
> > > 11:09:32 CDR GC: catalog recovery begin
> > > 11:09:32 CDR GC: server recovery complete
> > > 11:09:32 Checkpoint Completed: duration was 1 seconds.
> > > 11:09:36 CDR GC: replicate recovery complete
> > > 11:09:37 CDR GC: replicate pending state recovery complete
> > > 11:09:37 CDR GC: group recovery complete
> > > 11:09:37 CDR GC: group pending state recovery complete
> > > 11:09:37 CDR GC: delete table recovery complete
> > > 11:09:37 CDR GC: delete mapping recovery complete
> > > 11:09:37 CDR GC: catalog recovery complete
> > > 11:09:37 CDR queuer initialization complete
> > > 11:09:37 CDR NIF site: 3 <g_dbsvr2> clocks differ by 37 seconds
> > > 11:09:38 DDR Log Snooping - Snooping started in log 100
> > > 11:09:38 DDR Log Snooping - Catchup phase started, userthreads blocked
> > > 11:09:38 CDR NIF site: 3 <g_pinsvr1> clocks differ by 27 seconds
> > > 11:14:23 Checkpoint Completed: duration was 0 seconds.
> > > 11:29:13 Logical Log 117 Complete.
> > > 11:29:15 Logical Log 117 - Backup Started> > >
> > >
> #**************************************************************************
> > > #
> > > # INFORMIX SOFTWARE, INC.
> > > #
> > > # Title: onconfig.std
> > > # Description: Informix Dynamic Server Configuration Parameters
> > > #
> > >
> #**************************************************************************
> > >
> > > # Root Dbspace Configuration
> > >
> > > ROOTNAME rootdbs # Root dbspace name
> > > ROOTPATH /dev/links/vg00rl13 # Path for device containing root> dbspace
> > > ROOTOFFSET 0 # Offset of root dbspace into device
> (Kbytes)
> > > ROOTSIZE 200000 # Size of root dbspace (Kbytes)> > >
> > > # Disk Mirroring Configuration Parameters
> > >
> > > MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> > > MIRRORPATH # Path for device containing mirrored> root
> > > MIRROROFFSET 0 # Offset into mirrored device (Kbytes)> > >
> > > # Physical Log Configuration
> > >
> > > PHYSDBS physdbs # Location (dbspace) of physical log
> > > PHYSFILE 149894 # Physical log file size (Kbytes)> > >
> > > # Logical Log Configuration
> > >
> > > LOGFILES 33 # Number of logical log files
> > > LOGSIZE 500 # Logical log size (Kbytes)> > >
> > > # Diagnostics
> > >
> > > MSGPATH /usr/informix/online.log # System message log file path
> > > CONSOLE /dev/console # System console message path
> > > ALARMPROGRAM /usr/informix/etc/log_full.sh # Alarm program path> > > SYSALARMPROGRAM /usr/informix/etc/evidence.sh # System Alarm program
> path
> > > TBLSPACE_STATS 1> > >
> > > # System Archive Tape Device
> > >
> > > TAPEDEV /dev/null # Tape device path
> > > TAPEBLK 16 # Tape block size (Kbytes)
> > > TAPESIZE 10240 # Maximum
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g