Re: Replication: Understanding CDRGeval# threads
Posted in 2000
I can't go into all of the details, but we have to match all of the after/before
images for a given row.
Steve Romankiw wrote:
> Madison:
>
> You win the prize! I did not check onstat -g grp E but there was a LARGE
> transaction going against a replicated table. I am sure this was the cause. I
> worked with tech support to clear out my problem and all is better. I ended up
> dropping and recreating my replication environment.
>
> BTW, can you elaborate on :
>
> [snip]
> 2) unfortunatly the engineer that coded the original code for the
> compression phase did not fully understand the implication of a worst-case
> bubble sort logic.
> [snip]
>
> Q: I assume by worst-case, you mean the list is order already, but in reverse
> order. Ex: 54321 (unsorted list), 12345 (sorted list)
>
> Q: Just curious, what are you sorting?
>
> TIA
>
> Steve Romankiw
>
> Madison Pruet wrote:
>
> > Well - this sounds familar with pre-7.31 releases of ER.
> >
> > If you run 'onstat -g grp E' while this is happening, I'm going to bet that
> > the CDRGeval thread is 'compressing'.
> >
> > The original design of the grouper compression phase is really not very
> > efficient. We corrected the algorthem in 7.31.
> >
> > Basically the situation is:
> > 1) we have to compress the transaction so that we can maintain the local
> > delete shadow tables.
> > 2) unfortunatly the engineer that coded the original code for the
> > compression phase did not fully understand the implication of a worst-case
> > bubble sort logic.
> > 3) fortunatly we corrected the logic in 7.31
> > 4) unfortunatly the correction can not be backported
> >
> > I wrote and taught an ER-Internals class for advanced support a year ago
> > just as 7.31 was about to be released. One of my exercises was to
> > demonstrate the difference between the 7.30 and 7.31 grouper compression
> > phase. What I did was to create a simple table containing 10,000 rows of
> > (col1 int primary key, col2 int). I then added one to all of the col2 in
> > the table. All of the students were following the output of 'onstat -g grp
> > E', watching to see when the compression phase was ended. I then proceeded
> > with my lecture.
> >
> > Every few minutes or so, I would ask, "Is it still compressing?" Of course,
> > I knew that it would be. After 20 minutes of this, I killed the engine and
> > switched to 7.31. When I ran the same test, the compression phase was
> > finished in about 3 seconds. So I'm fairly certain that the problem is
> > corrected.
> >
> > In pre-7.30, the only way that this problem can be avoided is to ensure that
> > the transactions are fairly small.
> >
> > Oh yes - concerning the engineer that did the original work--- Well, he no
> > longer works for Informix. He was recruted by one of our competitors. You
> > know the one-- they have an office on the east side of highway 101 in San
> > Mateo CA - Seems like their name starts with Ora....
> >
> > I still don't know if he understands the performance impact of a bubble
> > sort. ;-)
> >
> > Steve Romankiw wrote:
> >
> > > I am noticing some strange behaviour with ER. The response time for
> > > sqlexec threads is very poor
> > > due to the fact that the CDRGeval# threads are not yielding very nicely.
> > >
> > > What are the CDRGeval# doing? I suspect that are trying to "catch up".
> > > I say this because nothing is running so what can they be doing?.
> > >
> > > Also notice the onstat -g seg, a lot of additional virtual segments are
> > > getting created. Is this a feature, bug, or lack of understanding?
> > >
> > > TIA
> > >
> > > Environment:
> > > Solaris 2.6
> > > Informix Dynamic Server Version 7.30.UC9
> > >
> > > onstat -g act
> > > -------------
> > > 7 11c87588 0 2 running 9tli> > > tlitcppoll
> > > 8 11c87a80 0 2 running 10tli
> > > tlitcppoll
> > > 9 11c8a128 0 2 running 11tli
> > > tlitcppoll
> > > 10 11c8a590 0 2 running 12tli
> > > tlitcppoll
> > > 665137 1214c990 11bd2760 2 running 1cpu
> > > CDRGeval0
> > > 665139 1214cf50 11bd0fdc 2 running 4cpu
> > > CDRGeval2
> > > 665140 1214d230 11bcf3a4 2 running 3cpu
> > > CDRGeval3
> > >
> > > onstat -g ath
> > > -------------
> > > [snip]
> > > 665126 11d39b40 11bd938c 2 sleeping forever 1cpu> > > CDRSchedMgr
> > > 665127 120b28f8 11bd51b4 2 cond wait CDRfanout 3cpu
> > > CDRN_CM
> > > 665135 1214c630 11bd22ac 2 ready 1cpu
> > > CDRGfan
> > > 665136 1214c768 11bd6484 2 cond wait CDRGclean 4cpu
> > > CDRGclean
> > > 665137 1214c990 11bd2760 2 running 3cpu
> > > CDRGeval0
> > > 665138 1214cc70 11bd72a0 2 running 4cpu
> > > CDRGeval1
> > > 665139 1214cf50 11bd0fdc 2 running 1cpu
> > > CDRGeval2
> > > 665140 1214d230 11bcf3a4 2 ready 3cpu
> > > CDRGeval3
> > > 665141 1214d510 11bd8ed8 3 sleeping secs: 1 3cpu
> > > ddr_snoopy
> > > 665142 1214d8e0 11bd0b28 3 sleeping secs: 1 4cpu
> > > ddr_log_io
> > > 950142 11d6d6c8 11bd4398 2 cond wait netnorm 1cpu
> > > sqlexec
> > > 953734 121b9258 11bd80bc 2 cond wait netnorm 4cpu
> > > srvinfx
> > > 954695 11db0678 11bd7754 2 cond wait netnorm 3cpu
> > > sqlexec
> > > 954711 11ee8ff0 11bdafc4 2 cond wait netnorm 1cpu
> > > sqlexec
> > > 954995 15874f68 11bd1df8 2 cond wait CDR connec 3cpu
> > > CDRNsT111
> > > 954996 15875518 11bdc748 2 sleeping forever 3cpu
> > > CDRNsA111
> > > 954998 11d394a8 11bd5668 2 cond wait netnorm 1cpu
> > > CDRNrA111
> > > 955015 12348f08 11bdcbfc 2 cond wait CDRDssleep 4cpu
> > > CDRD_0
> > > 955016 12176630 11bcfd0c 2 sleeping forever 1cpu
> > > ackTh111
> > > 955017 121768e0 11bd3ee4 2 sleeping secs: 173 3cpu
> > > replTh111
> > > 955840 120b2a90 11bd1490 2 cond wait netnorm 4cpu
> > > sqlexec
> > > 956862 11dbaa08 11bd0674 2 sleeping secs: 1 1cpu
> > > ontape
> > > [snip]> > >
> > > onstat -g seg
> > > -------------
> > > Informix Dynamic Server Version 7.30.UC9 -- On-Line -- Up 45 days
> > > 01:25:20 -- 458576 Kbytes> > > Segment Summary:
> > > id key addr size ovhd class blkused blkfree
> > > 500 1381386241 a000000 129466368 2844 R 15797 7
> > > 701 1381386242 11b78000 189120512 3484 V 23086 0
> > > 702 1381386243 1cfd4000 622592 608 M 68 8
> > > 169 1381386244 1d06c000 16777216 852 V 2048 0
> > > 170 1381386245 1e06c000 16777216 852 V 2048 0
> > > 171 1381386246 1f06c000 16777216 852 V 2048 0
> > > 172 1381386247 2006c000 16777216 852 V 2048 0
> > > 73 1381386248 2106c000 16777216 852 V 20