Translating with DrWatson… this can take a few seconds the first time.
This is a genuine, complex translation. DrWatson protects commands, error codes, and log output while naturally translating the surrounding text. It’s translated once and saved.
A single-direction CDR/ER setup on IDS 11.50FC7 went into 'RUN,BLOCK' on the source and stopped replicating when an application deleted ~600,000 rows in one transaction; the receive queue's first transaction didn't advance for over an hour, and only stopping/deleting the replicate restored things. Madison Pruet explained BLOCK is set when half the receive queue fills and suggested checking repeated 'onstat -g nif' output to see if bytes were still moving, plus possible stable-queue space exhaustion. Suggested workarounds: smaller transactions, 'begin work without replication', or more ER resources. The poster ultimately defined the replicate with '-D y' to skip DELETEs and ran the deletes on the target via cron.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
HI,
I config the single direction CDR system. IDS version is 11.50FC7
When the application execute below SQL to delete about 600000 records.
delete from japf10 where tdate='20141203'
the row size is 688 bytes for single row.
onstat -g nif show CDR is in 'RUN,BLOCK' state, and the replication doesn't
work.onstat -g rqm recvq show the 'commit time' doesn't change.onstat -d show the qhdr dbspace and qdata smart blobspace still has free space.
after I execute 'cdr stop repl japf10' and 'cdr delete repl japf10',the CDR
replication work fine.
Do anybody know why IDS being blocked by the big transaction? How to fixes or
avoid such kind of problem?
thanks for your time.
↪ replying to CHUAN LU
Madison Pruet — — source: IIUG Forums & Mailing Lists
The blocked state is a signal for the source node not to send any other=
transactions until the existing replicated transactions have been
processed. It it triggered when 1/2 of the existing receive queue spac=
e
has been filled. The transfer of the transaction currently being
transferred, however, is allowed to continue.
I suspect that in your case, the transaction was still in the process o=
f
being transmitted, possibly in the process of being executed, and/or th=
e
disk portion of the receive queue was full and the system was requestin=
g
that the stable queue be increased.
It is unfortunate that we do not have the diagnostics from the target
system.
From: "CHUAN LU" <luchuan@cn.ibm.com>
To: ids@iiug.org
Date: 12/03/2014 09:07 PM
Subject: CDR question on big transaction [34284]
Sent by: ids-bounces@iiug.org
HI,
I config the single direction CDR system. IDS version is 11.50FC7
When the application execute below SQL to delete about 600000 records.
delete from japf10 where tdate=3D'20141203'
the row size is 688 bytes for single row.
onstat -g nif show CDR is in 'RUN,BLOCK' state, and the replication doe=sn't
work.
onstat -g rqm recvq show the 'commit time' doesn't change.onstat -d show the qhdr dbspace and qdata smart blobspace still has fre=e
space.
after I execute 'cdr stop repl japf10' and 'cdr delete repl japf10',the=
CDR
replication work fine.
Do anybody know why IDS being blocked by the big transaction? How to fi=
xes
or
avoid such kind of problem?
thanks for your time.
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
HI,
But I execute 'onstat -g rqm recvq' in the target size, I see the first
transaction doesn't change for more than 1 hours.
↪ replying to CHUAN LU
Madison Pruet — — source: IIUG Forums & Mailing Lists
But what did onstat -g nif show?
From: "CHUAN LU" <luchuan@cn.ibm.com>
To: ids@iiug.org
Date: 12/04/2014 08:20 AM
Subject: Re: CDR question on big transaction [34286]
Sent by: ids-bounces@iiug.org
HI,
But I execute 'onstat -g rqm recvq' in the target size, I see the first=
transaction doesn't change for more than 1 hours.
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
In the source side, 'onstat -g nif' show the state is 'RUN,BLOCK';
In the target side, 'onstat -g nif' show the state is 'RUN';
↪ replying to CHUAN LU
Madison Pruet — — source: IIUG Forums & Mailing Lists
no -- you would have needed to have run several and then check to see i=
f
the number of bytes being transfered was increasing.
From: "CHUAN LU" <luchuan@cn.ibm.com>
To: ids@iiug.org
Date: 12/04/2014 08:53 AM
Subject: Re: CDR question on big transaction [34288]
Sent by: ids-bounces@iiug.org
In the source side, 'onstat -g nif' show the state is 'RUN,BLOCK';
In the target side, 'onstat -g nif' show the state is 'RUN';
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
I compare the record number in both side. the application are do 'delete ' sql,
But the row count in the target size doesn't decrease during 1 hours.
↪ replying to CHUAN LU
Madison Pruet — — source: IIUG Forums & Mailing Lists
The thing that is important is to determine if the large transaction wa=
s
being sent to the target. You need the NIF output for that. If the
transmission caused the stable queue disk space to fill up, then you sh=
ould
have gotten warnings that you needed to increase the size of stable
storage. If you did not increase stable storage, then that would have
caused the transfer to not be able to complete, and hence the transacti=
on
to not be applied.
Since you are an IBM employee, we should be communicating directly and =
not
via IIUG.
M.P.
From: "CHUAN LU" <luchuan@cn.ibm.com>
To: ids@iiug.org
Date: 12/04/2014 09:20 AM
Subject: Re: CDR question on big transaction [34290]
Sent by: ids-bounces@iiug.org
I compare the record number in both side. the application are do 'delet=
e '
sql,
But the row count in the target size doesn't decrease during 1 hours.
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
Big ER transaction causing ER block or slow is nothing new, agree?
Things can do,
--- make your transaction smaller
--- begin work without replication; delete .....; commit;
--- make ER more powerful
Thanks
Frank
On Wed, Dec 3, 2014 at 10:06 PM, CHUAN LU <luchuan@cn.ibm.com> wrote:
> HI,
>
> I config the single direction CDR system. IDS version is 11.50FC7
>
> When the application execute below SQL to delete about 600000 records.
>
> delete from japf10 where tdate='20141203'>
> the row size is 688 bytes for single row.
>
> onstat -g nif show CDR is in 'RUN,BLOCK' state, and the replication doesn't
> work.>
> onstat -g rqm recvq show the 'commit time' doesn't change.>
> onstat -d show the qhdr dbspace and qdata smart blobspace still has free
> space.>
> after I execute 'cdr stop repl japf10' and 'cdr delete repl japf10',the CDR
> replication work fine.
>
> Do anybody know why IDS being blocked by the big transaction? How to fixes
> or
> avoid such kind of problem?
>
> thanks for your time.
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--001a1132f08007f40e0509800d9e
'cdr define repl' command has "-D y " option that doesn't replicate the DELETE
operation. and I do the DELETE operatoin in the target side within the OS
crontab.
We use strictly necessary cookies to make this site work. With your
consent we’d also use optional cookies for analytics and marketing. You can accept all,
reject all, or choose. Read our Cookie Policy.