A bug or "expected behaviour"?
Posted in 2014
Topics: High Availability & Replication, Transactions, Locking & Isolation, Logging & Checkpoints, Clustering, Grid & MACH11
We are running HDR on 11.70.FC7 with UPDATABLE_SECONDARY.
An attempt was made on the primary to update 112 million rows in a table,
something like:
begin work;
lock table mytab in exclusive mode;
update mytab set col1="xxx" where 1=1;commit work;
No locks were accumulated on the primary.
On the secondary though, locks were added at a high rate, which resulted in
memory problems and eventually the secondary stopped.
19:53:53 Dynamically allocated new virtual shared memory segment (size
126492KB)
19:53:53 Memory sizes:resident:184832000 KB, virtual:7035808 KB, no SHMTOTAL
limit
19:53:53 dynamically allocated 1000000 locks
19:54:41 Logical Log 5460 Complete, timestamp: 0x62dc54ed.
19:56:16 Checkpoint Completed: duration was 1 seconds.
19:56:16 Sun Apr 28 - loguniq 5461, logpos 0x39ed084, timestamp: 0x62e2fb1f
Interval: 20235
The above happened quite often until:
05:30:18 Process exited with return code 126: /bin/sh /bin/sh -c
/opt/informix/etc/alarmprogram.sh 3 15 "Data Replication failure." "DR:
Received connection request from remote server when DR is not Off
[Local type: Secondary, Current state: FAILED]
[Remote type: Primary]
" "" 15004
05:30:20 DR: Terminating redirected write subsystem due to server disconnect.
All open redirected transactions will be rolled back.
05:30:21 Updates from secondary currently not allowed
05:30:22 Updates from secondary currently not allowed
05:30:23 Updates from secondary currently not allowed
05:30:24 Updates from secondary currently not allowed
05:30:25 Updates from secondary currently not allowed
05:30:26 Updates from secondary currently not allowed
05:30:27 Updates from secondary currently not allowed
05:30:28 Updates from secondary currently not allowed
05:30:28 DR: Turned off on secondary server
I'm running with unbuffered log and DRINTERVAL -1
So the "lock table xx in exclusive mode" is not used on secondary servers in
MACH11.
According to IBM this is currently "expected behavior". What is your opinion
on this? To my knowledge this limitation is not documented and this basically
means you have to take HDR offline while performing tasks like this.
Best regards,
-Snorri
I do remember seeing some PMR or technote or something referring to this
internally. But I don't remember the details. I think this may be
considered expected for the people who developed the code to "transport"
the locks from the primary to the secondaries. But It seems obvious to me
that this would never be expected by a customer.
The only workaround that I can think off for now is to split the
transactions into smaller blocks. Clearly not ideal, but in most cases it's
not that hard to do either... and I'd say much better that taking HDR off.
Regards.
On Fri, Oct 3, 2014 at 4:36 PM, SNORRI BERGMANN <snorri@init.is> wrote:
> We are running HDR on 11.70.FC7 with UPDATABLE_SECONDARY.
> An attempt was made on the primary to update 112 million rows in a table,
> something like:
>
> begin work;
> lock table mytab in exclusive mode;
> update mytab set col1="xxx" where 1=1;> commit work;
>
> No locks were accumulated on the primary.
>
> On the secondary though, locks were added at a high rate, which resulted in
> memory problems and eventually the secondary stopped.
>
> 19:53:53 Dynamically allocated new virtual shared memory segment (size
> 126492KB)
> 19:53:53 Memory sizes:resident:184832000 KB, virtual:7035808 KB, no
> SHMTOTAL> limit
> 19:53:53 dynamically allocated 1000000 locks
> 19:54:41 Logical Log 5460 Complete, timestamp: 0x62dc54ed.
> 19:56:16 Checkpoint Completed: duration was 1 seconds.
> 19:56:16 Sun Apr 28 - loguniq 5461, logpos 0x39ed084, timestamp: 0x62e2fb1f
> Interval: 20235
>
> The above happened quite often until:
>
> 05:30:18 Process exited with return code 126: /bin/sh /bin/sh -c
> /opt/informix/etc/alarmprogram.sh 3 15 "Data Replication failure." "DR:
> Received connection request from remote server when DR is not Off
>
> [Local type: Secondary, Current state: FAILED]
>
> [Remote type: Primary]
> " "" 15004
> 05:30:20 DR: Terminating redirected write subsystem due to server
> disconnect.
>
> All open redirected transactions will be rolled back.
> 05:30:21 Updates from secondary currently not allowed
> 05:30:22 Updates from secondary currently not allowed
> 05:30:23 Updates from secondary currently not allowed
> 05:30:24 Updates from secondary currently not allowed
> 05:30:25 Updates from secondary currently not allowed
> 05:30:26 Updates from secondary currently not allowed
> 05:30:27 Updates from secondary currently not allowed
> 05:30:28 Updates from secondary currently not allowed
> 05:30:28 DR: Turned off on secondary server
>
> I'm running with unbuffered log and DRINTERVAL -1
>
> So the "lock table xx in exclusive mode" is not used on secondary servers
> in
> MACH11.
> According to IBM this is currently "expected behavior". What is your
> opinion
> on this? To my knowledge this limitation is not documented and this
> basically
> means you have to take HDR offline while performing tasks like this.
>
> Best regards,
> -Snorri
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--20cf303636dbac2686050486b79c