Socket listener stuck behind non-yielding thread
Posted in 2016
While rebuilding ~1,900 foreign keys (several concurrent ALTER TABLE ... ADD CONSTRAINT jobs) on 11.70.FC8, new connections intermittently failed with -27001. onstat showed the soctcplst listener thread stuck 'ready' on a CPU VP that was monopolised by a long-running ALTER TABLE; connections resumed only when the ALTER finished. Suggestions included INFORMIXCONRETRY, moving poll threads to NET VPs (IBM noted listener threads can't run there, so it wouldn't help), reducing CPU VP load, and using a shared-memory connection. IBM's view was that the non-yielding ALTER thread is likely a defect, needing pstack/onmode -X stacks via a support case. No workaround or fix is confirmed in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Server Administration, Triggers, Constraints & Referential Integrity, Migration, Import/Export & Data Conversion, Versions, Editions & End-of-Life
I'm re-building all 1,900 foreign keys on a system, sometimes a many as 5 such
jobs are running concurrently.
Occasionally we are getting -27001 failures. I trapped one this morning:
dbschema -dhangs and eventually fails -27001 Read error occurred during connection
attempt.
Looking at the ready queue I see the Sockets TCP listener permanently waiting
on 8cpu.
IBM Informix Dynamic Server Version 11.70.FC8 -- On-Line -- Up 11:33:15 --224961056 Kbytes
Ready threads:
tid tcb rstcb prty status vp-class name
23 2ede9fe028 0 2 ready 8cpu* soctcplst
132 2edf9cea90 2edc8c18e8 3 ready 8cpu* memory
And looking at the active queue I see an application query running permanently
on that vcpu:
[informix@orbdb-1a<staging>:Migration_201601]$ onstat -g ses 19581
IBM Informix Dynamic Server Version 11.70.FC8 -- On-Line -- Up 11:36:46 --224961056 Kbytes
session effective #RSAM total used dynamic
id user user tty pid hostname threads memory memory explain
19581 xbinfinf - - 13063 orbdb-1a 1 151552 144392 off
Program :
/mount/informix/11.78/bin/dbaccess
tid name rstcb flags curstk status
53451 sqlexec 2ee8b29208 --BPR-- 32560 running-
Memory pools count 1
name class addr totalsize freesize #allocfrag #freefrag
19581 V 2ee0b2b040 151552 7160 250 11
name free used name free used
overhead 0 3288 scb 0 144
opentable 0 18400 filetable 0 4944
ru 0 600 log 0 16536
temprec 0 2248 keys 0 384
ralloc 0 46576 gentcb 0 1616
ostcb 0 2944 sqscb 0 21064
sql 0 72 hashfiletab 0 552
osenv 0 2704 sqtcb 0 8288
fragman 0 14032
sqscb info
scb sqscb optofc pdqpriority optcompind directives
2ede2091c0 2eedbb3028 0 0 2 1
Sess SQL Current Iso Lock SQL ISAM F.E.
Id Stmt type Database Lvl Mode ERR ERR Vers Explain
19581 ALTER TABLE xxx CR Not Wait 0 0 9.24 Off
Current SQL statement (3) :
alter table "... constraint (foreign key
(xxx) references xxx constraint
"yyyy
Once the ALTER TABLE finished 8cpu was freed up, the Sockets listener
disappeared from the ready queue, and the connections flowed again.
Can't think this is correct behaviour. We've seen a similar bug in 11.50. Any
easy workarounds (other than using a shared-memory INFORMIXSERVER?)
Thanks
Neil
Did you get a onstat -g stk several times to see what the session is doing?
Would INFORMIXCONRETRY help?
Or temporarily running NETTYPE NET rather than CPU?
Regards,
David.
> On 02 February 2016 at 12:51 NEIL TRUBY <neil.truby@ardenta.com> wrote:
>
>
> I'm re-building all 1,900 foreign keys on a system, sometimes a many as 5
such
> jobs are running concurrently.
>
> Occasionally we are getting -27001 failures. I trapped one this morning:
>
> dbschema -d> hangs and eventually fails -27001 Read error occurred during connection
> attempt.
>
> Looking at the ready queue I see the Sockets TCP listener permanently waiting
> on 8cpu.
> IBM Informix Dynamic Server Version 11.70.FC8 -- On-Line -- Up 11:33:15 --> 224961056 Kbytes
>
> Ready threads:
> tid tcb rstcb prty status vp-class name
> 23 2ede9fe028 0 2 ready 8cpu* soctcplst
> 132 2edf9cea90 2edc8c18e8 3 ready 8cpu* memory
>
> And looking at the active queue I see an application query running
permanently
> on that vcpu:
>
> [informix@orbdb-1a<staging>:Migration_201601]$ onstat -g ses 19581
>
> IBM Informix Dynamic Server Version 11.70.FC8 -- On-Line -- Up 11:36:46 --> 224961056 Kbytes
>
> session effective #RSAM total used dynamic
> id user user tty pid hostname threads memory memory explain
> 19581 xbinfinf - - 13063 orbdb-1a 1 151552 144392 off
>
> Program :
> /mount/informix/11.78/bin/dbaccess
>
> tid name rstcb flags curstk status
> 53451 sqlexec 2ee8b29208 --BPR-- 32560 running-
>
> Memory pools count 1
> name class addr totalsize freesize #allocfrag #freefrag
> 19581 V 2ee0b2b040 151552 7160 250 11
>
> name free used name free used
> overhead 0 3288 scb 0 144
> opentable 0 18400 filetable 0 4944
> ru 0 600 log 0 16536
> temprec 0 2248 keys 0 384
> ralloc 0 46576 gentcb 0 1616
> ostcb 0 2944 sqscb 0 21064
> sql 0 72 hashfiletab 0 552
> osenv 0 2704 sqtcb 0 8288
> fragman 0 14032
>
> sqscb info
> scb sqscb optofc pdqpriority optcompind directives
> 2ede2091c0 2eedbb3028 0 0 2 1
>
> Sess SQL Current Iso Lock SQL ISAM F.E.
> Id Stmt type Database Lvl Mode ERR ERR Vers Explain
> 19581 ALTER TABLE xxx CR Not Wait 0 0 9.24 Off
>
> Current SQL statement (3) :
> alter table "... constraint (foreign key
>
> (xxx) references xxx constraint
>
> "yyyy
>
> Once the ALTER TABLE finished 8cpu was freed up, the Sockets listener
> disappeared from the ready queue, and the connections flowed again.
>
> Can't think this is correct behaviour. We've seen a similar bug in 11.50. Any
> easy workarounds (other than using a shared-memory INFORMIXSERVER?)
>
> Thanks
> Neil
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Original post:
I'm re-building all 1,900 foreign keys on a system, sometimes a many as 5 such
jobs are running concurrently.
Occasionally we are getting -27001 failures. I trapped one this morning:
dbschema -dhangs and eventually fails -27001 Read error occurred during connection
attempt.
Looking at the ready queue I see the Sockets TCP listener permanently waiting
on 8cpu.
IBM Informix Dynamic Server Version 11.70.FC8 -- On-Line -- Up 11:33:15 --224961056 Kbytes
Ready threads:
tid tcb rstcb prty status vp-class name
23 2ede9fe028 0 2 ready 8cpu* soctcplst
132 2edf9cea90 2edc8c18e8 3 ready 8cpu* memory
And looking at the active queue I see an application query running permanently
on that vcpu:
[informix@orbdb-1a<staging>:Migration_201601]$ onstat -g ses 19581
IBM Informix Dynamic Server Version 11.70.FC8 -- On-Line -- Up 11:36:46 --224961056 Kbytes
session effective #RSAM total used dynamic
id user user tty pid hostname threads memory memory explain
19581 xbinfinf - - 13063 orbdb-1a 1 151552 144392 off
Program :
/mount/informix/11.78/bin/dbaccess
tid name rstcb flags curstk status
53451 sqlexec 2ee8b29208 --BPR-- 32560 running-
Memory pools count 1
name class addr totalsize freesize #allocfrag #freefrag
19581 V 2ee0b2b040 151552 7160 250 11
name free used name free used
overhead 0 3288 scb 0 144
opentable 0 18400 filetable 0 4944
ru 0 600 log 0 16536
temprec 0 2248 keys 0 384
ralloc 0 46576 gentcb 0 1616
ostcb 0 2944 sqscb 0 21064
sql 0 72 hashfiletab 0 552
osenv 0 2704 sqtcb 0 8288
fragman 0 14032
sqscb info
scb sqscb optofc pdqpriority optcompind directives
2ede2091c0 2eedbb3028 0 0 2 1
Sess SQL Current Iso Lock SQL ISAM F.E.
Id Stmt type Database Lvl Mode ERR ERR Vers Explain
19581 ALTER TABLE xxx CR Not Wait 0 0 9.24 Off
Current SQL statement (3) :
alter table "... constraint (foreign key
(xxx) references xxx constraint
"yyyy
Once the ALTER TABLE finished 8cpu was freed up, the Sockets listener
disappeared from the ready queue, and the connections flowed again.
Can't think this is correct behaviour. We've seen a similar bug in 11.50. Any
easy workarounds (other than using a shared-memory INFORMIXSERVER?)
Thanks
Neil
Response:
It would not be correct behavior. The thread doing the alter should likely
have some sort of yield check to make sure it doesn't run constantly without
yielding. As when threads do that, they tend to wreak havoc on other threads
that have to be bound to specific cpu vps...I think you opened a case already.
I would expect a request to get stack traces for that running thread...and
since it's running, you'll need to use something like procstack/pstack/onmode
-X to get the stack off the cpu vp as onstat -g stk is not a reliable
indication of what the thread is doing when it's running.
Jacques Renaut
IBM Informix APD Team
I agree with David. For my $$ best practice is to always run TCP listeners
in NET VPs for exactly this reason. Informix's thread model is
non-interrupting (aka cooperative) so a busy thread that doesn't have to
wait for IO or other resources can dominate a CPU VP.
Art
Art S. Kagel, President and Principal Consultant
ASK Database Management
www.askdbmgt.com
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on the IIUG, nor any other organization with which I am
associated either explicitly, implicitly, or by inference. Neither do
those opinions reflect those of other individuals affiliated with any
entity with which I am affiliated nor those of the entities themselves.
On Tue, Feb 2, 2016 at 7:51 AM, NEIL TRUBY <neil.truby@ardenta.com> wrote:
> I'm re-building all 1,900 foreign keys on a system, sometimes a many as 5
> such
> jobs are running concurrently.
>
> Occasionally we are getting -27001 failures. I trapped one this morning:
>
> dbschema -d> hangs and eventually fails -27001 Read error occurred during connection
> attempt.
>
> Looking at the ready queue I see the Sockets TCP listener permanently
> waiting
> on 8cpu.
> IBM Informix Dynamic Server Version 11.70.FC8 -- On-Line -- Up 11:33:15 --> 224961056 Kbytes
>
> Ready threads:
> tid tcb rstcb prty status vp-class name
> 23 2ede9fe028 0 2 ready 8cpu* soctcplst
> 132 2edf9cea90 2edc8c18e8 3 ready 8cpu* memory
>
> And looking at the active queue I see an application query running
> permanently
> on that vcpu:
>
> [informix@orbdb-1a<staging>:Migration_201601]$ onstat -g ses 19581
>
> IBM Informix Dynamic Server Version 11.70.FC8 -- On-Line -- Up 11:36:46 --> 224961056 Kbytes
>
> session effective #RSAM total used dynamic
> id user user tty pid hostname threads memory memory explain
> 19581 xbinfinf - - 13063 orbdb-1a 1 151552 144392 off
>
> Program :
> /mount/informix/11.78/bin/dbaccess
>
> tid name rstcb flags curstk status
> 53451 sqlexec 2ee8b29208 --BPR-- 32560 running-
>
> Memory pools count 1
> name class addr totalsize freesize #allocfrag #freefrag
> 19581 V 2ee0b2b040 151552 7160 250 11
>
> name free used name free used
> overhead 0 3288 scb 0 144
> opentable 0 18400 filetable 0 4944
> ru 0 600 log 0 16536
> temprec 0 2248 keys 0 384
> ralloc 0 46576 gentcb 0 1616
> ostcb 0 2944 sqscb 0 21064
> sql 0 72 hashfiletab 0 552
> osenv 0 2704 sqtcb 0 8288
> fragman 0 14032
>
> sqscb info
> scb sqscb optofc pdqpriority optcompind directives
> 2ede2091c0 2eedbb3028 0 0 2 1
>
> Sess SQL Current Iso Lock SQL ISAM F.E.
> Id Stmt type Database Lvl Mode ERR ERR Vers Explain
> 19581 ALTER TABLE xxx CR Not Wait 0 0 9.24 Off
>
> Current SQL statement (3) :
> alter table "... constraint (foreign key
>
> (xxx) references xxx constraint
>
> "yyyy
>
> Once the ALTER TABLE finished 8cpu was freed up, the Sockets listener
> disappeared from the ready queue, and the connections flowed again.
>
> Can't think this is correct behaviour. We've seen a similar bug in 11.50.
> Any
> easy workarounds (other than using a shared-memory INFORMIXSERVER?)
>
> Thanks
> Neil
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--089e013a22682d417d052ac9bd44
Original post: I agree with David. For my $$ best practice is to always run TCP listeners in NET VPs for exactly this reason. Informix's thread model is non-interrupting (aka cooperative) so a busy thread that doesn't have to wait for IO or other resources can dominate a CPU VP. Art Art S. Kagel, President and Principal Consultant ASK Database Management www.askdbmgt.com Blog: http://informix-myview.blogspot.com/ Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. Response: You can't run listener threads on net vps, you can only run poll threads on net vps, and having your poll threads on net vps will not help. Even a thread that doesn't have to wait on resource should still be checking to yield to prevent this sort of issue, and if it isn't, that would likely be considered a defect. Jacques Renaut IBM Informix APD Team
I only have to do this once, so I really need a workaround (or a version if 11.70FC8 that doesn't exhibit the behaviour), as any patch will even with Godspeed take many days. Any suggestions? Thanks Neil
It is not clear if the alter table is 1. not yielding at all 2. not yielding enough 3. 2 to lesser extent + cpu vp is just too busy with other work to get to the connection in time For 2 and 3 if we can reduce the load on the cpus vps that MAY help. Regards, David. > On 02 February 2016 at 13:56 JACQUES RENAUT <jrenaut@us.ibm.com> wrote: > > > Original post: > > I agree with David. For my $$ best practice is to always run TCP listeners > in NET VPs for exactly this reason. Informix's thread model is > non-interrupting (aka cooperative) so a busy thread that doesn't have to > wait for IO or other resources can dominate a CPU VP. > > Art > > Art S. Kagel, President and Principal Consultant > ASK Database Management > www.askdbmgt.com > > Blog: http://informix-myview.blogspot.com/ > > Disclaimer: Please keep in mind that my own opinions are my own opinions > and do not reflect on the IIUG, nor any other organization with which I am > associated either explicitly, implicitly, or by inference. Neither do > those opinions reflect those of other individuals affiliated with any > entity with which I am affiliated nor those of the entities themselves. > > Response: > > You can't run listener threads on net vps, you can only run poll threads on > net vps, and having your poll threads on net vps will not help. Even a thread > that doesn't have to wait on resource should still be checking to yield to > prevent this sort of issue, and if it isn't, that would likely be considered a > defect. > > Jacques Renaut > IBM Informix APD Team > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. >
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g