Thread Hang partially Instance - URGENT
Posted in 2011
Hi All
IFX 11.50 FC8 - AIX 6.1
Client - 4GL 7.32 / CSDK 2.9 - AIX 5.3 (other machine, connected from TCP)
I will try be objective.
This is the second time happen on our production and is happening just now.
On the first time, we need to bounce the instance... (with onclean + kill )
- The process of 4GL died for some reason (not KILL, probably the user close
the telnet session on the wrong way).
- The session on the database, "freeze" running on the CPU VP (with the isam
error -936 Error on remote connection connection-name.).
This session stay more part of time running on the CPU VP 1 (one)
Sometimes they migrate to other CPU VPs for some minutes and than back to CPU
VP1....
- When this session run on CPU VP1 , create a lot of administrative problems,
any ontape , onmode stop works
So , we isn't capable to run physical and logical backup any more... kill a
session or run a checkpoint (via onmode)...
The stack trace of this session is (I get this with procstack of AIX, because
the onstat -g stk don't return the stack) :
749612: oninit -v
0x0000000100040b18 mt_set_hang() + 0x4
0x00000001002170ac _iflushbuff() + 0xf4
0x0000000100218674 _iwrite() + 0x64
0x0000000100218800 _iputint() + 0x20
0x00000001003fda60 puttxstat() + 0xec
0x00000001007569b0 sqrollback() + 0x820
0x00000001003d8290 exec_sysdbproc() + 0x644
0x0000000100cef168 sqscb_cleanup() + 0x2cc
0x0000000100139a58 destroy_session() + 0xe8
0x00000001001f9890 sqsetconerr() + 0x84
0x000000010021726c asf_send() + 0x8c
0x000000010021705c _iflushbuff() + 0xa4
0x0000000100218674 _iwrite() + 0x64
0x00000001002183a4 _iputbuf() + 0x10
0x00000001004024fc puttuple() + 0xcc0
0x00000001004461f4 sql_scrollfetch() + 0x2e4
0x0000000100446604 sq_sfetch() + 0x110
0x0000000100224638 sqmain() + 0x91c
0x000000010037c15c listen_verify() + 0x498
0x000000010037a604 spawn_thread() + 0xe00
0x0000000100dd5288 startup() + 0xa8
0x0000000100dd51dc run_system() + 0x1a4
We already open PMR for both situations (24907,228,631 the actual and
24099,228,631what happen a few weeks)
Some details :
- Appear the user "close" the 4GL session exactly when a checkpoint occur (42
seconds of checkpoint where appear have some thread waits). I get this looking
the onstat -g ntt with the online.log messages...
My hypothesis is , ...maybe when the checkpoint occur, freeze the user
session, this user lost the patience and close the telnet window, when the
checkpoint release the session, the pid of this session don't exists any more.
The engine start rollback the transaction and freeze.
any tips , how kill this session?
onmoe -z , onmode -Z , onmode -H don't works.. just freeze.
task("onmode","z"...) too...
Curiosity... task("onmode", "c") works (the command don't return, appear
freeze, but the checkpoint occur)... onmode -c don't...