Re: Assert Fail - onmode hangs
Posted in 2005
Sorry it took so long to get back to these suggestions, but I hope the
new information helps a little.
david@smooth1.co.uk wrote:
> Yes the network socket IDS was listening on. Do netstat -a and this
> will be in time wait.
>
I will try that the next time it happens and view the results.
> Nothing you can do but wait until the TIME_WAIT state times out and
> releases the port.
>
> You should never need to run kill -9 on oninit processes. What is IDS
> doing?
>
This was actually a suggestion from IBM tech support, running a kill -9
on the VP1 process.
> We need the first 40-50 lines of the af file. The first 20 lines ends
> in a stack trace but we
> don't have all of it!
>
sidb1p:informix[/opt/prod/ids_coredumps/tlipita]-[17]% cat -n
af.7b0b5bbc | head -50
1 15:27:41
2 15:27:41 Informix Dynamic Server Version 9.30.UC1X6 Software
Serial Number AAD#J292950
3
4 15:27:41 Assert Failed: Exception Caught. Type: MT_EX_OS,
Context: mem
5 15:27:41 Who: Session(25312, informix@9d7989c7, -1,
452600008)
6 Thread(30499, sqlexec, 18173818, 6)
7 File: mtex.c Line: 368
8 15:27:41 Action: Please notify Informix Technical Support.
9 15:27:41 Stack for thread: 30499 sqlexec
10
11 base: 0x22c12000
12 len: 135168
13 pc: 0x006b25b8
14 tos: 0x22c31db0
15 state: running
16 vp: 6
17
18 0x006b1b24 (oninit)afhandler(0xa80000, 0xa09c9c, 0xa7fe48,
0xa0a0e4, 0x1, 0x7b0b5bbc)
19 0x006b1430 (oninit)affail_interface(0x22c324c8, 0x0,
0x1800a578, 0x0, 0xa0a0e4, 0x170)
20 0x006b5658 (oninit)mt_ex_throw_sig(0x91e0b0, 0xa07248, 0x0,
0xa0a0c4, 0xa6fce4, 0x94cce0)
21 0x00685d4c (oninit)afsig_handler(0xb, 0x22c32a78, 0x22c327c0,
0x0, 0x0, 0x0)
22 0x0028bb78 (oninit)find_opcinst(0x0, 0x0, 0x19516b48,
0x1adf86e0, 0x0, 0x1adf86dc)
23 0x001fc3fc (oninit)opinit (0x0, 0x1, 0x2000, 0x0, 0x0,
0x1adf86dc)
24 0x0021c9dc (oninit)op_opinit(0x21d565f8, 0x22c32d6c,
0x22c32d68, 0x22c32d78, 0x22c32d68, 0x22c32d6c)
25 0x001faf54 (oninit)sqoptim (0x21d565f8, 0xa709ec, 0xa8653c,
0x0, 0x1fa638, 0x0)
26 0x0031f664 (oninit)bldstructs(0xa86400, 0x20000, 0x1000,
0x1082, 0x0, 0x22c32e4c)
27 0x0031f3e0 (oninit)sqcmd (0xa86400, 0xa8653c, 0x1,
0x1fd3b018, 0x21d565f8, 0x0)
28 0x0031f01c (oninit)sq_cmnd (0xa86400, 0x0, 0x0, 0x30, 0xa8653c,
0xa709ec)
29 0x0039409c (oninit)sqmain (0x2, 0xa86400, 0x91cc00, 0xa79df4,
0x1, 0x9421f4)
30 0x00692b28 (oninit)startup (0xa801f0, 0x0, 0x0, 0x1800b8e0,
0x18682018, 0x1800b8a0)
31 0x00686d5c (oninit)idle_processor(0x0, 0x0, 0x0, 0x0, 0x0, 0x0)
32 0x00000000 (*nosymtab*)0x0
33
34
35 15:27:41 See Also:
/opt/prod/ids_coredumps/tlipita/af.7b0b5bbc
36
37 ---------------------------------
38 Begin System Alarm Program Output
39 ---------------------------------
40
41 Assertion Failure Type: FAILURE
42 Host Name: sidb1p
43 Database Server Name: tlipita
44 Time of failure: Wed Nov 9 15:27:41 EST 2005
45 AF file:
/opt/prod/ids_coredumps/tlipita/af.7b0b5bbc
46 Shared memory file: None
47 System Blocking: OFF
48
49
50 ===========------------- - - - - - -
sidb1p:informix[/opt/prod/ids_coredumps/tlipita]-[18]%
> 9.30 was an unstable version especially a UC1 version. Consider
> upgrading to
> at least 9.40.latest.
>
This is understood (and though it was before my time here, the reason
that we had an "X" version of 9.30). We are in the process of moving
to 9.40.UC4 as soon as we have verified all of our functionality in our
DEV environment.
>
> Run onstat -g stk all and exclude the sqlexec and btcleaner threads.
> What else is there?
>From the AF file, there was onstat -g stk data, including thread
entries for aslogflush, sapp_listener, onmode_mon, tlitcplst,
sm_discon, sm_listen, tlitcppoll, sm_poll, main_loop, msc vp, aio vp,
pio vp, lio vp, and several for kaio and flush_sub. The threads were
in various states, including cond wait, ready, running, and sleeping.
Is there anything specific that you would be looking for in the stk
data? I did notice a lot of the stacks ended with something similar to
the following:
0x00000000 (*nosymtab*)0x0
ScottishPoet, I assume you were talking about DUMPSHMEM, but it was
already set to 0.