Re: IDS 7.31.UD1 hangs
Posted in 2004
Topics: High Availability & Replication, SQL Development & Query Writing, Server Administration, Platform-Specific Issues, Versions, Editions & End-of-Life
Denis,
we need more info to diagnose the problem.
You said the engine hangs but it has been working before, Can you
identify any changes that has happened same time the engine started
hanging?
it doesn't allow connections.....What type SHM, TCP ?????? Have you
tried both the engine is hang?
onstat shows no problem, what onstat commands are you executing?
could you please post/send onstat -m , onstat -, onstat -g act, onstat
-F, onstat -R | grep dirty when the engine is hang? Any kind of
onstats to identify the problem.
your pstack information is good but you could get the same with onstat
-g stk . I would recomend you first to work with informix tools
rather than OS tools then if we need aditional information you may use
pstack, truss etc.
are you working with HDR?
have you openned a tech suppport case?
esteban.-
"Denis Melnikov" <dm'lnik@regent.ru> wrote in message news:<chhk87$h2a$1@n6.co.ru>...
> Hi, all!
>
> The engine hangs regularly: it stops calculations and doesn't
> allow new connections while 'onstat' shows no problem.
> This way it can hang indefinitely long.
> The only way to bring it up is to 'kill -9' and then 'oninit'.
>
> OS is SPARC Solaris 2.6. IDS 7.31.UD1.
>
> I've catched 'pstack' first time while hanging and once again
> when idle. The only difference that I've noticed is the stack
> of cpu vps (when hanging it run mt_spin_lock_wait()):
>
> hang> >>>>>> 15938: oninit cpu <<<<<<
> hang> lwp#1 ----------
> hang> ef537408 poll (0, 0, 1)
> hang> 00430cf0 mt_spin_lock_wait (1, 65b9f0, 430c00, 1388, 1f4, 2c0062cd) +
> e0
> hang> 00433c1c dospinlock (3870, 65bebc, 1, 433d68, 1, 134) + 618
> hang> 000b14f4 zero_profile (6d3c00, 6c1400, 7, f14c, 248b2, f0cc) + 12c
> hang> 0009febc onmode_monitor (0, a019684, a00a824, 2c56aaf8, a0115a4,
> 6c07b4) + 360
> hang> 00432388 startup (6d0210, 7, 0, 2c006958, 0, 2c006918) + b4
> hang> 0042c344 mt_poll_yield (0, 0, 0, 0, 0, 0) + 110
> hang> lwp#2 ----------
> hang> ef536f24 kaio (6, 2c094068, 68d338, 2c094068, 2c09403c, 118, 0)
> hang> ef536f24 _kaio (6, 2c094068, 68d338, 2c094068, 2c09403c, 118) + 4
> hang> ef539698 _lwp_wait (0, 0, 0, 0, 0, 0) + 1c
>
> idle> >>>>>> 27453: oninit cpu <<<<<<
> idle> lwp#1 ----------
> idle> ef53647c semsys (2, 70004, 2c5d1b90, 1, 12345678)
> idle> 0047dcac net_sm_poll_thread (2c5d2018, 6, 6d00e4, 3efc00, 65d42c, 1)
> + 228
> idle> 00432388 startup (6d0210, 7, 0, 2c006958, 0, 2c006918) + b4
> idle> 0046e23c net_startup_step (0, 0, 0, 0, 0, 0) + 234
> idle> lwp#2 ----------
> idle> ef536f24 kaio (6, 0, 0, 0, 0, 0, 0)
> idle> ef536f24 _kaio (6, 0, 0, 0, 0, 0) + 4
> idle> ef539698 _lwp_wait (0, 0, 0, 0, 0, 0) + 1c
> =======================================================================
> hang> >>>>>> 16155: oninit cpu <<<<<<
> hang> lwp#1 ----------
> hang> ef537408 poll (0, 0, 1)
> hang> 00430cf0 mt_spin_lock_wait (1, 65b9f0, 430c00, 1388, 1f4, 2d762018) +
> e0
> hang> 00444e0c mt_shm_free_blkpool (a005e00, 34b3b544, 34b3b540, 3b0d6000,
> fffffff8, 2e4e2000) + a8
> hang> 001099f8 hjfreemem (33e8d488, b000, 3b0d6000, 33e8d488, 3b0e1082, 1)
> + 28
> hang> 00108a68 hjoin_next (33e8d488, 2e75bbc0, 33e8d420, 33e8c210,
> 33e8d5a8, 3b0d6000) + 338
> hang> 0022dd18 join_next (1, 2e75bd60, 33e8d2f8, 80000000, 1000000,
> 40000000) + b0
> hang> 0010a6c8 insertla_next (6c1398, 0, 0, 1, 0, 2e75bdb0) + d8
> hang> 000f62fc doselinto (33e8d228, 2d926378, 1, 0, 2cc68fd8, 0) + 20c
> hang> 001c279c excommand (1226, 3, 117c, 0, 2e75bbc0, 2d926378) + 6e0
> hang> 001987f8 sq_exselect (6d3c00, 4, 6d3dd0, 2000, 20000, 0) + 68
> hang> 001ab168 sqmain (10, 6d3c00, 612400, 6c1398, 6c9fcc, 0) + 2f0
> hang> 00432388 startup (6d0210, 7, 3158c3e8, 2c006958, 0, 2c006918) + b4
> hang> 00477258 net_wait_for_io (0, 0, 0, 0, 0, 0) + 2c
> hang> lwp#2 ----------
> hang> ef536f24 kaio (6, 0, 0, 0, 0, 0, 0)
> hang> ef536f24 _kaio (6, 0, 0, 0, 0, 0) + 4
> hang> ef539698 _lwp_wait (0, 0, 0, 0, 0, 0) + 1c
>
> idle> >>>>>> 27455: oninit cpu <<<<<<
> idle> lwp#1 ----------
> idle> ef53647c semsys (2, 70008, 2c5ebb90, 1, 12345678)
> idle> 0047dcac net_sm_poll_thread (2c5d3cc0, 6, 6d00e4, 3efc00, 65d42c, 1)
> + 228
> idle> 00432388 startup (6d0210, 7, 0, 2c006958, 0, 2c006918) + b4
> idle> 0046e23c net_startup_step (0, 0, 0, 0, 0, 0) + 234
> idle> lwp#2 ----------
> idle> ef536f24 kaio (6, 0, 0, 0, 0, 0, 0)
> idle> ef536f24 _kaio (6, 0, 0, 0, 0, 0) + 4
> idle> ef539698 _lwp_wait (0, 0, 0, 0, 0, 0) + 1c
> ============================================================================
> hang> >>>>>> 16162: oninit cpu <<<<<<
> hang> lwp#1 ----------
> hang> 002cf378 btusearch (1, 0, 0, 10, 0, 2000) + 234
> hang> 002c2dac btadditem (3604ae44, 3604af54, 0, 11e638f4, 3604ae5e,
> 3604af64) + dc
> hang> 003f4aec wrtrecord (6d3db8, 6c13f0, 1000, 6c16ac, 7, 6c1628) + 5a4
> hang> 003f43a8 rswrite (44, 6c16ac, 315a9880, 3bbbf3e8, 3604b078,
> 315a9880) + d4
> hang> 004bdfe4 fmla_write (ffffffff, 0, ffffffff, 1, 3bbbf3e8, 315a9880) +
> 34c
> hang> 000f7d30 writetab (80000, 359d76f8, 359d7a28, 10000000, 359d7a28, 1)
> + 1e0
> hang> 000f5eb0 doopen (359d76f8, 359d7a28, 6c1398, 31fe0820, 359d7a28,
> 359d76f8) + 78
> hang> 0011fc9c doopen_allany (7, 7, 359d76f8, 3604b29c, 0, 3604b341) + 1c
> hang> 0011ffa4 exsubqm (3604b300, 0, 359d7a28, 3604b344, 7, 25) + 134
> hang> 001204a8 subqcmp (359d7a28, 3604b538, 7, 7, 3604b5f8, 25) + 64
> hang> 0010d28c geval (1, 0, 25, 3604b5f8, 359d7fb8, 25) + 8a0
> hang> 000f1838 gettupl (1000, 1264, 31fe0990, 4000000, 10000, 4000020) +
> 330
> hang> 000ef840 scan_next (315a9df0, 1, 6c1398, 315a9df0, 50, 2d3da85c) +
> 230
> hang> 0022dd18 join_next (1, 104, 5, 80000000, 1000000, 40000000) + b0
> hang> 002305b8 next_row (315a9c20, 31fe0820, 359d8080, 15, 20, 315a9c20) +
> 6c
> hang> 0023091c get_first_row_from_producer (315a9c20, 20, 393122a0, 0,
> 315a9658, 39d17898) + 30
> hang> 002311cc hash_process_all_groups (315a9c20, 39304000, 315a9c20,
> 315a9c20, 2, 39d17898) + 4
> hang> 0022f0b4 group_open (20000, 8000, 315a9c20, 39d17898, 2, 0) + 204
> hang> 00102ed0 sort_open (315a9698, 31fe0920, 315a9170, 0, 2, 315a9718) +
> 88
> hang> 000f69e0 prepselect (6d3dd0, 1226, 6c1398, 31fe0820, 31fe0c20,
> 31fe0830) + 600
> hang> 00198050 open_cursor (1000, 6d3dd0, 12, 31fe0830, 3604bb36, 0) + 5ac
> hang> 00197a80 sq_open (6d3c00, 12, 6d3dd0, 2000, 20000, 0) + c0
> hang> 001ab168 sqmain (6, 6d3c00, 612400, 6c1398, 6c9fcc, 0) + 2f0
> hang> 00432388 startup (6d0210, 7, 32f111a0, 2c006958, 0, 2c006918) + b4
> hang> 00477258 net_wait_for_io (0, 0, 0, 0, 0, 0) + 2c
> hang> lwp#2 ----------
> hang> ef536f24 kaio (6, 0, 0, 0, 0, 0, 0)
> hang> ef536f24 _kaio (6, 0, 0, 0, 0, 0) + 4
> hang> ef539698 _lwp_wait (0, 0, 0, 0, 0, 0) + 1c
>
> idle> >>>>>> 27456: oninit cpu <<<<<<
> idle> lwp#1 ----------
> idle> ef53647c semsys (2, 70000, 6c99b0, 1, a00b9a0)
> idle> 00425174 P (2c51e018, 6d0000, 7, 2c51e018, 1f4, 2c006918) + 24
> idle> 004281a8 idle_processor (c0, 6d3df8, 6d0234, 2c006958, 0, 2c0069
"Esteban Casuscelli" <estebanc00@yahoo.com.ar> ???????/???????? ? ???????? ?????????:
news:8a8f14b9.0409070551.3c59d1f0@posting.google.com...
> Denis,
>
> we need more info to diagnose the problem.
Esteban, thank you for your answer, I've almost lost my hope.
> You said the engine hangs but it has been working before, Can you
> identify any changes that has happened same time the engine started
> hanging?
It started hanging from the very beginning when I've upgraded UC7 to UD1
in 2001. It hanged approx. once in two months. Now we have more data,
dataspaces and users and running CDR (as primary). And now it hangs
approx. twice in a week.
> it doesn't allow connections.....What type SHM, TCP ?????? Have you
> tried both the engine is hang?
No matter what type. Two years ago we used SHM, now it is TCP mainly.
> onstat shows no problem, what onstat commands are you executing?
-m, -u, -F, -R
> could you please post/send onstat -m , onstat -, onstat -g act, onstat
> -F, onstat -R | grep dirty when the engine is hang? Any kind of
> onstats to identify the problem.
If I knew what can help...
######### onstat -m (one can see -25587 error at the end of listing,
######### however it plays no role)
Informix Dynamic Server Version 7.31.UD1 -- On-Line -- Up 2 days 23:30:23 -- 801376
Kbytes
Message Log File: /usr/informix/online.log
13:23:16 Checkpoint Completed: duration was 27 seconds.
13:23:16 Checkpoint loguniq 48102, logpos 0x3163018
13:33:01 Logical Log 48102 Complete.
13:33:02 Logical Log 48102 - Backup Started
13:34:01 Logical Log 48102 - Backup Completed
13:53:34 Checkpoint Completed: duration was 17 seconds.
13:53:34 Checkpoint loguniq 48103, logpos 0x1226018
14:23:56 Checkpoint Completed: duration was 23 seconds.
14:23:56 Checkpoint loguniq 48103, logpos 0x29c73fc
14:39:24 Logical Log 48103 Complete.
14:39:25 Logical Log 48103 - Backup Started
14:40:30 Logical Log 48103 - Backup Completed
14:54:14 Checkpoint Completed: duration was 17 seconds.
14:54:14 Checkpoint loguniq 48104, logpos 0xf28018
14:56:06 listener-thread: err = -25587: oserr = 0: errstr = : Network receive failed.
######### onstat -g act
Threads:
tid tcb rstcb prty status vp-class name
51 2d56a948 2d0568c4 4 running 1cpu onmode_mon
1384 2d53a138 2d05cc74 2 running 4cpu sqlexec
27090 2f6fa3c0 2d07156c 2 running 3cpu sqlexec
######### onstat -F
Fg Writes LRU Writes Chunk Writes
0 0 0
address flusher state data
2d050514 0 I 0 = 0X0
2d050a10 1 I 0 = 0X0
2d050f0c 2 I 0 = 0X0
2d051408 3 I 0 = 0X0
2d051904 4 I 0 = 0X0
2d051e00 5 I 0 = 0X0
2d0522fc 6 I 0 = 0X0
2d0527f8 7 I 0 = 0X0
2d052cf4 8 I 0 = 0X0
2d0531f0 9 I 0 = 0X0
2d0536ec 10 I 0 = 0X0
2d053be8 11 I 0 = 0X0
2d0540e4 12 I 0 = 0X0
2d0545e0 13 I 0 = 0X0
2d054adc 14 I 0 = 0X0
2d054fd8 15 I 0 = 0X0
states: Exit Idle Chunk Lru
######### onstat -R | grep dirty
10142 dirty, 254969 queued, 256000 total, 262144 hash buckets, 2048 buffer size
####################
-g lmx shows much locked mutexes while in normal state there is no ones.
> your pstack information is good but you could get the same with onstat
> -g stk .
As I know it dumps stack of thread. Can it show the process stack? Moreover,
Administrator's Reference says: "This option is not supported on all
platforms and is not always accurate." Is this true?
> I would recomend you first to work with informix tools
> rather than OS tools then if we need aditional information you may use
> pstack, truss etc.
OK, I will try :).
> are you working with HDR?
No, but CDR is the matter.
> have you openned a tech suppport case?
No.
> esteban.-
>
> "Denis Melnikov" <dm'lnik@regent.ru> wrote in message news:<chhk87$h2a$1@n6.co.ru>...
> > Hi, all!
> >
> > The engine hangs regularly: it stops calculations and doesn't
> > allow new connections while 'onstat' shows no problem.
> > This way it can hang indefinitely long.
> > The only way to bring it up is to 'kill -9' and then 'oninit'.
> >
> > OS is SPARC Solaris 2.6. IDS 7.31.UD1.
> >
> > I've catched 'pstack' first time while hanging and once again
> > when idle. The only difference that I've noticed is the stack
> > of cpu vps (when hanging it run mt_spin_lock_wait()):