Re: Server hangs with "mutex wait nsf.lock"
Posted in 2003
Topics: Server Administration, Platform-Specific Issues
The culprit is your NETTYPE entry for tcp connections. TCP listeners should
NEVER be set up in CPU VPs. This causes several problems. First the CPU VP
cannot practically block on the sockets waiting for new connections and
requests
from existing connections so it must poll the sockets, this burns CPU time to
no purpose and pushes the system call overhead way up tying up the OS system
call threads. Now, obviously, the CPU VP(s) have better things to do than poll
so they do so only when they have time at certain breakpoints in their runloop.
This means that if the CPU VPs are busy responding to queries they will not
poll the sockets and so you get the apparent lockups you are experiencing
especially if there are also shared memory connections.
Also I notice that your shared memory NETTYPE only specifies one listener poll
thread, shared memory listeners should always be run in ALL CPU VPs for best
responsiveness and to keep CPU VP#1 (which handles the first poll thread) from
running flat out and taking 90% of the work for itself.
Change the NETTYPE settings as follows:
NETTYPE ipcshm,3,50,CPU # Configure poll thread(s) for nettype
NETTYPE tlitcp,2,300,NET # Configure poll thread(s) for nettype
The problem will vanish.
Art S. Kagel
----- Original Message -----
From: Patrick McD.... <patmcd@intraware.com>
At: 1/30 18:51
> Long post..
>
> this occurs a few times a week. Shared memory connections work fine,
> tcp connections will "hang". So when this condition occurs, it seems as
> if all sessions no longer get served except for one. I have seen this
> occur with as little as 2 sessions. If I fire up a new dbaccess
> session, my window will show up blank and hang and will not show up in
> onstat -u.. The session that is doing the blocking usually clears in a> minute or two, but occasionally blocks longer ( most we let it go was 30
> minutes in a test). If it takes more than a couple of minutes, we
> bounce the engine. When bringing down, there is an error something
> like, "server not responding, timed out after 120 seconds" although the
> server is brought down.
>
> E450 with 4 procs.
> Informix Dynamic Server Version 9.30.UC3
> Solaris 2.6
>
> Basically a blob image file server. Blobs ranging from 1k to 1G are
> stored and served from machine.
>
> Has anyone experienced this "mutex wait nsf.lock" behavior? Informix
> tech support has been looking into this, but they have been unsuccessful
> in diagnosing.
>
> All shortened for space..
>
> onconfig
>
> ----------
> ## Tried several different NETTYPE configs.. currently
>
> NETTYPE ipcshm,1,50,CPU # Configure poll thread(s) for nettype
> #NETTYPE tlitcp,2,50,NET # Configure poll thread(s) for nettype
> NETTYPE tlitcp,2,300,CPU # Configure poll thread(s) for nettype
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 3 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> NOAGE 1 # Process aging> ## Messed with this. Tried AFF and no AFF and AFF 2 processors etc..
> same problem.
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors
> ---------->
> onstat -u
> ==== the sessions with "S" appear hung.
> ---------
> 14152b34 ---P--- 11338 testftp - 0 0 1 250 0
> 14154344 ---P--- 13400 testftp - 0 0 1 250 0
> 1415977c S--P--- 13396 testftp - 1417fdc8 0 0 23500 0
> 1415a384 ---P--- 13401 testftp - 0 0 1 250 0> 1415a988 S--P--- 13398 informix 2 1417fdc8 0 0 27500 0
> ---------
>
> onstat -g ath
> ---------> 11422 14d1ddc0 14152b34 2 sleeping forever 1cpu sqlexec
> 13480 14f2c778 1415977c 2 mutex wait nsf.lock 1cpu sqlexec
> 13482 159451b8 1415a988 2 mutex wait nsf.lock 1cpu srvinfx
> 13484 14dcc2b0 14154344 3 running 4cpu sqlexec
> 13485 14ce1ed0 1415a384 2 sleeping forever 1cpu sqlexec
> ---------
>
> onstat -g lmx:
> ----------> Locked mutexes:
> mid addr name holder lkcnt waiter waittime
> 3258 1417fdc8 nsf.lock 13484 0 13 986
> 13482 972
> 13480 757
> 5757 15287988 slcommon_ct 13480 0
> 6020 15945840 slcommon_ct 13482 0
>
> Number of mutexes on VP free lists: 568
> ----------
>
> onstat -g wmx:
> -----------> Mutexes with waiters:
> mid addr name holder lkcnt waiter waittime
> 3258 1417fdc8 nsf.lock 13484 0 13 986
> 13482 972
> 13480 757
>
> Number of mutexes on VP free lists: 568
> -----------
> onstat -k all
> ---------> Stack for thread: 13484 sqlexec <--looks like session blocking others.> base: 0x14ce6000
> len: 36864
> pc: 0x006995b8
> tos: 0x14cee848
> state: running
> vp: 4
>
> Stack for thread: 13482 srvinfx <--- looks like a "blocked" session> base: 0x14cc1000
> len: 36864
> pc: 0x006995b8
> tos: 0x14cc9778
> state: mutex wait
> vp: 1
>
> 0x00691580 (oninit)mt_lock_wait(0xa8b428, 0xa8b400, 0x2, 0x1417fdc8,
> 0x1, 0x0)
> 0x006916f8 (oninit)mt_lock (0x1417fdc8, 0x1417fdc8, 0x0, 0x1400a8e0,
> 0x0, 0x1400a8a0)
> 0x006e3fa0 (oninit)sl_lock (0x1457c6d0, 0x157b6d88, 0x7, 0x1400a8e0,
> 0x14cc99cc, 0x1400a8a0)
> 0x006e3350 (oninit)net_nsf_closefd(0x9581f8, 0x14ca5e30, 0x14cc9a3c,
> 0xa, 0x4, 0x14ca5e30)
> 0x00717640 (oninit)disctli (0x4, 0xa, 0x14d07d04, 0x14ca5e30,
> 0x15d95da8, 0x1)
> 0x007066b8 (oninit)tlDiscon(0x14cdf030, 0xa, 0x2, 0x2, 0x1, 0xa9f1b8)
> 0x007037fc (oninit)slSQIdiscon(0xa, 0xa, 0xffffffff, 0x147b7d1c, 0x0,
> 0x14d07d04)
> 0x006fe050 (oninit)pfDiscon(0x14cdf030, 0x147b7ca8, 0x14cc9d5c, 0x0,
> 0xa, 0x14d07d04)
> 0x006f5aac (oninit)cmDiscon(0x14cdf0a8, 0xa, 0x14cc9d5c, 0xa,
> 0x14cdf030, 0x14cc9c38)
> 0x006f2a08 (oninit)ascAbort(0x14d07d04, 0x14cdf0a8, 0x14cc9d58, 0x0,
> 0x14cdf0a8, 0xef508c10)
> 0x006d22a0 (oninit)asfExit (0x1, 0x0, 0x0, 0x14cdf0a8, 0x14cdf030,
> 0x14cc9d58)
> 0x006d2078 (oninit)ASF_Call(0xa0, 0x0, 0x14cdf0a8, 0xa, 0x14cc9d58,
> 0x14cc9d61)
> 0x003c5bc8 (oninit)sqscb_cleanup(0xa91400, 0x927400, 0xa85000, 0x1000,
> 0xa9174c, 0x1)
> 0x0069e034 (oninit)destroy_session(0x927400, 0xa7bc00, 0xa7bf2c,
> 0x14119950, 0x927438, 0x14119950
> )
> 0x0039c18c (oninit)sq_exit (0xa7b800, 0x0, 0x14d07018, 0x0, 0x1470f018,
> 0x4e)
> 0x0039810c (oninit)sqmain (0x38, 0xa91400, 0x927400, 0xa85004, 0x1,
> 0x94c6c4)
> 0x00699b38 (oninit)startup (0xa8b400, 0x0, 0x0, 0x1400a8e0, 0x1462a018,
> 0x1400a8a0)
> 0x0068dd6c (oninit)idle_processor(0x0, 0x0, 0x0, 0x0, 0x0, 0x0)
> 0x00000000 (*nosymtab*)0x0
>
> ---------
>
> Box is under very low load..
>
> ---------
> mpstat 1 5
> CPU minf mjf xcal intr ithr csw icsw migr smtx srw syscl usr sys wt
> idl
> 0 7 0 426 126 115 197 11 12 28 0 2017 17 7
> 0 76
> 1 6 0 560 408 200 146 10 10 20 0 2116 27 10
> 0 63
> 2 6 0 312 111 100 133 10 10 20 0 3120 27 11
> 0 61
> 3 6 0 269 265 259 85 6 6 19 0 1783 31 13
> 0 56
> CPU minf mjf xcal intr ithr csw icsw migr smtx srw syscl usr sys wt
> idl
> 0 0 0 0 100 100 18 0 1 0 0 7 0 0 0
> 100
> 1 3 0 312 400 200 34 0 0 1 0 39 0 0 0
> 100
> 2 0 0 0 101 100 15 0 2 0 0 16 0 0 0
> 100@@NL@
Art S. Kagel wrote:
> Also I notice that your shared memory NETTYPE only specifies
> one listener poll thread, shared memory listeners should always
> be run in ALL CPU VPs for best responsiveness and to keep
> CPU VP#1 (which handles the first poll thread) from
> running flat out and taking 90% of the work for itself.
>
> Change the NETTYPE settings as follows:
>
> NETTYPE ipcshm,3,50,CPU # Configure poll thread(s) for nettype
> NETTYPE tlitcp,2,300,NET # Configure poll thread(s) for nettype
What about an engine that has almost no shared memory connections
(dedicated database server) and about 70-80 concurrent connections
over TCP/IP? Machine has 2 processors (Intel PIII 933MHz) running
under Linux.
Regards, Richard
--
+-----------------------------+--------------------------------------+
| Dr. med Richard Spitz | Mail: spitz@ana.med.uni-muenchen.de |
| Klinik für Anaesthesiologie | Tel : +49-89-7095-6110 |
| Klinikum der Univ. München | FAX : +49-89-7095-6420 |
| 81366 München, Germany | GSM : +49-172-8933578 |
+-----------------------------+--------------------------------------+