Mutex wait
Posted in 2011
Vikas reported heavy mutex waits (nsf.lock, pt_*, hash, netnorm, <unknown>) and near outages on IDS 11.50.FC8W2 on HP-UX/Itanium with 350+ sessions. Responders explained the mutex names (netnorm = waiting on client input, pt_* = partition/extent growth, nsf.lock = likely culprit, possibly tunable via NUMFDSERVERS), suggested checking onstat -g lmx, buffer/LRU contention and chunk service times, and pointed to APAR IC75072 (ontape/sqlexec deadlock, fixed in FC9). Vikas noted the problem disappeared when the IBM Guardium monitoring tool was stopped. No confirmed resolution is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning
Hello All,
IDS 11.50.FC8W2 on HP-Unix
Reposting this issue as I did not got any response for my earlier thread and
yesterday we again saw a near outage due to mutex waits in the database.
We experinced some dip in our performance(load jobs) and while checking the
user I saw quite many sessions with S- flag indicating these were waiting on
mutex.
I checked in onstat -g ath and found the following different waits:
8815514 c000000a2ee36030 c000000a8fa7db48 1 mutex wait pt_940000a 12cpu sqlexec
9376796 c000000a54e18c10 c000000a865a1550 1 mutex wait nsf.lock 11cpu* sqlexec
9331804 c000000a2e8c1720 c000000a7ddc4888 1 mutex wait hash 0 175 10cpu sqlexec
8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait <unknown> 9cpu sqlexec
8897413 c000000a511d3ae0 c000000a3777d370 1 mutex wait netnorm 1cpu sqlexec
8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait hash 0 286 5cpu sqlexec
Not all of these are waiting for nsf.lock! and I dont know what the others
mean(hash 0 175, pt_940000a) ?
Drilling down from ath - users - ses I saw the SQLs waiting for mutex, however
these are normal SQLs running almost all the time while loading jobs run.
onstat -lmx shows some holding and lots of waiters.
Any comments and suggestions, why this could be happening in my system?
Regards,
Vikas
Ignore the netnorm waits. That just means that the thread is waiting for
input from the client application.
pt__ are locks on the partition. Normally that means that threads are
expanding a tablespace.
nsf.lock is probably going to be the problem child.
From: "VIKAS HIVARKAR" <vikas.hivarkar@gmail.com>
To: ids@iiug.org
Date: 09/20/2011 11:01 PM
Subject: Mutex wait [24974]
Sent by: ids-bounces@iiug.org
Hello All,
IDS 11.50.FC8W2 on HP-Unix
Reposting this issue as I did not got any response for my earlier thread
and
yesterday we again saw a near outage due to mutex waits in the database.
We experinced some dip in our performance(load jobs) and while checking the
user I saw quite many sessions with S- flag indicating these were waiting
on
mutex.
I checked in onstat -g ath and found the following different waits:
8815514 c000000a2ee36030 c000000a8fa7db48 1 mutex wait pt_940000a 12cpu
sqlexec
9376796 c000000a54e18c10 c000000a865a1550 1 mutex wait nsf.lock 11cpu*
sqlexec
9331804 c000000a2e8c1720 c000000a7ddc4888 1 mutex wait hash 0 175 10cpu
sqlexec
8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait <unknown> 9cpu
sqlexec
8897413 c000000a511d3ae0 c000000a3777d370 1 mutex wait netnorm 1cpu sqlexec
8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait hash 0 286 5cpu
sqlexec
Not all of these are waiting for nsf.lock! and I dont know what the others
mean(hash 0 175, pt_940000a) ?
Drilling down from ath - users - ses I saw the SQLs waiting for mutex,
however
these are normal SQLs running almost all the time while loading jobs run.
onstat -lmx shows some holding and lots of waiters.
Any comments and suggestions, why this could be happening in my system?
Regards,
Vikas
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
How many concurrent sessions do you normally have?
How many LRU queues are configured?
How many physical processor cores on the system? How fast are they?
Itanium or PA-RISC?
How many CPU VPs are configured?
What are your settings for:
VPCLASS cpu
MULTIPROCESSOR
RESIDENT
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Tue, Sep 20, 2011 at 11:58 PM, VIKAS HIVARKAR
<vikas.hivarkar@gmail.com>wrote:
> Hello All,
>
> IDS 11.50.FC8W2 on HP-Unix
>
> Reposting this issue as I did not got any response for my earlier thread
> and
> yesterday we again saw a near outage due to mutex waits in the database.
>
> We experinced some dip in our performance(load jobs) and while checking the
> user I saw quite many sessions with S- flag indicating these were waiting
> on
> mutex.
>
> I checked in onstat -g ath and found the following different waits:
> 8815514 c000000a2ee36030 c000000a8fa7db48 1 mutex wait pt_940000a 12cpu
> sqlexec
> 9376796 c000000a54e18c10 c000000a865a1550 1 mutex wait nsf.lock 11cpu*
> sqlexec
> 9331804 c000000a2e8c1720 c000000a7ddc4888 1 mutex wait hash 0 175 10cpu
> sqlexec
> 8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait <unknown> 9cpu
> sqlexec
> 8897413 c000000a511d3ae0 c000000a3777d370 1 mutex wait netnorm 1cpu sqlexec
> 8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait hash 0 286 5cpu
> sqlexec
>
> Not all of these are waiting for nsf.lock! and I dont know what the others
> mean(hash 0 175, pt_940000a) ?
>
> Drilling down from ath - users - ses I saw the SQLs waiting for mutex,
> however
> these are normal SQLs running almost all the time while loading jobs run.
>
> onstat -lmx shows some holding and lots of waiters.>
> Any comments and suggestions, why this could be happening in my system?
>
> Regards,
> Vikas
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--000e0cd489c498dad804ad6ca1e3
we are running on AIX and ran into this issue on FC8W2 ... maybe it is the
same issue?
IC75072: RACE CONDITION BETWEEN ONTAPE AND SQLEXEC THREAD CAN CAUSE DEADLOCK
ON DBS MUTEX.
Support informed us that it was fixed in 11.50.FC9
________________________________________
From: ids-bounces@iiug.org [ids-bounces@iiug.org] On Behalf Of Art Kagel
[art.kagel@gmail.com]
Sent: Wednesday, September 21, 2011 12:13 AM
To: ids@iiug.org
Subject: Re: Mutex wait [24977]
How many concurrent sessions do you normally have?
How many LRU queues are configured?
How many physical processor cores on the system? How fast are they?
Itanium or PA-RISC?
How many CPU VPs are configured?
What are your settings for:
VPCLASS cpu
MULTIPROCESSOR
RESIDENT
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Tue, Sep 20, 2011 at 11:58 PM, VIKAS HIVARKAR
<vikas.hivarkar@gmail.com>wrote:
> Hello All,
>
> IDS 11.50.FC8W2 on HP-Unix
>
> Reposting this issue as I did not got any response for my earlier thread
> and
> yesterday we again saw a near outage due to mutex waits in the database.
>
> We experinced some dip in our performance(load jobs) and while checking the
> user I saw quite many sessions with S- flag indicating these were waiting
> on
> mutex.
>
> I checked in onstat -g ath and found the following different waits:
> 8815514 c000000a2ee36030 c000000a8fa7db48 1 mutex wait pt_940000a 12cpu
> sqlexec
> 9376796 c000000a54e18c10 c000000a865a1550 1 mutex wait nsf.lock 11cpu*
> sqlexec
> 9331804 c000000a2e8c1720 c000000a7ddc4888 1 mutex wait hash 0 175 10cpu
> sqlexec
> 8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait <unknown> 9cpu
> sqlexec
> 8897413 c000000a511d3ae0 c000000a3777d370 1 mutex wait netnorm 1cpu sqlexec
> 8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait hash 0 286 5cpu
> sqlexec
>
> Not all of these are waiting for nsf.lock! and I dont know what the others
> mean(hash 0 175, pt_940000a) ?
>
> Drilling down from ath - users - ses I saw the SQLs waiting for mutex,
> however
> these are normal SQLs running almost all the time while loading jobs run.
>
> onstat -lmx shows some holding and lots of waiters.>
> Any comments and suggestions, why this could be happening in my system?
>
> Regards,
> Vikas
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--000e0cd489c498dad804ad6ca1e3
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
It would be interesting to see the "onstat -g lmx" output. Which mutexs have
waiters and how long are they waiting?
nsf.lock is kind of a "classic". There are several things that can be done
to solve it.
The number of new connections per second is important.
The cpu contention is also an important factor
Some new configuration parameters can help
Regards
On Wed, Sep 21, 2011 at 4:58 AM, VIKAS HIVARKAR
<vikas.hivarkar@gmail.com>wrote:
> Hello All,
>
> IDS 11.50.FC8W2 on HP-Unix
>
> Reposting this issue as I did not got any response for my earlier thread
> and
> yesterday we again saw a near outage due to mutex waits in the database.
>
> We experinced some dip in our performance(load jobs) and while checking the
> user I saw quite many sessions with S- flag indicating these were waiting
> on
> mutex.
>
> I checked in onstat -g ath and found the following different waits:
> 8815514 c000000a2ee36030 c000000a8fa7db48 1 mutex wait pt_940000a 12cpu
> sqlexec
> 9376796 c000000a54e18c10 c000000a865a1550 1 mutex wait nsf.lock 11cpu*
> sqlexec
> 9331804 c000000a2e8c1720 c000000a7ddc4888 1 mutex wait hash 0 175 10cpu
> sqlexec
> 8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait <unknown> 9cpu
> sqlexec
> 8897413 c000000a511d3ae0 c000000a3777d370 1 mutex wait netnorm 1cpu sqlexec
> 8660524 c000000a3511a030 c000000a7ddf58d8 1 mutex wait hash 0 286 5cpu
> sqlexec
>
> Not all of these are waiting for nsf.lock! and I dont know what the others
> mean(hash 0 175, pt_940000a) ?
>
> Drilling down from ath - users - ses I saw the SQLs waiting for mutex,
> however
> these are normal SQLs running almost all the time while loading jobs run.
>
> onstat -lmx shows some holding and lots of waiters.>
> Any comments and suggestions, why this could be happening in my system?
>
> Regards,
> Vikas
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--0016364ef0962ec5c204ad6ed548
Hello All,
One important observation:
When Guardium (monitoring tool from IBM) is stoped I do not see this bad mutex
behaviour!
@Madison Pruet:
Yes, nsf.lock waits are a problem, however in my case If I see 200 mutex waits
(onstat -g ath |grep mutex / onstat -g lmx)then only 3-5 are nsf.lock hence
not sure if the performance issue is only because of these 3 to 4 nsf.locks.
https://www-304.ibm.com/support/docview.wss?uid=swg21145897
says that if we see lots of nsf.locks then setting NUMFDSERVERS to 1 for
11.50. could resolve the issue.
I see large number of other types of mutex waits, should I try setting this
parameter( it requires instance restart)?
@Art:
concurrent sessions: 350+
BUFFERPOOL
size=2K,buffers=7200000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=5.000000
BUFFERPOOL
size=4K,buffers=3000000,lrus=512,lru_min_dirty=10.000000,lru_max_dirty=20.000000
BUFFERPOOL
size=16K,buffers=750000,lrus=512,lru_min_dirty=10.000000,lru_max_dirty=20.000000
AUTO_LRU_TUNING 1
Number of CPUs = 16
Clock speed = 1598 MHz
Bus speed = 533 MT/s
processor family: 32 Intel(R) Itanium 2 9000 series
VPCLASS cpu,num=12,noage
RESIDENT 1
MULTIPROCESSOR 1
@Joe:
We take the backup during night time, during day time only logical log backup
is running, so will check the bug you shared.
@Fernando:
I do see couple of threads and many waiter threads against each holder,
however I am not able to trace them further as the old onces are replaced by
new onces faster than my tracking speed.
Any suggestions for further investigation.
Regards,
Vikas
40GB of buffers for 350 users? Are they all being used? In the onstat -P
and onstat -g buf reports, are the lots of buffers in the 'other' column for
partnum zero?
Separate issue: I was suspecting LRU contention, but with 512 LRUs in each
of three buffer pools, that's unlikely. Just to be sure, what are your
Bufwaits Ratio and Buffer Turnover Rates (for each buffer pool separately if
you can)? Maybe there is contention for a very small subset of the buffer
pool? Slow IO perhaps. Run onstat -g iof or onstat -g ppf and look at the
service times for each chunk and partition to check that. Service times
should be under 10ms ideally.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Wed, Sep 21, 2011 at 5:20 AM, VIKAS HIVARKAR
<vikas.hivarkar@gmail.com>wrote:
> Hello All,
>
> One important observation:
> When Guardium (monitoring tool from IBM) is stoped I do not see this bad
> mutex
> behaviour!
>
> @Madison Pruet:
> Yes, nsf.lock waits are a problem, however in my case If I see 200 mutex
> waits
> (onstat -g ath |grep mutex / onstat -g lmx)then only 3-5 are nsf.lock hence
> not sure if the performance issue is only because of these 3 to 4
> nsf.locks.
>
> https://www-304.ibm.com/support/docview.wss?uid=swg21145897
> says that if we see lots of nsf.locks then setting NUMFDSERVERS to 1 for
> 11.50. could resolve the issue.
>
> I see large number of other types of mutex waits, should I try setting this
> parameter( it requires instance restart)?
>
> @Art:
> concurrent sessions: 350+
>
> BUFFERPOOL>
>
size=2K,buffers=7200000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=5.000000
> BUFFERPOOL>
>
size=4K,buffers=3000000,lrus=512,lru_min_dirty=10.000000,lru_max_dirty=20.000000
> BUFFERPOOL>
>
size=16K,buffers=750000,lrus=512,lru_min_dirty=10.000000,lru_max_dirty=20.000000
> AUTO_LRU_TUNING 1
>
> Number of CPUs = 16
> Clock speed = 1598 MHz
> Bus speed = 533 MT/s
> processor family: 32 Intel(R) Itanium 2 9000 series
>
> VPCLASS cpu,num=12,noage
> RESIDENT 1
> MULTIPROCESSOR 1>
> @Joe:
> We take the backup during night time, during day time only logical log
> backup
> is running, so will check the bug you shared.
>
> @Fernando:
> I do see couple of threads and many waiter threads against each holder,
> however I am not able to trace them further as the old onces are replaced
> by
> new onces faster than my tracking speed.
>
> Any suggestions for further investigation.
>
> Regards,
> Vikas
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--00151773d9f2cf6f4704ad725f82
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g