Mutex spins, polling. and performance - Help Please?
Posted in 2000
Topics: Performance & Tuning, Server Administration, Platform-Specific Issues, Clustering, Grid & MACH11
Let me start off by saying that I am a System Admin - not a programmer
and not a DBA. Please try to take it easy on me if I ask something
stupid ;).
I have a Sun 6000 with 14 250MHz CPUs, 7GB RAM, running Solaris 2.6 with
a recent patch cluster, and Informix 7.30UC3-1.
Performance is so-so - no complaints, but no praise either. In
preparing to move it to a new 6500 (6000 is a lease), I took a cursory
look at performance. I noticed that, even though the machine is
consistently over 50% idle, the load average is usually around 5 to 7+.
Further investigation shows gobs of mutex spins - from 2500 to over 10K
per second! Lockstat indicates that these are on plocks and that the
callers are always poll or polldel from oninit processes.
Reading the Informix OnLine Dynamic Server Admin Guide and some other
posts on this group led me to suspect onconfig issues, but I get out of
my element pretty quickly here. Basically, I want to be able to speak
semi-intelligently when I talk to the DBAs. These are the points that I
wonder about:
The entry:
NETTYPE tlitcp,10,250,NET
I thought I might find this set as a CPU entry, which would prevent it
from blocking and force it to poll. But this isn't the case. Am I on
the money here?
The entries:
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 8 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vpsto one
As I said, the box has 14 CPUs - shouldn't NUMCPUVPS be 14? This has
probably just not been updated after addition of CPUs, but will it make
that big a difference?
Anywhere else I should look? Thanks a bunch.
Wally Blackburn
Sr. Unix Admin
Sent via Deja.com http://www.deja.com/
Before you buy.
In article <8o415q$g3j$1@nnrp1.deja.com>, <wallyblackburn@hotmail.com> wrote: >I have a Sun 6000 with 14 250MHz CPUs, 7GB RAM, running Solaris 2.6 with >a recent patch cluster, and Informix 7.30UC3-1. > >Performance is so-so - no complaints, but no praise either. In >preparing to move it to a new 6500 (6000 is a lease), I took a cursory >look at performance. I noticed that, even though the machine is >consistently over 50% idle, the load average is usually around 5 to 7+. I don't have any comments on the other stuff you mentioned, but the fact that the load average is 7 on a system with 14 CPUs isn't a problem, if I understand load averages correctly. The load average is average of the number of processes (or maybe of LWPs) that are eligible to be run at any given time. So, on a single CPU system, a load average of 0 means that no processes are eligible to run, i.e. that the system is completely idle. A load average of 1 means that at any given time, one process is ready to run. So, if the machine is ideally loaded (exactly one process that's runnable, and it's always runnable and always running), then the load average will be 1. On an ideally-loaded system with 14 processors, you'd have 14 processes, each of which is runnable all the time and each of which is continuously running on a processor. In that case, there is never a case where a process is runnable but can't access the CPU. Yet, 14 processes are runnable, so the load average would be 14. My point is that, a 14 processor machine is just as busy when its load average is 14 as a 1 processor machine is when its load average is 1. Of course, what Unix really needs in my opinion is not something that shows the average load over a period of time, but something that shows the peak load. The load average could easily be less than the number of processors but there still might be moments when processes are runnable but processors aren't available. In that case, you'd like to know the peak so that you can detect that. A nice alternative would be for there to be some way to keep track (on a process-by-process basis) of how many times or what percentage of the time a process was runnable but a processor wasn't available. (There may even be a tool I'm not aware of that already does this...) - Logan [1] or are in short-term I/O wait, but let's ignore that for now.
A lot on the number of cpuvps depends on the number of threads in the ready
queue. If the ready queue has several threads on it, then you might
consider adding another cpuvp. (onstat -g rea).
Where are your client programs running? If the clients are running on the
same physical machine, then it might make sense to not use all of the
processors for the cpuvps. If they are connecting to the database via tcp,
then you might as well up the number of cpuvps.
If your clients are running on a different machine and are connecting via
TCP, then yes - 8 CPUVPS on a 14 CPUVP would max out at 57%, which sounds
like your situation.
The poll commands could be IO completion as well as NET. That might mean
that you are starving the buffer pool.
I don't know for sure. You didn't submit an onstat -p, so I really can't
tell what your basic profile looks like.
It might be worth it to submit the output of onstat -p, your onconfig,
onstat -g rea (many), so that other folks within the news group can maketheir comments. There's quite a few folks here that are fairly
knowledgable.
wallyblackburn@hotmail.com wrote:
> Let me start off by saying that I am a System Admin - not a programmer
> and not a DBA. Please try to take it easy on me if I ask something
> stupid ;).
>
> I have a Sun 6000 with 14 250MHz CPUs, 7GB RAM, running Solaris 2.6 with
> a recent patch cluster, and Informix 7.30UC3-1.
>
> Performance is so-so - no complaints, but no praise either. In
> preparing to move it to a new 6500 (6000 is a lease), I took a cursory
> look at performance. I noticed that, even though the machine is
> consistently over 50% idle, the load average is usually around 5 to 7+.
>
> Further investigation shows gobs of mutex spins - from 2500 to over 10K
> per second! Lockstat indicates that these are on plocks and that the
> callers are always poll or polldel from oninit processes.
>
> Reading the Informix OnLine Dynamic Server Admin Guide and some other
> posts on this group led me to suspect onconfig issues, but I get out of
> my element pretty quickly here. Basically, I want to be able to speak
> semi-intelligently when I talk to the DBAs. These are the points that I
> wonder about:
>
> The entry:
> NETTYPE tlitcp,10,250,NET>
> I thought I might find this set as a CPU entry, which would prevent it
> from blocking and force it to poll. But this isn't the case. Am I on
> the money here?
>
> The entries:
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 8 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> As I said, the box has 14 CPUs - shouldn't NUMCPUVPS be 14? This has
> probably just not been updated after addition of CPUs, but will it make
> that big a difference?
>
> Anywhere else I should look? Thanks a bunch.
>
> Wally Blackburn
> Sr. Unix Admin
>
> Sent via Deja.com http://www.deja.com/
> Before you buy.
wallyblackburn@hotmail.com wrote in message <8o415q$g3j$1@nnrp1.deja.com>...
>Let me start off by saying that I am a System Admin - not a programmer
>and not a DBA. Please try to take it easy on me if I ask something
>stupid ;).
>
>I have a Sun 6000 with 14 250MHz CPUs, 7GB RAM, running Solaris 2.6 with
>a recent patch cluster, and Informix 7.30UC3-1.
>
Move to at 7.31.UC4. 7.31 has a lot of performance
improvements
>Performance is so-so - no complaints, but no praise either. In
>preparing to move it to a new 6500 (6000 is a lease), I took a cursory
>look at performance. I noticed that, even though the machine is
>consistently over 50% idle, the load average is usually around 5 to 7+.
>
>Further investigation shows gobs of mutex spins - from 2500 to over 10K
>per second! Lockstat indicates that these are on plocks and that the
>callers are always poll or polldel from oninit processes.
>
>Reading the Informix OnLine Dynamic Server Admin Guide and some other
>posts on this group led me to suspect onconfig issues, but I get out of
>my element pretty quickly here. Basically, I want to be able to speak
>semi-intelligently when I talk to the DBAs. These are the points that I
>wonder about:
>
>The entry:
>NETTYPE tlitcp,10,250,NET>
>I thought I might find this set as a CPU entry, which would prevent it
>from blocking and force it to poll. But this isn't the case. Am I on
>the money here?
>
Yes, check the FAQ at www.smooth1.demon.co.uk.
>The entries:
>MULTIPROCESSOR 1 # 0 for single-processor, 1 for>multi-processor
>NUMCPUVPS 8 # Number of user (cpu) vps
>SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps>to one
>
>As I said, the box has 14 CPUs - shouldn't NUMCPUVPS be 14?
13 - leave one for the kernel to run on. Normally the first CPU
since under some OS's and on certain hardware CPU 0 handles
more interrupts than other CPUs.
This has
>probably just not been updated after addition of CPUs, but will it make
>that big a difference?
>
>Anywhere else I should look? Thanks a bunch.
>
>Wally Blackburn
>Sr. Unix Admin
Look for a tool called Virtual Adrian, it can be use to monitor
Sun Servers.
>
>
>Sent via Deja.com http://www.deja.com/
>Before you buy.
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g