Re: RE: onstat -g glo: cpu efficiency
Posted in 2016
Thread about poor CPU efficiency (~60%) and long spins on an IBM Power 750 LPAR with 16 cores, SMT4 (64 logical CPUs), 54 CPU VPs and 2200 connections, with occasional 60-second SQL timeouts. Advice: turn SMT down to 2 (or off), since Informix runs better with fewer hardware threads, and cut the soctcp poll threads drastically (each handles 350-500 connections, so ~5-7 NET VPs, not 20), judging health via onstat -g rea/act rather than spin counts. Long spins were said to be normal and not necessarily bad, often rising when CPUs are faster. The poster moved to SMT2 and doubled cores, saw fewer active-queue threads but more long spins; no definitive final outcome on the timeouts is recorded, though another user confirmed SMT2 fixed similar random performance problems.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning
We have an IBM Power 750 (Power 7+) with 16 cores defined running SMT so that
the cores appears to the OS as 64 cores.
We have 54 VPCLASS cpu
Our average efficiency looks to be in the low 60% range.
We are experience about 80 lngspins per day with an peek connections of about
2200.
NETTYPE ipcshm,2,50,CPU
NETTYPE soctcp,20,150,NET
Would you suggest lower the cpus to maybe 48, or does the 60% efficiency rule
out any issues with the cpu VPCLASS configuration?
I've been increasing the number of soctcp polling threads and decreasing the
number of connections and that seems to have helped some of our performance
issues, but I was thinking about moving up to 25,120,NET.
Thoughts?
Turn off SMT - in all the testing I have done then Informix will run quicker
The nice thing it is easy to turn on and off so there is a negative impact
it close to instant to turn SMT back on
Cheers
Paul
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
ANTHONY LANDRY
Sent: Monday, September 12, 2016 3:52 PM
To: ids@iiug.org
Subject: Re: RE: onstat -g glo: cpu efficiency [37800]
We have an IBM Power 750 (Power 7+) with 16 cores defined running SMT so
that
the cores appears to the OS as 64 cores.
We have 54 VPCLASS cpu
Our average efficiency looks to be in the low 60% range.
We are experience about 80 lngspins per day with an peek connections of
about
2200.
NETTYPE ipcshm,2,50,CPU
NETTYPE soctcp,20,150,NET
Would you suggest lower the cpus to maybe 48, or does the 60% efficiency
rule
out any issues with the cpu VPCLASS configuration?
I've been increasing the number of soctcp polling threads and decreasing the
number of connections and that seems to have helped some of our performance
issues, but I was thinking about moving up to 25,120,NET.
Thoughts?
****************************************************************************
***
Forum Note: Use "Reply" to post a response in the discussion forum.
Really? My immediate recommendation would be to chnage the NETTYPE for
soctcp... lower the threads and increase the connections.... Too many poll
threads tend to consume a lot of system CPU time.
Are you sure that helped your performance issues? Didn't you change
anything else?
I usually use around 300 connections per thread.... maybe 200 which would
mean a lot less poll threads.
What performance issues were you addressing?
On Mon, Sep 12, 2016 at 10:51 PM, ANTHONY LANDRY <tony@clerk.org> wrote:
> We have an IBM Power 750 (Power 7+) with 16 cores defined running SMT so
> that
> the cores appears to the OS as 64 cores.
>
> We have 54 VPCLASS cpu
>
> Our average efficiency looks to be in the low 60% range.
>
> We are experience about 80 lngspins per day with an peek connections of
> about
> 2200.
>
> NETTYPE ipcshm,2,50,CPU
> NETTYPE soctcp,20,150,NET>
> Would you suggest lower the cpus to maybe 48, or does the 60% efficiency
> rule
> out any issues with the cpu VPCLASS configuration?
>
> I've been increasing the number of soctcp polling threads and decreasing
> the
> number of connections and that seems to have helped some of our performance
> issues, but I was thinking about moving up to 25,120,NET.
>
> Thoughts?
>
>
> ************************************************************
> *******************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--94eb2c034022079d8b053c55f674
My recommendation would be to disable the four threads per core and reduce
that to 2 threads per core (or turn them off entirely). You will get better
throughput from the CPU VPs that way and so fewer spins waits.
You already have many more NET VPs configured than you need. IBM says that
each listener/poll thread can handle between 350 and 500 connections. If
you typically have about 2200 TCP connections at a time then you can live
with 5 to 7 NET VPs for the TCP connections.
If the shared memory connections (NETTYPE ipcshm) are only used for DBA
operations, diagnostics, and maintenance then two poll threads are
sufficient. However, if there are many connections that are local through
shared memory then you should ipcshm poll threads running in every CPU VP.
Finally, to me the most important thing to watch to determine if you have
sufficient CPU VP resources available is the onstat -g rea (ready queue
report) and onstat -g act (active threads). The ready queue should nearly
always be empty with not thread living there more than a second in general
and active threads should be running and not in iowait status.
Art
Art S. Kagel, President and Principal Consultant
ASK Database Management
www.askdbmgt.com
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on the IIUG, nor any other organization with which I am
associated either explicitly, implicitly, or by inference. Neither do
those opinions reflect those of other individuals affiliated with any
entity with which I am affiliated nor those of the entities themselves.
On Mon, Sep 12, 2016 at 4:51 PM, ANTHONY LANDRY <tony@clerk.org> wrote:
> We have an IBM Power 750 (Power 7+) with 16 cores defined running SMT so
> that
> the cores appears to the OS as 64 cores.
>
> We have 54 VPCLASS cpu
>
> Our average efficiency looks to be in the low 60% range.
>
> We are experience about 80 lngspins per day with an peek connections of
> about
> 2200.
>
> NETTYPE ipcshm,2,50,CPU
> NETTYPE soctcp,20,150,NET>
> Would you suggest lower the cpus to maybe 48, or does the 60% efficiency
> rule
> out any issues with the cpu VPCLASS configuration?
>
> I've been increasing the number of soctcp polling threads and decreasing
> the
> number of connections and that seems to have helped some of our performance
> issues, but I was thinking about moving up to 25,120,NET.
>
> Thoughts?
>
>
> ************************************************************
> *******************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--001a11470622f414e9053c55f799
Thank you for your reply.
When I said performance issues, I guess I should have defined what I was
trying to solve better.
I was getting SQL timeouts (over 60 seconds) from clients (about 2 to 6 a day)
on a query that is such a low cost and returns so fast that it doesn't make
sense.
To me it felt like that SQL request wasn't getting picked up and serviced by
the engine, so I took a stab at increasing the polling threads.
After increasing the polling threads I didn't get any SQL timeouts for a week,
but I didn't get 2 yesterday.
Additionally, last night we doubled our physical CPUs and moved from SMT4 to
SMT2 on that LPAR as a test to see if that would address our long spins (we
were having about 30 by 10am, and as of right now we have over 60 at 9:30am.
So why more CPU and SMT2 is creating more long spins is interesting.
I did notice using onstat -g glo that the system created more aio processes,
which I would think is good.
The results of doubling our physical cores plus moving to SMT2 (instead of
SMT4) resulting in more long spins seem somewhat counter intuitive...
More CPU power will give you more long spins. I believe that's normal and
expected although this is a grey area...
Let's start by understanding what is a longspin. When a thread tries to
hold on to a mutex/latch it starts by "spinning". That is, trying in loop
to get the mutex.
If it isn't able to do it, it will be put on a waiting list. The move from
the waiting list, and hten from the waiting list to the ready list, and
finnaly to the executing list, can be expensive. That's why spinning is
done.
If you don't have enough CPU resources, your threads take longer to be on
the executing list and have less opportunities to spin.
Now... are longspins a bad thing? Yes and no... as I mention above,
spinning can be less costly that the alternative. I do remember a famous
informix issue (the infamous nsf.lock) where in part it happened because
the threads were not spinning long enough, given the fact that the "long
spin" definition was defined many years ago for much slower CPUS (so recent
CPUs would "long spin" too early.
This caused the htreads to be put on the ready queue, and these threads had
to move to the running queue, get the mutex, use it, release it before a
new one could get it (snowball effect).
So, my advice would be to take long spins with some care... they aren't
necessarily evil...
Regards.
On Fri, Sep 16, 2016 at 2:40 PM, ANTHONY LANDRY <tony@clerk.org> wrote:
> Thank you for your reply.
>
> When I said performance issues, I guess I should have defined what I was
> trying to solve better.
>
> I was getting SQL timeouts (over 60 seconds) from clients (about 2 to 6 a
> day)
> on a query that is such a low cost and returns so fast that it doesn't make
> sense.
>
> To me it felt like that SQL request wasn't getting picked up and serviced
> by
> the engine, so I took a stab at increasing the polling threads.
>
> After increasing the polling threads I didn't get any SQL timeouts for a
> week,
> but I didn't get 2 yesterday.
>
> Additionally, last night we doubled our physical CPUs and moved from SMT4
> to
> SMT2 on that LPAR as a test to see if that would address our long spins (we
> were having about 30 by 10am, and as of right now we have over 60 at
> 9:30am.
> So why more CPU and SMT2 is creating more long spins is interesting.
>
> I did notice using onstat -g glo that the system created more aio
> processes,
> which I would think is good.
>
> The results of doubling our physical cores plus moving to SMT2 (instead of
> SMT4) resulting in more long spins seem somewhat counter intuitive...
>
>
> ************************************************************
> *******************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--001a1143e2084d7372053ca05cf6
Thank you for your reply! I took your advice and moved to SMT2 last night... the results (somewhat unintuitive in my opinion) are posted above. I've also been monitoring active threads and before any changes to our system I saw that list empty about 50% of the time and when it wasn't empty it almost always had only one active thread. After the changes the results lean more toward about 80% of the time the list is empty.
Depends on which long spins you are getting. What is the name of the longspins? Regards, David. > On 16 September 2016 at 15:03 ANTHONY LANDRY <tony@clerk.org> wrote: > > > Thank you for your reply! > > I took your advice and moved to SMT2 last night... the results (somewhat > unintuitive in my opinion) are posted above. > > I've also been monitoring active threads and before any changes to our system > I saw that list empty about 50% of the time and when it wasn't empty it almost > > always had only one active thread. After the changes the results lean more > toward about 80% of the time the list is empty. > > > ******************************************************************************* > > Forum Note: Use "Reply" to post a response in the discussion forum. >
I'm going to agree with Fernando. Spins are not a good metric for
determining the efficiency of your configuration. If the oninits are
working faster because you switched to SMT2 that's a good thing and will
often result in more spins. The real question is: How are those query
timeouts behaving now? Still seeing them or are they gone?
Art
Art S. Kagel, President and Principal Consultant
ASK Database Management
www.askdbmgt.com
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on the IIUG, nor any other organization with which I am
associated either explicitly, implicitly, or by inference. Neither do
those opinions reflect those of other individuals affiliated with any
entity with which I am affiliated nor those of the entities themselves.
On Fri, Sep 16, 2016 at 9:40 AM, ANTHONY LANDRY <tony@clerk.org> wrote:
> Thank you for your reply.
>
> When I said performance issues, I guess I should have defined what I was
> trying to solve better.
>
> I was getting SQL timeouts (over 60 seconds) from clients (about 2 to 6 a
> day)
> on a query that is such a low cost and returns so fast that it doesn't make
> sense.
>
> To me it felt like that SQL request wasn't getting picked up and serviced
> by
> the engine, so I took a stab at increasing the polling threads.
>
> After increasing the polling threads I didn't get any SQL timeouts for a
> week,
> but I didn't get 2 yesterday.
>
> Additionally, last night we doubled our physical CPUs and moved from SMT4
> to
> SMT2 on that LPAR as a test to see if that would address our long spins (we
> were having about 30 by 10am, and as of right now we have over 60 at
> 9:30am.
> So why more CPU and SMT2 is creating more long spins is interesting.
>
> I did notice using onstat -g glo that the system created more aio
> processes,
> which I would think is good.
>
> The results of doubling our physical cores plus moving to SMT2 (instead of
> SMT4) resulting in more long spins seem somewhat counter intuitive...
>
>
> ************************************************************
> *******************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--001a114b2e0a4ebe6e053ca0cc30
I agree with Art and Fernando - this is exactly what we had to do for a client that moved to AIX 7.1 on P7 - we turned the number of threads down to 2 from 4 per core. This was surprising but it seemed to be the overall culprit in our somewhat random performance issues. Long ready queues (200+) followed by the typical issues from there. We would watch the cpu threads - the 3 and 4 thread would get very little time if any, whereas thread 1 and 2 got a lot of time. Thanks - Mark Scranton The Mark Scranton Group mark@markscranton.com
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g