how to balance CPU VP
Posted in 2008
Poster saw very uneven CPU usage across 10 CPU VPs in 'onstat -g glo' and asked how to balance the load. Respondents (Art Kagel, Madison Pruet, Mark Scranton) explained that Informix doesn't load-balance: threads run on the first available CPU VP, so a descending 'inverted pyramid' of CPU times is normal and actually indicates spare capacity. Advice: keep TCP poll threads in NET VPs (e.g. NETTYPE soctcp,3,200,NET), roughly one per 200 connections, rather than in CPU VPs; with shared-memory connections, match ipcshm poll threads to CPU VPs. No change was needed in this case.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
I set onconfig parameter
NUMCPUVPS 10
AFF_SPROC 2
AFF_NPROCS 10and run ' onstat -g glo'
vp pid class usercpu syscpu total
1 2863646 cpu 9598.41 149.49 9747.90
2 155844 adm 0.76 1.87 2.63
3 1962506 cpu 14388.87 125.14 14514.01
4 3281730 cpu 9848.78 108.75 9957.53
5 512092 cpu 6566.11 75.35 6641.46
6 3023388 cpu 5093.21 62.39 5155.60
7 585780 cpu 2972.64 40.21 3012.85
8 1327988 cpu 2214.82 26.54 2241.36
9 614874 cpu 1057.71 14.46 1072.17
10 704684 cpu 582.57 6.83 589.40
11 86596 cpu 341.86 3.86 345.72
My question is :
How to balance loading on the 10 CPU VPS ?
roger@star2000.com.tw wrote:
> I set onconfig parameter
> NUMCPUVPS 10
> AFF_SPROC 2
> AFF_NPROCS 10> and run ' onstat -g glo'
> vp pid class usercpu syscpu total
> 1 2863646 cpu 9598.41 149.49 9747.90
> 2 155844 adm 0.76 1.87 2.63
> 3 1962506 cpu 14388.87 125.14 14514.01
> 4 3281730 cpu 9848.78 108.75 9957.53
> 5 512092 cpu 6566.11 75.35 6641.46
> 6 3023388 cpu 5093.21 62.39 5155.60
> 7 585780 cpu 2972.64 40.21 3012.85
> 8 1327988 cpu 2214.82 26.54 2241.36
> 9 614874 cpu 1057.71 14.46 1072.17
> 10 704684 cpu 582.57 6.83 589.40
> 11 86596 cpu 341.86 3.86 345.72
>
> My question is :
> How to balance loading on the 10 CPU VPS ?
You'll never get nearly even usage out of all CPU VPs, however, if you
are using shared memory connections, make sure that the NETTYPE
parameter for ipcshm has the same number of poll threads as CPU VPs and
that they are configured to run in the CPU VPs. Similarly, you should
have at least two, but as many TCP poll threads configured in NET VPs so
that no one poll thread is handling more than 200 connections. This
configuration will get you as close to a balanced CPU load as possible.
Actually, the figures above don't show a particularly poorly tuned
instance for this one tunable.
Art S. Kagel
Oninit
===========================================================================================
Please access the attached hyperlink for an important electronic communications disclaimer:
http://www.oninit.com/home/disclaimer.php
===========================================================================================
> You'll never get nearly even usage out of all CPU VPs, however, if you > are using shared memory connections, make sure that the NETTYPE > parameter for ipcshm has the same number of poll threads as CPU VPs and > that they are configured to run in the CPU VPs. Similarly, you should > have at least two, but as many TCP poll threads configured in NET VPs so > that no one poll thread is handling more than 200 connections. This > configuration will get you as close to a balanced CPU load as possible. > > Actually, the figures above don't show a particularly poorly tuned > instance for this one tunable. > > Art S. Kagel > Oninit Dear Art, All client software is Delphi , and connect DB with tcp/ip. We can't use shared memory ! Any other way I can do to solve my problem? Can I set NETTYPE soctcp,10,200,TCP or NETTYPE soctcp,10,200,CPU to dispatch the CPU VP evenly?
On Feb 19, 4:11 am, ro...@star2000.com.tw wrote:
> I set onconfig parameter
> NUMCPUVPS 10
> AFF_SPROC 2
> AFF_NPROCS 10> and run ' onstat -g glo'
> vp pid class usercpu syscpu total
> 1 2863646 cpu 9598.41 149.49 9747.90
> 2 155844 adm 0.76 1.87 2.63
> 3 1962506 cpu 14388.87 125.14 14514.01
> 4 3281730 cpu 9848.78 108.75 9957.53
> 5 512092 cpu 6566.11 75.35 6641.46
> 6 3023388 cpu 5093.21 62.39 5155.60
> 7 585780 cpu 2972.64 40.21 3012.85
> 8 1327988 cpu 2214.82 26.54 2241.36
> 9 614874 cpu 1057.71 14.46 1072.17
> 10 704684 cpu 582.57 6.83 589.40
> 11 86596 cpu 341.86 3.86 345.72
>
> My question is :
> How to balance loading on the 10 CPU VPS ?
what problem do you think "balancing up" the CPU VPs will solve? Why
do you believe the current workloads indicates you have a problem?
IDS will use any CPU available, not necessarlity the one that has been
used least
you may get better "balance" if you reduce the number of CPU VPs but
i am not sure that will solve whatever business problem you are
seeking to fix.
roger@star2000.com.tw wrote:
>> You'll never get nearly even usage out of all CPU VPs, however, if you
>> are using shared memory connections, make sure that the NETTYPE
>> parameter for ipcshm has the same number of poll threads as CPU VPs and
>> that they are configured to run in the CPU VPs. Similarly, you should
>> have at least two, but as many TCP poll threads configured in NET VPs so
>> that no one poll thread is handling more than 200 connections. This
>> configuration will get you as close to a balanced CPU load as possible.
>>
>> Actually, the figures above don't show a particularly poorly tuned
>> instance for this one tunable.
>>
>> Art S. Kagel
>> Oninit
>>
>
> Dear Art,
> All client software is Delphi , and connect DB with tcp/ip.
> We can't use shared memory !
> Any other way I can do to solve my problem?
> Can I set NETTYPE soctcp,10,200,TCP
> or NETTYPE soctcp,10,200,CPU to dispatch the CPU VP evenly?
OK, so then then the shm polling is fine as is. First, the correct
NETTYPE to use NET VPs for TCP poll threads would be:
NETTYPE soctcp,3,200,NET
Next, it's a bad idea to place TCP poll threads in the CPU VPs. That's
because the CPU VPs cannot block on the sockets, as the NET VPs can,
since they have other work to do, so they have to poll the sockets
periodically when they are between other work and at certain defined
breakpoints in the code where the active threads can release the VP for
polling. This will indeed increase the CPU usage of all of the CPU VPs
running poll threads, however, it will be just wasted cycles and not any
increase in work volume, performance, or responsiveness. Keep the TCP
poll threads in the NET VPs configuring, as I said, one for every 200
expected connections. Third, as I think I hinted, there's nothing
wrong with the CPU utilization pattern you are seeing currently. It
just happens that the lower numbered CPU VPs tend to take most of the
work onto themselves simply because they are least likely to be asleep
or even swapped out when a new request comes in since they are actively
servicing existing requests. Thus often by the time another CPU VP
awakens to poll the ready queue some lower numbered VP has already
drained the queue and scheduled the work. Only when the lower numbered
CPU VPs are too busy to take on more work or are in the middle of an
uninterruptable (word?) section of code will the lower order CPU VPs
take on a job. So the 'inverted triangle' of CPU times you are seeing
in the onstat -g glo listing is normal and indicated a well tuned
instance that has lots of free cycles available for more work when it's
presented.
My only point about the shared memory poll threads is that IF you only
have one CPU VP out of many polling shared memory that will by the first
CPU VP, and as described, that's typically the most busy CPU VP in
versions earlier than 10 (the second CPU VP tends to be busier in 10.00
and later due to some restructuring of overhead task code), since it
mainly polls when it's not busy it also tends to take even more of the
work onto itself than is the case for TCP connection work and so gets
even busier, polling less often. Thus, having only the one poll thread
tends to reduce responsiveness if shared memory connections are common
amd tends to skew the CPU usage triangle even more than normal. Having
SHM poll threads in all CPU VPs reduced the effect and spreads the work
a bit more evenly. I'm not saying the work needs to be spread for any
reason other than that user requests will tend to be processed more
quickly since they spend less time waiting to be picked up out of the
shared memory request buffer and end up more likely to be picked up
immediately by an otherwise idle CPU VP than to be placed in the ready
queue to wait again for a free VP.
Art S. Kagel
Oninit
===========================================================================================
Please access the attached hyperlink for an important electronic communications disclaimer:
http://www.oninit.com/home/disclaimer.php
===========================================================================================
roger@star2000.com.tw wrote:
> I set onconfig parameter
> NUMCPUVPS 10
> AFF_SPROC 2
> AFF_NPROCS 10> and run ' onstat -g glo'
> vp pid class usercpu syscpu total
> 1 2863646 cpu 9598.41 149.49 9747.90
> 2 155844 adm 0.76 1.87 2.63
> 3 1962506 cpu 14388.87 125.14 14514.01
> 4 3281730 cpu 9848.78 108.75 9957.53
> 5 512092 cpu 6566.11 75.35 6641.46
> 6 3023388 cpu 5093.21 62.39 5155.60
> 7 585780 cpu 2972.64 40.21 3012.85
> 8 1327988 cpu 2214.82 26.54 2241.36
> 9 614874 cpu 1057.71 14.46 1072.17
> 10 704684 cpu 582.57 6.83 589.40
> 11 86596 cpu 341.86 3.86 345.72
>
> My question is :
> How to balance loading on the 10 CPU VPS ?
>
>
The CPUVP usage is based on the NEED to use that CPUVP, not simply on
the fact that a CPUVP exists. When a thread is able to perform
processing work, it will be run on which ever CPUVP is available,
starting from the top down. The First CPUVP has some additional
responsibilities, expecially if it is also running the network poll
threads, so it might appear that it is not doing as much work as the
other VPs, but it really is. So in general when you see an inverted
pyramid, you are looking at a system which has spare CPU horse power.
>
On Feb 19, 10:21 am, Madison Pruet <mpru...@verizon.net> wrote:
> ro...@star2000.com.tw wrote:
> > I set onconfig parameter
> > NUMCPUVPS 10
> > AFF_SPROC 2
> > AFF_NPROCS 10> > and run ' onstat -g glo'
> > vp pid class usercpu syscpu total
> > 1 2863646 cpu 9598.41 149.49 9747.90
> > 2 155844 adm 0.76 1.87 2.63
> > 3 1962506 cpu 14388.87 125.14 14514.01
> > 4 3281730 cpu 9848.78 108.75 9957.53
> > 5 512092 cpu 6566.11 75.35 6641.46
> > 6 3023388 cpu 5093.21 62.39 5155.60
> > 7 585780 cpu 2972.64 40.21 3012.85
> > 8 1327988 cpu 2214.82 26.54 2241.36
> > 9 614874 cpu 1057.71 14.46 1072.17
> > 10 704684 cpu 582.57 6.83 589.40
> > 11 86596 cpu 341.86 3.86 345.72
>
> > My question is :
> > How to balance loading on the 10 CPU VPS ?
>
> The CPUVP usage is based on the NEED to use that CPUVP, not simply on
> the fact that a CPUVP exists. When a thread is able to perform
> processing work, it will be run on which ever CPUVP is available,
> starting from the top down. The First CPUVP has some additional
> responsibilities, expecially if it is also running the network poll
> threads, so it might appear that it is not doing as much work as the
> other VPs, but it really is. So in general when you see an inverted
> pyramid, you are looking at a system which has spare CPU horse power.
>
>
I get this question often in my travels....the idea of "balancing the
workload" between CPUVPS. As Madison pointed out (and others), the
engine doesn't attempt to load balance - there's really no reason to
do so. A thread running on a VP yields for one of about 3 reasons - no
more work to do, waiting on a resource, or hit an mt_yield() call - I
always forget the 4th one (or so I thought there was a 4th one...there
may be more)). The engine doesn't do pre-emption, scheduling, or
timeslicing in this context (all pun intended!). So as Madison stated,
usage of a CPUVP is on a "need to use" basis. When a thread is to
yield, he'll do the "switching" (used very broadly here) needed to get
the next thread on a ready queue to take a turn on that cpuvp, for
example. Wanna know more? I have a couple of presentations on our
multi-threading I can direct you to if you contact me directly. Been
dragging them around for years.
HTH -
Mark Scranton
Xtivia Inc.