CPUvps
Posted in 2000
Topics: Performance & Tuning, Server Administration, Versions, Editions & End-of-Life
Hi,
Got the following on my IDS 7.24UC4X1 on HPUX10.20 (5 CPUs, 9 CPUvps)
pri:informix 23> onstat -g glo|grep cpu
class vps usercpu syscpu total
cpu 9 192149.07 15658.73 207807.80
vp pid class usercpu syscpu total
1 20113 cpu 36889.68 3181.98 40071.66
3 20129 cpu 46050.40 2896.56 48946.96
4 20130 cpu 28517.67 1699.41 30217.08
5 20131 cpu 18292.99 1215.59 19508.58
6 20132 cpu 11601.74 748.16 12349.90
7 20133 cpu 9118.92 1071.51 10190.43
8 20135 cpu 17394.67 1923.77 19318.44
9 20137 cpu 11906.31 1386.56 13292.87
10 20138 cpu 12376.69 1535.19 13911.88
pri:informix 24> onstat -g sch|grep cpu
vp pid class semops busy waits spins/wait
1 20113 cpu 168 235 879
3 20129 cpu 43 43 1001
4 20130 cpu 7 7 1001
5 20131 cpu 6 6 1001
6 20132 cpu 8 8 1001
7 20133 cpu 8 8 1001
8 20135 cpu 475 716 847
9 20137 cpu 277 396 845
10 20138 cpu 423 638 867
anyone know of where there seems to be discrepancies in the results under
"onstat -g sch" for vp 3-7?
What I understand under "spins/wait" would translate that these vps seems to
be idling more.
Also, under "onstat -g seg" the vps does not seems to have an equal
distribution of workload. would this means that if there is a way to
distribute more evenly, perhaps I can get more performance out of them?
btw, previously when I declared 4 cpuvps, and then used "onmode -p +5 cpu",
the results under "onstat -g sch" seems to be more evenly distributed than
what I am getting now.
thanks
Steven
I think that Mark Stock has already raised the question of 'more CPUVPs than
physical CPUs'. I'm going to attempt to address CPUVP scheduling.
We don't attempt to equalize the load of the CPUvps because unlike disk IO there
is nothing to be gained by that attempt. Each of the physical processors in a
given machine runs at the same clock speed and really is always running, even if
it is in the CPU idle loop waiting to see if there is anything in the process
ready queue.
What we do is very similar except in the engine there is no idle loop. Instead
when the CPUVP has nothing to do, it will either block on a select() call (if it
is using KAIO or if it has an inline poll thread) or it will block on a
semaphore. There is a ready queue in the engine consisting of thread contexts
waiting for access to a CPUVP. If the CPUVP discovers that there is somthing in
the ready queue, then it will process that thread rather than go into a blocked
state. When a thread is spawned or the poll thread has received new input for a
thread, if it is discovered that all of the active CPUVPs are busy, then one of
the blocked threads is activated by simply unlocking the semaphore that it is
blocked on. That CPUVP will then discover the ready thread an process it.
Since the CPUVPs are scanned from the first to the last to determine which to
activate, that will tend to make the first CPUVPs appear to be carrying the bulk
of the work. If the poll thread is running on the first CPUVP, then it will
tend to make the second CPUVP to be doing just a bit more work than the first.
This is because when input comes from a client on the poll thread running on
CPUVP-0, it will try to activate CPUVP-1 to do the work. Then CPUVP-0 will
check to see if there is anything in the ready queue.
Steven wrote:
> Hi,
>
> Got the following on my IDS 7.24UC4X1 on HPUX10.20 (5 CPUs, 9 CPUvps)
>
> pri:informix 23> onstat -g glo|grep cpu
> class vps usercpu syscpu total
> cpu 9 192149.07 15658.73 207807.80
> vp pid class usercpu syscpu total
> 1 20113 cpu 36889.68 3181.98 40071.66
> 3 20129 cpu 46050.40 2896.56 48946.96
> 4 20130 cpu 28517.67 1699.41 30217.08
> 5 20131 cpu 18292.99 1215.59 19508.58
> 6 20132 cpu 11601.74 748.16 12349.90
> 7 20133 cpu 9118.92 1071.51 10190.43
> 8 20135 cpu 17394.67 1923.77 19318.44
> 9 20137 cpu 11906.31 1386.56 13292.87
> 10 20138 cpu 12376.69 1535.19 13911.88
> pri:informix 24> onstat -g sch|grep cpu
> vp pid class semops busy waits spins/wait
> 1 20113 cpu 168 235 879
> 3 20129 cpu 43 43 1001
> 4 20130 cpu 7 7 1001
> 5 20131 cpu 6 6 1001
> 6 20132 cpu 8 8 1001
> 7 20133 cpu 8 8 1001
> 8 20135 cpu 475 716 847
> 9 20137 cpu 277 396 845
> 10 20138 cpu 423 638 867
>
> anyone know of where there seems to be discrepancies in the results under
> "onstat -g sch" for vp 3-7?
> What I understand under "spins/wait" would translate that these vps seems to
> be idling more.
>
> Also, under "onstat -g seg" the vps does not seems to have an equal
> distribution of workload. would this means that if there is a way to
> distribute more evenly, perhaps I can get more performance out of them?
>
> btw, previously when I declared 4 cpuvps, and then used "onmode -p +5 cpu",
> the results under "onstat -g sch" seems to be more evenly distributed than
> what I am getting now.
>
> thanks
> Steven
Madison Pruet wrote: > > I think that Mark Stock has already raised the question of 'more CPUVPs than > physical CPUs'. I'm going to attempt to address CPUVP scheduling. > > We don't attempt to equalize the load of the CPUvps because unlike disk IO there > is nothing to be gained by that attempt. Each of the physical processors in a > given machine runs at the same clock speed and really is always running, even if > it is in the CPU idle loop waiting to see if there is anything in the process > ready queue. > > What we do is very similar except in the engine there is no idle loop. Instead > when the CPUVP has nothing to do, it will either block on a select() call (if it > is using KAIO or if it has an inline poll thread) or it will block on a > semaphore. There is a ready queue in the engine consisting of thread contexts > waiting for access to a CPUVP. If the CPUVP discovers that there is somthing in > the ready queue, then it will process that thread rather than go into a blocked > state. When a thread is spawned or the poll thread has received new input for a > thread, if it is discovered that all of the active CPUVPs are busy, then one of > the blocked threads is activated by simply unlocking the semaphore that it is > blocked on. That CPUVP will then discover the ready thread an process it. > Since the CPUVPs are scanned from the first to the last to determine which to > activate, that will tend to make the first CPUVPs appear to be carrying the bulk > of the work. If the poll thread is running on the first CPUVP, then it will > tend to make the second CPUVP to be doing just a bit more work than the first. > This is because when input comes from a client on the poll thread running on > CPUVP-0, it will try to activate CPUVP-1 to do the work. Then CPUVP-0 will > check to see if there is anything in the ready queue. Just two things to add. You can improve the distribution of work and improve responsiveness if you have shm poll threads on all CPU VPs. On the subject of CPU VP > CPUs, there is some evidence that for systems with processors faster then about 400MHZ, on HP/PA and Sun/ Sparc at least, having up to 2 CPU VPs per CPU can improve throughput. I do not know Steven's model or CPU speed but he may be trying to take advantage of this, so far anecdotal, evidence. In general Mark is correct CPU VPs should be #CPUs - 1. -- Art S. Kagel & Family kagel@erols.com
In article <38E26B6D.C3B089BA@erols.com>, Art S. Kagel & Family <kagel@erols.com> writes >Just two things to add. You can improve the distribution of work and >improve responsiveness if you have shm poll threads on all CPU VPs. > >On the subject of CPU VP > CPUs, there is some evidence that for >systems with processors faster then about 400MHZ, on HP/PA and Sun/ >Sparc at least, having up to 2 CPU VPs per CPU can improve throughput. >I do not know Steven's model or CPU speed but he may be trying to take >advantage of this, so far anecdotal, evidence. In general Mark is >correct CPU VPs should be #CPUs - 1. > I have found that PSORT_NPROCS > CPUs does improve performance even on Sun 170Mhz processors!! -- David Williams
David Williams wrote: > > In article <38E26B6D.C3B089BA@erols.com>, Art S. Kagel & Family > <kagel@erols.com> writes > >Just two things to add. You can improve the distribution of work and > >improve responsiveness if you have shm poll threads on all CPU VPs. > > > >On the subject of CPU VP > CPUs, there is some evidence that for > >systems with processors faster then about 400MHZ, on HP/PA and Sun/ > >Sparc at least, having up to 2 CPU VPs per CPU can improve throughput. > >I do not know Steven's model or CPU speed but he may be trying to take > >advantage of this, so far anecdotal, evidence. In general Mark is > >correct CPU VPs should be #CPUs - 1. > > > > I have found that PSORT_NPROCS > CPUs does improve performance > even on Sun 170Mhz processors!! Maxing out PSORT_NPROCS always helps, yes. That is a separate issue from the value for NUMCPUVPS. -- Art S. Kagel & Family kagel@erols.com
In article <38E8CAE6.41F85229@erols.com>, Art S. Kagel & Family <kagel@erols.com> writes >David Williams wrote: >> >> In article <38E26B6D.C3B089BA@erols.com>, Art S. Kagel & Family >> <kagel@erols.com> writes >> >Just two things to add. You can improve the distribution of work and >> >improve responsiveness if you have shm poll threads on all CPU VPs. >> > >> >On the subject of CPU VP > CPUs, there is some evidence that for >> >systems with processors faster then about 400MHZ, on HP/PA and Sun/ >> >Sparc at least, having up to 2 CPU VPs per CPU can improve throughput. >> >I do not know Steven's model or CPU speed but he may be trying to take >> >advantage of this, so far anecdotal, evidence. In general Mark is >> >correct CPU VPs should be #CPUs - 1. >> > >> >> I have found that PSORT_NPROCS > CPUs does improve performance >> even on Sun 170Mhz processors!! > >Maxing out PSORT_NPROCS always helps, yes. That is a separate issue >from the value for NUMCPUVPS. > Not really since it shows that having Number of active processes > Number of physical CPUS is not always a bad thing and can even improve performance. -- David Williams
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g