Re: Poor performance for a $750,000 box
Posted in 1999
Topics: Performance & Tuning
"Art S. Kagel" wrote: > It > started when I noticed that 90% of the CPU time my IDS engines were using > was being used by CPU VP #1 and that the CPU that I had pinned that VP to > was running at 98% ALL the time. Uhm why not turn off processor affinity? Then see what happens. ;-) (Which platform and OS are you running?) -Mikey
Mike Segel wrote: > > "Art S. Kagel" wrote: > > > It > > started when I noticed that 90% of the CPU time my IDS engines were using > > was being used by CPU VP #1 and that the CPU that I had pinned that VP to > > was running at 98% ALL the time. > > Uhm why not turn off processor affinity? > Then see what happens. ;-) > (Which platform and OS are you running?) Good question Mikey, I ... never mind. We run most of our servers on Data General Aviion which are Numa architecture mulitprocessor boxes. If a process starts on one processor board it's memory stays on the physical memory owned by that 4 CPU board so if the process migrates off that card its memory accesses are remoted across the bus. Too slow! Besides, all that would have done would be to allow that VP to migrate to CPUs which were less utilized and since it was hogging 90% of the work that would be any CPU it was not on so it would just stay where it was an hog that one CPU anyway unless someother process snuck (sneeked?) in while it was in wait state. Anyway I am much happier having 40% of 28 processors free than 2% or two or three processors and 80% of the rest. We were getting application delays waiting before the "load balancing" was performed. No more. Art S. Kagel
Art, the problem is that by forcing processor affinity, you limit the vp to run only on that CPU, however you don't stop other system processes from using that CPU. Having said that, depending on the hardware/OS they may be doing some load balancing between the processess. As to a "generic" view of SMP, all memory is shared therefore it doesn't matter if your process comes off the stack and into a different CPU, at least in theory. I was just curious if you had benchmarked the performance without CPU affinity on.... -Mikey "Art S. Kagel" wrote: > Mike Segel wrote: > > > > "Art S. Kagel" wrote: > > > > > It > > > started when I noticed that 90% of the CPU time my IDS engines were using > > > was being used by CPU VP #1 and that the CPU that I had pinned that VP to > > > was running at 98% ALL the time. > > > > Uhm why not turn off processor affinity? > > Then see what happens. ;-) > > (Which platform and OS are you running?) > > Good question Mikey, I ... never mind. > > We run most of our servers on Data General Aviion which are Numa > architecture mulitprocessor boxes. If a process starts on one processor > board it's memory stays on the physical memory owned by that 4 CPU board > so if the process migrates off that card its memory accesses are remoted > across the bus. Too slow! Besides, all that would have done would be to > allow that VP to migrate to CPUs which were less utilized and since it > was hogging 90% of the work that would be any CPU it was not on so it > would just stay where it was an hog that one CPU anyway unless someother > process snuck (sneeked?) in while it was in wait state. > > Anyway I am much happier having 40% of 28 processors free than 2% or two > or three processors and 80% of the rest. We were getting application > delays waiting before the "load balancing" was performed. No more. > > Art S. Kagel
OH yes! We do not use the Informix affinity but affine ALL of the VPs,
round robin within class (ie aio, cpu, tli, etc) across the processors
using a script I developed and a utility from DG advanced support. First
we affine the script's session to all 28 CPUs so that its children's
memory is not immediately locked down, then start up the engine. Once the
engine is online the script runs onstat -g glo and gathers the pids and
classes for all VPs, it then executes a subscript for each class that
affines the VPs for that class round robin with an option that allows the
process to float within the 4 processor group that owns the memory image
if the affined processor is too busy. AIO VPs are affined across the CPUs
in reverse order (ie #27, #26, #25, etc) while CPU and TLI VPs are affined
in natural order (ie #4, #5, #6, etc [CPUs 0-3 are reserved]). This is
because I have seen that even with the load balancing that having shared
memory poll threads in all the CPU VPs gives us the first 10 or 15 CPU VPs
are still getting the bulk of the work as are the first half of the AIO
VPs so by reversing the affinity order we balance the CPU load better
still. This plan works well, BTW, as the CPU load monitor (cpustat) shows
that all 28 of these CPUs run within 2% of each other in load on any
machine.
Performance is 50% better than either non-affined or affinity using the
Informix affinit parameters. In either case CPU utilization is choppy
with cpustat showing some processors at 100% and others at 40%.
Of course I test my looney ideas Mikey. I had the same, perfectly
reasonable objections to the plan when the DG technicians first suggested
that we try it and explained the memory image problems. Remember that a
Numa architecture box is NOT true SMP, it is more like a collection of
cooperating SMP boxes attached by a fast interconnect bus. While memory
images on one CPU card IS accessible by a CPU on another board that access
is over the interconnect which is an order of magnitude slower than local
access by CPUs on the same board. As I stated if a process migrates to a
CPU on a different board than the one that contains its memory image those
pages are NOT copied to local memory on the new processor board (though
the local cache is updated as needed but the update is slow). Actually
the memory (up to 2GB) are owned by a 4 processor card and 2 cards share a
level 3 cache on the same board and have CPU bus speed access to the other
cards memory. Migrating from one card to the other on the same board
carries a small penalty but this is minor compared to migrating off the 8
processor board. Sun would have the same problem with their E6000 and
E10000 systems if they did not have a massively fast interconnect that is
actually faster than the CPU bus.
Art S. Kagel
Mike Segel wrote:
>
> Art, the problem is that by forcing processor affinity, you limit the vp to run
> only on that CPU, however you don't stop other system processes from using that
> CPU. Having said that,
> depending on the hardware/OS they may be doing some load balancing between the
> processess.
>
> As to a "generic" view of SMP, all memory is shared therefore it doesn't matter
> if your process comes off the stack and into a different CPU, at least in theory.
>
> I was just curious if you had benchmarked the performance without CPU affinity
> on....
>
> -Mikey
>
> "Art S. Kagel" wrote:
>
> > Mike Segel wrote:
> > >
> > > "Art S. Kagel" wrote:
> > >
> > > > It
> > > > started when I noticed that 90% of the CPU time my IDS engines were using
> > > > was being used by CPU VP #1 and that the CPU that I had pinned that VP to
> > > > was running at 98% ALL the time.
> > >
> > > Uhm why not turn off processor affinity?
> > > Then see what happens. ;-)
> > > (Which platform and OS are you running?)
> >
> > Good question Mikey, I ... never mind.
> >
> > We run most of our servers on Data General Aviion which are Numa
> > architecture mulitprocessor boxes. If a process starts on one processor
> > board it's memory stays on the physical memory owned by that 4 CPU board
> > so if the process migrates off that card its memory accesses are remoted
> > across the bus. Too slow! Besides, all that would have done would be to
> > allow that VP to migrate to CPUs which were less utilized and since it
> > was hogging 90% of the work that would be any CPU it was not on so it
> > would just stay where it was an hog that one CPU anyway unless someother
> > process snuck (sneeked?) in while it was in wait state.
> >
> > Anyway I am much happier having 40% of 28 processors free than 2% or two
> > or three processors and 80% of the rest. We were getting application
> > delays waiting before the "load balancing" was performed. No more.
> >
> > Art S. Kagel
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g