Re: AW: NETTYPE
Posted in 2003
Dan,
I've researched and tested this one to death long ago and it still holds. To
explain. We originally went with the recommendations in the admin guide and
had only one CPU listener for shared memory connections and a couple for TCP
connections. Since most connections here are shared memory (all user accesses
are through middleware servers running local to the engine) response was OK
most
of the time, though we did notice that CPU VP #1 was using several times the
CPU time as the other 23 CPU VPs and some peak periods simple requests seemed
to
hang.
However, the SA's complained that once the Informix engine was started up the
system call overhead on the machines went from about 2000/sec on similarly
loaded machine not running IDS to 20-25000/sec on the IDS server machines. This
was on DG/UX on early Motorolla 88000 based Aviion platforms which could not
easily handle system call overhead beyond about 20000/sec without dogging out
so
we had a problem. My normal approach to problem solving is to figure out how
the problem could be caused and prove or disprove the various proposals.
I spoke to senior Informix techs about how and when the engine might be using
excessive system calls and eventually the discussion got around to
polling/listening. It turns out that quite naturally if the TCP litener threads
are running in NET VPs which have nothing else to do they set up a select()
system call with all of the connection listening ports and all of the ports
assigned to existing connections and block on that call until a request for a
new connection or for service of an existing one comes in. One system call, all
is cool. On the other hand if the thread is running in the CPU VP it cannot
block because the CPU VP has other work to do and remember IDS threading is
done
in the IDS library not the OS scheduler so the system call block will block the
entire VP not just the one thread. So instead of sitting on the select call
the thread will poll all active ports in a short timer loop when it is not
otherwise busy and at certain break points during query processing. The polling
involves a non-blocking read() system call to each active port. MASSIVE system
call overload results. Plus if the CPU VP is very busy it will be polling
relatively seldom from a new requests point of view causing response delays
when
the load goes up.
Now conversely, for ipcshm connections there is no blocking call possible.
There is only reading the shared memory locations for each allocated connection
(active or not) in a loop. Each read requires a system call to latch the
location and another to reset the mutex. If poll thread is assigned to a NET VP
which has nothing else to do the poll loop will run almost continuously until
the VPs timeslice expires. Again massive system call overload and a CPU that is
burning cycles for nothing. (However, as the admin guide suggests shared
memory responsiveness is great!) Not place that ipcshm poll thread into a CPU
VP and the poll loop can only run when the VP is quiet or at those designated
breakpoints and then only for a few iterations. Which is why my shared memory
responsiveness was down at peak. It also explained why the 1st CPU VP was
getting to do most of the work. Remember that the point in time when the CPU VP
has the most time to poll is when it is not processing requests or is waiting
for IOs. So when it picks up a new request it is likely to assign the new user
thread to itself instead of putting into the queue for another CPU VP to pick
it
up. This causes the one listening VP to be even busier and less responsive.
So based on this research and thought-experiment I began to test the obvious
and
less obvious solutions. The best course was to run shared memory poll threads
in EVERY CPU VP to spread the load and improve responsiveness and to run TCP
listeners in NET VPs so they could block when possible. The results? Our CPU
VPs now run a smooth curve with the first 3rd taking on about 50% of the work
the second 3rd of the VPs taking about 30% and the remaining 20% on the
remaining CPU VPs and shared memory response is instantaneous under all load
conditions since some CPU VP is always polling or about to poll. Meanwhile
system call overhead on the server host was cut to under 10000/sec. As a bonus
CPU usage on the CPU to which the first CPU VP was affined drop from 80% to
40%.
After that I began to advocate these rules of thumb for connections:
TCP connections ONLY in NET VPs only and as few as needed to maintain
responsiveness. Shared memory connections ONLY in CPU VPs and running in EVERY
CPU VP always.
Art S. Kagel
----- Original Message -----
From: Dan.Michael.... <dan.michaelis@verizon.com>
At: 2/21 14:07
>
> Art,
>
> Help me understand why we never want to put TCP poll threads on CPU VP's.
> I might well agree that we'd not want to put any poll threads on CPU VP's,
> since we want CPU VP's pretty much exclusively doing CPU VP kind of stuff,
> rather than listening for connections. What I don't understand is why it
> is any better to have a shared memory poll thread on the CPU VP than it is
> to have a TCP poll thread.
>
> My understanding of the connection process is that the VP listening for
> whichever type of connection you're coming in on (TCP or shared memory)
> picks up the request, it forces the creation of an sqlexec thread, and then
> hands off the session to the sqlexec thread, then goes back to listening.
> I understand from the documentation (and from some real life experiences
> that others I know have had) that the CPU VP's make faster connections, but
> given that the functionality of the poll threads is the same, even that
> doesn't make much theoretical sense. Is it because the CPU VP's have fewer
> yield points than the NET VP's? I'm not sure where/why you'd yield in the
> connection process anyway, so that doesn't make much sense, but it's the
> only thing I can think of.
>
> In Edmund's case, given that the existing configuration has 4 tli poll
> threads, my assumption is that most users are connecting via TCP, so I
> don't know that I'd decrease the number of VP's handling those connections.
> I'd probably go with something like this:
>
> NETTYPE ipcshm 1,50,NET
> NETTYPE tlitcp 1,200,CPU
> NETTYPE tlitcp 3,200,NET>
> I'm guessing that the only shared memory connections made for this instance
> are maintenance ones (and the additional shared memory listener is the
> result of earlier misunderstandings -- I've done that in the past), so
> you'd only need 1 poll thread, and it doesn't have to be lightning fast.
> The 1 TCP poll thread that listens on the CPU VP would handle most of the
> connections, and do so more quickly than a NET VP would, but in cases
> where the instance was slammed, the NET TCP poll threads would pick up the
> slack at a slightly slower rate. Alternatively, increase the tlitcp/NET VP
> threads by 1, and get rid of the CPU VP poll threads altogether.
>
> If I'm off base, please let me, and the rest of us know... I know that
> there are TONS of misunderstandings about the poll threads and VP classes
> out there...
>
> Thanks.
>
> Dan Michaelis
> 813.978.6534 (office)
> 813.303.3225 (pager)
> dan.michaelis@verizon.com@@N