Re: Listener threads and CPU VP imbalance
Posted in 2001
Posted on behalf of Art S. Kagel, who had the good grace to answer a
directly mailed question.
----- Original Message -----
> Hi Art,
>
> I saw in the FAQ that way back (11th Dec 1997) you wrote something about
> insufficient number of listener threads causing skewed CPU VP activity,
and
> recommended increasing the number of listeners.
It is not the number of listeners, strictly speaking, but that there be a
listener thread assigned to each and every CPU VP.
> Is this still a problem with the latest engines?
As far as I can tell yes. The reduction in redundant system calls in the
7.24, and again in the 7.30, code base theoretically may either have
improved
the situation by increasing the possibility that the polling thread would
find its VP busy and so assign the request to another VP and also by making
each polling CPU VP faster and so more responsive to new requests or it may
have exacerbated the situation by making the first few CPU VPs even busier
thus less responsive. My experience, when someone has insisted that all
those poll threads MUST be hurting performance and insisted we try reducing
them yet again, has been that load balancing under 7.31 is better than it
was
under 7.1x or 7.2x but still not good and that response time suffers when
the number of shared memory listener (poll) threads is less than the number
of CPU VPs (note that this all assumes that shared memory listener threads
are being set up in CPU VPs using "NETTYPE ipcshm,...,CPU").
> Is there any way to see service time statistics from the clients - ie to
> show that the poll and/or listener threads are struggling, or does it come
> down to noticing the load skew on the VP's?
It is actually not easy to gather service time stats to diagnose this
problem
and so I rely on the load balancing as shown in onstat -g glo making sure
that the total cpu usage of all of the CPU VPs (assuming a busy system of
course) are relatively equal to diagnose the problem. What I have found is
that you can force a test using many processes or threads performing the
same
VERY simple memory based query to limit any effect to response time only and
not disk latency or IO queuing. The test program should issue the request
repeatedly timing the response using gettimeofday() for microsecond accuracy
(of course actual accuracy will depend on the system clock rate). A second
test that is useful is engine connect time but only relevant if you run many
jobs which connect, perform quick queries/updates and exit, otherwise
connection response is pretty irrelevant. What you will see with 4 or more
CPU VPs and only one shared memory listener (obviously the test clients must
be shared memory clients) is a fairly smooth falloff in the CPU usage from
one CPU VP to another. Service times for the test queries will be erratic.
Try the same test again with a long DSS style query (or even some Cartesian
product) running on CPU VP #1 (onstat -g ses <sid> & onstat -g ath to track)
and you should see the problem more clearly. Now rerun the tests after
restarting with listeners in all CPU VPs to see the difference. If you have
a CPU monitor, like DG's cpu_stat, you will be able to see the uneven stress
on the physical CPUs as well.
> Do you know of other reasons that may cause the skew instead of listener
> imbalance? I'm thinking here about possible false-alerts and a day
wasted...
Well you will see CPU usage skewing if you simply are not needing all of the
CPU VPS all of the time. If your load is periodic there will be period when
only one or two CPU VPS are actually needed to get the work done.
> Probably worth a reply to c.d.i.
I cannot get to CDI right now. Feel free to post this reply if you like.
Art S. Kagel
> TIA