Gathering EXTREME amounts of time in %sys CPU
Answered: green (hollow confidence) — The original asker reports that turning off the Patrol monitoring agent made the runaway %sys CPU problem go away, after ruling out semaphores, memory leaks, PDQ, and NET/CPU VP shared-memory-listener contention raised by other repliers; the asker himself is not fully certain Patrol was the cause since another change happened around the same time.
Advisory only.
Posted in 1999
A 16-CPU NCR MP-RAS box running Informix 7.24 began accumulating huge %sys CPU time after an OS/Informix/Top End upgrade, climbing day by day from ~35% to ~80% until the machine locked up, even when the system was essentially idle. Suggestions covered semaphores, memory leaks/onmode -F, PDQ settings, onstat diagnostics, and NETTYPE/sqlhosts issues (shared-memory and TCP sharing a service name, or listeners in the wrong VP type, causing CPU spinning). A system dump analysed by NCR showed the Patrol monitoring tool running on all CPUs; disabling Patrol made the problem disappear, though the poster noted another change was made at the same time.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: General Discussion
Configuration: NCR 3555 with 16 P133 Intel processors 2 Gb RAM NCR MPRAS 3.02 Informix 7.24 UC5 Mirrored Raid 5 Symbios Disk Arrays containing about 190 Gb Informix ODS database. Disk Arrays are dual hosted to a pair of small NCR boxes which are attached to a Robotic Tape Library. They only go active when large the host is passive. Quad SCSI controllers running in an active/active configuration. Top End 2.05.01 Networking is via 16 Mb Token Ring for most traffic, private ethernet to our NCR 5100. We loaded the previously named versions of UNIX, Informix, and Top End onto our operational data store in early December. We have suffered numerous outages due to a number of factors that we have slowly whittled down. We have achieved some degree of stability, and are now able to run for several days without interruption. (Several days, I've got to say that's pitiful). However, in a dramatic shift from our previous configuration, we are gathering %sys CPU time in extreme amounts. If we IPL in the morning, the end of day curve is as follows (days end with machine mostly idle) Days Since IPL Old Config New Config 1 4% 35% 2 5% 48% 3 4% 58% 4 4% 64% 5 5% lock up around 78 to 80% Has anyone any ideas ? Have you seen anything like this ? We are open to any and all suggestions for research possibilitys. Isn't Y2K compliance fun ? Richard Thomas Harrah's Entertainment, Inc. rthomas@harrahs.com
Richard, I recognize Harrah's as one of the sites we've worked on in the past. Have you reported an incident on this situation ? If you have an incident number, please send it and I'll see what I can do to move it along. If you haven't opened an incident yet, please do so. This issue sounds somewhat familiar and the General Purpose Computing Global Support Center should be able to provide some assistance in explaining the cause of the numbers. Thanks, Greg Staska Director, Sustaining Engineering - GPCGSC (NCR) Richard Thomas wrote: > Configuration: > NCR 3555 with 16 P133 Intel processors > 2 Gb RAM > NCR MPRAS 3.02 > Informix 7.24 UC5 > Mirrored Raid 5 Symbios Disk Arrays containing about 190 Gb Informix ODS > database. > Disk Arrays are dual hosted to a pair of small NCR boxes which are > attached to a Robotic Tape Library. They only go active when > large the host is passive. > Quad SCSI controllers running in an active/active configuration. > Top End 2.05.01 > Networking is via 16 Mb Token Ring for most traffic, private ethernet > to our NCR 5100. > > We loaded the previously named versions of UNIX, Informix, > and Top End onto our operational data store in early December. > > We have suffered numerous outages due to a number of factors > that we have slowly whittled down. > > We have achieved some degree of stability, and are now able to > run for several days without interruption. (Several days, I've got > to say that's pitiful). > > However, in a dramatic shift from our previous configuration, > we are gathering %sys CPU time in extreme amounts. > > If we IPL in the morning, the end of day curve is as follows > (days end with machine mostly idle) > Days Since IPL Old Config New Config > 1 4% 35% > 2 5% 48% > 3 4% 58% > 4 4% 64% > 5 5% lock up around 78 > to 80% > > Has anyone any ideas ? > Have you seen anything like this ? > > We are open to any and all suggestions for research possibilitys. > > Isn't Y2K compliance fun ? > > Richard Thomas > Harrah's Entertainment, Inc. > rthomas@harrahs.com
It would be interesting to know if you are attempting to do non-DB work on the computer. If so, I suspect that 16 P133s can be maxed out in a hurry. "Richard Thomas" <rthomas@harrahs.com> wrote: >Configuration: >NCR 3555 with 16 P133 Intel processors >2 Gb RAM >NCR MPRAS 3.02 >Informix 7.24 UC5 >Mirrored Raid 5 Symbios Disk Arrays containing about 190 Gb Informix ODS >database. >Disk Arrays are dual hosted to a pair of small NCR boxes which are >attached to a Robotic Tape Library. They only go active when >large the host is passive. >Quad SCSI controllers running in an active/active configuration. >Top End 2.05.01 >Networking is via 16 Mb Token Ring for most traffic, private ethernet >to our NCR 5100. > >We loaded the previously named versions of UNIX, Informix, >and Top End onto our operational data store in early December. > >We have suffered numerous outages due to a number of factors >that we have slowly whittled down. > >We have achieved some degree of stability, and are now able to >run for several days without interruption. (Several days, I've got >to say that's pitiful). > >However, in a dramatic shift from our previous configuration, >we are gathering %sys CPU time in extreme amounts. > >If we IPL in the morning, the end of day curve is as follows >(days end with machine mostly idle) >Days Since IPL Old Config New Config >1 4% 35% >2 5% 48% >3 4% 58% >4 4% 64% >5 5% lock up around 78 >to 80% > >Has anyone any ideas ? >Have you seen anything like this ? > >We are open to any and all suggestions for research possibilitys. > >Isn't Y2K compliance fun ? > >Richard Thomas >Harrah's Entertainment, Inc. >rthomas@harrahs.com > > > >
In article <369E3978.E04AF9AC@ColumbiaSC.NCR.COM>, Greg Staska
<greg.staska@ColumbiaSC.NCR.COM> writes
>Richard,
>
>I recognize Harrah's as one of the sites we've worked on in the past. Have
>you reported an incident on this situation ? If you have an incident number,
>please send it and I'll see what I can do to move it along.
>
>If you haven't opened an incident yet, please do so. This issue sounds
>somewhat familiar and the General Purpose Computing Global Support Center
>should be able to provide some assistance in explaining the cause of the
>numbers.
>
>Thanks,
>Greg Staska
>Director, Sustaining Engineering - GPCGSC (NCR)
>
>Richard Thomas wrote:
>
>> Configuration:
>> NCR 3555 with 16 P133 Intel processors
>> 2 Gb RAM
>> NCR MPRAS 3.02
>> Informix 7.24 UC5
>> Mirrored Raid 5 Symbios Disk Arrays containing about 190 Gb Informix ODS
>> database.
>> Disk Arrays are dual hosted to a pair of small NCR boxes which are
>> attached to a Robotic Tape Library. They only go active when
>> large the host is passive.
>> Quad SCSI controllers running in an active/active configuration.
>> Top End 2.05.01
>> Networking is via 16 Mb Token Ring for most traffic, private ethernet
>> to our NCR 5100.
>>
>> We loaded the previously named versions of UNIX, Informix,
>> and Top End onto our operational data store in early December.
>>
>> We have suffered numerous outages due to a number of factors
>> that we have slowly whittled down.
>>
>> We have achieved some degree of stability, and are now able to
>> run for several days without interruption. (Several days, I've got
>> to say that's pitiful).
>>
>> However, in a dramatic shift from our previous configuration,
>> we are gathering %sys CPU time in extreme amounts.
>>
>> If we IPL in the morning, the end of day curve is as follows
>> (days end with machine mostly idle)
>> Days Since IPL Old Config New Config
>> 1 4% 35%
>> 2 5% 48%
>> 3 4% 58%
>> 4 4% 64%
>> 5 5% lock up around 78
>> to 80%
>>
>> Has anyone any ideas ?
Several...
>> Have you seen anything like this ?
Not exactly, the only time I saw a system freeze that hard was an old
7.1x running against Pyramid which hung solid waiting for a semaphore.
Check you have plenty of semaphores configured.
System CPU could well be swapping. There are Online bugs related to
memory leaks.
One per day (in cron) do
onstat -g seg #
onmode -F # free unused shared memory
onstat -g seg #
does it increase?
Also you need to know what sessions are doing. Is PDQ in use or have
you set MAXPDQPRIORITY to 0 in ONCONFIG, and have no environment
variables to do with PDQ set?
Do
onstat -z
sleep 10
onstat -p # What is online doing (in general)?
onstat -u # What sessions are active?
onstat -D # Where is Disk I/O happening?
onstat -g mem # What memory are sessions using?
onstat -g glo # What are VPs doing?
For active sessions run
onstat -g ses <session-id>
several times
what large sql's are the sessions running?
>>
>> We are open to any and all suggestions for research possibilitys.
>>
>> Isn't Y2K compliance fun ?
>>
>> Richard Thomas
>> Harrah's Entertainment, Inc.
>> rthomas@harrahs.com
>
--
David Williams
<snip - delete a lot of old text>
David Williams wrote in message <0mcD2LAtTeo2EwkJ@smooth1.demon.co.uk>...
> Several...
>>> Have you seen anything like this ?
>
> Not exactly, the only time I saw a system freeze that hard was an old
> 7.1x running against Pyramid which hung solid waiting for a semaphore.
>
> Check you have plenty of semaphores configured.
Plenty of semaphores. Not within 35% of limit.
>
> System CPU could well be swapping. There are Online bugs related to
> memory leaks.
>
Disk I/O on swap devices and sar reports do not indicate this condition.
> One per day (in cron) do
>
> onstat -g seg #
> onmode -F # free unused shared memory
> onstat -g seg #>
We nail shared memory, and then constantly monitor it. We never run out,
or come within 15% of maximum. We have operators and programmers
constantly watching for this condition.
> does it increase?
>
> Also you need to know what sessions are doing. Is PDQ in use or have
> you set MAXPDQPRIORITY to 0 in ONCONFIG, and have no environment
> variables to do with PDQ set?
The programs that run on this box are all handcrafted e-sql, running with
a PDQPRIORITY of 0.
>
> Do
>
> onstat -z
> sleep 10
> onstat -p # What is online doing (in general)?
> onstat -u # What sessions are active?
> onstat -D # Where is Disk I/O happening?
> onstat -g mem # What memory are sessions using?
> onstat -g glo # What are VPs doing?
I'll talk to the DBA about this.
>
> For active sessions run
>
> onstat -g ses <session-id>>
> several times
>
> what large sql's are the sessions running?
About midnight, there is basically nothing happening. We have programs
connected to Informix, but they are waiting on msg_recv() for Top End
to pass in some read-only request. The output from on onstat -g <ses>
shows that On-Line is waiting on a semaphore operation from the connected
programs. In other words, the sessions are idle.
More or less, all database releated programs are either waiting to be passed
a message from Top End, or they are servicing a short lived transaction,
usually
less than 3 seconds elapsed tim in the database engine.
Thanks for your input.
>--
>David Williams
Fran Akridge wrote in message <36a09d97.997649475@news.mindspring.com>...
>It would be interesting to know if you are attempting to do non-DB
>work on the computer. If so, I suspect that 16 P133s can be maxed out
>in a hurry.
>
Before the code load, we instituted a freeze on both applications code loads
and non-critical ad-hoc work.
If anything, the work load has lightened up because of this prohibition.
We do a lot of file I/O using the fopen() family of functions, but that
basic
workload hasn't changed for a long time.
This application, while relatively new, and nowhere near the "mature system"
stage
in the software lifecycle, has been running for a couple of years without
exhibiting
this type of behavior.
It is NOT disturbing to see a high % in %sys when all programs are running
and
processing data. When you load new system level software, you should expect
to see some basic fundamental changes in how the system is utilized.
Otherwise,
why load it ?
It IS disturbing that when there is more or less no processing occuring, and
98%
of the programs on the machine are waiting on either a msg_recv()
(applications)
or a semop() (Informix - oninit), that the %sys is so high.
Thanks for your input.
Richard
>"Richard Thomas" <rthomas@harrahs.com> wrote:
>
>>Configuration:
>>NCR 3555 with 16 P133 Intel processors
>>2 Gb RAM
>>NCR MPRAS 3.02
>>Informix 7.24 UC5
>>Mirrored Raid 5 Symbios Disk Arrays containing about 190 Gb Informix ODS
>>database.
>>Disk Arrays are dual hosted to a pair of small NCR boxes which are
>>attached to a Robotic Tape Library. They only go active when
>>large the host is passive.
>>Quad SCSI controllers running in an active/active configuration.
>>Top End 2.05.01
>>Networking is via 16 Mb Token Ring for most traffic, private ethernet
>>to our NCR 5100.
>>
>>We loaded the previously named versions of UNIX, Informix,
>>and Top End onto our operational data store in early December.
>>
>>We have suffered numerous outages due to a number of factors
>>that we have slowly whittled down.
>>
>>We have achieved some degree of stability, and are now able to
>>run for several days without interruption. (Several days, I've got
>>to say that's pitiful).
>>
>>However, in a dramatic shift from our previous configuration,
>>we are gathering %sys CPU time in extreme amounts.
>>
>>If we IPL in the morning, the end of day curve is as follows
>>(days end with machine mostly idle)
>>Days Since IPL Old Config New Config
>>1 4% 35%
>>2 5% 48%
>>3 4% 58%
>>4 4% 64%
>>5 5% lock up around 78
>>to 80%
>>
>>Has anyone any ideas ?
>>Have you seen anything like this ?
>>
>>We are open to any and all suggestions for research possibilitys.
>>
>>Isn't Y2K compliance fun ?
>>
>>Richard Thomas
>>Harrah's Entertainment, Inc.
>>rthomas@harrahs.com
>>
>>
>>
>>
>
all of the following have been done and they look "normal":
onstat -pcached reads and writes percetages are in upper the upper 80 percentile and
upper 50's to mid sixty percentile
bufwaits are lower than expected for the activity that runs on this system.
onstat -g gloalthought there is unbalanced CPU vp activity ( NOAGE is set but we do no
proccessor binding at the informix level nor the os level, which if we did it at
the OS level, it would remedy that ), these exact conditions existed before we
did the upgrade.... ( UPDATE STATISTICS would also help a whole lot, BUT since
we just converted this database over and we expect 99% availability, running
upstats on all 358 tables with an average rowcount of 3.5 million rows and 395
indexes would probably take about as much time for us to test and roll forward
into 7.3uc6 and move to a bigger hardware platform )
onstat -D ( sar -d, sar -b )( I know that we can improve on some of these data layouts and index layouts, we
are also probably doing a lot of page compression because of the size of our
extents and whatnot, we have plans to move this database, fixing all of these
and reconfiguring this system to run DAMN GOOD. )
As richard said, it all looks good.
it's a real stinker.....
Carlos Bolden
Database Administrator
Harrah's Entertainment..
Richard Thomas wrote:
> <snip - delete a lot of old text>
> David Williams wrote in message <0mcD2LAtTeo2EwkJ@smooth1.demon.co.uk>...
>
> > Several...
> >>> Have you seen anything like this ?
> >
> > Not exactly, the only time I saw a system freeze that hard was an old
> > 7.1x running against Pyramid which hung solid waiting for a semaphore.
> >
> > Check you have plenty of semaphores configured.
>
> Plenty of semaphores. Not within 35% of limit.
>
> >
> > System CPU could well be swapping. There are Online bugs related to
> > memory leaks.
> >
>
> Disk I/O on swap devices and sar reports do not indicate this condition.
>
> > One per day (in cron) do
> >
> > onstat -g seg #
> > onmode -F # free unused shared memory
> > onstat -g seg #> >
>
> We nail shared memory, and then constantly monitor it. We never run out,
> or come within 15% of maximum. We have operators and programmers
> constantly watching for this condition.
>
> > does it increase?
> >
> > Also you need to know what sessions are doing. Is PDQ in use or have
> > you set MAXPDQPRIORITY to 0 in ONCONFIG, and have no environment
> > variables to do with PDQ set?
>
> The programs that run on this box are all handcrafted e-sql, running with
> a PDQPRIORITY of 0.
>
> >
> > Do
> >
> > onstat -z
> > sleep 10
> > onstat -p # What is online doing (in general)?
> > onstat -u # What sessions are active?
> > onstat -D # Where is Disk I/O happening?
> > onstat -g mem # What memory are sessions using?
> > onstat -g glo # What are VPs doing?>
> I'll talk to the DBA about this.
>
> >
> > For active sessions run
> >
> > onstat -g ses <session-id>> >
> > several times
> >
> > what large sql's are the sessions running?
>
> About midnight, there is basically nothing happening. We have programs
> connected to Informix, but they are waiting on msg_recv() for Top End
> to pass in some read-only request. The output from on onstat -g <ses>
> shows that On-Line is waiting on a semaphore operation from the connected
> programs. In other words, the sessions are idle.
>
> More or less, all database releated programs are either waiting to be passed
> a message from Top End, or they are servicing a short lived transaction,
> usually
> less than 3 seconds elapsed tim in the database engine.
>
> Thanks for your input.
>
> >--
> >David Williams
In article <36A5800B.48312933@midsouth.rr.com>, Carlos Bolden
<cbolden1@midsouth.rr.com> writes
>all of the following have been done and they look "normal":
>
>onstat -p>cached reads and writes percetages are in upper the upper 80 percentile and
>upper 50's to mid sixty percentile
>bufwaits are lower than expected for the activity that runs on this system.
>
But doing an onstat -z first shows which operations are being
performed in that time period, reads? write? idx-RA? commits? rollbacks?
what is online doing?
>onstat -g glo>althought there is unbalanced CPU vp activity ( NOAGE is set but we do no
>proccessor binding at the informix level nor the os level, which if we did it at
>the OS level, it would remedy that ), these exact conditions existed before we
>did the upgrade.... ( UPDATE STATISTICS would also help a whole lot, BUT since
>we just converted this database over and we expect 99% availability, running
>upstats on all 358 tables with an average rowcount of 3.5 million rows and 395
>indexes would probably take about as much time for us to test and roll forward
>into 7.3uc6 and move to a bigger hardware platform )
>
?? Is the which VP's have their time increasing?
>onstat -D ( sar -d, sar -b )>( I know that we can improve on some of these data layouts and index layouts, we
>are also probably doing a lot of page compression because of the size of our
>extents and whatnot, we have plans to move this database, fixing all of these
>and reconfiguring this system to run DAMN GOOD. )
>As richard said, it all looks good.
>
Are disk I/Os going up?
>it's a real stinker.....
>
>
>Carlos Bolden
>Database Administrator
>Harrah's Entertainment..
>
>Richard Thomas wrote:
>
>> <snip - delete a lot of old text>
>> David Williams wrote in message <0mcD2LAtTeo2EwkJ@smooth1.demon.co.uk>...
>>
>> > Several...
>> >>> Have you seen anything like this ?
>> >
>> > Not exactly, the only time I saw a system freeze that hard was an old
>> > 7.1x running against Pyramid which hung solid waiting for a semaphore.
>> >
>> > Check you have plenty of semaphores configured.
>>
>> Plenty of semaphores. Not within 35% of limit.
>>
>> >
>> > System CPU could well be swapping. There are Online bugs related to
>> > memory leaks.
>> >
>>
>> Disk I/O on swap devices and sar reports do not indicate this condition.
>>
>> > One per day (in cron) do
>> >
>> > onstat -g seg #
>> > onmode -F # free unused shared memory
>> > onstat -g seg #>> >
>>
>> We nail shared memory, and then constantly monitor it. We never run out,
>> or come within 15% of maximum. We have operators and programmers
>> constantly watching for this condition.
>>
>> > does it increase?
>> >
>> > Also you need to know what sessions are doing. Is PDQ in use or have
>> > you set MAXPDQPRIORITY to 0 in ONCONFIG, and have no environment
>> > variables to do with PDQ set?
>>
>> The programs that run on this box are all handcrafted e-sql, running with
>> a PDQPRIORITY of 0.
>>
>> >
>> > Do
>> >
>> > onstat -z
>> > sleep 10
>> > onstat -p # What is online doing (in general)?
>> > onstat -u # What sessions are active?
>> > onstat -D # Where is Disk I/O happening?
>> > onstat -g mem # What memory are sessions using?
>> > onstat -g glo # What are VPs doing?>>
>> I'll talk to the DBA about this.
>>
>> >
>> > For active sessions run
>> >
>> > onstat -g ses <session-id>>> >
>> > several times
>> >
>> > what large sql's are the sessions running?
>>
>> About midnight, there is basically nothing happening. We have programs
>> connected to Informix, but they are waiting on msg_recv() for Top End
>> to pass in some read-only request. The output from on onstat -g <ses>
>> shows that On-Line is waiting on a semaphore operation from the connected
>> programs. In other words, the sessions are idle.
>>
>> More or less, all database releated programs are either waiting to be passed
>> a message from Top End, or they are servicing a short lived transaction,
>> usually
>> less than 3 seconds elapsed tim in the database engine.
>>
>> Thanks for your input.
>>
>> >--
>> >David Williams
>
--
David Williams
In article <36a48dd0@hwilkins.harrahs.com>, Richard Thomas
<rthomas@harrahs.com> writes
>Fran Akridge wrote in message <36a09d97.997649475@news.mindspring.com>...
>>It would be interesting to know if you are attempting to do non-DB
>>work on the computer. If so, I suspect that 16 P133s can be maxed out
>>in a hurry.
>>
>
>Before the code load, we instituted a freeze on both applications code loads
>and non-critical ad-hoc work.
>
>If anything, the work load has lightened up because of this prohibition.
>
>We do a lot of file I/O using the fopen() family of functions, but that
>basic
>workload hasn't changed for a long time.
>
>This application, while relatively new, and nowhere near the "mature system"
>stage
>in the software lifecycle, has been running for a couple of years without
>exhibiting
>this type of behavior.
>
>It is NOT disturbing to see a high % in %sys when all programs are running
>and
>processing data. When you load new system level software, you should expect
>to see some basic fundamental changes in how the system is utilized.
>Otherwise,
>why load it ?
>
>It IS disturbing that when there is more or less no processing occuring, and
>98%
>of the programs on the machine are waiting on either a msg_recv()
>(applications)
>or a semop() (Informix - oninit), that the %sys is so high.
>
>Thanks for your input.
>Richard
>
%sys will only be that high if online is doing something...
Hold on what are your NETTYPEs set to, check the FAQ
at www.smooth1.demon.co.uk Section 6.35
<h2><a name="6.35">6.35 Any known issues with shared memory
connections?</a></h2>
<p>Yes, in yuor sqlhosts file shared memory and tcp connections should
not use the same service name</p>
<p>On 16th Jun 1998 kagel@bloomberg.net (Art S. Kagel) wrote:-</p>
<p>Yes, as I said, there will be a noticable increase in system calls
per
second if you do this. With only one CPU VP and one NET VP you may not
be able to descern the effect but we run 8 NET VPs and have shared
memory listeners in all 28 CPU VPs that we run. In our case, not only
is there a noticable, and quantifiable, increase in system calls when
the network and shared memory connection share a service, there is a
noticeable slowdown in system responsiveness. This is true even though
we have 4 out of 32 CPUs reserved for UNIX services. I'd call that a
drawback!</p>
Check that!!
--
David Williams
David Williams wrote: > In article <36a48dd0@hwilkins.harrahs.com>, Richard Thomas > <rthomas@harrahs.com> writes > >Fran Akridge wrote in message <36a09d97.997649475@news.mindspring.com>... [Dialogue SNIPPED] > %sys will only be that high if online is doing something... > Hold on what are your NETTYPEs set to, check the FAQ > at www.smooth1.demon.co.uk Section 6.35 Related to this see my comment below. > <h2><a name="6.35">6.35 Any known issues with shared memory > connections?</a></h2> > <p>Yes, in yuor sqlhosts file shared memory and tcp connections should > not use the same service name</p> [Details SNIPPED] > Check that!! Also note that using NET VPs for shared memory listeners or CPU VPs for TCP connections will send the CPU time for the corresponding VPs through the roof. Shared memory must be polled, you cannot block on a shared memory read, so shared memory listeners in NET VPs spin the CPU. The CPU VPs only poll the shared memory communications areas between other operations making for a much quieter system. Conversely, NET VPs can use select() to listen on the TCP port(s) for connections and requests and so use almost no system time when the system is not being flooded with requests whereas if the CPU VP has to listen for TCP connections and requests it must add additional poll loops for the TCP listening port and for each user connection port because it cannot block since it has other work to do. Art S. Kagel
It looks like we found our problem. We run Patrol to monitor our system automatically. When we took a dump tape and sent it to NCR for analysis, the only thing that they found that was running on all cpu's was Patrol. We subsequently turned Patrol off, and the %sys cpu problem has gone away. We are not writing in stone that the problem was with Patrol, as one other change was made to the system at about the same point in time, but at this moment, it appears to be the culprit. Thanks for you assistance. Richard Art S. Kagel wrote in message <36A7633B.141D@bloomberg.net>... >David Williams wrote: > >> In article <36a48dd0@hwilkins.harrahs.com>, Richard Thomas >> <rthomas@harrahs.com> writes >> >Fran Akridge wrote in message <36a09d97.997649475@news.mindspring.com>... >[Dialogue SNIPPED] > >> %sys will only be that high if online is doing something... > >> Hold on what are your NETTYPEs set to, check the FAQ >> at www.smooth1.demon.co.uk Section 6.35 > >Related to this see my comment below. > >> <h2><a name="6.35">6.35 Any known issues with shared memory >> connections?</a></h2> > >> <p>Yes, in yuor sqlhosts file shared memory and tcp connections should >> not use the same service name</p> >[Details SNIPPED] > >> Check that!! > >Also note that using NET VPs for shared memory listeners or CPU VPs for >TCP connections will send the CPU time for the corresponding VPs >through the roof. Shared memory must be polled, you cannot block on a >shared memory read, so shared memory listeners in NET VPs spin the CPU. >The CPU VPs only poll the shared memory communications areas between >other operations making for a much quieter system. Conversely, NET VPs >can use select() to listen on the TCP port(s) for connections and >requests and so use almost no system time when the system is not being >flooded with requests whereas if the CPU VP has to listen for TCP >connections and requests it must add additional poll loops for the TCP >listening port and for each user connection port because it cannot >block since it has other work to do. > >Art S. Kagel