onstat -g spi and vp_proc lock
Posted in 2000
Topics: Performance & Tuning, Connectivity: ESQL/C, 4GL & Embedded SQL, Server Administration, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
I've been following Art Kagel's advice and trying to reduce bufwaits
on our servers. Several questions arise...Deep breath...
We are running IDS 7.31.UC4-1 on Solaris 2.5.1 / Solaris 2.6.
Hardware is either single or dual CPU E450 with CPU speeds varying
from ~200-~300 Mhz.
I've increases LRUS and CLEANERS noticed checkpoint times
come down (users were complaining the system froze, onstat's showed
nothing odd but checkpoint times were high).
Art can checkpoints stop queries?? This were seeing a bug you
mentioned before here.
Checkpoint stats are:-
8 LRUS 1 CLEANER
Average checkpoint 15-20 seconds max checkpoint 120 seconds
64 LRUS 32 CLEANERS
Averages checkpoints were 15-17 seconds.
127 LRUS 48 CLEANERS
Average checkpoint 3-5 seconds max checkpoint 8-10 seconds.
I'm trying to get a standard config for our servers since they all
run the same 4gl application. We notice that on the server with
the most disks (14 on 2 controllers) that above 45 cleaners that
any more are not used (onstat -u gives all 0's) so I though add
a few more for peak periods.
Is this ok having so many LRUS and CLEANERS on a single/dual CPU
machine?? Unused CLEANERS seem to show all 0's in onstat -u I've
seen a bug non the Techinfo site where cleaners were woken up
unnecessarily
Defect 110622 SLOW CHECKPOINTS IN 7.3/7.31/9.2 DUE TO REDUNDANT
WAKEUP OF CLEANERS AND UNNEEDED VISITATION OF MODIFIED LRUS (CASE
#842680)
but this is fixed in 7.31.UC4-1. Also onstat -g spi shows a few
spin on LRU queues E.g. <100 per queue with >1,000,000 bufreads
per day.
We have notice erratic performance for one program on this server.
This program is a real pig, it does loads of selects,deletes,updates,
inserts, makes heavy use of temp tables, the works.
Normally the program completed in 30 minutes but sometimes it takes
2 1/2 hours and it only 1/4 done! Stop the program and restart it
from within the SAME session i.e. same environment and WITHOUT an
reboot of online and it completes in 1/2 hour again.
Since we only run the program once per week this is hard to debug.
onstat's give little info but I did notice that onstat -g spi
showed
******************************************************************
vp_proc lock increasing by ~10,000 per seconds and of all the long
spins 1 million out of 1.3 million were on the two vp_proc locks!!
******************************************************************
Yikes, definate bottleneck. Informix Tech Support said looked at
the onstat output+ onconfig and said the server seemed ok??
I'm stumped. I'll look at the sql again but any ideas???
Art, Help!!!
PS We have 10,000 buffers and 1,000,000 bufreads in 11 hours.
Computing the buffer turnover figure from Art's forumlae a while
back gives 41 buffer turnovers per hour!
Thanks,
David.
--
David Williams
David Williams wrote:
>
> I've been following Art Kagel's advice and trying to reduce bufwaits
> on our servers. Several questions arise...Deep breath...
>
> We are running IDS 7.31.UC4-1 on Solaris 2.5.1 / Solaris 2.6.
> Hardware is either single or dual CPU E450 with CPU speeds varying
> from ~200-~300 Mhz.
>
> I've increases LRUS and CLEANERS noticed checkpoint times
> come down (users were complaining the system froze, onstat's showed
> nothing odd but checkpoint times were high).
>
> Art can checkpoints stop queries?? This were seeing a bug you
> mentioned before here.
There was an old 'BUG' that Informix insisted was just the way things were
designed. Since the checkpoint thread runs in CPU VP #1 and will only
relinquish the VP to other threads if there are no CLEANERS available when
it has gathered a bunch of buffers to be flushed, which with 48 CLEANERS
will just not happen, any query whose thread is running on CPU VP #1 at the
time the checkpoint begins will hang until all of the dirty buffers have
been flushed, yes. This was going on in 7.2x. Informix assured me that
they were redesigning the entire checkpointing process in 7.3x and 9.2x
and the problem should not exist there. Honestly I have not had an
opportunity to test it out. The new machines we are running 7.31 on are
just so much faster than the DG M88K machines that noone is complaining so
I have left the issue lay. I had suggested they either move the checkpoint
thread to one of the admin VPs or push any user threads in VP #1 back onto
the ready queue before starting it so other VPs, if any, could pick them
up. I do not know how they resolved the issue. Obviously in 9.2x this is
no longer an issue as Fuzzy Checkpointing does not flush dirty buffers.
> Checkpoint stats are:-
>
> 8 LRUS 1 CLEANER
> Average checkpoint 15-20 seconds max checkpoint 120 seconds
>
> 64 LRUS 32 CLEANERS
> Averages checkpoints were 15-17 seconds.
A bit better.
> 127 LRUS 48 CLEANERS
> Average checkpoint 3-5 seconds max checkpoint 8-10 seconds.
There you go.
> I'm trying to get a standard config for our servers since they all
> run the same 4gl application. We notice that on the server with
> the most disks (14 on 2 controllers) that above 45 cleaners that
> any more are not used (onstat -u gives all 0's) so I though add
> a few more for peak periods.
I do like to keep LRUS = CLEANERS but if they are not being used....
> Is this ok having so many LRUS and CLEANERS on a single/dual CPU
> machine?? Unused CLEANERS seem to show all 0's in onstat -u I've
> seen a bug non the Techinfo site where cleaners were woken up
> unnecessarily
>
> Defect 110622 SLOW CHECKPOINTS IN 7.3/7.31/9.2 DUE TO REDUNDANT
> WAKEUP OF CLEANERS AND UNNEEDED VISITATION OF MODIFIED LRUS (CASE
> #842680)
>
> but this is fixed in 7.31.UC4-1. Also onstat -g spi shows a few
> spin on LRU queues E.g. <100 per queue with >1,000,000 bufreads
> per day.
>
> We have notice erratic performance for one program on this server.
> This program is a real pig, it does loads of selects,deletes,updates,
> inserts, makes heavy use of temp tables, the works.
>
> Normally the program completed in 30 minutes but sometimes it takes
> 2 1/2 hours and it only 1/4 done! Stop the program and restart it
> from within the SAME session i.e. same environment and WITHOUT an
> reboot of online and it completes in 1/2 hour again.
Weird. It's like it got into a funk and could not get out again. But maybe
it's just that at the time it started the buffer pool was busy so it was
constantly fighting to reload pages it had already read in once an when it
is run with the system quieter it flies?!?!?
> Since we only run the program once per week this is hard to debug.
> onstat's give little info but I did notice that onstat -g spi
> showed
>
> ******************************************************************
> vp_proc lock increasing by ~10,000 per seconds and of all the long
> spins 1 million out of 1.3 million were on the two vp_proc locks!!
> ******************************************************************
>
> Yikes, definate bottleneck. Informix Tech Support said looked at
> the onstat output+ onconfig and said the server seemed ok??
>
> I'm stumped. I'll look at the sql again but any ideas???
> Art, Help!!!
>
> PS We have 10,000 buffers and 1,000,000 bufreads in 11 hours.
> Computing the buffer turnover figure from Art's forumlae a while
> back gives 41 buffer turnovers per hour!
Um, I get only 9.09 turnovers per hour based on the bufreads:
1000000/11 = 90909 reads/hour
90909/10000 = 9.0909 turnovers/hour
or about every 6.6 minutes between complete buffer turnover.
But what were the pagreads this is the key to turnover?
Nothing to sneeze at non-the-less and I would increase to 100,000 buffers
easily and see how it goes even without knowing pagreads. You should see
that errant 4GL program fly.
Art S. Kagel