Re: HELP: MEANING OF onstat -g sch output
Posted in 1997
In article <5ds4qr$hg4@cssun.mathcs.emory.edu>, Rick Ward
<wardr@gusco.com> writes
>Edward Bidinotto wrote:
>>
>> Can anyone explain what the output on onstat -g sch means especially the
>> column
>> "busy waits"? Does a high number in that column mean that the vp is
>> staying busy
>> or that it is staying idle?
>>
>> We need to find out from the following output if we need to add cpuvps
>> or delete
>> also if we need to delete aiovps or add them.
>>
>> It is so confusing and Informix Tech Support is no help at all......
>>
>> Thanks in advance.
>>
>> INFORMIX-OnLine Version 7.11.UC1 -- On-Line -- Up 01:13:41 -- 578280
>> Kbytes
>>
>> VP Scheduler Statistics:
>>
>> vp pid class semops busy waits
>> spins/wait
>> 1 2476 cpu 43651 44160
>> 9952
>> 2 2510 adm 0
>> 0 0
>> 3 2513 cpu 41192 41627
>> 9955
>> 4 2514 cpu 39627 40069
>> 9953
>> 5 2516 cpu 37104 37458
>> 9964
>> 6 2522 lio 939
>> 0 0
>> 9 2523 msc 4369
>> 0 0
>> 8 2524 aio 21318
>> 0 0
>> 10 2525 aio 17915
>> 0 0
>> 7 2526 pio 333
>> 0 0
>> 13 2527 shm 17 19
>> 918
>> 11 2528 aio 12682
>> 0 0
>> 12 2529 aio 9731
>> 0 0
>> 14 2530 shm 8 10
>> 865
>> 15 2531 soc 0
>> 0 0
>> 16 2532 soc 0
>> 0 0
>> 17 2533 soc 0
>> 0 0
>> 18 2534 soc 0
>> 0 0
>
>I asked this same question of our Tech Support here in the UK. It took
>them a few days (but it wasn't high priority anyway), but they came
>back with this:
>
>"When a CPU VP finds no thread to run, the process performs a 'bus
>wait'. It is a loop of instructions designed to wait a short time.
Correct - sort of.. Remember the SPINLOCK parameter in Online 5.x?
Well it exists as an undocumented parameter in v7. This means I assume
that spins/wait is similar in botb versions. i.e. if a thread is
waiting for a latch then it will 'spin' e.g. somthing like this in C.
int counter;
counter = (SPINLOCK parameter)
while ((counter-->0) and latch_is_locked(latch))
{
{long l; for (l=0;l<10000;l++);}
}
if latch_is_locked(latch);
waitfor(latch); /* semaphore operation */
The idea behind this is that it is better (from the point of view of
memory caching/context switching) to stay running until the latch is
freed. However if this thread runs it does not let other threads run
on that VP and it not doing anything useful whilst spinning.
The parameters above therefore mean
a) spins/wait =
number of times while loops gets executed and latch is locked (spins)
---------------------------------------------------------------------
b) busy waits = number of times while loop is entered at least once
(waits)
c) sema ops = number of times even spinning (going around the
while loop did) not help and when the while loop was exited the
latch was still locked.
How do you reduce the number of times a latch is locked?
Well latches are locked when :-
creating an entry in the lock table -taking a lock.
accessing a buffer on an LRU queue.
logical/physical log buffers have associated latchs.
The Online Admin Guide 7.1 says that
"A large number of latch waits typically results from a high volume of
processing activity, in which OnLine is logging most of the
transactions."
There i would review the application to make sure
a) Indexes are being used - onstat -u and look for sessions which have
a relatively large number of reads. I usually assume reads>3000
is large and also monitor twice 1 second apart and watch which
sessions accumlates reads. Then get the process id for the session.
Get the process name (x.4ge), put SET EXPLAIN ON for it and
check the sqexplain.out file generated to see if it is
sequential scanning a large table.
b) That temporary tables are created with no log.
c) Check how many LRU queues you have. There generally
should be more LRU queues than page cleaners.
The next thing to do is
a) Monitor CPU VP's
Use onstat -g glo to monitor CPU VP's.
Do this twice about 6 seconds apart and see if they are
all accumulating ~6 seconds of CPU time. If so then
add a CPU VP.
Tools like top and mpstat can also give CPU utilisation on a
per process basis (don't be suprised if oninits are always near
the top) but if you have N CPU VP's and N physical CPUs are
always busy then adding CPU VP's may help.
Always keep number of CPU VPs less then the number of physical CPUs.
This is to allow one physical CPU for other processes to run on the
machine.
b) Monitor AIO VP's
First check your release notes $INFORMIXDIR/release/ONLINE_7.1 to
see if kernel AIO is available on your platform and if so how
to enable it. UNIX patches may be required or special devices
need to be created. Check it is runing by doing onstat -g ath
and looking for a kaio thread.
If it is running AND you are using ONLY RAW chunks then set AIO
VP's = 1. They would only be used for to writing to online.log
and the console.
If it is not available/running then use onstat -g ioq to monitor
AIO I/O queues (OnLine Admin Guide 7.1, Volume 2 Page 33-33).
If they gow over time then add an AIO VP.
>idea is that 1) a thread may become eligible to run very soon, sooner
>than any system call to wait could return, and 2) there is no other more
>useful work to perform. "Spins" refers to the number of loops in a busy
>wait.
>
>So the busy wait count is the number of times the VP was looking for a
>thread and found none. Spins is the number of times through the busy
>wait loop. Note that if a thread becomes runnable before the busy wait
>loop count expires, then the VP breaks out of the loop and runs the
>eligible thread.
>
>If the busy wait loop is executed n times, then the VP waits using a
>semaphore. Only when a runnable thread is available will the VP be
>awakened to run it. In onstat -g sch, the semops record the number
>of times this final step is taken."
>
>So, from your output I would say your system isn't too busy.
>
> vp pid class semops busy waits spins/wait
> 1 2476 cpu 43651