RE: Serious performance issues
Posted in 1999
Hi Bill
My $0.02...
I notice in your "onstat -g ioq" output, all your kaio queues are showing
pretty high (up to 85 in once case!) maxlen's. According to my Informix
instructor, a maxlen > 25 indicates that disk requests are not being
serviced fast enough. The first step in remedying this would be to log some
disk performance metrics (something along the lines of "iostat -xtc") and
check the service time and utilization (busyness) of each disk, and the
average occupancy of the wait and active queues. You're using striping so at
least it should be more or less consistent across each RAID set. Ensure
you've tuned your stripe size to enable truly "across drives" striping, and
to avoid unnecessary writes. Also remember that when using striping, it is
of critical importance that you distribute the drives across multiple
controllers, to avoid saturating a single bus which has to service many
logical requests from one physical request.
You lokwaits (in "onstat -p") to lock requests ratio is currently about
0.07%. Keep as low as possible, and start to think about rewriting some
code, or examining lock levels, if gets over 1%.
One other thing is that your sequential scans is shudderingly high (almost
1M). Update stats ?
Cheers!
Andrew
> ----------
> From: Bill Weaver[SMTP:billw@fscorp.com]
> Sent: Saturday, June 05, 1999 2:06 AM
> To: informix-list@iiug.org; Murray Wood
> Subject: RE: Serious performance issues
>
> <<File: onstats>>
> Database is more relational and table driven than before. It was quite an
> extensive redesign
> BUT the actual tables are similiar to the previous design so logically
> there wasn't much that
> changed.
>
> We have a Sequent Symmetry SE40 with 8 processors and 1 Gig of memory.
> There are lots of disk
> which are mirrored stripped pairs - 43 of these stripes are used for
> Informix (all 2 Gig stripes
> so there are 43 dbspaces with 1 chunk per space). OS is Dynix/ptx 4.4.2.
> IDS 7.30.UC3, 4GL
> 7.20.UD1X1.
>
> We've done some pretty extensive analysis on indicies, especially on
> problem programs, and
> everything looks good there. I won't say it is perfect yet, but most
> queries are using an index
> path.
>
> Right now, there are 252 users - 149 using 4gl applications over shared
> memory connections and
> 103 using a client/server application over network connections. This can
> be as much as the low
> 300's for total users.
>
> Yes we are using KAIO. Yes we've updated statistics (I use Art Kagel's
> dostats program to
> update statistics on a regular basis). Attached are all the onstats you> requested.
>
>
> As an example, we rebooted the box last night to try and help things out.
> After the reboot,
> there were 5 major processes that we started up to run overnight. Those 5
> processes brought the
> system to 0% idle time which NEVER happened before.
>
> --- On Fri, 4 Jun 1999 10:49:26 +1200 Murray Wood <murray@quanta.co.nz>
> wrote:
>
> What is the effect of the application / database redesign?
>
> You dont say what hardware configuration, OS, Informix versions .... If
> IDS, have you run
> update statistics medium?>
> Do you now have the right indexes for the new application?
> Can you trace a slow sql from onstat -g ses {number} and put it / them
> through set explain on?
>
> Number of users?
> KAIO?
>
> Post: onstat -p onstat -g seg onstat -g ioq
> onstat -F
> onstat -R> onconfig onstat -d
>
> Regards
> Murray Wood
>
>
>
> -----Original Message-----
> From: Bill Weaver [SMTP:billw@fscorp.com]
> Sent: Friday, June 04, 1999 3:54 AM
> To: informix-list@iiug.org
> Subject: Serious performance issues
>
> In the past few weeks, we went through a major upgrade of our database
> where we completely
> redesigned a large portion of the database structure, migrated the data,
> and updated the
> applications. Since that period, we've had serious performance problems.
> We went from a system
>
> that averaged 30-40% cpu idle time to one that stays 0-5% idle (more often
> at 0% idle during the
>
> day). I have to believe the problem is inefficencies in the new database
> structure as that is
> what changed. However, I'm at a total loss as to how to find where these
> inefficencies are!
> I/O is more evenly spread out on this new system than on the old so it
> isn't an I/O problem
> (which is also supported by the sudden increase in cpu utilization). The
> applications (4GL
> code) are the same applications, the only changes made to them were to
> support the new structure
>
> - no major logic changes were made - yet they themselves perform worse
> under the new system.
>
> Anybody have any advice as to how to track down what's causing this and
> eliminate it? Any areas
>
> that I need to be looking at? One addition we made to this new structure
> that wasn't in the old
>
> was that we added referrential integrity utilizing foreign keys. Can the
> additional overhead
> from referrential integrity checks cause or contribute to our performance
> woes? I've considered
>
> eliminating the foreign keys (one key table is referenced by almost every
> other table in the
> system for example) and replacing them with regular indexes. Would this
> help?
>
> One interesting thing I've noticed is that even during periods of 0% idle
> time, one of my cpuvps
>
> is still performing busy waits and semops. There is stuff sitting on the
> ready queue (onstat -g
>
> rea) but for some reason it isn't picking it up. And, it is always the
> same processor that has
> these busy waits/semops. The other 5 cpuvps have 0 in busy waits and
> semops so they are staying
>
> fulling utilized (I have an 8 processor system - 6 are dedicated and
> affinitied to a production
> instance of Informix, 1 to a development instance of Informix, and the
> last one is left free for
>
> OS - it actually is the first physical processor). Following is the
> output of onstat -g sch:
>
> vp pid class semops busy waits spins/wait
> 1 22836 cpu 0 0 0
> 2 22937 adm 0 0 0
> 3 22938 cpu 0 0 0
> 4 22939 cpu 0 0 0
> 5 22940 cpu 0 0 0
> 6 22941 cpu 0 0 0
> 7 22942 cpu 6923 8395 9116
> 8 22944 lio 0 0 0
> 9 22945 pio 0 0 0
> 10 22946 aio 0 0 0
> 11 22947 msc 2004 0 0
> 12 22956 aio 0 0 0
> 13 22957 tli 18 27 786
>
> --------------------------------------------------------
> Name: Bill Weaver
> E-mail: Bill Weaver <billw@fscorp.com>
> Date: 06/03/99
> Time: 10:53:46
>
> (Retrospectively realizes there is no future in hindsight)@@