Re: Puzzling performance issue?
Posted in 2003
Topics: Performance & Tuning, Server Administration, Logging & Checkpoints, Platform-Specific Issues
Thanks for pointing out the issue of OS swaping. I didnt realize
what/how AIX was utilizing memory and paging. A little research into
this gave a quick fix to the paging issue. What I used was vmstat -s
and lsps -s Output below shows the before and after my tuning of the
onconfig to decrease the memory footprint of the instance. The "now"
output is after running for several days and has had no paging space
ins or outs.
============================= Then =========================
Tue Nov 18 13:33:13 PST 2003
# vmstat -s
2615215 total address trans. faults
971298 page ins
333974 page outs
230925 paging space page ins
201494 paging space page outs
0 total reclaims
990150 zero filled pages faults
11549 executable filled pages faults
5555107 pages examined by clock
44 revolutions of the clock hand
974844 pages freed by the clock
18227 backtracks
0 lock misses
691 free frame waits
0 extend XPT waits
343309 pending I/O waits
711645 start I/Os
711639 iodones
185630777 cpu context switches
167837162 device interrupts
0 software interrupts
0 traps
266313581 syscalls
$ lsps -sTotal Paging Space Percent Used
800MB 53%
=================== Now =========================
Tue Nov 25 06:36:02 PST 2003
$ vmstat -s
1307910 total address trans. faults
11367 page ins
80001 page outs
0 paging space page ins
0 paging space page outs
0 total reclaims
571241 zero filled pages faults
3887 executable filled pages faults
0 pages examined by clock 0 revolutions of the clock hand
0 pages freed by the clock
11289 backtracks
0 lock misses
0 free frame waits
0 extend XPT waits
8888 pending I/O waits
77573 start I/Os
77573 iodones
176493850 cpu context switches
149339879 device interrupts
0 software interrupts
0 traps
232470895 syscalls
$ lsps -sTotal Paging Space Percent Used
800MB 1%
========== Informix stats
==============================================
Although performance is not exactly what the customer desires I
imagine this may be one of those times where performance and
expectation may never be on the same page. The Informix instance
profile shows a good % for cached reads and writes.
IMHO, BR and BTR are something I can't truly address due to the lack
of resources on the box. (2x112Mhz Process) (512 RAM)
I have still to play with MULTIPROCESSOR, NUMCPUVPS , SINGLE_CPU_VP
and possibly the affinity parameters.
RA_PAGES & RA_THRESHOLD caused more more IO overhead than performance
gain when modified to higher values.
Values tried respectively: (32, 28) (26, 22) (18, 14) (12, 8) (4,2)
(0,0) where 4,2 gave the best percentage over a 2 hour period and
honestly I think I'm ok with a read ahead utilization of 99.94 %
This time around I can't address/rewrite any of the SQL. Yet I
completely identify with your comments. I also don't expect a magic
parameter, only to give the best 'shine-ola' to the outside that can
occur knowing I'm not addressing the issue at the granularity it
deserves. Rewriting the SQL will have to be addressed by the creators
of the application and I believe that issue is being addressed for
future releases. However, a machines with more resources are doing
fine running the same applications under a much heavier client
load/use of the same sql functions.
Checkpoint duration seems quite acceptable at 0,1 and peaks of 5
during what would be defined by the site as bulk inserts, especially
since CPs are currently driven by CKPTINTVL. (The site does not have a
DBA.) I've debated about instructing the client how the physical log
works but not sure it would be maintained. I would like comments from
the field on this issue or for that matter, any of the text here.
Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
17:31:46 -- 315104 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
46273344 68611369 2524475315 98.17 993026 1750300 117515173 99.15
isamtot open start read write rewrite delete commit
rollbk
1997713079 12818214 191846714 1153172649 54959419 390200 233788
584182 151
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 145154.16 23452.51 1362 2724
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress
seqscans
17118603 850 776660865 1 0 703 145888 255490
ixda-RA idx-RA da-RA RA-pgsused lchwaits
32765140 538792 1269482 34552373 1271182
(UR): 99.9391 %
(BR): 9.19729 %
(BTR): 25.2224/HR
$ onstat -g iov
Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
17:32:11 -- 315104 Kbytes
AIO I/O vps:
class/vp s io/s totalops dskread dskwrite dskcopy wakeups io/wup
errors
kio 0 i 41.3 16885025 16653904 231121 0 35555023 0.5
0
kio 1 i 42.5 17380026 17159670 220356 0 36570219 0.5
0
msc 0 i 0.9 380480 0 0 0 379868 1.0
0
aio 0 i 0.0 34 0 0 0 30 1.1
0
aio 1 i 0.0 21 4 0 0 21 1.0
0
aio 2 i 0.0 18 9 0 0 18 1.0
0
aio 3 i 0.0 12 5 0 0 12 1.0
0
pio 0 i 0.0 0 0 0 0 1 0.0
0
lio 0 i 0.0 0 0 0 0 1 0.0
0
$ onstat -g ioq
Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
17:32:14 -- 315104 Kbytes
AIO I/O queues:
q name/id len maxlen totalops dskread dskwrite dskcopy
kio 0 0 16 23845771 23359842 485929 0
kio 1 0 19 24442378 23916804 525574 0
adt 0 0 0 0 0 0 0
msc 0 0 2 380480 0 0 0
aio 0 0 3 85 18 0 0
pio 0 0 0 0 0 0 0
lio 0 0 0 0 0 0 0
gfd 3 0 0 0 0 0 0
gfd 4 0 0 0 0 0 0
gfd 5 0 0 0 0 0 0
gfd 6 0 0 0 0 0 0
gfd 7 0 0 0 0 0 0
gfd 8 0 0 0 0 0 0
gfd 9 0 0 0 0 0 0
gfd 10 0 0 0 0 0 0
gfd 11 0 0 0 0 0
On Tue, 25 Nov 2003 11:14:40 -0500, Eric wrote:
Eric, comments and suggestions below:
<SNIP>
> Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days 17:31:46
> -- 315104 Kbytes
>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached 46273344
> 68611369 2524475315 98.17 993026 1750300 117515173 99.15
>
> isamtot open start read write rewrite delete commit
> rollbk
> 1997713079 12818214 191846714 1153172649 54959419 390200 233788 584182 151
>
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs 0 0 0
> 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes 0 0
> 0 145154.16 23452.51 1362 2724
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 17118603 850 776660865 1 0 703 145888 255490
Some sequential scans, not outrageous but there's some things you can do to help
some of these queries, see below.
> ixda-RA idx-RA da-RA RA-pgsused lchwaits 32765140 538792 1269482
> 34552373 1271182
>
> (UR): 99.9391 %
> (BR): 9.19729 %
> (BTR): 25.2224/HR
As you note, the BR and BTR are not great, but maybe you can help without too
much additional resources.
<SNIP>
> NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype NETTYPE
> ipcshm,1,100,CPU
You have 2 CPU VPs, the ipcshm should have listeners in both for best
responsiveness, change to:
NETTYPE ipcshm,2,100,CPU # You can make it 50 instead of 100 if you want.
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed
> env.
> RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
Making RESIDENT set to 1 or -1 will insure that other things are swapped out not
IDS's memory (1==> Resident segment marked not swappable, -1 RESIDENT and
initial VIRTUAL segments are marked non-swappable)
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 2 # Number of user (cpu) vps SINGLE_CPU_VP 0
> BUFFERS 65000 # 60000 Maximum number of shared buffers NUMAIOVPS 4
> #
Increasing this just a little to say 75000 may help the BTR a bit. Also, take a
BTR reading across peak load (ie a full day or full week) and, separately,
during a significant period that does not include peak load to see if this is
affecting normal processing or only bulk loading/peak loads.
> CLEANERS 31 # Number of buffer cleaner processes LRUS 31
> # Number of LRU queues
Increasing LRUS and cleaners will reduce the BR below the critical 7% level and
requires very little resources. Try 63 (CLEANERS should match LRUS) never use
64, though!
> # Read Ahead Variables
> RA_PAGES 4 # Number of pages to attempt to read ahead
> RA_THRESHOLD 2 # Number of pages left before next group
Good catch. I've always held that with fast disks and a caching controller the
RA_ parameters are no longer effective. I like 8 & 4 because 8 pages is a 'big'
read that the server will execute in a single operation, but these #s seem to be
working for you.
> DBSPACETEMP temp_dbs # Default temp dbspaces
Most sequential scans result in a sort to satisfy an ORDER BY or GROUP BY clause
so if you have three or more temp dbspaces listed above (or if you set the
PSORT_NPROCS=12 (or more) and PSORT_DBTEMP=<list of 3-6 filesystems> in the
environment) the improved sort speed can dramatically improve user performance
perceptions. Also it's a good idea to list at least one 'normal' (not temp)
dbspace in DBSPACETEMP so that logged temp tables have a place to go instead of
the ROOTDB space which is critical to performance. List a low load dbspace,
preferably on a different disk structure than the ROOTDB, logical & physical
logs and other high activity chunks, that is NOT a temp space for this.
> LBU_PRESERVE 0 # Preserve last log for log backup
This should always be 1 (at least in versions <9.4) so that you cannot get
caught without log space to record a logical log backup, in case the logs fill
up.
> TBLSPACE_STATS 1
This is expensive. Set to zero except when trying to diagnose performance
problems.
Art S. Kagel
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g