Puzzling performance issue?
Posted in 2003
Topics: Performance & Tuning, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration
Hello all,
I'm inherited a small Informix system that I can't seem to tune.
I've had good results with similar systems running same versions of
OS, Informix and the same application. I do not consider myself an
expert but normally have good luck with this.
Without going into detail I'm just going to present the onstats below.
The general complaint by the users is that the system has a slow
response time. I can post more data if requested. My main interest
is to gather opinions about Informix's health / behavior based on the
numbers below.
I've been playing with LRU/CLEANERS, BUFFERS but beginning to run out
of resources.. Possibly the box is just small to have expectation?
Update Statistics is run nightly using dostats...
I'd really appreciate any thoughts or suggestions you may have...
Thanks,
Eric
PS. I'm aware of the pathing issues /dev/ I've advised the site of
the problems that can arrise in the event of disk replacments, etc..
However, this shouldnt be a problem with performance.
AIX 4.3.3
Informix 7.31.UD4
Number of users: 10-15 including automated report jobs, etc..
Application is a power builder App and the unit is functioning as OLTP
system.
Hardware is comprised of 2x112 Mhz with 512 RAM.
----------------------------------------------------------------------------
Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
08:15:53 -- 473136 Kbytes
Segment Summary:
id key addr size ovhd class blkused
blkfree
1048577 1381386241 30000000 386187264 25596 R 47135 7
786434 1381386242 50000000 98304000 2104 V 5181 6819
Total: - - 484491264 - - 52316 6826
(* segment locked in memory)
Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
08:14:43 -- 473136 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
29867337 47431597 2068415140 98.56 1046285 4817223 87704954 98.81
isamtot open start read write rewrite delete commit
rollbk
1712137397 20607320 152843649 1027261174 38801477 459235 291320 820482
271
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 125145.33 17885.58 1207 2430
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
10290628 1756 660635403 1 0 1020 230061 403729
ixda-RA idx-RA da-RA RA-pgsused lchwaits
19777532 350325 1157981 21269907 1156324
Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
08:14:53 -- 473136 Kbytes
AIO I/O vps:
class/vp s io/s totalops dskread dskwrite dskcopy wakeups io/wup
errors
kio 0 i 28.0 10513541 10130121 383420 0 22065593 0.5 0
kio 1 i 33.5 12568142 12244994 323148 0 26403231 0.5 0
msc 0 i 2.3 849118 0 0 0 847745 1.0 0
aio 0 i 0.0 64 2 0 0 62 1.0 0
aio 1 i 0.0 42 11 0 0 43 1.0 0
aio 2 i 0.0 28 13 0 0 28 1.0 0
aio 3 i 0.0 18 7 0 0 18 1.0 0
aio 4 i 0.0 6 0 0 0 6 1.0 0
pio 0 i 0.0 0 0 0 0 1 0.0 0
lio 0 i 0.0 0 0 0 0 1 0.0 0
Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
08:15:18 -- 473136 Kbytes
AIO global files:
gfd pathname totalops dskread dskwrite io/s
3 /dev/rroot_chunk 11037 4487 6550 0.0
4 /dev/rlog_chunk1 167759 18962 148797 0.4
5 rprod_chunk1 6272525 6224455 48070 16.7
6 phys_chunk1 21543 23 21520 0.1
7 test_chunk1 299121 249873 49248 0.8
8 temp_chunk1 368053 308946 59107 1.0
9 rprod_chunk3 4758108 4691315 66793 12.7
10 test_chunk2 192611 147595 45016 0.5
11 rprod_chunk2 3582750 3493799 88951 9.5
12 /dev/rlog_chunk2 12 11 1 0.0
13 prod2_chunk1 28 26 2 0.0
14 prod_chunk4 5091721 5034669 57052 13.6
15 rprod_chunk5 2320098 2204739 115359 6.2
Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
08:15:43 -- 473136 Kbytes
Configuration File: /digi/informix/etc/onconfig
#**************************************************************************
#
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig.std
# Description: INFORMIX-OnLine Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace nameROOTPATH /dev/rroot_chunk # Path for device containing root
dbspace
ROOTOFFSET 100 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 48000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirroredroot
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS phys_dbs # Location (dbspace) of physical log
PHYSFILE 100000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 19 #19 # Number of logicallog files
LOGSIZE 10000 #10000 # Logical log size
(Kbytes)
# Diagnostics
MSGPATH /digi/informix/online.log # System message log file
path
CONSOLE /digi/informix/console.log # System console message
path
ALARMPROGRAM /digi/informix/etc/no_log.sh # Alarm program path
# System Archive Tape Device
TAPEDEV /dev/rmt0 # Tape device path
TAPEBLK 1024 # Tape block size (Kbytes)
TAPESIZE 4800000 # Maximum amount of data to put on
tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV /dev/rmt0 # Log tape device path
LTAPEBLK 1024 # Log tape block size (Kbytes)
LTAPESIZE 4800000 # Max amount of data to put on log
tape (Kbytes)
# Optical
STAGEBLOB # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLineinstance
DBSERVERNAME worx_server # Name of default database server
DBSERVERALIASES worx_srv # List of alternate dbservernames
NETTYPE soctcp,2,100,NET # Configure poll thread(s) fornettype
DEADLOCK_TIMEOUT 60 # Max time to wait of lock indistributed env.
RESIDENT 0 # Forced residency flag (Yes = 1, No
= 0)
MULTIPROCESSOR 1 # 0 for
Eric wrote:
> Hello all,
> I'm inherited a small Informix system that I can't seem to tune.
> I've had good results with similar systems running same versions of
> OS, Informix and the same application. I do not consider myself an
> expert but normally have good luck with this.
> Without going into detail I'm just going to present the onstats below.
> The general complaint by the users is that the system has a slow
> response time. I can post more data if requested. My main interest
> is to gather opinions about Informix's health / behavior based on the
> numbers below.
>
> I've been playing with LRU/CLEANERS, BUFFERS but beginning to run out
> of resources.. Possibly the box is just small to have expectation?
>
> Update Statistics is run nightly using dostats...>
> I'd really appreciate any thoughts or suggestions you may have...
>
> Thanks,
> Eric
>
> PS. I'm aware of the pathing issues /dev/ I've advised the site of
> the problems that can arrise in the event of disk replacments, etc..
> However, this shouldnt be a problem with performance.
>
>
> AIX 4.3.3
> Informix 7.31.UD4
> Number of users: 10-15 including automated report jobs, etc..
> Application is a power builder App and the unit is functioning as OLTP
> system.
> Hardware is comprised of 2x112 Mhz with 512 RAM.
>
> ----------------------------------------------------------------------------
> Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
> 08:15:53 -- 473136 Kbytes
>
> Segment Summary:
> id key addr size ovhd class blkused
> blkfree
> 1048577 1381386241 30000000 386187264 25596 R 47135 7
> 786434 1381386242 50000000 98304000 2104 V 5181 6819
> Total: - - 484491264 - - 52316 6826
>
> (* segment locked in memory)
>
>
>
> Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
> 08:14:43 -- 473136 Kbytes
>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 29867337 47431597 2068415140 98.56 1046285 4817223 87704954 98.81
>
> isamtot open start read write rewrite delete commit
> rollbk
> 1712137397 20607320 152843649 1027261174 38801477 459235 291320 820482
> 271
>
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 0 0 0 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 0 125145.33 17885.58 1207 2430
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 10290628 1756 660635403 1 0 1020 230061 403729
>
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 19777532 350325 1157981 21269907 1156324
>
>
> Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
> 08:14:53 -- 473136 Kbytes
>
> AIO I/O vps:
> class/vp s io/s totalops dskread dskwrite dskcopy wakeups io/wup
> errors
> kio 0 i 28.0 10513541 10130121 383420 0 22065593 0.5 0
> kio 1 i 33.5 12568142 12244994 323148 0 26403231 0.5 0
> msc 0 i 2.3 849118 0 0 0 847745 1.0 0
> aio 0 i 0.0 64 2 0 0 62 1.0 0
> aio 1 i 0.0 42 11 0 0 43 1.0 0
> aio 2 i 0.0 28 13 0 0 28 1.0 0
> aio 3 i 0.0 18 7 0 0 18 1.0 0
> aio 4 i 0.0 6 0 0 0 6 1.0 0
> pio 0 i 0.0 0 0 0 0 1 0.0 0
> lio 0 i 0.0 0 0 0 0 1 0.0 0
>
> Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
> 08:15:18 -- 473136 Kbytes
>
> AIO global files:
> gfd pathname totalops dskread dskwrite io/s
> 3 /dev/rroot_chunk 11037 4487 6550 0.0
> 4 /dev/rlog_chunk1 167759 18962 148797 0.4
> 5 rprod_chunk1 6272525 6224455 48070 16.7
> 6 phys_chunk1 21543 23 21520 0.1
> 7 test_chunk1 299121 249873 49248 0.8
> 8 temp_chunk1 368053 308946 59107 1.0
> 9 rprod_chunk3 4758108 4691315 66793 12.7
> 10 test_chunk2 192611 147595 45016 0.5
> 11 rprod_chunk2 3582750 3493799 88951 9.5
> 12 /dev/rlog_chunk2 12 11 1 0.0
> 13 prod2_chunk1 28 26 2 0.0
> 14 prod_chunk4 5091721 5034669 57052 13.6
> 15 rprod_chunk5 2320098 2204739 115359 6.2
>
>
> Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up 4 days
> 08:15:43 -- 473136 Kbytes
>
> Configuration File: /digi/informix/etc/onconfig
> #**************************************************************************
> #
> # INFORMIX SOFTWARE, INC.
> #
> # Title: onconfig.std
> # Description: INFORMIX-OnLine Configuration Parameters
> #
> #**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name> ROOTPATH /dev/rroot_chunk # Path for device containing root
> dbspace
> ROOTOFFSET 100 # Offset of root dbspace into device
> (Kbytes)
> ROOTSIZE 48000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH # Path for device containing mirrored> root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS phys_dbs # Location (dbspace) of physical log
> PHYSFILE 100000 # Physical log file size (Kbytes)>
>
> # Logical Log Configuration
>
> LOGFILES 19 #19 # Number of logical> log files
> LOGSIZE 10000 #10000 # Logical log size
> (Kbytes)>
> # Diagnostics
>
> MSGPATH /digi/informix/online.log # System message log file
> path
> CONSOLE /digi/informix/console.log # System console message
> path
> ALARMPROGRAM /digi/informix/etc/no_log.sh # Alarm program path
>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt0 # Tape device path
> TAPEBLK 1024 # Tape block size (Kbytes)
> TAPESIZE 4800000 # Maximum amount of data to put on
> tape (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /dev/rmt0 # Log tape device path
> LTAPEBLK 1024 # Log tape block size (Kbytes)
> LTAPESIZE 4800000 # Max amount of data to put on log
> tape (Kbytes)>
> # Optical
>
> STAGEBLOB # INFORMIX-OnLine/Optical staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a OnLine> instance
> DBSERVERNAME worx_server # Name of default database server>
Random thoughts:
Given an uptime of 4 days, you appear to have clocked quite a bit of
user cpu time. If your hardware assessment is correct, a 2-way 112mhz
processor, aren't you pretty cpu-bound? What does your AIX level
statistics tell you about overall cpu utilization?
Quite a few buffer-level operations with not much I/O showing in the
onstat -D. Reset your stats and look at smaller snapshot. Maybeyou're not seeing all of the I/O because those number "roll over".
512 mb of memory, are you swapping? You appear to be very close at
484491264.
Is response time always slow, or does it degrade during certain times?
Long checkpoints?
Eric wrote:
>
> Hardware is comprised of 2x112 Mhz with 512 RAM.
>
> ----------------------------------------------------------------------
> ------ Informix Dynamic Server Version 7.31.UD4 -- On-Line -- Up
> 4 days 08:15:53 -- 473136 Kbytes
Not just maybe, but definitely this machine is swapping heavily. Find out
how to measure swap activity, and reduce IDS memory consumption until
swapping is all but stopped. Look more at active swapping rather than volume
left on swap, 'cos unused daemons could be swapped out and comfortably stay
there.
Or, get another 256M or more put into the box.
(lack of) memory and swapping is probably the core of your problems.
> AIO I/O vps:
> class/vp s io/s totalops dskread dskwrite dskcopy wakeups io/wup
> errors
> kio 0 i 28.0 10513541 10130121 383420 0 22065593 0.5 0
> kio 1 i 33.5 12568142 12244994 323148 0 26403231 0.5 0
> msc 0 i 2.3 849118 0 0 0 847745 1.0 0
> aio 0 i 0.0 64 2 0 0 62 1.0 0
> aio 1 i 0.0 42 11 0 0 43 1.0 0
> aio 2 i 0.0 28 13 0 0 28 1.0 0
> aio 3 i 0.0 18 7 0 0 18 1.0 0
> aio 4 i 0.0 6 0 0 0 6 1.0 0
> pio 0 i 0.0 0 0 0 0 1 0.0 0
> lio 0 i 0.0 0 0 0 0 1 0.0 0
If using kio on all chunks, then you don't need those aio at all. Just one
or two is plenty. Actually, since they appear to be doing sod-all activity,
it looks like you do indeed have kaio on all chunks.
> AIO global files:
> gfd pathname totalops dskread dskwrite io/s
> 3 /dev/rroot_chunk 11037 4487 6550 0.0
> 4 /dev/rlog_chunk1 167759 18962 148797 0.4
> 5 rprod_chunk1 6272525 6224455 48070 16.7
> 6 phys_chunk1 21543 23 21520 0.1
> 7 test_chunk1 299121 249873 49248 0.8
> 8 temp_chunk1 368053 308946 59107 1.0
> 9 rprod_chunk3 4758108 4691315 66793 12.7
> 10 test_chunk2 192611 147595 45016 0.5
> 11 rprod_chunk2 3582750 3493799 88951 9.5
> 12 /dev/rlog_chunk2 12 11 1 0.0
> 13 prod2_chunk1 28 26 2 0.0
> 14 prod_chunk4 5091721 5034669 57052 13.6
> 15 rprod_chunk5 2320098 2204739 115359 6.2
Activity doesn't appear to be well-distributed. Can u shuffle tables around
so that each chunk tends towards the same amount of activity?
> NETTYPE soctcp,2,100,NET # Configure poll thread(s) for
Pointless. If only 16 users, change to 1,100. Remember the old rule too -
one tcp listener can handle upto about 200(?check-bad memory) connections,
so using 2,100 is pointless even if you need that menay connectors.
Also, local processes will get better results using shmem connectors. Since
you have 2 CPUVPs (but see below) then allocate
NETTYPE ipcshm,2,50,CPU
and get them users connecting using that.
> RESIDENT 0
With the memory shortage, it will indeed be impossible to set this. But I'd
call it a mandatory setting that can only be solved when you solve the
memory problems.
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 2 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> NOAGE 1 # Process aging
D'Oh! This is telling the processes to hog CPU without regard for other
demand. Since you've allocated 2 CPUVPs on a 2 CPU machine, this is a recipe
for disaster. You'd be better off either by switching off NOAGE, or
MULTIPROCESSOR 1
NUMCPUVPS 1
SINGLE_CPU_VP 1
NOAGE 1
AFF_SPROC 1
AFF_NPROCS 1
I don't know AIX intimately, but find out if the processors are numbered
starting from 0 or 1, and make sure you direct the single CPUVP to the
second processor, since on a lot of platforms, the first physical CPU often
does specific jobs for the hardware and/or kernel, so you don't want to
conflict with that.
NOTE: going to 1 CPUVP means that the shm connector will need to be:
NETTYPE ipcshm,1,100,CPU
> LOCKS 50000 # Maximum number of locks
> BUFFERS 90000 # Maximum number of shared buffers
> NUMAIOVPS 5 # Number of IO vps
As mentioned, drop aio's to 1 or 2
> PHYSBUFF 1024 # Physical log buffer size (Kbytes)
> LOGBUFF 512 # Logical log buffer size (Kbytes)
Yoiks! Get everything balanced, and go back to fairly normal settings for
these fellas.
> CLEANERS 111 # Number of buffer cleaner processes
Way too many cleaners. Since you have about 12 chunks, set cleaners to 12
(or whatever) Ditto LRU's. Maxing out LRU's doesn't actually help on a small
system. It's all about balance.
> SHMVIRTSIZE 96000
> SHMADD 32768 # Size of new shared memory segments
Since memory is tight, and the virtual segment doesn't appear to be getting
much use, back this off to 32M/8M and keep an eye on the segments after
that. If another seg is allocated with my smaller suggestions, then adjust
carefully to suit.
> LRUS 111 # Number of LRU queues
As mentioned, adjust to approx 12 (ie match cleaners)
> # Read Ahead Variables
> RA_PAGES 4
> RA_THRESHOLD 2
Read-ahead is your friend. Try 32/24 for these. When read-ahead happens, you
want it to be effective. May as well use the fact that the disk heads are in
the right place at the right time...
> DBSPACETEMP temp_dbs # Default temp dbspaces
Try a logged AND an unlogged temp space. Then pay attention to which one
gets the most use, and maybe adjust the sizes accordingly if disk space it
tight.
Did anyone mention that memory and swapping is probably your major problem?
Related threads
- IDS 10 table-level restore
- Informix Development Webinar December 11, 2007
- ontape -p/r with changed ROOTPATH
- Migrate from HP PA-RISC to HP ITANIUM by ontape