RE: performance problems due to lock waits?
Posted in 2003
Topics: High Availability & Replication, Performance & Tuning, Installation, Setup & Upgrades, Storage & Space Management, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues
Try turning on the "NOAGE" in the config file......
Wayne Martin
Database Administrator
Kmart Corporation
-----Original Message-----
From: Jay [mailto:remove:jay@td.ca]
Sent: Monday, July 21, 2003 9:49 AM
To: informix-list@iiug.org
Subject: performance problems due to lock waits?
Hi,
I'm having a major decrease in performance, and I'm wondering if the number
of spin locks are contributing to this.
Enviroment
Informix 7.24 (I know we must upgrade, but unfortunatly it's not my descion
to as when), AIX 4.3.3, 4 CPUs, 1GB RAM, 8 disks no mirror or stripes and
volitile tables fragmented by round robin across 4 disks (indexes for these
tables are on seperate a disk).
I am really hoping someone can help me!
Thanks in advance
J.
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace nameROOTPATH /dev/prod2_rootdbs # Path for device containing root
dbspace
ROOTOFFSET 0 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 1024000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS rootdbs # Location (dbspace) of physical log
PHYSFILE 80000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 6 # Number of logical log files
LOGSIZE 45000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /isgprod2/inf7dev/online.log # System message log file path
CONSOLE /isgprod2/inf7dev/console.log # System console message path
ALARMPROGRAM /isgprod2/inf7dev/etc/no_log.sh # Alarm program path
# System Archive Tape Device
TAPEDEV /dev/rmt1 # Tape device path
TAPEBLK 512 # Tape block size (Kbytes)
TAPESIZE 20971520 #Maximum amount of data to put on tape
(kbytes)
# Log Archive Tape Device
LTAPEDEV /dev/null # Log tape device path
LTAPEBLK 512 # Log tape block size (Kbytes)
LTAPESIZE 12582412 # Max amount of data to put on log tape
(Kbytes)
# Optical
STAGEBLOB # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLineinstance
DBSERVERNAME aixprod2_on # Name of default database server
DBSERVERALIASES rt_aixprod2_on # List of alternate dbservernames
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed
env.
RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 4 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps toone
NOAGE 0 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 10000 # Maximum number of locks#BUFFERS 8000 # Maximum number of shared buffers
BUFFERS 100000
NUMAIOVPS 6 # Number of IO vps
PHYSBUFF 64 # Physical log buffer size (Kbytes)
LOGBUFF 64 # Logical log buffer size (Kbytes)LOGSMAX 6 # Maximum number of logical log files
CLEANERS 6 # Number of buffer cleaner processes
SHMBASE 0x30000000 # Shared memory base address
SHMVIRTSIZE 98304 # initial virtual shared memory segment size
SHMADD 16384 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 32 # Number of LRU queues
LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water markpercentage
LTXEHWM 60 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 64 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline workerthreads
ON_RECVRY_THREADS 1 # Default number of online worker threads
# Data Replication Variables
# DRAUTO: 0 manual, 1 retain type, 2 reverse type
DRAUTO 0 # DR automatic switchover
DRINTERVAL 30 # DR max time between DR buffer flushes (in
sec)
DRTIMEOUT 30 # DR network timeout (in sec)DRLOSTFOUND /isgprod2/inf7dev/etc/dr.lostfound # DR lost+found file path
# CDR Variables
CDR_LOGBUFFERS 2048 # size of log reading buffer pool (Kbytes)
CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-vp,additional)
CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR queue
(Kbytes)
# Backup/Restore variables
BAR_ACT_LOG /tmp/bar_act.log
BAR_MAX_BACKUP 0
BAR_RETRY 1
BAR_NB_XPORT_COUNT 10
BAR_XFER_BUF_SIZE 31
# Read Ahead Variables
RA_PAGES 32 # Number of pages to attempt to read ahead
RA_THRESHOLD 30 # Number of pages left before next group
# DBSPACETEMP:
# OnLine equivalent of DBTEMP for SE. This is the list of dbspaces
# that the OnLine SQL Engine will use to create temp tables etc.
# If specified it must be a colon separated list of dbspaces that exist
# when the OnLine system is brought online. If not specified, or if
# all dbspaces specified are invalid, various ad hoc queries will create
# temporary files in /tmp instead.
DBSPACETEMP tempdbs1:tempdbs2:tempdbs3 # Default temp dbspaces
#DBSPACETEMP # Default temp dbspaces
# DUMP*:
# The following parameters control the type of diagnostics information which
# is preserved when an unanticipated error condition (assertion failure)
occurs
# during OnLine operations.
# For DUMPSHMEM, DUMPGCORE and DUMPCORE 1 means Yes, 0 means No.
DUMPDIR /tmp # Preserve diagnostics in this directory
DUMPSHMEM 0
Martin, Wayne E. wrote:
> Try turning on the "NOAGE" in the config file......
But only in careful balance with some other perhaps necessary changes! This
one has the potential to hog all CPU's and thus cripple performance. When
switching on NOAGE (a good thing) it's important to make sensible decisions
about what you want to do with your physical CPU's.
If N = number of physical CPU's, then your best typical starting point is
MULTIPROCESSOR 1 # always set for multi-cpu box
NUMCPUVPS N-1 # do the math and write it in
SINGLE_CPU_VP 0
NOAGE 1
AFF_SPROC 1 # or 2 if your architecture counts from 1 not zero
AFF_NPROCS N-1 # do the math...
You should also set RESIDENCY -1 because there is absolutely no point in
letting the engine buffers swap out.
If -1 is not supported in your 7.24 (I don't remember - check the release
notes files) then use onstat -g seg to count what you've got and put that
number in.
For an 8 disk machine, you've got way too few AIOVPs unless you are using
KAIO - which should be preferable anyway. If not using KAIO, count the
number of chunks on your system and then allocate C * 1.5 AIOVPS and
CLEANERS. Monitor with onstat -g iov. If any are showing > 1.2 io/wup then
you probably need more. If many are showing < 1.0 io/wup, or generally
appear lazy, then you probably need less.
When LRU writes kick in, the long queues are allocated to a cleaner. If all
your LRU's tend to need LRU writes between checkpoints, then each LRU needs
a cleaner. IF you have more LRUS than cleaners, then some of the LRU queues
will not start to flush until other ones are finished and a cleaner becomes
available. So reduce or increase LRUS to match CLEANERS if you get a lot of
LRU writes. I couldn't see that in your message unless I'm going blind (no
comments from you, Ob) What is in the first two information lines from
onstat -F? eg
Fg Writes LRU Writes Chunk Writes
0 6258 23171
comes out of a little linux development box here.
I'm dubious about the small sizing of your physical and logical logs, but
that's not necessarily a performance problem, unless it's provoking too many
checkpoints before they are really due. What does your MSGFILE say about
checkpoints? What's an average AND worst-case checkpoint duration? Anything
over 2-3 seconds will be both noticed by users and perhaps start to irritate
them. Anything over 5 seconds will probably definitely irritate them. Do
your checkpoints really appear to trigger every 300 seconds (5 mins?)
I see you have your LRU percents really low. That's almost certainly making
them queue up for service from the cleaners, and the cleaners will in turn
be keeping the AIOVP's busy and interfering with user activity. I'd be
surprised if u didn't get benefit from the above changes...
Apart from that, it's too early and wintery here to go into any more
investigation, but that stuff above is all the basics IMHO so call that
round 1 %^)
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix