A question of load . . .
Posted in 2006
A site running IDS 9.40 on an HP9000/HP-UX 11i reported an hour-long slowdown under unusually heavy load (300-400 sessions, three oninit processes pegging CPU, simple queries taking minutes), with no onstat output captured at the time. Replies stressed that diagnosis is impossible without onstat/OS data, and suggested checking PDQ gating (onstat -g mgm), saving onstat -o snapshots and ps/vmstat/iostat data for next time. Art Kagel recommended computing bufwaits ratio, buffer turnover and read-ahead utilization (ratios.ksh), raising LRUS (and CLEANERS to match) from 8, likely increasing BUFFERS, and checking KAIO/AIO VPs via onstat -g iov. No confirmed root cause or resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Performance & Tuning, Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration
We have an HP9000/L2000 with 4 64bit PA-RISC chips running at 440Mhz and
5 Gigs of ram. We're running Inforimx 9.40.HC5 on HPUX 11i. We ran
into an issue on Monday where our database engine seemed to grind to a
halt. The machine itself responded reasonably to everything but
database queries. It seemed like anything having to do with extracting
or querying data in the database took exceptionally long to respond.
Interacting with onstat did not. Unfortunately, no one thought to save
off a copy of what onstat dumped during the incident.
Our application is a touch heterogeneous, we have everything from cognos
impromptu, MS Access "applications", a massive 4gl based application
using shm to connect, to perl dbi based connections running off the same
database. The majority of users probably connect using shm connections
through the 4gl application but, MS Access makes an excessive number of
connections per user of a given application. Everything but the MS
Access and Congos Impromptu reports/applications are running from the
same machine as the engine. This machine also has an apache web server
to handle some cgi scripts, mostly perl code and the source of most of
our perl dbi connections.
When our problem occured, we had a reasonably heavy load from all
around. The machine had approximatly 560 or so proccess running total,
some 300-400 connections appearing in onstat at any one instant, and
three oninit processes eating up 90% of the cpu time solid for about an
hour. Simple queries would take 2 to 3 minutes to complete. There
were no apparent harware problems with the equipment, no excessive
numbers of locks or full logical logs, no single process or small group
of processes that looked like they were doing anything out of the
ordinary.
Fortunately, everything seemed to calm down again about an hour latter.
As a note, that day represented an exceptionally high load in comparison
to our normal operations and this could all be related to having maxed
out what our hardware could handle.
Is there any tuning that might be suggested for our system? I've
attached a copy of our config file. ( I've sense increased the frequency of my
logging cron job )
CONFIG FILE:
#**************************************************************************
#
# IBM Corporation
#
# Title: onconfig.std
# Description: IBM Informix Dynamic Server Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME root # Root dbspace nameROOTPATH /opt/informix/dev/root.1 # Path for device containing root
dbspace
ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 1024000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH /opt/informix/dev/root.1-m # Path for device containing
mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS root # Location (dbspace) of physical log
PHYSFILE 50000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 60 # Number of logical log files
LOGSIZE 10000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /opt/informix/Logs/cars.log # System message log file path
CONSOLE /dev/console # System console message path
# To automatically backup logical logs, edit alarmprogram.sh and set
# BACKUPLOGS=Y
ALARMPROGRAM /opt/informix/etc/log_full.sh # Alarm program path
TBLSPACE_STATS 1 # Maintain tblspace statistics
# System Archive Tape Device
TAPEDEV /dev/rmt/1m # Tape device path
TAPEBLK 32 # Tape block size (Kbytes)
TAPESIZE 0 # Maximum amount of data to put on tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV /dev/rmt/0m # Log tape device path
LTAPEBLK 32 # Log tape block size (Kbytes)
LTAPESIZE 0 # Max amount of data to put on log tape (Kbytes)
# Optical
STAGEBLOB # Informix Dynamic Server staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLine instance
DBSERVERNAME paul # Name of default database server
DBSERVERALIASES carsitcp # List of alternate dbservernames
NETTYPE ipcshm,2,500,CPU # Configure poll thread(s) for nettype
NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 3 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vpsto one
NOAGE 1 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 256000 # Maximum number of locks
BUFFERS 90000 # Maximum number of shared buffers
NUMAIOVPS 8 # Number of IO vps
PHYSBUFF 32 # Physical log buffer size (Kbytes)
LOGBUFF 32 # Logical log buffer size (Kbytes)LOGSMAX 120 # Maximum number of logical log files
CLEANERS 8 # Number of buffer cleaner processes
SHMBASE 0x0 # Shared memory base address
SHMVIRTSIZE 196608 # initial virtual shared memory segment size
SHMADD 16384 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
CKPTINTVL 900 # Check point interval (in sec)
LRUS 8 # Number of LRU queues
LRU_MAX_DIRTY 4 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 64 # Stack size (Kbytes)
# Dynamic Logging
# DYNAMIC_LOGS:
# 2 : server automatically add a new logical log when necessary. (ON)
# 1 : notify DBA to add new logical logs when necessary. (ON)
# 0 : cannot add logical log on the fly. (OFF)
#
# When dynamic logging is on, we can have higher values for
LTXHWM/LTXEHWM,
# because the server can add new logical logs during long transaction
rollback.
# However, to limit the number of new logical logs being added,
LTXHWM/LTXEHWM
# can be set to smaller values.
#
# If dynamic logging is off, LTXHWM/LTXEHWM need to be set to smaller
values
# to avoid long transaction rollback hanging the server due to lack of
logical
# log space, i.e. 50/60 or lower.
DYNAMIC_LOGS 0
LTXHWM 45
LTXEHWM 55
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline worker threads
ON_RECVRY_THREADS 1 # Default number of online worker threads
# Data Replication Variables
DRINTERVAL 30 # DR max time between DR buffer flushes (in sec)
DRTIMEOUT 30 # DR network timeout (in sec)
DRLOSTFOUND /opt/informix/etc/dr.lostfound # DR lost+found file path
# CDR Variables
CDR_EVA
MAX_PDQPRIORITY 100
DId you check to see if there were any PDQ queries running and to see if
anything was gated? onstat -g mgm.
j.
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org]On Behalf Of
CHRIS SALCH
Sent: Thursday, August 31, 2006 1:18 PM
To: ids@iiug.org
Subject: A question of load . . . [7374]
We have an HP9000/L2000 with 4 64bit PA-RISC chips running at 440Mhz and
5 Gigs of ram. We're running Inforimx 9.40.HC5 on HPUX 11i. We ran
into an issue on Monday where our database engine seemed to grind to a
halt. The machine itself responded reasonably to everything but
database queries. It seemed like anything having to do with extracting
or querying data in the database took exceptionally long to respond.
Interacting with onstat did not. Unfortunately, no one thought to save
off a copy of what onstat dumped during the incident.
Our application is a touch heterogeneous, we have everything from cognos
impromptu, MS Access "applications", a massive 4gl based application
using shm to connect, to perl dbi based connections running off the same
database. The majority of users probably connect using shm connections
through the 4gl application but, MS Access makes an excessive number of
connections per user of a given application. Everything but the MS
Access and Congos Impromptu reports/applications are running from the
same machine as the engine. This machine also has an apache web server
to handle some cgi scripts, mostly perl code and the source of most of
our perl dbi connections.
When our problem occured, we had a reasonably heavy load from all
around. The machine had approximatly 560 or so proccess running total,
some 300-400 connections appearing in onstat at any one instant, and
three oninit processes eating up 90% of the cpu time solid for about an
hour. Simple queries would take 2 to 3 minutes to complete. There
were no apparent harware problems with the equipment, no excessive
numbers of locks or full logical logs, no single process or small group
of processes that looked like they were doing anything out of the
ordinary.
Fortunately, everything seemed to calm down again about an hour latter.
As a note, that day represented an exceptionally high load in comparison
to our normal operations and this could all be related to having maxed
out what our hardware could handle.
Is there any tuning that might be suggested for our system? I've
attached a copy of our config file. ( I've sense increased the frequency of
my
logging cron job )
CONFIG FILE:
#**************************************************************************
#
# IBM Corporation
#
# Title: onconfig.std
# Description: IBM Informix Dynamic Server Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME root # Root dbspace nameROOTPATH /opt/informix/dev/root.1 # Path for device containing root
dbspace
ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 1024000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH /opt/informix/dev/root.1-m # Path for device containing
mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS root # Location (dbspace) of physical log
PHYSFILE 50000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 60 # Number of logical log files
LOGSIZE 10000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /opt/informix/Logs/cars.log # System message log file path
CONSOLE /dev/console # System console message path
# To automatically backup logical logs, edit alarmprogram.sh and set
# BACKUPLOGS=Y
ALARMPROGRAM /opt/informix/etc/log_full.sh # Alarm program path
TBLSPACE_STATS 1 # Maintain tblspace statistics
# System Archive Tape Device
TAPEDEV /dev/rmt/1m # Tape device path
TAPEBLK 32 # Tape block size (Kbytes)
TAPESIZE 0 # Maximum amount of data to put on tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV /dev/rmt/0m # Log tape device path
LTAPEBLK 32 # Log tape block size (Kbytes)
LTAPESIZE 0 # Max amount of data to put on log tape (Kbytes)
# Optical
STAGEBLOB # Informix Dynamic Server staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLine instance
DBSERVERNAME paul # Name of default database server
DBSERVERALIASES carsitcp # List of alternate dbservernames
NETTYPE ipcshm,2,500,CPU # Configure poll thread(s) for nettype
NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 3 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vpsto one
NOAGE 1 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 256000 # Maximum number of locks
BUFFERS 90000 # Maximum number of shared buffers
NUMAIOVPS 8 # Number of IO vps
PHYSBUFF 32 # Physical log buffer size (Kbytes)
LOGBUFF 32 # Logical log buffer size (Kbytes)LOGSMAX 120 # Maximum number of logical log files
CLEANERS 8 # Number of buffer cleaner processes
SHMBASE 0x0 # Shared memory base address
SHMVIRTSIZE 196608 # initial virtual shared memory segment size
SHMADD 16384 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
CKPTINTVL 900 # Check point interval (in sec)
LRUS 8 # Number of LRU queues
LRU_MAX_DIRTY 4 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 64 # Stack size (Kbytes)
# Dynamic Logging
# DYNAMIC_LOGS:
# 2 : server automatically add a new logical log when necessary. (ON)
# 1 : notify DBA to add new logical logs when necessary. (ON)
# 0 : cannot add logical log on the fly. (OFF)
#
# When dynamic logging is on, we can have higher values for
LTXHWM/LTXEHWM,
# because the server can add new logical logs during long transaction
rollback.
# However, to limit the number of new logical logs being added,
LTXHWM/LTXEHWM
# can be set to smaller values.
#
# If dynamic logging is off, LTXHWM/LTXEHWM need to be set to smaller
values
# to avoid long transaction rollback hanging the server due to lack of
logical
# log space, i.e. 50/60 or lower.
DYNAMIC_LOGS 0
LTXHWM 45
LTXEHWM 55
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online resto
I would have to say without information from tools like "top", "ps",
vmstat, and other OS related tools as well as an array of Informix
onstats I wouldn't even know where to offer a suggestion.
Because of the same type of issue in the past we have made a habit
(for good or bad) to take an onstat -o filename.out. This allows us
to go back to that point in time to look at informix at that time.
Note: this requires disk space as large as the memory foot print of
ids. We have a space just to save a bunch of these in times of need.
Luckily we don't need them often.
I would suggest looking for orphaned or runaway processes... Anything
accumulating CPU time beyond what is normal... Ask your UNIX SA to
look for this on regular bases... "ps -ef" Also watch for processes
that have a really low priority... Not sure about HPUX but normally
low means more resources to be given.
But it sounds like you have covered most of the bases
On 8/31/06, CHRIS SALCH <chrissalch@letu.edu> wrote:
>
> We have an HP9000/L2000 with 4 64bit PA-RISC chips running at 440Mhz and
> 5 Gigs of ram. We're running Inforimx 9.40.HC5 on HPUX 11i. We ran
> into an issue on Monday where our database engine seemed to grind to a
> halt. The machine itself responded reasonably to everything but
> database queries. It seemed like anything having to do with extracting
> or querying data in the database took exceptionally long to respond.
> Interacting with onstat did not. Unfortunately, no one thought to save
> off a copy of what onstat dumped during the incident.
>
> Our application is a touch heterogeneous, we have everything from cognos
> impromptu, MS Access "applications", a massive 4gl based application
> using shm to connect, to perl dbi based connections running off the same
> database. The majority of users probably connect using shm connections
> through the 4gl application but, MS Access makes an excessive number of
> connections per user of a given application. Everything but the MS
> Access and Congos Impromptu reports/applications are running from the
> same machine as the engine. This machine also has an apache web server
> to handle some cgi scripts, mostly perl code and the source of most of
> our perl dbi connections.
>
> When our problem occured, we had a reasonably heavy load from all
> around. The machine had approximatly 560 or so proccess running total,
> some 300-400 connections appearing in onstat at any one instant, and
> three oninit processes eating up 90% of the cpu time solid for about an
> hour. Simple queries would take 2 to 3 minutes to complete. There
> were no apparent harware problems with the equipment, no excessive
> numbers of locks or full logical logs, no single process or small group
> of processes that looked like they were doing anything out of the
> ordinary.
>
> Fortunately, everything seemed to calm down again about an hour latter.
> As a note, that day represented an exceptionally high load in comparison
> to our normal operations and this could all be related to having maxed
> out what our hardware could handle.
>
> Is there any tuning that might be suggested for our system? I've
> attached a copy of our config file. ( I've sense increased the frequency of
my
> logging cron job )
>
> CONFIG FILE:
>
> #**************************************************************************
> #
> # IBM Corporation
> #
> # Title: onconfig.std
> # Description: IBM Informix Dynamic Server Configuration Parameters
> #
> #**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME root # Root dbspace name> ROOTPATH /opt/informix/dev/root.1 # Path for device containing root
> dbspace
> ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
> ROOTSIZE 1024000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH /opt/informix/dev/root.1-m # Path for device containing
> mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS root # Location (dbspace) of physical log
> PHYSFILE 50000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 60 # Number of logical log files
> LOGSIZE 10000 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /opt/informix/Logs/cars.log # System message log file path
> CONSOLE /dev/console # System console message path>
> # To automatically backup logical logs, edit alarmprogram.sh and set
> # BACKUPLOGS=Y
> ALARMPROGRAM /opt/informix/etc/log_full.sh # Alarm program path
> TBLSPACE_STATS 1 # Maintain tblspace statistics>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt/1m # Tape device path
> TAPEBLK 32 # Tape block size (Kbytes)
> TAPESIZE 0 # Maximum amount of data to put on tape (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /dev/rmt/0m # Log tape device path
> LTAPEBLK 32 # Log tape block size (Kbytes)
> LTAPESIZE 0 # Max amount of data to put on log tape (Kbytes)>
> # Optical
>
> STAGEBLOB # Informix Dynamic Server staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a OnLine instance
> DBSERVERNAME paul # Name of default database server
> DBSERVERALIASES carsitcp # List of alternate dbservernames
> NETTYPE ipcshm,2,500,CPU # Configure poll thread(s) for nettype
> NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 3 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 256000 # Maximum number of locks
> BUFFERS 90000 # Maximum number of shared buffers
> NUMAIOVPS 8 # Number of IO vps
> PHYSBUFF 32 # Physical log buffer size (Kbytes)
> LOGBUFF 32 # Logical log buffer size (Kbytes)> LOGSMAX 120 # Maximum number of logical log files
> CLEANERS 8 # Number of buffer cleaner processes
> SHMBASE 0x0 # Shared memory base address
> SHMVIRTSIZE 196608 # initial virtual shared memory segment size
> SHMADD 16384 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
> 0=>unlimited
> CKPTINTVL 900 # Check point interval (in sec)
> LRUS 8 # Number of LRU queues
> LRU_MAX_DIRTY 4 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 64 # Stack size (Kbytes)>
> # Dynamic Logging
> # DYNAMIC_LOGS:
> # 2 : server automatically add a new logical log when necessary. (ON)
> # 1 : notify DBA to add new logical logs when necessary. (ON)
> # 0 : cannot add logical log on the fly. (OFF)
> #
> # When dynamic logging is on, we can have higher values for
( Unfortunatly, none of us thought to save any of that data when things
started slowing down, so all we've got is what we can remember. )
In answer to several:
3 oninit processes were the only things really eating cpu time. There
were a few other processes that would show up using a small amount of
cpu time but nothing excessive. We were getting a load average just
topping 2 at the time ( this has happened before without any significant
impact to the system )
iostat was showing about 40kbps - 60kbps across all the drives,
significantly lower than what I see now or even after everything calmed
back down and our scheduled backup was started.
Here's an onstat from shortly after if it means much, (The backup
process was started after the whole mess).
IBM Informix Dynamic Server Version 9.40.HC5 -- On-Line -- Up 13
days 09:21:20 -- 426136 Kbytes
Userthreads
address flags sessid user tty wait tout locks nreads
nwrites
d2767018 ---P--D 1 root - 0 0 0 633
4202
d276763c ---P--F 0 root - 0 0 0 0
77695
d2767c60 ---P--F 0 root - 0 0 0 0
11181
d2768284 ---P--F 0 root - 0 0 0 0
7118
d27688a8 ---P--F 0 root - 0 0 0 0
18300
d2768ecc ---P--F 0 root - 0 0 0 0
9348
d27694f0 ---P--F 0 root - 0 0 0 0
7618
d2769b14 ---P--F 0 root - 0 0 0 0
6462
d276a138 ---P--F 0 root - 0 0 0 0
4758
d276a75c ---P--- 13 root - 0 0 0 0
24
d276ad80 ---P--B 14 root - 0 0 0 65074
480
d276b9c8 Y--P--- 352236 carsu tC d8375840 0 1 1019
604
d276bfec ---P--D 17 root - 0 0 0 0
0
d276c610 Y--P--D 26 root - c4fd3b6c 0 0 0
0
d276cc34 Y--P--- 27 roedelm ASTRASER d3201420 0 1 0
0
d276d258 Y--P--- 28 roedelm ASTRASER d30a6358 0 1 0
0
d276d87c Y--P--- 384521 zijlstre tg d9e4c658 0 1 9
0
d276dea0 Y--P--- 390938 carsu tO d7df4560 0 4 0
0
d276eae8 Y--P--- 392296 erpeldic JOHN d70e6b20 0 1 0
0
d276fd54 Y--P--- 389877 carsu tkb d4f0b150 0 1 132
170
d2770378 Y--P--- 386225 sgps JOHN d9961308 0 1 1
0
d277099c ---PR-- 392439 academus - 0 0 1 10
0
d2770fc0 Y--P--- 391698 gullettj tnb d9e3a278 0 1 30
0
d27715e4 Y--P--- 392236 bakerl LINDABAK d53d67f8 0 1 0
0
d2771c08 Y--P--- 392438 academus - d502aa20 0 1 0
0
d277222c Y--P--- 26513 academus - d4865868 0 1 0
0
d2772850 Y--P--- 392379 sgps JOHN d9bcf078 0 1 0
0
d2772e74 Y--P--- 344045 carsu tG d71e18b8 0 1 8179
0
d2773498 Y--P--- 381962 trogdond t3 d43afe78 0 1 1
0
d2774704 Y--P--- 383934 hoffpowj teb d7e04de0 0 1 10
0
d2774d28 Y--P--- 387046 bakerl LINDABAK d600ad58 0 1 1
0
d277534c Y--P--- 389228 burrowsr t0 d7332a28 0 1 104
0
d2775970 Y--P--- 343854 murleyt TOBIMURL d35e5818 0 1 319
36
d2775f94 Y--P--- 390673 childree tI d8d7adf8 0 1 32
0
d2776bdc Y--P--- 352174 sackettp PAMSACKE d78af750 0 1 173
22
d2777200 Y--P--- 385510 zijlstre tg dc89a9b0 0 2 134
0
d2777824 Y--P--- 392149 sgps JOHN d93b61f8 0 1 0
0
d2777e48 Y--P--- 344941 holleyc DAL-REP1 d5aad900 0 1 581
46
d277846c Y--P--- 773 carsu tf d45ae1b0 0 1 0
0
d27790b4 Y--P--- 344192 carsu tt d74e7bc8 0 1 1795
0
d27796d8 Y--P--- 375850 hyattr RENEEHYA d48798b8 0 1 1548
54
d277b58c Y--P--- 357910 hueyt TERESAHU d3b50968 0 1 1467
134
d277bbb0 Y--P--- 387128 carsu tl d66aef80 0 1 225
0
d277c1d4 Y--P--- 390897 holleyc DAL-REP1 d75015b0 0 1 1
0
d277da64 Y--P--- 343646 vestalt tq d4df9da8 0 1 2
0
d277e6ac Y--P--- 155417 carsu tD d59b3a20 0 1 0
0
d277f2f4 Y--P--- 343715 sarverl tu d4331c58 0 1 1221
0
d277ff3c Y--P--- 392116 sgps JOHN d761ce30 0 1 0
0
d2780560 Y--P--- 343793 cowartj JULIANCO d39af2b8 0 1 1494
722
d2781df0 Y--P--- 392277 smeetond - d64dec78 0 1 0
0
d2782a38 ---PR-- 384779 wolffd - 0 0 3 16752
7880
d278305c Y--P--- 390841 johnsonh tO d548c7c0 0 1 3
0
d27842c8 Y--P--- 390978 carsu t8 d3682c60 0 1 0
0
d27848ec Y--P--- 390741 carsu t8 d8b67fa8 0 1 157
0
d2785b58 ------- 391626 informix tty2a11 0 0 0 312628
0
d278617c Y--P--- 382952 carsu tk d91fb7d8 0 1 251
0
d27867a0 Y--P--- 252531 carsu t4 d9534b28 0 2 0
0
d2786dc4 Y--P--- 376405 gullettj tnb d70c1558 0 1 7
0
d27873e8 Y--P--- 381469 wolffd - d600acf8 0 1 2
0
d2787a0c Y--P--- 353053 meltonc t9 d7057fb8 0 1 1
0
d2788654 Y--P--- 392449 carsu td d478b828 0 1 869
0
d278929c Y--P--- 392295 erpeldic JOHN d66c4c78 0 1 0
0
d278a508 Y--P--- 390298 carsu tk d8b814f8 0 1 2738
0
d278ab2c Y--P--- 346340 wessonc tl dc89aad0 0 1 132
0
d278b150 Y--P--- 389623 sgps JOHN d39afb58 0 1 0
0
d278b774 Y--P--- 392066 keilersv ti d4e43ca0 0 1 4
0
d278bd98 Y--P--- 387915 carsu teb d91fbea0 0 1 1195
0
d278d004 Y--P--- 348675 bradfork tgb d70c17b8 0 1 3
0
d278dc4c Y--P--- 344032 pattersk tG d9be5680 0 1 2
0
d278eeb8 Y--P--- 344035 carsu tG d7123220 0 1 13
0
d2790124 Y--P--- 390760 carsu t8 d4c13748 0 1 3255
0
d2790748 Y--P--- 388516 carsu tnb d676d478 0 1 2992
0
d27919b4 Y--P--- 384549 carsu tg d3255478 0 4 6536
40
d2791fd8 Y--P--- 344477 bakerl tR da0f4e88 0 1 10
0
d27925fc Y--P--- 343677 vestalt tq d42de718 0 2 2182
0
d2792c20 Y--P--- 390425 sgps JOHN d4f87150 0 1 0
0
d2793244 Y--P--- 391011 suarezj - d94736f8 0 2 251
4
d27944b0 Y--P--- 117944 carsu th d4787428 0 1 0
0
d279571c Y--P--- 344043 carsu tG d73f1bd0 0 1 10003
634
d2796364 Y--P--- 385067 kunderp tv d4e31d40 0 1 7
0
d2796988 Y--P--- 392168 bakerl LINDABAK d3b64c98 0 1 11
0
d2796fac Y--P--- 745 hammd tf d5adaf68 0 1 0
0
d27975d0 Y--P--- 746 hammd tf d5af5ca0 0 1 0
0
d5c30ecc Y--P--- 392385 carsu tg d9f50e18 0 1 167
0
d5c31b14 Y--P--- 344484 carsu tR d9b80588 0 1 4042
58
d5c32138 Y--P--- 383938 carsu teb d4f872e8 0 4 6761
0
d5c3275c Y--P--- 392440 academus - d9961018 0 1 0
0
d5c339c8 Y--P--- 346791 mcgheeb tbb d70c1cb0 0 1 5
0
d5c33fec Y--P--- 384998 carsu teb d4df9558 0 1 0
0
d5c34c34 Y--P--- 376219 carsu tkb d8237630 0 1 9807
0
d5c35258 Y--P--- 390653 childree tI d77f2f98 0 1 0
0
d5c35ea0 Y--P--- 344041 carsu tG d624f140 0 1 7429
4780
d5c36ae8 ---P--- 392464 roedelm - 0 0 1 0
0
d5c3710c Y--P--- 792 carsu tf d4726018 0 1 0
0
d5c37730 Y--P--- 390893 ardend tJ d8b281b0 0 1 24
0
d5c38378 Y--P--- 390057 barnesc t8 d7d89930 0 1 1
0
d5c3899c Y--P--- 122431 jonesl tp d31a08d8 0 2 0
0
d5c38fc0 Y--P--- 346800 mcgheeb tbb d8ee0210 0 2 2471
140
d5c395e4 Y--P--- 392245 bakerl LINDABAK d93b6c48 0 1 2
0
d5c3a22c Y--P--- 348610 carsu tab d7e08d80 0 1 477
0
d5c3ae74 Y--P--- 390740 jonesmi MICHAELJ d4589a70 0 1 1
0
d5c3c704 Y--P--- 346765 hallv tab d93b6ac8 0 1 6
0
d5c3d34c ------- 391626 informix tty2a11 0 0 0 0
1472
d5c3e5b8 Y--P--- 389097 hueyt tvb d9b80078 0 2 112
2
d5c3ebdc Y--P--- 363978 carsu tgb d8d7a858 0 1 380
80
d5c3f824 Y--P--- 389756 jordanm MONIQUEJ d7e08f30 0 1 509
16
d5c40a90 Y--P--- 353058 meltonc t9 d4682d58 0 1 463
2
d5c410b4 Y--P--- 391922 childree tI d556d078 0 2 89
0
d5c416d8 Y--P--- 384710 sgps JOHN d39af3d8 0 1 0
0
d5c41cfc Y--P--- 376203 bullardk tkb d645f9b0 0 1 3
0
d5c42320 Y--P--- 343654 carsu tq d895c9c0 0 1 2605@
As already pointed out, without the onstat output it's very difficult to
diagnose what went on. However, there are some things you can look at in
hindsight that I will detail. Also there are a couple of suspect ONCONFIG
values I can point you at (the suggested tests will help sort these out also).
If you have not yet zero'd out your server stats (onstat -z) DON'T DO IT YET!
First calculate my three basic metrics:
Bufwaits Ratio (BR), Buffer Turnover Rate (BTR), and ReadAhead Utilization
(RAU). You can download the script ratios.ksh from the IIUG site to do this for
you. These are the formula:
BR = ((bufwaits/(pagreads + bufwrits)) * 100)
BTR = ((bufwrits+pagreads)/BUFFERS)/<fractional hours since stats were cleared>
RAU = ((ixda-RA+idx-RA+da-RA) / RA-pgsused) * 100)
BR is % of buffer accesses that had to wait for a buffer or for an LRU queue
BR <= 7 server's cruising
7 < BR <= 10 server's struggling
BR > 10 users are complaining, server's hosed
BTR is an approximation of how often your server turns over the entire cache
(in
practice high BTR more often tells you that a smaller number of buffers is
being churned very quickly).
BTR should be small single digits - I like to keep mine at 6 or less
RAU is the percentage of readahead pages that are actually accessed. This value
should be as close to 100% as possible (I worry when it falls below 99.5%!)
Note that you have only 8 LRUS configured with 300 users that means that over
40
users are trying to access each LRU concurrently during peak load. I would be
surprised if your BR value isn't double digits. Perhaps as high as 25 or so!
If this is so, increase LRUS to the maximum value (128 for 32bit versions, 256
for 64bit versions). You should always set CLEANERS >= LRUS so increase that
value as well.
BUFFERS = 90000 - Unless every one of those 300 users is accessing the same
data this looks low. BTR will start you in the right direction to see that.
Also look at onstat -P over time during peak loads to see what partnums are
hogging the cache and to see if a small number of partnums are trading buffers
back an forth between them. All signs that you need more buffers.
Are you using KAIO? If not monitor onstat -g iov to make sure that at least one
AIO VP has an io/wup value below 1.0. If not you can take advantage of more
aio vps. These are VERY low overhead (they sleep if they are not busy) so add a
bunch if needed and cull them latter (decrease by the number of vps with io.wup
< 0.5 or so).
Art S. Kagel
----- Original Message -----
From: Chris Salch <ids@iiug.org>
At: 8/31 13:23:45
We have an HP9000/L2000 with 4 64bit PA-RISC chips running at 440Mhz and
5 Gigs of ram. We're running Inforimx 9.40.HC5 on HPUX 11i. We ran
into an issue on Monday where our database engine seemed to grind to a
halt. The machine itself responded reasonably to everything but
database queries. It seemed like anything having to do with extracting
or querying data in the database took exceptionally long to respond.
Interacting with onstat did not. Unfortunately, no one thought to save
off a copy of what onstat dumped during the incident.
Our application is a touch heterogeneous, we have everything from cognos
impromptu, MS Access "applications", a massive 4gl based application
using shm to connect, to perl dbi based connections running off the same
database. The majority of users probably connect using shm connections
through the 4gl application but, MS Access makes an excessive number of
connections per user of a given application. Everything but the MS
Access and Congos Impromptu reports/applications are running from the
same machine as the engine. This machine also has an apache web server
to handle some cgi scripts, mostly perl code and the source of most of
our perl dbi connections.
When our problem occured, we had a reasonably heavy load from all
around. The machine had approximatly 560 or so proccess running total,
some 300-400 connections appearing in onstat at any one instant, and
three oninit processes eating up 90% of the cpu time solid for about an
hour. Simple queries would take 2 to 3 minutes to complete. There
were no apparent harware problems with the equipment, no excessive
numbers of locks or full logical logs, no single process or small group
of processes that looked like they were doing anything out of the
ordinary.
Fortunately, everything seemed to calm down again about an hour latter.
As a note, that day represented an exceptionally high load in comparison
to our normal operations and this could all be related to having maxed
out what our hardware could handle.
Is there any tuning that might be suggested for our system? I've
attached a copy of our config file. ( I've sense increased the frequency of my
logging cron job )
CONFIG FILE:
#**************************************************************************
#
# IBM Corporation
#
# Title: onconfig.std
# Description: IBM Informix Dynamic Server Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME root # Root dbspace nameROOTPATH /opt/informix/dev/root.1 # Path for device containing root
dbspace
ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 1024000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH /opt/informix/dev/root.1-m # Path for device containing
mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS root # Location (dbspace) of physical log
PHYSFILE 50000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 60 # Number of logical log files
LOGSIZE 10000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /opt/informix/Logs/cars.log # System message log file path
CONSOLE /dev/console # System console message path
# To automatically backup logical logs, edit alarmprogram.sh and set
# BACKUPLOGS=Y
ALARMPROGRAM /opt/informix/etc/log_full.sh # Alarm program path
TBLSPACE_STATS 1 # Maintain tblspace statistics
# System Archive Tape Device
TAPEDEV /dev/rmt/1m # Tape device path
TAPEBLK 32 # Tape block size (Kbytes)
TAPESIZE 0 # Maximum amount of data to put on tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV /dev/rmt/0m # Log tape device path
LTAPEBLK 32 # Log tape block size (Kbytes)
LTAPESIZE 0 # Max amount of data to put on log tape (Kbytes)
# Optical
STAGEBLOB # Informix Dynamic Server staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLine instance
DBSERVERNAME paul # Name of default database server
DBSERVERALIASES carsitcp # List of alternate dbservernames
NETTYPE ipcshm,2,500,CPU # Configure poll thread(s) for nettype
NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 3 # Number of user (cpu) vps
OK, I've calculated the metrics from the onstat -p output below. It's not
particularly useful unless you've zerod the stats more recently than server
startup 13 days ago, but here they are:
Pagreads: 7908428
Bufwrits: 35477827
Bufwaits: 1443280
BUFFERS: 90000
Time (hours) since reset: 321.33
ixda-RA: 2346196
idx-RA: 86301
da-RA: 4646343
RA-pgsused: 7071512
BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
These look fine, but a shorter accumulation period would give more reliable
numbers. For best results you should be saving the underlying values at least
daily (some save them before and after peak load periods each day) and clearing
the stats at least weekly. That will allow you to recalculate the metrics over
several time spans and during peak periods.
Art S. Kagel
----- Original Message -----
From: Chris Salch <ids@iiug.org>
At: 8/31 14:42:38
( Unfortunatly, none of us thought to save any of that data when things
started slowing down, so all we've got is what we can remember. )
In answer to several:
3 oninit processes were the only things really eating cpu time. There
were a few other processes that would show up using a small amount of
cpu time but nothing excessive. We were getting a load average just
topping 2 at the time ( this has happened before without any significant
impact to the system )
iostat was showing about 40kbps - 60kbps across all the drives,
significantly lower than what I see now or even after everything calmed
back down and our scheduled backup was started.
Here's an onstat from shortly after if it means much, (The backup
process was started after the whole mess).
IBM Informix Dynamic Server Version 9.40.HC5 -- On-Line -- Up 13
days 09:21:20 -- 426136 Kbytes
Userthreads
address flags sessid user tty wait tout locks nreads
nwrites
d2767018 ---P--D 1 root - 0 0 0 633
4202
d276763c ---P--F 0 root - 0 0 0 0
77695
d2767c60 ---P--F 0 root - 0 0 0 0
11181
d2768284 ---P--F 0 root - 0 0 0 0
7118
d27688a8 ---P--F 0 root - 0 0 0 0
18300
d2768ecc ---P--F 0 root - 0 0 0 0
9348
d27694f0 ---P--F 0 root - 0 0 0 0
7618
d2769b14 ---P--F 0 root - 0 0 0 0
6462
d276a138 ---P--F 0 root - 0 0 0 0
4758
d276a75c ---P--- 13 root - 0 0 0 0
24
d276ad80 ---P--B 14 root - 0 0 0 65074
480
d276b9c8 Y--P--- 352236 carsu tC d8375840 0 1 1019
604
d276bfec ---P--D 17 root - 0 0 0 0
0
d276c610 Y--P--D 26 root - c4fd3b6c 0 0 0
0
d276cc34 Y--P--- 27 roedelm ASTRASER d3201420 0 1 0
0
d276d258 Y--P--- 28 roedelm ASTRASER d30a6358 0 1 0
0
d276d87c Y--P--- 384521 zijlstre tg d9e4c658 0 1 9
0
d276dea0 Y--P--- 390938 carsu tO d7df4560 0 4 0
0
d276eae8 Y--P--- 392296 erpeldic JOHN d70e6b20 0 1 0
0
d276fd54 Y--P--- 389877 carsu tkb d4f0b150 0 1 132
170
d2770378 Y--P--- 386225 sgps JOHN d9961308 0 1 1
0
d277099c ---PR-- 392439 academus - 0 0 1 10
0
d2770fc0 Y--P--- 391698 gullettj tnb d9e3a278 0 1 30
0
d27715e4 Y--P--- 392236 bakerl LINDABAK d53d67f8 0 1 0
0
d2771c08 Y--P--- 392438 academus - d502aa20 0 1 0
0
d277222c Y--P--- 26513 academus - d4865868 0 1 0
0
d2772850 Y--P--- 392379 sgps JOHN d9bcf078 0 1 0
0
d2772e74 Y--P--- 344045 carsu tG d71e18b8 0 1 8179
0
d2773498 Y--P--- 381962 trogdond t3 d43afe78 0 1 1
0
d2774704 Y--P--- 383934 hoffpowj teb d7e04de0 0 1 10
0
d2774d28 Y--P--- 387046 bakerl LINDABAK d600ad58 0 1 1
0
d277534c Y--P--- 389228 burrowsr t0 d7332a28 0 1 104
0
d2775970 Y--P--- 343854 murleyt TOBIMURL d35e5818 0 1 319
36
d2775f94 Y--P--- 390673 childree tI d8d7adf8 0 1 32
0
d2776bdc Y--P--- 352174 sackettp PAMSACKE d78af750 0 1 173
22
d2777200 Y--P--- 385510 zijlstre tg dc89a9b0 0 2 134
0
d2777824 Y--P--- 392149 sgps JOHN d93b61f8 0 1 0
0
d2777e48 Y--P--- 344941 holleyc DAL-REP1 d5aad900 0 1 581
46
d277846c Y--P--- 773 carsu tf d45ae1b0 0 1 0
0
d27790b4 Y--P--- 344192 carsu tt d74e7bc8 0 1 1795
0
d27796d8 Y--P--- 375850 hyattr RENEEHYA d48798b8 0 1 1548
54
d277b58c Y--P--- 357910 hueyt TERESAHU d3b50968 0 1 1467
134
d277bbb0 Y--P--- 387128 carsu tl d66aef80 0 1 225
0
d277c1d4 Y--P--- 390897 holleyc DAL-REP1 d75015b0 0 1 1
0
d277da64 Y--P--- 343646 vestalt tq d4df9da8 0 1 2
0
d277e6ac Y--P--- 155417 carsu tD d59b3a20 0 1 0
0
d277f2f4 Y--P--- 343715 sarverl tu d4331c58 0 1 1221
0
d277ff3c Y--P--- 392116 sgps JOHN d761ce30 0 1 0
0
d2780560 Y--P--- 343793 cowartj JULIANCO d39af2b8 0 1 1494
722
d2781df0 Y--P--- 392277 smeetond - d64dec78 0 1 0
0
d2782a38 ---PR-- 384779 wolffd - 0 0 3 16752
7880
d278305c Y--P--- 390841 johnsonh tO d548c7c0 0 1 3
0
d27842c8 Y--P--- 390978 carsu t8 d3682c60 0 1 0
0
d27848ec Y--P--- 390741 carsu t8 d8b67fa8 0 1 157
0
d2785b58 ------- 391626 informix tty2a11 0 0 0 312628
0
d278617c Y--P--- 382952 carsu tk d91fb7d8 0 1 251
0
d27867a0 Y--P--- 252531 carsu t4 d9534b28 0 2 0
0
d2786dc4 Y--P--- 376405 gullettj tnb d70c1558 0 1 7
0
d27873e8 Y--P--- 381469 wolffd - d600acf8 0 1 2
0
d2787a0c Y--P--- 353053 meltonc t9 d7057fb8 0 1 1
0
d2788654 Y--P--- 392449 carsu td d478b828 0 1 869
0
d278929c Y--P--- 392295 erpeldic JOHN d66c4c78 0 1 0
0
d278a508 Y--P--- 390298 carsu tk d8b814f8 0 1 2738
0
d278ab2c Y--P--- 346340 wessonc tl dc89aad0 0 1 132
0
d278b150 Y--P--- 389623 sgps JOHN d39afb58 0 1 0
0
d278b774 Y--P--- 392066 keilersv ti d4e43ca0 0 1 4
0
d278bd98 Y--P--- 387915 carsu teb d91fbea0 0 1 1195
0
d278d004 Y--P--- 348675 bradfork tgb d70c17b8 0 1 3
0
d278dc4c Y--P--- 344032 pattersk tG d9be5680 0 1 2
0
d278eeb8 Y--P--- 344035 carsu tG d7123220 0 1 13
0
d2790124 Y--P--- 390760 carsu t8 d4c13748 0 1 3255
0
d2790748 Y--P--- 388516 carsu tnb d676d478 0 1 2992
0
d27919b4 Y--P--- 384549 carsu tg d3255478 0 4 6536
40
d2791fd8 Y--P--- 344477 bakerl tR da0f4e88 0 1 10
0
d27925fc Y--P--- 343677 vestalt tq d42de718 0 2 2182
0
d2792c20 Y--P--- 390425 sgps JOHN d4f87150 0 1 0
0
d2793244 Y--P--- 391011 suarezj - d94736f8 0 2 251
4
d27944b0 Y--P--- 117944 carsu th d4787428 0 1 0
0
d279571c Y--P--- 344043 carsu tG d73f1bd0 0 1 10003
634
d2796364 Y--P--- 385067 kunderp tv d4e31d40 0 1 7
0
d2796988 Y--P--- 392168 bakerl LINDABAK d3b64c98 0 1 11
0
d2796fac Y--P--- 745 hammd tf d5adaf68 0 1 0
0
d27975d0 Y--P--- 746 hammd tf d5af5ca0 0 1 0
0
d5c30ecc Y--P--- 392385 carsu tg d9f50e18 0 1 167
0
d5c31b14 Y--P--- 344484 carsu tR d9b80588 0 1 4042
58
d5c32138 Y--P--- 383938 carsu teb d4f872e8 0 4 6761
0
d5c3275c Y--P--- 392440 academus - d9961018 0 1 0
0
d5c339c8 Y--P--- 346791 mcgheeb tbb d70c1cb0 0 1 5
0
d5c33fec Y--P--- 384998 carsu teb d4df9558 0 1 0
0
d5c34c34 Y--P--- 376219 carsu tkb d8237630 0 1 9807
0
d5c35258 Y--P--- 390653 childree tI d77f2f98 0 1 0
0
d5c35ea0 Y--P--- 344041 carsu tG d624f140 0 1 7429
4780
d5c36ae8 ---P--- 392464 roedelm - 0 0 1 0
0@@NL
Hola Chris!
No estas usando los dbspaces temporales.
DBSPACETEMP temp0:temp1 # Default temp dbspaces
Cambiar los 2 puntos ":" por coma ","
DBSPACETEMP temp0,temp1 # Default temp dbspaces
Saludos!
----- Original Message -----
From: "CHRIS SALCH" <chrissalch@letu.edu>
To: <ids@iiug.org>
Sent: Thursday, August 31, 2006 12:17 PM
Subject: A question of load . . . [7374]
>
> We have an HP9000/L2000 with 4 64bit PA-RISC chips running at 440Mhz and
>
> 5 Gigs of ram. We're running Inforimx 9.40.HC5 on HPUX 11i. We ran
> into an issue on Monday where our database engine seemed to grind to a
> halt. The machine itself responded reasonably to everything but
> database queries. It seemed like anything having to do with extracting
> or querying data in the database took exceptionally long to respond.
> Interacting with onstat did not. Unfortunately, no one thought to save
> off a copy of what onstat dumped during the incident.
>
> Our application is a touch heterogeneous, we have everything from cognos
>
> impromptu, MS Access "applications", a massive 4gl based application
> using shm to connect, to perl dbi based connections running off the same
>
> database. The majority of users probably connect using shm connections
> through the 4gl application but, MS Access makes an excessive number of
> connections per user of a given application. Everything but the MS
> Access and Congos Impromptu reports/applications are running from the
> same machine as the engine. This machine also has an apache web server
> to handle some cgi scripts, mostly perl code and the source of most of
> our perl dbi connections.
>
> When our problem occured, we had a reasonably heavy load from all
> around. The machine had approximatly 560 or so proccess running total,
> some 300-400 connections appearing in onstat at any one instant, and
> three oninit processes eating up 90% of the cpu time solid for about an
> hour. Simple queries would take 2 to 3 minutes to complete. There
> were no apparent harware problems with the equipment, no excessive
> numbers of locks or full logical logs, no single process or small group
> of processes that looked like they were doing anything out of the
> ordinary.
>
> Fortunately, everything seemed to calm down again about an hour latter.
> As a note, that day represented an exceptionally high load in comparison
>
> to our normal operations and this could all be related to having maxed
> out what our hardware could handle.
>
> Is there any tuning that might be suggested for our system? I've
> attached a copy of our config file. ( I've sense increased the frequency
> of my
> logging cron job )
>
> CONFIG FILE:
>
> #***********************************************************************
> ***
> #
> # IBM Corporation
> #
> # Title: onconfig.std
> # Description: IBM Informix Dynamic Server Configuration Parameters
> #
> #***********************************************************************
> ***
>
> # Root Dbspace Configuration
>
> ROOTNAME root # Root dbspace name> ROOTPATH /opt/informix/dev/root.1 # Path for device containing root
> dbspace
> ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
> ROOTSIZE 1024000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH /opt/informix/dev/root.1-m # Path for device containing
> mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS root # Location (dbspace) of physical log
> PHYSFILE 50000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 60 # Number of logical log files
> LOGSIZE 10000 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /opt/informix/Logs/cars.log # System message log file path
> CONSOLE /dev/console # System console message path>
> # To automatically backup logical logs, edit alarmprogram.sh and set
> # BACKUPLOGS=Y
> ALARMPROGRAM /opt/informix/etc/log_full.sh # Alarm program path
> TBLSPACE_STATS 1 # Maintain tblspace statistics>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt/1m # Tape device path
> TAPEBLK 32 # Tape block size (Kbytes)
> TAPESIZE 0 # Maximum amount of data to put on tape (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /dev/rmt/0m # Log tape device path
> LTAPEBLK 32 # Log tape block size (Kbytes)
> LTAPESIZE 0 # Max amount of data to put on log tape (Kbytes)>
> # Optical
>
> STAGEBLOB # Informix Dynamic Server staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a OnLine instance
> DBSERVERNAME paul # Name of default database server
> DBSERVERALIASES carsitcp # List of alternate dbservernames
> NETTYPE ipcshm,2,500,CPU # Configure poll thread(s) for nettype
> NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 3 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 256000 # Maximum number of locks
> BUFFERS 90000 # Maximum number of shared buffers
> NUMAIOVPS 8 # Number of IO vps
> PHYSBUFF 32 # Physical log buffer size (Kbytes)
> LOGBUFF 32 # Logical log buffer size (Kbytes)> LOGSMAX 120 # Maximum number of logical log files
> CLEANERS 8 # Number of buffer cleaner processes
> SHMBASE 0x0 # Shared memory base address
> SHMVIRTSIZE 196608 # initial virtual shared memory segment size
> SHMADD 16384 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
> 0=>unlimited
> CKPTINTVL 900 # Check point interval (in sec)
> LRUS 8 # Number of LRU queues
> LRU_MAX_DIRTY 4 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 64 # Stack size (Kbytes)>
> # Dynamic Logging
> # DYNAMIC_LOGS:
> # 2 : server automatically add a new logical log when necessary. (ON)
> # 1 : notify DBA to add new logical logs when necessary. (ON)
> # 0 : cannot add logical log on the fly. (OFF)
> #
> # When dynamic logging is on, we can have higher values for
> LTXHWM/LTXEHWM,
> # because the server can add new logical logs during long transaction
> rollback.
> # However, to limit the number of new logical logs being added,
> LTXHWM/LTXEHWM
> # can be set to smaller values.
> #
> # If dynamic logging is off, LTXHWM/LTXEHWM need to be set to smaller
> values
> # to avoid long transaction rollback hanging the server due to lack of
> logical
> # log space, i.e. 50/60 or lower.
>
> DYNAMIC_LOGS 0
> LTXHWM 45
> LTXEHWM 55>
> # System Page Size@@NL
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Norberto Valverde LLanos
Sent: Thursday, August 31, 2006 3:38 PM
To: ids@iiug.org
Subject: Re: A question of load . . . [7382]
Hola Chris!
No estas usando los dbspaces temporales.
DBSPACETEMP temp0:temp1 # Default temp dbspaces
Cambiar los 2 puntos ":" por coma ","
DBSPACETEMP temp0,temp1 # Default temp dbspaces
Saludos!
----- Original Message -----
From: "CHRIS SALCH" <chrissalch@letu.edu>
To: <ids@iiug.org>
Sent: Thursday, August 31, 2006 12:17 PM
Subject: A question of load . . . [7374]
>
> We have an HP9000/L2000 with 4 64bit PA-RISC chips running at 440Mhz
and
>
> 5 Gigs of ram. We're running Inforimx 9.40.HC5 on HPUX 11i. We ran
> into an issue on Monday where our database engine seemed to grind to a
> halt. The machine itself responded reasonably to everything but
> database queries. It seemed like anything having to do with extracting
> or querying data in the database took exceptionally long to respond.
> Interacting with onstat did not. Unfortunately, no one thought to save
> off a copy of what onstat dumped during the incident.
>
> Our application is a touch heterogeneous, we have everything from
cognos
>
> impromptu, MS Access "applications", a massive 4gl based application
> using shm to connect, to perl dbi based connections running off the
same
>
> database. The majority of users probably connect using shm connections
> through the 4gl application but, MS Access makes an excessive number
of
> connections per user of a given application. Everything but the MS
> Access and Congos Impromptu reports/applications are running from the
> same machine as the engine. This machine also has an apache web server
> to handle some cgi scripts, mostly perl code and the source of most of
> our perl dbi connections.
>
> When our problem occured, we had a reasonably heavy load from all
> around. The machine had approximatly 560 or so proccess running total,
> some 300-400 connections appearing in onstat at any one instant, and
> three oninit processes eating up 90% of the cpu time solid for about
an
> hour. Simple queries would take 2 to 3 minutes to complete. There
> were no apparent harware problems with the equipment, no excessive
> numbers of locks or full logical logs, no single process or small
group
> of processes that looked like they were doing anything out of the
> ordinary.
>
> Fortunately, everything seemed to calm down again about an hour
latter.
> As a note, that day represented an exceptionally high load in
comparison
>
> to our normal operations and this could all be related to having maxed
> out what our hardware could handle.
>
> Is there any tuning that might be suggested for our system? I've
> attached a copy of our config file. ( I've sense increased the
frequency
> of my
> logging cron job )
>
> CONFIG FILE:
>
>
#***********************************************************************
> ***
> #
> # IBM Corporation
> #
> # Title: onconfig.std
> # Description: IBM Informix Dynamic Server Configuration Parameters
> #
>
#***********************************************************************
> ***
>
> # Root Dbspace Configuration
>
> ROOTNAME root # Root dbspace name> ROOTPATH /opt/informix/dev/root.1 # Path for device containing root
> dbspace
> ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
> ROOTSIZE 1024000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH /opt/informix/dev/root.1-m # Path for device containing
> mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS root # Location (dbspace) of physical log
> PHYSFILE 50000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 60 # Number of logical log files
> LOGSIZE 10000 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /opt/informix/Logs/cars.log # System message log file path
> CONSOLE /dev/console # System console message path>
> # To automatically backup logical logs, edit alarmprogram.sh and set
> # BACKUPLOGS=Y
> ALARMPROGRAM /opt/informix/etc/log_full.sh # Alarm program path
> TBLSPACE_STATS 1 # Maintain tblspace statistics>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt/1m # Tape device path
> TAPEBLK 32 # Tape block size (Kbytes)
> TAPESIZE 0 # Maximum amount of data to put on tape (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /dev/rmt/0m # Log tape device path
> LTAPEBLK 32 # Log tape block size (Kbytes)
> LTAPESIZE 0 # Max amount of data to put on log tape (Kbytes)>
> # Optical
>
> STAGEBLOB # Informix Dynamic Server staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a OnLine instance
> DBSERVERNAME paul # Name of default database server
> DBSERVERALIASES carsitcp # List of alternate dbservernames
> NETTYPE ipcshm,2,500,CPU # Configure poll thread(s) for nettype
> NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 3 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 256000 # Maximum number of locks
> BUFFERS 90000 # Maximum number of shared buffers
> NUMAIOVPS 8 # Number of IO vps
> PHYSBUFF 32 # Physical log buffer size (Kbytes)
> LOGBUFF 32 # Logical log buffer size (Kbytes)> LOGSMAX 120 # Maximum number of logical log files
> CLEANERS 8 # Number of buffer cleaner processes
> SHMBASE 0x0 # Shared memory base address
> SHMVIRTSIZE 196608 # initial virtual shared memory segment size
> SHMADD 16384 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
> 0=>unlimited
> CKPTINTVL 900 # Check point interval (in sec)
> LRUS 8 # Number of LRU queues
> LRU_MAX_DIRTY 4 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 64 # Stack size (Kbytes)>
> # Dynamic Logging
> # DYNAMIC_LOGS:
> # 2 : server automatically add a new logical log when necessary. (ON)
> # 1 : notify DBA to add new logical logs when necessary. (ON)
> # 0 : cannot add logical log on the fly. (OFF)
> #
> # When dynamic logging is on, we can have higher values for
> LTXHWM/LTXEHWM,
> # because the server can add new logical logs during long transaction
> rollback.
> # However, to limit the number of new logical logs being added,
> LTXHWM/LTXEHW
Either comma (,) or colon (:) are accepted. These lists mimic PATH lists and
since UNIX uses colon and MS OSes use comma, Informix accespts either.
Art S. Kagel
----- Original Message -----
From: Andrew G Ford <ids@iiug.org>
At: 8/31 15:49:16
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Norberto Valverde LLanos
Sent: Thursday, August 31, 2006 3:38 PM
To: ids@iiug.org
Subject: Re: A question of load . . . [7382]
Hola Chris!
No estas usando los dbspaces temporales.
DBSPACETEMP temp0:temp1 # Default temp dbspaces
Cambiar los 2 puntos ":" por coma ","
DBSPACETEMP temp0,temp1 # Default temp dbspaces
Saludos!
----- Original Message -----
From: "CHRIS SALCH" <chrissalch@letu.edu>
To: <ids@iiug.org>
Sent: Thursday, August 31, 2006 12:17 PM
Subject: A question of load . . . [7374]
>
> We have an HP9000/L2000 with 4 64bit PA-RISC chips running at 440Mhz
and
>
> 5 Gigs of ram. We're running Inforimx 9.40.HC5 on HPUX 11i. We ran
> into an issue on Monday where our database engine seemed to grind to a
> halt. The machine itself responded reasonably to everything but
> database queries. It seemed like anything having to do with extracting
> or querying data in the database took exceptionally long to respond.
> Interacting with onstat did not. Unfortunately, no one thought to save
> off a copy of what onstat dumped during the incident.
>
> Our application is a touch heterogeneous, we have everything from
cognos
>
> impromptu, MS Access "applications", a massive 4gl based application
> using shm to connect, to perl dbi based connections running off the
same
>
> database. The majority of users probably connect using shm connections
> through the 4gl application but, MS Access makes an excessive number
of
> connections per user of a given application. Everything but the MS
> Access and Congos Impromptu reports/applications are running from the
> same machine as the engine. This machine also has an apache web server
> to handle some cgi scripts, mostly perl code and the source of most of
> our perl dbi connections.
>
> When our problem occured, we had a reasonably heavy load from all
> around. The machine had approximatly 560 or so proccess running total,
> some 300-400 connections appearing in onstat at any one instant, and
> three oninit processes eating up 90% of the cpu time solid for about
an
> hour. Simple queries would take 2 to 3 minutes to complete. There
> were no apparent harware problems with the equipment, no excessive
> numbers of locks or full logical logs, no single process or small
group
> of processes that looked like they were doing anything out of the
> ordinary.
>
> Fortunately, everything seemed to calm down again about an hour
latter.
> As a note, that day represented an exceptionally high load in
comparison
>
> to our normal operations and this could all be related to having maxed
> out what our hardware could handle.
>
> Is there any tuning that might be suggested for our system? I've
> attached a copy of our config file. ( I've sense increased the
frequency
> of my
> logging cron job )
>
> CONFIG FILE:
>
>
#***********************************************************************
> ***
> #
> # IBM Corporation
> #
> # Title: onconfig.std
> # Description: IBM Informix Dynamic Server Configuration Parameters
> #
>
#***********************************************************************
> ***
>
> # Root Dbspace Configuration
>
> ROOTNAME root # Root dbspace name> ROOTPATH /opt/informix/dev/root.1 # Path for device containing root
> dbspace
> ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
> ROOTSIZE 1024000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH /opt/informix/dev/root.1-m # Path for device containing
> mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS root # Location (dbspace) of physical log
> PHYSFILE 50000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 60 # Number of logical log files
> LOGSIZE 10000 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /opt/informix/Logs/cars.log # System message log file path
> CONSOLE /dev/console # System console message path>
> # To automatically backup logical logs, edit alarmprogram.sh and set
> # BACKUPLOGS=Y
> ALARMPROGRAM /opt/informix/etc/log_full.sh # Alarm program path
> TBLSPACE_STATS 1 # Maintain tblspace statistics>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt/1m # Tape device path
> TAPEBLK 32 # Tape block size (Kbytes)
> TAPESIZE 0 # Maximum amount of data to put on tape (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /dev/rmt/0m # Log tape device path
> LTAPEBLK 32 # Log tape block size (Kbytes)
> LTAPESIZE 0 # Max amount of data to put on log tape (Kbytes)>
> # Optical
>
> STAGEBLOB # Informix Dynamic Server staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a OnLine instance
> DBSERVERNAME paul # Name of default database server
> DBSERVERALIASES carsitcp # List of alternate dbservernames
> NETTYPE ipcshm,2,500,CPU # Configure poll thread(s) for nettype
> NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 3 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 256000 # Maximum number of locks
> BUFFERS 90000 # Maximum number of shared buffers
> NUMAIOVPS 8 # Number of IO vps
> PHYSBUFF 32 # Physical log buffer size (Kbytes)
> LOGBUFF 32 # Logical log buffer size (Kbytes)> LOGSMAX 120 # Maximum number of logical log files
> CLEANERS 8 # Number of buffer cleaner processes
> SHMBASE 0x0 # Shared memory base address
> SHMVIRTSIZE 196608 # initial virtual shared memory segment size
> SHMADD 16384 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
> 0=>unlimited
> CKPTINTVL 900 # Check point interval (in sec)
> LRUS 8 # Number of LRU queues
> LRU_MAX_DIRTY 4 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 64 # Stack size (Kbytes)>
> # Dynamic Logging
> # DYNAMIC_LOGS:
> # 2 : server automatically add a new logical log when necessary. (ON)
> # 1 : notify DBA to add new logical logs when necessary. (ON)
> # 0 : cannot add logical log on the fly. (OFF)
Those stats are reset every sunday at 1:05 am. So that BTR is a
significantly higher number than what you have shown. That would be a
BTR of about 12 ouch! (by your info)
On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
wrote:
> OK, I've calculated the metrics from the onstat -p output below. It's not
> particularly useful unless you've zerod the stats more recently than server
> startup 13 days ago, but here they are:
>
> Pagreads: 7908428
> Bufwrits: 35477827
> Bufwaits: 1443280
> BUFFERS: 90000
> Time (hours) since reset: 321.33
> ixda-RA: 2346196
> idx-RA: 86301
> da-RA: 4646343
> RA-pgsused: 7071512
>
> BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
> BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
> RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
>
> These look fine, but a shorter accumulation period would give more reliable
> numbers. For best results you should be saving the underlying values at least
> daily (some save them before and after peak load periods each day) and
> clearing
> the stats at least weekly. That will allow you to recalculate the metrics
over
> several time spans and during peak periods.
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Chris Salch <ids@iiug.org>
> At: 8/31 14:42:38
>
> ( Unfortunatly, none of us thought to save any of that data when things
> started slowing down, so all we've got is what we can remember. )
>
> In answer to several:
>
> 3 oninit processes were the only things really eating cpu time. There
> were a few other processes that would show up using a small amount of
> cpu time but nothing excessive. We were getting a load average just
> topping 2 at the time ( this has happened before without any significant
> impact to the system )
>
> iostat was showing about 40kbps - 60kbps across all the drives,
> significantly lower than what I see now or even after everything calmed
> back down and our scheduled backup was started.
>
> Here's an onstat from shortly after if it means much, (The backup
> process was started after the whole mess).
>
> IBM Informix Dynamic Server Version 9.40.HC5 -- On-Line -- Up 13
> days 09:21:20 -- 426136 Kbytes>
> Userthreads
> address flags sessid user tty wait tout locks nreads
> nwrites
> d2767018 ---P--D 1 root - 0 0 0 633
> 4202
> d276763c ---P--F 0 root - 0 0 0 0
> 77695
> d2767c60 ---P--F 0 root - 0 0 0 0
> 11181
> d2768284 ---P--F 0 root - 0 0 0 0
> 7118
> d27688a8 ---P--F 0 root - 0 0 0 0
> 18300
> d2768ecc ---P--F 0 root - 0 0 0 0
> 9348
> d27694f0 ---P--F 0 root - 0 0 0 0
> 7618
> d2769b14 ---P--F 0 root - 0 0 0 0
> 6462
> d276a138 ---P--F 0 root - 0 0 0 0
> 4758
> d276a75c ---P--- 13 root - 0 0 0 0
> 24
> d276ad80 ---P--B 14 root - 0 0 0 65074
> 480
> d276b9c8 Y--P--- 352236 carsu tC d8375840 0 1 1019
> 604
> d276bfec ---P--D 17 root - 0 0 0 0
> 0
> d276c610 Y--P--D 26 root - c4fd3b6c 0 0 0
> 0
> d276cc34 Y--P--- 27 roedelm ASTRASER d3201420 0 1 0
> 0
> d276d258 Y--P--- 28 roedelm ASTRASER d30a6358 0 1 0
> 0
> d276d87c Y--P--- 384521 zijlstre tg d9e4c658 0 1 9
> 0
> d276dea0 Y--P--- 390938 carsu tO d7df4560 0 4 0
> 0
> d276eae8 Y--P--- 392296 erpeldic JOHN d70e6b20 0 1 0
> 0
> d276fd54 Y--P--- 389877 carsu tkb d4f0b150 0 1 132
> 170
> d2770378 Y--P--- 386225 sgps JOHN d9961308 0 1 1
> 0
> d277099c ---PR-- 392439 academus - 0 0 1 10
> 0
> d2770fc0 Y--P--- 391698 gullettj tnb d9e3a278 0 1 30
> 0
> d27715e4 Y--P--- 392236 bakerl LINDABAK d53d67f8 0 1 0
> 0
> d2771c08 Y--P--- 392438 academus - d502aa20 0 1 0
> 0
> d277222c Y--P--- 26513 academus - d4865868 0 1 0
> 0
> d2772850 Y--P--- 392379 sgps JOHN d9bcf078 0 1 0
> 0
> d2772e74 Y--P--- 344045 carsu tG d71e18b8 0 1 8179
> 0
> d2773498 Y--P--- 381962 trogdond t3 d43afe78 0 1 1
> 0
> d2774704 Y--P--- 383934 hoffpowj teb d7e04de0 0 1 10
> 0
> d2774d28 Y--P--- 387046 bakerl LINDABAK d600ad58 0 1 1
> 0
> d277534c Y--P--- 389228 burrowsr t0 d7332a28 0 1 104
> 0
> d2775970 Y--P--- 343854 murleyt TOBIMURL d35e5818 0 1 319
> 36
> d2775f94 Y--P--- 390673 childree tI d8d7adf8 0 1 32
> 0
> d2776bdc Y--P--- 352174 sackettp PAMSACKE d78af750 0 1 173
> 22
> d2777200 Y--P--- 385510 zijlstre tg dc89a9b0 0 2 134
> 0
> d2777824 Y--P--- 392149 sgps JOHN d93b61f8 0 1 0
> 0
> d2777e48 Y--P--- 344941 holleyc DAL-REP1 d5aad900 0 1 581
> 46
> d277846c Y--P--- 773 carsu tf d45ae1b0 0 1 0
> 0
> d27790b4 Y--P--- 344192 carsu tt d74e7bc8 0 1 1795
> 0
> d27796d8 Y--P--- 375850 hyattr RENEEHYA d48798b8 0 1 1548
> 54
> d277b58c Y--P--- 357910 hueyt TERESAHU d3b50968 0 1 1467
> 134
> d277bbb0 Y--P--- 387128 carsu tl d66aef80 0 1 225
> 0
> d277c1d4 Y--P--- 390897 holleyc DAL-REP1 d75015b0 0 1 1
> 0
> d277da64 Y--P--- 343646 vestalt tq d4df9da8 0 1 2
> 0
> d277e6ac Y--P--- 155417 carsu tD d59b3a20 0 1 0
> 0
> d277f2f4 Y--P--- 343715 sarverl tu d4331c58 0 1 1221
> 0
> d277ff3c Y--P--- 392116 sgps JOHN d761ce30 0 1 0
> 0
> d2780560 Y--P--- 343793 cowartj JULIANCO d39af2b8 0 1 1494
> 722
> d2781df0 Y--P--- 392277 smeetond - d64dec78 0 1 0
> 0
> d2782a38 ---PR-- 384779 wolffd - 0 0 3 16752
> 7880
> d278305c Y--P--- 390841 johnsonh tO d548c7c0 0 1 3
> 0
> d27842c8 Y--P--- 390978 carsu t8 d3682c60 0 1 0
> 0
> d27848ec Y--P--- 390741 carsu t8 d8b67fa8 0 1 157
> 0
> d2785b58 ------- 391626 informix tty2a11 0 0 0 312628
> 0
> d278617c Y--P--- 382952 carsu tk d91fb7d8 0 1 251
> 0
> d27867a0 Y--P--- 252531 carsu t4 d9534b28 0 2 0
> 0
> d2786dc4 Y--P--- 376405 gullettj tnb d70c1558 0 1 7
> 0
> d27873e8 Y--P--- 381469 wolffd - d600acf8 0 1 2
> 0
> d2787a0c Y--P--- 353053 meltonc t9 d7057fb8 0 1 1
> 0
> d2788654 Y--P--- 392449 carsu td d478b828 0 1 869
> 0
> d278929c Y--P--- 392295 erpeldic JOHN d66c4c78 0 1 0
> 0
> d278a508 Y--P--- 390298 carsu tk d8b814f8 0 1 2738
> 0
> d278ab2c Y--P--- 346340 wessonc tl dc89aad0 0 1 132
> 0
> d278b150 Y--P--- 389623 sgps JOHN d39afb58 0 1 0
> 0
> d278b774 Y--P--- 392066 keilersv ti d4e43ca0 0 1 4
> 0
> d278bd98 Y--P--- 387915 carsu teb d91fbea0 0 1 1195
> 0
> d278d004 Y--P--- 348675 bradfork tgb d70c17b8 0 1 3
> 0
> d278dc4c Y--P--- 344032 pattersk tG d9be5680 0 1 2
> 0
> d278eeb8 Y--P--- 344035 carsu tG d7123220 0 1 13
> 0
> d2790124 Y--P--- 390760 carsu t8 d4c13748 0 1 3255
> 0
> d2790748 Y--P--- 388516 carsu tnb d676d478 0 1 2992
> 0
> d27919b4 Y--P--- 384549 carsu tg d3255478 0 4 6536
> 40
> d2791fd8 Y--P--- 344477 bakerl tR da0f4e88 0 1 10
> 0
> d27925fc Y--P--- 343677 vestalt tq d42de718 0 2 2182
> 0
> d2792c20 Y--P--- 390425 sgps JOHN d4f87150 0 1 0
> 0
> d2793244 Y--P--- 391011 suarezj - d94736f8 0 2 251
> 4
> d27944b0 Y--P--- 117944 carsu th d4787428 0 1 0
> 0
> d279571c Y--P--- 344043 carsu tG d73f1bd0 0 1 10003
> 634
> d2796364 Y--P--- 385067 kunderp tv d4e31d40 0 1 7
> 0
> d2796988 Y--P--- 392168 bakerl LINDABAK d3b64c98 0 1 11
> 0
> d2796fac Y--P--- 745 hamm
And the ratios script shows:
pgsused/(ixda-RA+idx-RA+da-RA)*100
Read-Ahead Util (RAU): 99.90% 23922139/(6713181+244596+16988224)*100
bufwaits*100 / (pagreads+bufwrits)
Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=90000)
Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
cached)))/(BUFFERS)
BT Period-max: 19.90/hr every 3.01 minutes.
BT Period-min: 3.11/hr every 19.29 minutes.
Does this look as off as I think it does?
On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> Those stats are reset every sunday at 1:05 am. So that BTR is a
> significantly higher number than what you have shown. That would be a
> BTR of about 12 ouch! (by your info)
>
> On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
> wrote:
> > OK, I've calculated the metrics from the onstat -p output below. It's not
> > particularly useful unless you've zerod the stats more recently than server
> > startup 13 days ago, but here they are:
> >
> > Pagreads: 7908428
> > Bufwrits: 35477827
> > Bufwaits: 1443280
> > BUFFERS: 90000
> > Time (hours) since reset: 321.33
> > ixda-RA: 2346196
> > idx-RA: 86301
> > da-RA: 4646343
> > RA-pgsused: 7071512
> >
> > BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
> > BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
> > RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
> >
> > These look fine, but a shorter accumulation period would give more reliable
> > numbers. For best results you should be saving the underlying values at
> least
> > daily (some save them before and after peak load periods each day) and
> > clearing
> > the stats at least weekly. That will allow you to recalculate the metrics
> over
> > several time spans and during peak periods.
> >
> > Art S. Kagel
> >
> > ----- Original Message -----
> > From: Chris Salch <ids@iiug.org>
> > At: 8/31 14:42:38
> >
> > ( Unfortunatly, none of us thought to save any of that data when things
> > started slowing down, so all we've got is what we can remember. )
> >
> > In answer to several:
> >
> > 3 oninit processes were the only things really eating cpu time. There
> > were a few other processes that would show up using a small amount of
> > cpu time but nothing excessive. We were getting a load average just
> > topping 2 at the time ( this has happened before without any significant
> > impact to the system )
> >
> > iostat was showing about 40kbps - 60kbps across all the drives,
> > significantly lower than what I see now or even after everything calmed
> > back down and our scheduled backup was started.
> >
> > Here's an onstat from shortly after if it means much, (The backup
> > process was started after the whole mess).
> >
> > IBM Informix Dynamic Server Version 9.40.HC5 -- On-Line -- Up 13
> > days 09:21:20 -- 426136 Kbytes> >
> > Userthreads
> > address flags sessid user tty wait tout locks nreads
> > nwrites
> > d2767018 ---P--D 1 root - 0 0 0 633
> > 4202
> > d276763c ---P--F 0 root - 0 0 0 0
> > 77695
> > d2767c60 ---P--F 0 root - 0 0 0 0
> > 11181
> > d2768284 ---P--F 0 root - 0 0 0 0
> > 7118
> > d27688a8 ---P--F 0 root - 0 0 0 0
> > 18300
> > d2768ecc ---P--F 0 root - 0 0 0 0
> > 9348
> > d27694f0 ---P--F 0 root - 0 0 0 0
> > 7618
> > d2769b14 ---P--F 0 root - 0 0 0 0
> > 6462
> > d276a138 ---P--F 0 root - 0 0 0 0
> > 4758
> > d276a75c ---P--- 13 root - 0 0 0 0
> > 24
> > d276ad80 ---P--B 14 root - 0 0 0 65074
> > 480
> > d276b9c8 Y--P--- 352236 carsu tC d8375840 0 1 1019
> > 604
> > d276bfec ---P--D 17 root - 0 0 0 0
> > 0
> > d276c610 Y--P--D 26 root - c4fd3b6c 0 0 0
> > 0
> > d276cc34 Y--P--- 27 roedelm ASTRASER d3201420 0 1 0
> > 0
> > d276d258 Y--P--- 28 roedelm ASTRASER d30a6358 0 1 0
> > 0
> > d276d87c Y--P--- 384521 zijlstre tg d9e4c658 0 1 9
> > 0
> > d276dea0 Y--P--- 390938 carsu tO d7df4560 0 4 0
> > 0
> > d276eae8 Y--P--- 392296 erpeldic JOHN d70e6b20 0 1 0
> > 0
> > d276fd54 Y--P--- 389877 carsu tkb d4f0b150 0 1 132
> > 170
> > d2770378 Y--P--- 386225 sgps JOHN d9961308 0 1 1
> > 0
> > d277099c ---PR-- 392439 academus - 0 0 1 10
> > 0
> > d2770fc0 Y--P--- 391698 gullettj tnb d9e3a278 0 1 30
> > 0
> > d27715e4 Y--P--- 392236 bakerl LINDABAK d53d67f8 0 1 0
> > 0
> > d2771c08 Y--P--- 392438 academus - d502aa20 0 1 0
> > 0
> > d277222c Y--P--- 26513 academus - d4865868 0 1 0
> > 0
> > d2772850 Y--P--- 392379 sgps JOHN d9bcf078 0 1 0
> > 0
> > d2772e74 Y--P--- 344045 carsu tG d71e18b8 0 1 8179
> > 0
> > d2773498 Y--P--- 381962 trogdond t3 d43afe78 0 1 1
> > 0
> > d2774704 Y--P--- 383934 hoffpowj teb d7e04de0 0 1 10
> > 0
> > d2774d28 Y--P--- 387046 bakerl LINDABAK d600ad58 0 1 1
> > 0
> > d277534c Y--P--- 389228 burrowsr t0 d7332a28 0 1 104
> > 0
> > d2775970 Y--P--- 343854 murleyt TOBIMURL d35e5818 0 1 319
> > 36
> > d2775f94 Y--P--- 390673 childree tI d8d7adf8 0 1 32
> > 0
> > d2776bdc Y--P--- 352174 sackettp PAMSACKE d78af750 0 1 173
> > 22
> > d2777200 Y--P--- 385510 zijlstre tg dc89a9b0 0 2 134
> > 0
> > d2777824 Y--P--- 392149 sgps JOHN d93b61f8 0 1 0
> > 0
> > d2777e48 Y--P--- 344941 holleyc DAL-REP1 d5aad900 0 1 581
> > 46
> > d277846c Y--P--- 773 carsu tf d45ae1b0 0 1 0
> > 0
> > d27790b4 Y--P--- 344192 carsu tt d74e7bc8 0 1 1795
> > 0
> > d27796d8 Y--P--- 375850 hyattr RENEEHYA d48798b8 0 1 1548
> > 54
> > d277b58c Y--P--- 357910 hueyt TERESAHU d3b50968 0 1 1467
> > 134
> > d277bbb0 Y--P--- 387128 carsu tl d66aef80 0 1 225
> > 0
> > d277c1d4 Y--P--- 390897 holleyc DAL-REP1 d75015b0 0 1 1
> > 0
> > d277da64 Y--P--- 343646 vestalt tq d4df9da8 0 1 2
> > 0
> > d277e6ac Y--P--- 155417 carsu tD d59b3a20 0 1 0
> > 0
> > d277f2f4 Y--P--- 343715 sarverl tu d4331c58 0 1 1221
> > 0
> > d277ff3c Y--P--- 392116 sgps JOHN d761ce30 0 1 0
> > 0
> > d2780560 Y--P--- 343793 cowartj JULIANCO d39af2b8 0 1 1494
> > 722
> > d2781df0 Y--P--- 392277 smeetond - d64dec78 0 1 0
> > 0
> > d2782a38 ---PR-- 384779 wolffd - 0 0 3 16752
> > 7880
> > d278305c Y--P--- 390841 johnsonh tO d548c7c0 0 1 3
> > 0
> > d27842c8 Y--P--- 390978 carsu t8 d3682c60 0 1 0
> > 0
> > d27848ec Y--P--- 390741 carsu t8 d8b67fa8 0 1 157
> > 0
> > d2785b58 ------- 391626 informix tty2a11 0 0 0 312628
> > 0
> > d278617c Y--P--- 382952 carsu tk d91fb7d8 0 1 251
> > 0
> > d27867a0 Y--P--- 252531 carsu t4 d9534b28 0 2 0
> > 0
> > d2786dc4 Y--P--- 376405 gullettj tnb d70c1558 0 1 7
> > 0
> > d27873e8 Y--P--- 381469 wolffd - d600acf8 0 1 2
> > 0
> > d2787a0c Y--P--- 353053 meltonc t9 d7057fb8 0 1 1
> > 0
> > d2788654 Y--P--- 392449 carsu td d478b828 0 1 869
> > 0
> > d278929c Y--P--- 392295 erpeldic JOHN d66c4c78 0 1 0
> > 0
> > d278a508 Y--P--- 390298 carsu tk d8b814f8 0 1 2738
> > 0
> > d278ab2c Y--P--- 346340 wessonc tl dc89aad0 0 1 132
> > 0
> > d278b150 Y--P--- 389623 sgps JOHN d39afb58 0 1 0
> > 0
> > d278b774 Y--P--- 392066 keilersv ti d4e43ca0 0 1 4
>
19, yes, that's more like what I expected to see. Ignore the min - It was an attempt by me to filter out the effects of churning a subset of the cache. It turns out after experimentation that it's only a reliable number if the server is relatively quiet. Ahh well. Art S. Kagel
The next time it happens look to see if you have db transactions rolling
back, locking issues, long chk point times, take snapshots of sql operations in
process etc. it may give you a clue.
if you have hp glance - look at cpu, disk i/o rates when you have high cpu
usage, you may see oninit processes with high cpu usage bouncing from cpu to
cpu. I've had some success locking oninits to specific cpus, also playing
around
with config params on # of cpu vps has an effect on HP machines. Informix n-1
rule doesn't always work well with hp. Also look at shared memory segments,
sometimes you get extra ones that didn't get released when the engine was
bounced - you can tie up a lot of real memory. See if extra memory segments are
being added dynamically, this will slow you down - I'll told this is an HP
architecture issue - # mapping registers.
Ken Moskowitz
Chris
90000 buffers on a 2K page size O/S means 180 Mb of buffer space! With
5Gb memory, even with what else is happening on the machine, I would
try 350000 buffers taking up 700 Mb. You will need to increase LRUs
(try 127) and CLEANERS (127 to match LRUs). Also watch checkpoint
times and be prepared to decrease the interval or LRU_MAX and LRU_MIN.
I also noticed your LOG_BUFF and PHYS_BUFF are still at default
values. There may be some milage in increasing these (even as far as
128 or 256) which can help as well.
Keith
On 31/08/06, Chris Salch <chrissalch@letu.edu> wrote:
>
> And the ratios script shows:
>
> pgsused/(ixda-RA+idx-RA+da-RA)*100
> Read-Ahead Util (RAU): 99.90% 23922139/(6713181+244596+16988224)*100
>
> bufwaits*100 / (pagreads+bufwrits)
> Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
> Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=90000)
> Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
> cached)))/(BUFFERS)
>
> BT Period-max: 19.90/hr every 3.01 minutes.
>
> BT Period-min: 3.11/hr every 19.29 minutes.
>
> Does this look as off as I think it does?
>
> On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> > Those stats are reset every sunday at 1:05 am. So that BTR is a
> > significantly higher number than what you have shown. That would be a
> > BTR of about 12 ouch! (by your info)
> >
> > On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
> > wrote:
> > > OK, I've calculated the metrics from the onstat -p output below. It's not
> > > particularly useful unless you've zerod the stats more recently than
> server
> > > startup 13 days ago, but here they are:
> > >
> > > Pagreads: 7908428
> > > Bufwrits: 35477827
> > > Bufwaits: 1443280
> > > BUFFERS: 90000
> > > Time (hours) since reset: 321.33
> > > ixda-RA: 2346196
> > > idx-RA: 86301
> > > da-RA: 4646343
> > > RA-pgsused: 7071512
> > >
> > > BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
> > > BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
> > > RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
> > >
> > > These look fine, but a shorter accumulation period would give more
> reliable
> > > numbers. For best results you should be saving the underlying values at
> > least
> > > daily (some save them before and after peak load periods each day) and
> > > clearing
> > > the stats at least weekly. That will allow you to recalculate the metrics
> > over
> > > several time spans and during peak periods.
> > >
> > > Art S. Kagel
> > >
> > > ----- Original Message -----
> > > From: Chris Salch <ids@iiug.org>
> > > At: 8/31 14:42:38
> > >
> > > ( Unfortunatly, none of us thought to save any of that data when things
> > > started slowing down, so all we've got is what we can remember. )
<SNIPPED>
CHRIS SALCH wrote:
> We have an HP9000/L2000 with 4 64bit PA-RISC chips running at 440Mhz and
> 5 Gigs of ram. We're running Inforimx 9.40.HC5 on HPUX 11i. We ran
> into an issue on Monday where our database engine seemed to grind to a
> halt. The machine itself responded reasonably to everything but
> database queries. It seemed like anything having to do with extracting
> or querying data in the database took exceptionally long to respond.
> Interacting with onstat did not. Unfortunately, no one thought to save
> off a copy of what onstat dumped during the incident.
>
> Our application is a touch heterogeneous, we have everything from cognos
> impromptu, MS Access "applications", a massive 4gl based application
> using shm to connect, to perl dbi based connections running off the same
> database. The majority of users probably connect using shm connections
> through the 4gl application but, MS Access makes an excessive number of
> connections per user of a given application. Everything but the MS
> Access and Congos Impromptu reports/applications are running from the
> same machine as the engine. This machine also has an apache web server
> to handle some cgi scripts, mostly perl code and the source of most of
> our perl dbi connections.
>
> When our problem occured, we had a reasonably heavy load from all
> around. The machine had approximatly 560 or so proccess running total,
> some 300-400 connections appearing in onstat at any one instant, and
> three oninit processes eating up 90% of the cpu time solid for about an
> hour. Simple queries would take 2 to 3 minutes to complete. There
> were no apparent harware problems with the equipment, no excessive
> numbers of locks or full logical logs, no single process or small group
> of processes that looked like they were doing anything out of the
> ordinary.
>
> Fortunately, everything seemed to calm down again about an hour latter.
> As a note, that day represented an exceptionally high load in comparison
> to our normal operations and this could all be related to having maxed
> out what our hardware could handle.
>
> Is there any tuning that might be suggested for our system? I've
> attached a copy of our config file. ( I've sense increased the frequency of
my
> logging cron job )
>
> CONFIG FILE:
>
> #**************************************************************************
> #
> # IBM Corporation
> #
> # Title: onconfig.std
> # Description: IBM Informix Dynamic Server Configuration Parameters
> #
> #**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME root # Root dbspace name> ROOTPATH /opt/informix/dev/root.1 # Path for device containing root
> dbspace
> ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
> ROOTSIZE 1024000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH /opt/informix/dev/root.1-m # Path for device containing
> mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS root # Location (dbspace) of physical log
> PHYSFILE 50000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 60 # Number of logical log files
> LOGSIZE 10000 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /opt/informix/Logs/cars.log # System message log file path
> CONSOLE /dev/console # System console message path>
> # To automatically backup logical logs, edit alarmprogram.sh and set
> # BACKUPLOGS=Y
> ALARMPROGRAM /opt/informix/etc/log_full.sh # Alarm program path
> TBLSPACE_STATS 1 # Maintain tblspace statistics>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt/1m # Tape device path
> TAPEBLK 32 # Tape block size (Kbytes)
> TAPESIZE 0 # Maximum amount of data to put on tape (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /dev/rmt/0m # Log tape device path
> LTAPEBLK 32 # Log tape block size (Kbytes)
> LTAPESIZE 0 # Max amount of data to put on log tape (Kbytes)>
> # Optical
>
> STAGEBLOB # Informix Dynamic Server staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a OnLine instance
> DBSERVERNAME paul # Name of default database server
> DBSERVERALIASES carsitcp # List of alternate dbservernames
> NETTYPE ipcshm,2,500,CPU # Configure poll thread(s) for nettype
> NETTYPE soctcp,1,100,NET # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 3 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 256000 # Maximum number of locks
> BUFFERS 90000 # Maximum number of shared buffers
> NUMAIOVPS 8 # Number of IO vps
> PHYSBUFF 32 # Physical log buffer size (Kbytes)
> LOGBUFF 32 # Logical log buffer size (Kbytes)> LOGSMAX 120 # Maximum number of logical log files
> CLEANERS 8 # Number of buffer cleaner processes
> SHMBASE 0x0 # Shared memory base address
> SHMVIRTSIZE 196608 # initial virtual shared memory segment size
> SHMADD 16384 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
> 0=>unlimited
> CKPTINTVL 900 # Check point interval (in sec)
> LRUS 8 # Number of LRU queues
> LRU_MAX_DIRTY 4 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 64 # Stack size (Kbytes)>
> # Dynamic Logging
> # DYNAMIC_LOGS:
> # 2 : server automatically add a new logical log when necessary. (ON)
> # 1 : notify DBA to add new logical logs when necessary. (ON)
> # 0 : cannot add logical log on the fly. (OFF)
> #
> # When dynamic logging is on, we can have higher values for
> LTXHWM/LTXEHWM,
> # because the server can add new logical logs during long transaction
> rollback.
> # However, to limit the number of new logical logs being added,
> LTXHWM/LTXEHWM
> # can be set to smaller values.
> #
> # If dynamic logging is off, LTXHWM/LTXEHWM need to be set to smaller
> values
> # to avoid long transaction rollback hanging the server due to lack of
> logical
> # log space, i.e. 50/60 or lower.
>
> DYNAMIC_LOGS 0
> LTXHWM 45
> LTXEHWM 55>
> # System Page Size
> # BUFFSIZE - OnLine no longer supports this configuration parameter.
> # To determine the page size used by OnLine on your platform
> # see the last line of output from the command, 'onstat -b'.
>
> # Recovery Variables
> # OFF_RECVRY_THREADS:
> # Number of parallel worker threads during fast recovery or an offline
> restore.
> # ON_RECVRY_THREADS:
> # Number of parallel worker threads during an online restore.
>
> OFF_RECVRY_THREADS 10 # Default number of offl
I will submit that there is no good reason to increase the size of the logical
log buffer or the physical log buffer. It just puts more data at risk in a
BUFFERED LOG environment and uses up logical log space faster in an UNBUFFERED
LOG environment for little or no performance improvement.
Art S. Kagel
----- Original Message -----
From: Keith Simmons <ids@iiug.org>
At: 9/01 5:54:13
Chris
90000 buffers on a 2K page size O/S means 180 Mb of buffer space! With
5Gb memory, even with what else is happening on the machine, I would
try 350000 buffers taking up 700 Mb. You will need to increase LRUs
(try 127) and CLEANERS (127 to match LRUs). Also watch checkpoint
times and be prepared to decrease the interval or LRU_MAX and LRU_MIN.
I also noticed your LOG_BUFF and PHYS_BUFF are still at default
values. There may be some milage in increasing these (even as far as
128 or 256) which can help as well.
Keith
On 31/08/06, Chris Salch <chrissalch@letu.edu> wrote:
>
> And the ratios script shows:
>
> pgsused/(ixda-RA+idx-RA+da-RA)*100
> Read-Ahead Util (RAU): 99.90% 23922139/(6713181+244596+16988224)*100
>
> bufwaits*100 / (pagreads+bufwrits)
> Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
> Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=90000)
> Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
> cached)))/(BUFFERS)
>
> BT Period-max: 19.90/hr every 3.01 minutes.
>
> BT Period-min: 3.11/hr every 19.29 minutes.
>
> Does this look as off as I think it does?
>
> On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> > Those stats are reset every sunday at 1:05 am. So that BTR is a
> > significantly higher number than what you have shown. That would be a
> > BTR of about 12 ouch! (by your info)
> >
> > On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
> > wrote:
> > > OK, I've calculated the metrics from the onstat -p output below. It's
not
> > > particularly useful unless you've zerod the stats more recently than
> server
> > > startup 13 days ago, but here they are:
> > >
> > > Pagreads: 7908428
> > > Bufwrits: 35477827
> > > Bufwaits: 1443280
> > > BUFFERS: 90000
> > > Time (hours) since reset: 321.33
> > > ixda-RA: 2346196
> > > idx-RA: 86301
> > > da-RA: 4646343
> > > RA-pgsused: 7071512
> > >
> > > BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
> > > BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
> > > RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
> > >
> > > These look fine, but a shorter accumulation period would give more
> reliable
> > > numbers. For best results you should be saving the underlying values at
> > least
> > > daily (some save them before and after peak load periods each day) and
> > > clearing
> > > the stats at least weekly. That will allow you to recalculate the
metrics
> > over
> > > several time spans and during peak periods.
> > >
> > > Art S. Kagel
> > >
> > > ----- Original Message -----
> > > From: Chris Salch <ids@iiug.org>
> > > At: 8/31 14:42:38
> > >
> > > ( Unfortunatly, none of us thought to save any of that data when things
> > > started slowing down, so all we've got is what we can remember. )
<SNIPPED>
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Although perhaps so critical nowadays,with Sans and large buffers
external to the engine, I think there is still some milage in bigger
log buffers, particularly on direct connect disks. The recommendation
comes straight from a set of Informix Perormance Tuning Course Notes
(with such worthys as John Mille and Mark Scranton on the authors
list!). It suggests (particularly for OLTP) that to minimise I/O the
buffers should be about 75% full when they are flushed. (any more
produces excessive I/O, any less potentially wastes memory. Monitor
using onstat -l |head -18. May not give much improvement (and may be
offset by disadvantages) but is worth consideration
Keith
On 01/09/06, ART KAGEL, BLOOMBERG/ 731 LEXIN <kagel@bloomberg.net> wrote:
>
> I will submit that there is no good reason to increase the size of the
logical
> log buffer or the physical log buffer. It just puts more data at risk in a
> BUFFERED LOG environment and uses up logical log space faster in an
UNBUFFERED
> LOG environment for little or no performance improvement.
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Keith Simmons <ids@iiug.org>
> At: 9/01 5:54:13
>
> Chris
>
> 90000 buffers on a 2K page size O/S means 180 Mb of buffer space! With
> 5Gb memory, even with what else is happening on the machine, I would
> try 350000 buffers taking up 700 Mb. You will need to increase LRUs
> (try 127) and CLEANERS (127 to match LRUs). Also watch checkpoint
> times and be prepared to decrease the interval or LRU_MAX and LRU_MIN.
> I also noticed your LOG_BUFF and PHYS_BUFF are still at default
> values. There may be some milage in increasing these (even as far as
> 128 or 256) which can help as well.
>
> Keith
>
> On 31/08/06, Chris Salch <chrissalch@letu.edu> wrote:
> >
> > And the ratios script shows:
> >
> > pgsused/(ixda-RA+idx-RA+da-RA)*100
> > Read-Ahead Util (RAU): 99.90% 23922139/(6713181+244596+16988224)*100
> >
> > bufwaits*100 / (pagreads+bufwrits)
> > Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
> > Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=90000)
> > Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
> > cached)))/(BUFFERS)
> >
> > BT Period-max: 19.90/hr every 3.01 minutes.
> >
> > BT Period-min: 3.11/hr every 19.29 minutes.
> >
> > Does this look as off as I think it does?
> >
> > On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> > > Those stats are reset every sunday at 1:05 am. So that BTR is a
> > > significantly higher number than what you have shown. That would be a
> > > BTR of about 12 ouch! (by your info)
> > >
> > > On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
> > > wrote:
> > > > OK, I've calculated the metrics from the onstat -p output below. It's
> not
> > > > particularly useful unless you've zerod the stats more recently than
> > server
> > > > startup 13 days ago, but here they are:
> > > >
> > > > Pagreads: 7908428
> > > > Bufwrits: 35477827
> > > > Bufwaits: 1443280
> > > > BUFFERS: 90000
> > > > Time (hours) since reset: 321.33
> > > > ixda-RA: 2346196
> > > > idx-RA: 86301
> > > > da-RA: 4646343
> > > > RA-pgsused: 7071512
> > > >
> > > > BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
> > > > BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
> > > > RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
> > > >
> > > > These look fine, but a shorter accumulation period would give more
> > reliable
> > > > numbers. For best results you should be saving the underlying values at
> > > least
> > > > daily (some save them before and after peak load periods each day) and
> > > > clearing
> > > > the stats at least weekly. That will allow you to recalculate the
> metrics
> > > over
> > > > several time spans and during peak periods.
> > > >
> > > > Art S. Kagel
> > > >
> > > > ----- Original Message -----
> > > > From: Chris Salch <ids@iiug.org>
> > > > At: 8/31 14:42:38
> > > >
> > > > ( Unfortunatly, none of us thought to save any of that data when things
> > > > started slowing down, so all we've got is what we can remember. )
> <SNIPPED>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
Here's my reasoning:
If you are using BUFFERED LOG databases then logical log buffers are not
flushed
to the logs on disk until they fill. For OLTP typical transaction size tends
to be small, a few KB is normal. With the default logical log buffer size (32K)
that means that on average 4-16 committed transactions are at risk for
undetacted rollback after committing successfully if the server crashes before
the log buffer fills and flushes. If you increase the buffer size to 256 K that
increases the number of transactions at risk to from 64-128. Note when you
increase the size of the buffer 8-fold you also increase the time needed to
flush that buffer to disk by 8 times. All unacceptable risk.
If you are using UNBUFFERED LOG databases (highly recommended for OLTP) then
the
logical log buffer is flushed as soon as a commit/rollback record is written to
it or it fills whichever comes first. In this case from 25-75% of the buffer
is unused during normal processing on an OLTP system with default 32K logical
log buffers. If you increase the logical log buffer size to 256K then on
average from 75-98.5% of the logical log buffer space is unused. This is in
constrast to the recommendation you quote to try to use >75% of the buffers
before flushing. Data risk is not significantly affected here however, but the
expanded buffers accomplish nothing.
Physical log buffering is another matter, though I didn't address it separately
before. Here you can make a case since the physical log is not strictly
neccessary for recovery or data integrity. There are only a small number of
scenarios under which the restoration of physical log pages after a crash
actually improves the data integrity. OK, for the physical log buffers I will
admit a case can be made that increasing the buffer size, based on observation
of the onstat -l header section, is reasonable.
Art S. Kagel
----- Original Message -----
From: Keith Simmons <ids@iiug.org>
At: 9/01 10:20:59
Although perhaps so critical nowadays,with Sans and large buffers
external to the engine, I think there is still some milage in bigger
log buffers, particularly on direct connect disks. The recommendation
comes straight from a set of Informix Perormance Tuning Course Notes
(with such worthys as John Mille and Mark Scranton on the authors
list!). It suggests (particularly for OLTP) that to minimise I/O the
buffers should be about 75% full when they are flushed. (any more
produces excessive I/O, any less potentially wastes memory. Monitor
using onstat -l |head -18. May not give much improvement (and may be
offset by disadvantages) but is worth consideration
Keith
On 01/09/06, ART KAGEL, BLOOMBERG/ 731 LEXIN <kagel@bloomberg.net> wrote:
>
> I will submit that there is no good reason to increase the size of the
logical
> log buffer or the physical log buffer. It just puts more data at risk in a
> BUFFERED LOG environment and uses up logical log space faster in an
UNBUFFERED
> LOG environment for little or no performance improvement.
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Keith Simmons <ids@iiug.org>
> At: 9/01 5:54:13
>
> Chris
>
> 90000 buffers on a 2K page size O/S means 180 Mb of buffer space! With
> 5Gb memory, even with what else is happening on the machine, I would
> try 350000 buffers taking up 700 Mb. You will need to increase LRUs
> (try 127) and CLEANERS (127 to match LRUs). Also watch checkpoint
> times and be prepared to decrease the interval or LRU_MAX and LRU_MIN.
> I also noticed your LOG_BUFF and PHYS_BUFF are still at default
> values. There may be some milage in increasing these (even as far as
> 128 or 256) which can help as well.
>
> Keith
>
> On 31/08/06, Chris Salch <chrissalch@letu.edu> wrote:
> >
> > And the ratios script shows:
> >
> > pgsused/(ixda-RA+idx-RA+da-RA)*100
> > Read-Ahead Util (RAU): 99.90% 23922139/(6713181+244596+16988224)*100
> >
> > bufwaits*100 / (pagreads+bufwrits)
> > Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
> > Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=90000)
> > Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
> > cached)))/(BUFFERS)
> >
> > BT Period-max: 19.90/hr every 3.01 minutes.
> >
> > BT Period-min: 3.11/hr every 19.29 minutes.
> >
> > Does this look as off as I think it does?
> >
> > On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> > > Those stats are reset every sunday at 1:05 am. So that BTR is a
> > > significantly higher number than what you have shown. That would be a
> > > BTR of about 12 ouch! (by your info)
> > >
> > > On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
> > > wrote:
> > > > OK, I've calculated the metrics from the onstat -p output below. It's
> not
> > > > particularly useful unless you've zerod the stats more recently than
> > server
> > > > startup 13 days ago, but here they are:
> > > >
> > > > Pagreads: 7908428
> > > > Bufwrits: 35477827
> > > > Bufwaits: 1443280
> > > > BUFFERS: 90000
> > > > Time (hours) since reset: 321.33
> > > > ixda-RA: 2346196
> > > > idx-RA: 86301
> > > > da-RA: 4646343
> > > > RA-pgsused: 7071512
> > > >
> > > > BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
> > > > BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
> > > > RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
> > > >
> > > > These look fine, but a shorter accumulation period would give more
> > reliable
> > > > numbers. For best results you should be saving the underlying values
at
> > > least
> > > > daily (some save them before and after peak load periods each day) and
> > > > clearing
> > > > the stats at least weekly. That will allow you to recalculate the
> metrics
> > > over
> > > > several time spans and during peak periods.
> > > >
> > > > Art S. Kagel
> > > >
> > > > ----- Original Message -----
> > > > From: Chris Salch <ids@iiug.org>
> > > > At: 8/31 14:42:38
> > > >
> > > > ( Unfortunatly, none of us thought to save any of that data when
things
> > > > started slowing down, so all we've got is what we can remember. )
> <SNIPPED>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
As a side note to that, IDS 10 seems to require that PHYS log buffer
and LOGICAL log buffer be the same size.
And if you don't have them the same, it changes them for you!
Perhaps JMiller could address that.
Bob Roussey
Unix / Informix Administration
Spirit Airlines
Robert.Roussey@SpiritAir.com
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
ART KAGEL, BLOOMBERG/ 731 LEXIN
Sent: Friday, September 01, 2006 10:35 AM
To: ids@iiug.org
Subject: Re: A question of load . . . [7399]
Here's my reasoning:
If you are using BUFFERED LOG databases then logical log buffers are not
flushed
to the logs on disk until they fill. For OLTP typical transaction size
tends
to be small, a few KB is normal. With the default logical log buffer
size
(32K)
that means that on average 4-16 committed transactions are at risk for
undetacted rollback after committing successfully if the server crashes
before
the log buffer fills and flushes. If you increase the buffer size to 256
K
that
increases the number of transactions at risk to from 64-128. Note when
you
increase the size of the buffer 8-fold you also increase the time needed
to
flush that buffer to disk by 8 times. All unacceptable risk.
If you are using UNBUFFERED LOG databases (highly recommended for OLTP)
then
the
logical log buffer is flushed as soon as a commit/rollback record is
written
to
it or it fills whichever comes first. In this case from 25-75% of the
buffer
is unused during normal processing on an OLTP system with default 32K
logical
log buffers. If you increase the logical log buffer size to 256K then on
average from 75-98.5% of the logical log buffer space is unused. This is
in
constrast to the recommendation you quote to try to use >75% of the
buffers
before flushing. Data risk is not significantly affected here however,
but the
expanded buffers accomplish nothing.
Physical log buffering is another matter, though I didn't address it
separately
before. Here you can make a case since the physical log is not strictly
neccessary for recovery or data integrity. There are only a small number
of
scenarios under which the restoration of physical log pages after a
crash
actually improves the data integrity. OK, for the physical log buffers I
will
admit a case can be made that increasing the buffer size, based on
observation
of the onstat -l header section, is reasonable.
Art S. Kagel
----- Original Message -----
From: Keith Simmons <ids@iiug.org>
At: 9/01 10:20:59
Although perhaps so critical nowadays,with Sans and large buffers
external to the engine, I think there is still some milage in bigger
log buffers, particularly on direct connect disks. The recommendation
comes straight from a set of Informix Perormance Tuning Course Notes
(with such worthys as John Mille and Mark Scranton on the authors
list!). It suggests (particularly for OLTP) that to minimise I/O the
buffers should be about 75% full when they are flushed. (any more
produces excessive I/O, any less potentially wastes memory. Monitor
using onstat -l |head -18. May not give much improvement (and may be
offset by disadvantages) but is worth consideration
Keith
On 01/09/06, ART KAGEL, BLOOMBERG/ 731 LEXIN <kagel@bloomberg.net>
wrote:
>
> I will submit that there is no good reason to increase the size of the
logical
> log buffer or the physical log buffer. It just puts more data at risk
in a
> BUFFERED LOG environment and uses up logical log space faster in an
UNBUFFERED
> LOG environment for little or no performance improvement.
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Keith Simmons <ids@iiug.org>
> At: 9/01 5:54:13
>
> Chris
>
> 90000 buffers on a 2K page size O/S means 180 Mb of buffer space! With
> 5Gb memory, even with what else is happening on the machine, I would
> try 350000 buffers taking up 700 Mb. You will need to increase LRUs
> (try 127) and CLEANERS (127 to match LRUs). Also watch checkpoint
> times and be prepared to decrease the interval or LRU_MAX and LRU_MIN.
> I also noticed your LOG_BUFF and PHYS_BUFF are still at default
> values. There may be some milage in increasing these (even as far as
> 128 or 256) which can help as well.
>
> Keith
>
> On 31/08/06, Chris Salch <chrissalch@letu.edu> wrote:
> >
> > And the ratios script shows:
> >
> > pgsused/(ixda-RA+idx-RA+da-RA)*100
> > Read-Ahead Util (RAU): 99.90% 23922139/(6713181+244596+16988224)*100
> >
> > bufwaits*100 / (pagreads+bufwrits)
> > Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
> > Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=90000)
> > Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
> > cached)))/(BUFFERS)
> >
> > BT Period-max: 19.90/hr every 3.01 minutes.
> >
> > BT Period-min: 3.11/hr every 19.29 minutes.
> >
> > Does this look as off as I think it does?
> >
> > On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> > > Those stats are reset every sunday at 1:05 am. So that BTR is a
> > > significantly higher number than what you have shown. That would
be a
> > > BTR of about 12 ouch! (by your info)
> > >
> > > On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
> > > wrote:
> > > > OK, I've calculated the metrics from the onstat -p output below.
It's
> not
> > > > particularly useful unless you've zerod the stats more recently
than
> > server
> > > > startup 13 days ago, but here they are:
> > > >
> > > > Pagreads: 7908428
> > > > Bufwrits: 35477827
> > > > Bufwaits: 1443280
> > > > BUFFERS: 90000
> > > > Time (hours) since reset: 321.33
> > > > ixda-RA: 2346196
> > > > idx-RA: 86301
> > > > da-RA: 4646343
> > > > RA-pgsused: 7071512
> > > >
> > > > BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
> > > > BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
> > > > RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
> > > >
> > > > These look fine, but a shorter accumulation period would give
more
> > reliable
> > > > numbers. For best results you should be saving the underlying
values
at
> > > least
> > > > daily (some save them before and after peak load periods each
day) and
> > > > clearing
> > > > the stats at least weekly. That will allow you to recalculate
the
> metrics
> > > over
> > > > several time spans and during peak periods.
> > > >
> > > > Art S. Kagel
> > > >
> > > > ----- Original Message -----
> > > > From: Chris Salch <ids@iiug.org>
> > > > At: 8/31 14:42:38
> > > >
> > > > ( Unfortunatly, none of us thought to save any of that data when
things
> > > > started slowing down, so all we've got is what we can remember.
)
> <SNIPPED>
>
>
>
************************************************************************
*******
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
******************************************************
Sounds like an undocumented feature :-)
Paul Watson
Tel: +44 1414161772
Mob: +44 7818003457
Web: www.oninit.com
GO FURTHER with DB2
GET THERE FASTER with Informix.
Attend the IDUG 2006 European Conference.
Vienna, Austria. 2-6 October 2006
Visit http://www.iiug.org/conf for more information.
> -----Original Message-----
> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On
> Behalf Of Robert Roussey(MIS)
> Sent: 01 September 2006 09:45
> To: ids@iiug.org
> Subject: RE: A question of load . . . [7401]
>
>
> As a side note to that, IDS 10 seems to require that PHYS log
> buffer and LOGICAL log buffer be the same size.
> And if you don't have them the same, it changes them for you!
> Perhaps JMiller could address that.
>
> Bob Roussey
> Unix / Informix Administration
> Spirit Airlines
> Robert.Roussey@SpiritAir.com
>
> -----Original Message-----
> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On
> Behalf Of ART KAGEL, BLOOMBERG/ 731 LEXIN
> Sent: Friday, September 01, 2006 10:35 AM
> To: ids@iiug.org
> Subject: Re: A question of load . . . [7399]
>
> Here's my reasoning:
>
> If you are using BUFFERED LOG databases then logical log
> buffers are not
>
> flushed
> to the logs on disk until they fill. For OLTP typical
> transaction size tends to be small, a few KB is normal. With
> the default logical log buffer size
> (32K)
> that means that on average 4-16 committed transactions are at
> risk for undetacted rollback after committing successfully if
> the server crashes before the log buffer fills and flushes.
> If you increase the buffer size to 256 K that increases the
> number of transactions at risk to from 64-128. Note when you
> increase the size of the buffer 8-fold you also increase the
> time needed to flush that buffer to disk by 8 times. All
> unacceptable risk.
>
> If you are using UNBUFFERED LOG databases (highly recommended
> for OLTP) then the logical log buffer is flushed as soon as a
> commit/rollback record is written to it or it fills whichever
> comes first. In this case from 25-75% of the buffer is unused
> during normal processing on an OLTP system with default 32K
> logical log buffers. If you increase the logical log buffer
> size to 256K then on
>
> average from 75-98.5% of the logical log buffer space is
> unused. This is in constrast to the recommendation you quote
> to try to use >75% of the buffers before flushing. Data risk
> is not significantly affected here however, but the expanded
> buffers accomplish nothing.
>
> Physical log buffering is another matter, though I didn't
> address it separately before. Here you can make a case since
> the physical log is not strictly neccessary for recovery or
> data integrity. There are only a small number of scenarios
> under which the restoration of physical log pages after a
> crash actually improves the data integrity. OK, for the
> physical log buffers I will admit a case can be made that
> increasing the buffer size, based on observation of the
> onstat -l header section, is reasonable.>
> Art S. Kagel
>
> ----- Original Message -----
> From: Keith Simmons <ids@iiug.org>
> At: 9/01 10:20:59
>
> Although perhaps so critical nowadays,with Sans and large
> buffers external to the engine, I think there is still some
> milage in bigger log buffers, particularly on direct connect
> disks. The recommendation comes straight from a set of
> Informix Perormance Tuning Course Notes (with such worthys as
> John Mille and Mark Scranton on the authors list!). It
> suggests (particularly for OLTP) that to minimise I/O the
> buffers should be about 75% full when they are flushed. (any
> more produces excessive I/O, any less potentially wastes
> memory. Monitor using onstat -l |head -18. May not give much
> improvement (and may be offset by disadvantages) but is worth
> consideration
>
> Keith
>
> On 01/09/06, ART KAGEL, BLOOMBERG/ 731 LEXIN <kagel@bloomberg.net>
> wrote:
> >
> > I will submit that there is no good reason to increase the
> size of the
>
> logical
> > log buffer or the physical log buffer. It just puts more
> data at risk
> in a
> > BUFFERED LOG environment and uses up logical log space faster in an
> UNBUFFERED
> > LOG environment for little or no performance improvement.
> >
> > Art S. Kagel
> >
> > ----- Original Message -----
> > From: Keith Simmons <ids@iiug.org>
> > At: 9/01 5:54:13
> >
> > Chris
> >
> > 90000 buffers on a 2K page size O/S means 180 Mb of buffer
> space! With
>
> > 5Gb memory, even with what else is happening on the
> machine, I would
> > try 350000 buffers taking up 700 Mb. You will need to increase LRUs
> > (try 127) and CLEANERS (127 to match LRUs). Also watch checkpoint
> > times and be prepared to decrease the interval or LRU_MAX
> and LRU_MIN.
>
> > I also noticed your LOG_BUFF and PHYS_BUFF are still at default
> > values. There may be some milage in increasing these (even as far as
> > 128 or 256) which can help as well.
> >
> > Keith
> >
> > On 31/08/06, Chris Salch <chrissalch@letu.edu> wrote:
> > >
> > > And the ratios script shows:
> > >
> > > pgsused/(ixda-RA+idx-RA+da-RA)*100
> > > Read-Ahead Util (RAU): 99.90%
> 23922139/(6713181+244596+16988224)*100
>
> > >
> > > bufwaits*100 / (pagreads+bufwrits)
> > > Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
> > > Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=90000)
> > > Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
> > > cached)))/(BUFFERS)
> > >
> > > BT Period-max: 19.90/hr every 3.01 minutes.
> > >
> > > BT Period-min: 3.11/hr every 19.29 minutes.
> > >
> > > Does this look as off as I think it does?
> > >
> > > On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> > > > Those stats are reset every sunday at 1:05 am. So that BTR is a
> > > > significantly higher number than what you have shown. That would
> be a
> > > > BTR of about 12 ouch! (by your info)
> > > >
> > > > On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL,
> BLOOMBERG/ 731 LEXIN
>
> > > > wrote:
> > > > > OK, I've calculated the metrics from the onstat -p
> output below.
> It's
> > not
> > > > > particularly useful unless you've zerod the stats
> more recently
> than
> > > server
> > > > > startup 13 days ago, but here they are:
> > > > >
> > > > > Pagreads: 7908428
> > > > > Bufwrits: 35477827
> > > > > Bufwaits: 1443280
> > > > > BUFFERS: 90000
> > > > > Time (hours) since reset: 321.33
> > > > > ixda-RA: 2346196
> > > > > idx-RA: 86301
> > > > > da-RA: 4646343
> > > > > RA-pgsused: 7071512
> > > > >
> > > > > BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200 BTR =
> > > > > (((35477827 + 7908428) / 90000) / 321.33) = 1.5002 RAU =
> > > > > (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
> > > > >
> > > > > These look fine, but a shorter accumulation period would give
> more
> > > reliable
> > > > > numbers. For best results you should be saving the underlying
> values
> at
>
Most definitely thank you all. Now I've got to see what our application
vendor, thinks about this little tidbit and see where to go from there.
I'll post what happens as soon as I know :)
On Fri, 2006-09-01 at 10:35 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
wrote:
> Here's my reasoning:
>
> If you are using BUFFERED LOG databases then logical log buffers are not
> flushed
> to the logs on disk until they fill. For OLTP typical transaction size tends
> to be small, a few KB is normal. With the default logical log buffer size
> (32K)
> that means that on average 4-16 committed transactions are at risk for
> undetacted rollback after committing successfully if the server crashes
before
> the log buffer fills and flushes. If you increase the buffer size to 256 K
> that
> increases the number of transactions at risk to from 64-128. Note when you
> increase the size of the buffer 8-fold you also increase the time needed to
> flush that buffer to disk by 8 times. All unacceptable risk.
>
> If you are using UNBUFFERED LOG databases (highly recommended for OLTP) then
> the
> logical log buffer is flushed as soon as a commit/rollback record is written
> to
> it or it fills whichever comes first. In this case from 25-75% of the buffer
> is unused during normal processing on an OLTP system with default 32K logical
> log buffers. If you increase the logical log buffer size to 256K then on
> average from 75-98.5% of the logical log buffer space is unused. This is in
> constrast to the recommendation you quote to try to use >75% of the buffers
> before flushing. Data risk is not significantly affected here however, but
the
> expanded buffers accomplish nothing.
>
> Physical log buffering is another matter, though I didn't address it
> separately
> before. Here you can make a case since the physical log is not strictly
> neccessary for recovery or data integrity. There are only a small number of
> scenarios under which the restoration of physical log pages after a crash
> actually improves the data integrity. OK, for the physical log buffers I will
> admit a case can be made that increasing the buffer size, based on
observation
> of the onstat -l header section, is reasonable.
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Keith Simmons <ids@iiug.org>
> At: 9/01 10:20:59
>
> Although perhaps so critical nowadays,with Sans and large buffers
> external to the engine, I think there is still some milage in bigger
> log buffers, particularly on direct connect disks. The recommendation
> comes straight from a set of Informix Perormance Tuning Course Notes
> (with such worthys as John Mille and Mark Scranton on the authors
> list!). It suggests (particularly for OLTP) that to minimise I/O the
> buffers should be about 75% full when they are flushed. (any more
> produces excessive I/O, any less potentially wastes memory. Monitor
> using onstat -l |head -18. May not give much improvement (and may be
> offset by disadvantages) but is worth consideration
>
> Keith
>
> On 01/09/06, ART KAGEL, BLOOMBERG/ 731 LEXIN <kagel@bloomberg.net> wrote:
> >
> > I will submit that there is no good reason to increase the size of the
> logical
> > log buffer or the physical log buffer. It just puts more data at risk in a
> > BUFFERED LOG environment and uses up logical log space faster in an
> UNBUFFERED
> > LOG environment for little or no performance improvement.
> >
> > Art S. Kagel
> >
> > ----- Original Message -----
> > From: Keith Simmons <ids@iiug.org>
> > At: 9/01 5:54:13
> >
> > Chris
> >
> > 90000 buffers on a 2K page size O/S means 180 Mb of buffer space! With
> > 5Gb memory, even with what else is happening on the machine, I would
> > try 350000 buffers taking up 700 Mb. You will need to increase LRUs
> > (try 127) and CLEANERS (127 to match LRUs). Also watch checkpoint
> > times and be prepared to decrease the interval or LRU_MAX and LRU_MIN.
> > I also noticed your LOG_BUFF and PHYS_BUFF are still at default
> > values. There may be some milage in increasing these (even as far as
> > 128 or 256) which can help as well.
> >
> > Keith
> >
> > On 31/08/06, Chris Salch <chrissalch@letu.edu> wrote:
> > >
> > > And the ratios script shows:
> > >
> > > pgsused/(ixda-RA+idx-RA+da-RA)*100
> > > Read-Ahead Util (RAU): 99.90% 23922139/(6713181+244596+16988224)*100
> > >
> > > bufwaits*100 / (pagreads+bufwrits)
> > > Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
> > > Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=90000)
> > > Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
> > > cached)))/(BUFFERS)
> > >
> > > BT Period-max: 19.90/hr every 3.01 minutes.
> > >
> > > BT Period-min: 3.11/hr every 19.29 minutes.
> > >
> > > Does this look as off as I think it does?
> > >
> > > On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> > > > Those stats are reset every sunday at 1:05 am. So that BTR is a
> > > > significantly higher number than what you have shown. That would be a
> > > > BTR of about 12 ouch! (by your info)
> > > >
> > > > On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXIN
> > > > wrote:
> > > > > OK, I've calculated the metrics from the onstat -p output below. It's
> > not
> > > > > particularly useful unless you've zerod the stats more recently than
> > > server
> > > > > startup 13 days ago, but here they are:
> > > > >
> > > > > Pagreads: 7908428
> > > > > Bufwrits: 35477827
> > > > > Bufwaits: 1443280
> > > > > BUFFERS: 90000
> > > > > Time (hours) since reset: 321.33
> > > > > ixda-RA: 2346196
> > > > > idx-RA: 86301
> > > > > da-RA: 4646343
> > > > > RA-pgsused: 7071512
> > > > >
> > > > > BR = (1443280 / (35477827 + 7908428)) * 100.00 = 3.3200
> > > > > BTR = (((35477827 + 7908428) / 90000) / 321.33) = 1.5002
> > > > > RAU = (7071512/(2346196+86301+4646343)) * 100.00 = 99.8900
> > > > >
> > > > > These look fine, but a shorter accumulation period would give more
> > > reliable
> > > > > numbers. For best results you should be saving the underlying values
> at
> > > > least
> > > > > daily (some save them before and after peak load periods each day)
and
> > > > > clearing
> > > > > the stats at least weekly. That will allow you to recalculate the
> > metrics
> > > > over
> > > > > several time spans and during peak periods.
> > > > >
> > > > > Art S. Kagel
> > > > >
> > > > > ----- Original Message -----
> > > > > From: Chris Salch <ids@iiug.org>
> > > > > At: 8/31 14:42:38
> > > > >
> > > > > ( Unfortunatly, none of us thought to save any of that data when
> things
> > > > > started slowing down, so all we've got is what we can remember. )
> > <SNIPPED>
> >
> >
> >
>
>
*******************************************************************************
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
> >
>
>
*******************************************************************************
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
>
>
************
I would like to add a comment to Art's response. The logging attribute=
to
be buffered or unbuffered logging is not controlled by the create datab=
ase
statement, but the create database statement sets the default mode for =
the
database. Each user has control over the logging attributed buffered
or non-buffered. So for people who are running in a non-buffered loggi=
ng
mode but have batch or restartable programs changing the logging
attribute for just their session to buffered logging can be done.
I was unable to reproduce the problem mention below. I tried starting=
my
version 10 server with a physical log of 96 and my logical log buffer a=
t 32
and it started just fine.
Physical Logging
Buffer bufused bufsize numpages numwrits pages/io
P-1 0 48 0 0 0.00
phybegin physize phypos phyused %used
1:263 3500 697 0 0.00
Logical Logging
Buffer bufused bufsize numrecs numpages numwrits recs/pages
pages/io
L-2 0 16 0 0 0 0.0 0.=
0
John
=
"Robert =
Roussey\\\\(MIS\\\\)" =
<Robert.Roussey@S =
To
piritAir.com> ids@iiug.org =
Sent by: =
cc
ids-bounces@iiug. =
org Subj=
ect
RE: A question of load . . . [74=
01]
=
09/01/2006 07:45 =
AM =
=
=
Please respond to =
ids@iiug.org =
=
=
As a side note to that, IDS 10 seems to require that PHYS log buffer
and LOGICAL log buffer be the same size.
And if you don't have them the same, it changes them for you!
Perhaps JMiller could address that.
Bob Roussey
Unix / Informix Administration
Spirit Airlines
Robert.Roussey@SpiritAir.com
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
ART KAGEL, BLOOMBERG/ 731 LEXIN
Sent: Friday, September 01, 2006 10:35 AM
To: ids@iiug.org
Subject: Re: A question of load . . . [7399]
Here's my reasoning:
If you are using BUFFERED LOG databases then logical log buffers are no=
t
flushed
to the logs on disk until they fill. For OLTP typical transaction size
tends
to be small, a few KB is normal. With the default logical log buffer
size
(32K)
that means that on average 4-16 committed transactions are at risk for
undetacted rollback after committing successfully if the server crashes=
before
the log buffer fills and flushes. If you increase the buffer size to 25=
6
K
that
increases the number of transactions at risk to from 64-128. Note when
you
increase the size of the buffer 8-fold you also increase the time neede=
d
to
flush that buffer to disk by 8 times. All unacceptable risk.
If you are using UNBUFFERED LOG databases (highly recommended for OLTP)=
then
the
logical log buffer is flushed as soon as a commit/rollback record is
written
to
it or it fills whichever comes first. In this case from 25-75% of the
buffer
is unused during normal processing on an OLTP system with default 32K
logical
log buffers. If you increase the logical log buffer size to 256K then o=
n
average from 75-98.5% of the logical log buffer space is unused. This i=
s
in
constrast to the recommendation you quote to try to use >75% of the
buffers
before flushing. Data risk is not significantly affected here however,
but the
expanded buffers accomplish nothing.
Physical log buffering is another matter, though I didn't address it
separately
before. Here you can make a case since the physical log is not strictly=
neccessary for recovery or data integrity. There are only a small numbe=
r
of
scenarios under which the restoration of physical log pages after a
crash
actually improves the data integrity. OK, for the physical log buffers =
I
will
admit a case can be made that increasing the buffer size, based on
observation
of the onstat -l header section, is reasonable.
Art S. Kagel
----- Original Message -----
From: Keith Simmons <ids@iiug.org>
At: 9/01 10:20:59
Although perhaps so critical nowadays,with Sans and large buffers
external to the engine, I think there is still some milage in bigger
log buffers, particularly on direct connect disks. The recommendation
comes straight from a set of Informix Perormance Tuning Course Notes
(with such worthys as John Mille and Mark Scranton on the authors
list!). It suggests (particularly for OLTP) that to minimise I/O the
buffers should be about 75% full when they are flushed. (any more
produces excessive I/O, any less potentially wastes memory. Monitor
using onstat -l |head -18. May not give much improvement (and may be
offset by disadvantages) but is worth consideration
Keith
On 01/09/06, ART KAGEL, BLOOMBERG/ 731 LEXIN <kagel@bloomberg.net>
wrote:
>
> I will submit that there is no good reason to increase the size of th=
e
logical
> log buffer or the physical log buffer. It just puts more data at risk=
in a
> BUFFERED LOG environment and uses up logical log space faster in an
UNBUFFERED
> LOG environment for little or no performance improvement.
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Keith Simmons <ids@iiug.org>
> At: 9/01 5:54:13
>
> Chris
>
> 90000 buffers on a 2K page size O/S means 180 Mb of buffer space! Wit=
h
> 5Gb memory, even with what else is happening on the machine, I would
> try 350000 buffers taking up 700 Mb. You will need to increase LRUs
> (try 127) and CLEANERS (127 to match LRUs). Also watch checkpoint
> times and be prepared to decrease the interval or LRU_MAX and LRU_MIN=
.
> I also noticed your LOG_BUFF and PHYS_BUFF are still at default
> values. There may be some milage in increasing these (even as far as
> 128 or 256) which can help as well.
>
> Keith
>
> On 31/08/06, Chris Salch <chrissalch@letu.edu> wrote:
> >
> > And the ratios script shows:
> >
> > pgsused/(ixda-RA+idx-RA+da-RA)*100
> > Read-Ahead Util (RAU): 99.90% 23922139/(6713181+244596+16988224)*10=
0
> >
> > bufwaits*100 / (pagreads+bufwrits)
> > Bufwaits Ratio (BWR): 2.39% (4819402*100) / (30380617+171168299)
> > Buffer Turnover(Max):2239.43 (pagreads+bufwrits)/(BUFFERS=3D90000)
> > Buffer Turnover(Min): 349.92 (pagreads+(bufwrits*(1-%
> > cached)))/(BUFFERS)
> >
> > BT Period-max: 19.90/hr every 3.01 minutes.
> >
> > BT Period-min: 3.11/hr every 19.29 minutes.
> >
> > Does this look as off as I think it does?
> >
> > On Thu, 2006-08-31 at 18:22 -0400, Chris Salch wrote:
> > > Those stats are reset every sunday at 1:05 am. So that BTR is a
> > > significantly higher number than what you have shown. That would
be a
> > > BTR of about 12 ouch! (by your info)
> > >
> > > On Thu, 2006-08-31 at 14:57 -0400, ART KAGEL, BLOOMBERG/ 731 LEXI=
N
> > > wrote:
> > > > OK, I've calculated the metrics from the onstat -p output below=
.
It's
> not
> > > > particularly useful unless you've zerod the stats more recently=
than
> > server
> > > > startup 13 days ago, but here
We had similiar problems too. What happen was one 'power user' was running
Cognos Impromptu query, using non-standard catalog to create her own report,
therefore, generating excess number of 'outer join' and suck out all CPU.
Next time happen, you can run 'onstat -g act' to see what is active, 'onstat
-u' to see what session is running that, 'onstat -g ses XXXXX' will show you
the sql. -- You maybe surprised. One Impromptu sql I caught has 8 pages of sql
statement, all under one 'select' statement!
Frank Lai
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix