Poor performance on replacement server
Posted in 2005
Simon Coyne moved an ~80GB IDS 9.3 instance from a 2-CPU DS25 to a 4-CPU Alpha ES45 (Tru64 5.1B) and saw severe degradation at ~55 users: CPU VPs pinned at 100% while disk I/O dropped to near zero, performance either fine or stalled. Replies suggested lock contention/cascading lock waits, OS/hardware config differences, fewer disks per SCSI channel, and ONCONFIG changes (MULTIPROCESSOR=1, more CPUVPS, RESIDENT, NOAGE, fewer AIO VPs with KAIO, smaller PHYSBUFF). Simon had posted the old server's ONCONFIG by mistake, said multiprocessor settings made no difference, and management had already reverted to the DS25, so no onstat diagnostics could be gathered. No resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Backup & Restore, Performance & Tuning, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration
Dear Informix users
The firm I work for have recently purchased a 4 processor ES45 Alpha server to
run our practice management (sales ledger) system. This is replacing a 2
processor DS25 Alpha server (which struggles at period end). Outline
specification is
4x1.2Ghz Alpha processors
8Gb RAM
5304 4 Channel Ultra SCSI controller
2xMA300(I think) disk cages populated with 36Gb disks and configured as RAID
1+0 for both OS and Informix raw chunks.
Tru 64 5.1B
Informix 9.3.FC3
The instance is about 80Gb total.
Unfortunately, we are experiencing crippling performance problems. When we get
to about 55 green screen users on the system, performance within Informix
becomes terrible. The 3 CPU VPs go to 100% and disk I/O drops to almost 0
across every one of the Informix disks. On occassion, when CPU utilisation
dips below ~85%, disk I/O goes back to normal levels and performance is great.
At all times, the disk performance for O/S based operations is more than
acceptable. Informix performance is almost binary; it's either working or
stopped. There doesn't seem to be an inbetween position where performance
starts degrading.
Has anyone seen anything like this at all or have any helpful suggestions?
(Update statistics runs daily.)
Thanks in advance
Simon Coyne
Infrastructure Specialist
DLA Piper Rudnick Gray Cary LLP
ONconfig follows
#**************************************************************************
#
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig.std
# Description: Informix Dynamic Server 2000 Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace name
ROOTPATH /dev/rootdbs # Path for device containing root dbspace
ROOTOFFSET 500 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 2000000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS physdbs # Location (dbspace) of physical log
PHYSFILE 1000000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 1034 # Number of logical log files
LOGSIZE 1500 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /usr/informix/online.log # System message log file path
CONSOLE /usr/informix/console.log # System console message path
ALARMPROGRAM /usr/informix/etc/log_full.sh # Alarm program path
TBLSPACE_STATS 1 # Maintain tblspace statistics
# System Archive Tape Device
TAPEDEV /devices/tape/tape17c # Tape device path
#TAPEDEV /devices/tape/tape16c # FS Backup device
#TAPEDEV /dev/null # Black Hole
#TAPEDEV /nfsdump/aristadbbackup.img
TAPEBLK 1024 # Tape block size (Kbytes)
TAPESIZE 85000000 # Maximum amount of data to put on tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV /devices/tape/tape16
LTAPEBLK 1024 # Log tape block size (Kbytes)
LTAPESIZE 35000000 # Max amount of data to put on log tape (Kbytes)
# Optical
STAGEBLOB # Informix Dynamic Server 2000 staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLine instance
DBSERVERNAME arista # Name of default database server
DBSERVERALIASES arista_tcp # List of alternate dbservernames
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 0 # 0 for single-processor, 1 for multi-processor
NUMCPUVPS 2 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
NOAGE 0 # Process aging# NRW changed proc affinity to second cpu to help system performance
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 1000000 # Maximum number of locks#[was] BUFFERS 300000 # Maximum number of shared buffers
#[was] BUFFERS 400000 # Maximum number of shared buffers
BUFFERS 500000 # Maximum number of shared buffers
NUMAIOVPS 16 # Number of IO vps
#[was] PHYSBUFF 655 # Physical log buffer size (Kbytes)
PHYSBUFF 864 # Physical log buffer size (Kbytes)
#[was] LOGBUFF 32 # Logical log buffer size (Kbytes)
LOGBUFF 38 # Logical log buffer size (Kbytes)LOGSMAX 2048 # Maximum number of logical log files
CLEANERS 16 # Number of buffer cleaner processes
SHMBASE 0x200000000 # Shared memory base address#[was] SHMVIRTSIZE 155648 # initial virtual shared memory segment size
#[was] SHMVIRTSIZE 188416 # initial virtual shared memory segment size
SHMVIRTSIZE 753644 # initial virtual shared memory segment size#[was] SHMADD 16000 # Size of new shared memory segments (Kbytes)
#[was] SHMADD 18841 # Size of new shared memory segments (Kbytes)
SHMADD 75366 # Size of new shared memory segments (Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 128 # Number of LRU queues
LRU_MAX_DIRTY 3 # LRU percent dirty begin cleaning limit# [was] LRU_MAX_DIRTY 10 # LRU percent dirty begin cleaning limit
# [was] LRU_MIN_DIRTY 5 # LRU percent dirty end cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water mark percentage
LTXEHWM 60 # Long transaction high water mark (exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline worker threads
ON_RECVRY_THREADS 1 # Default number of online worker threads
# Data Replication Variables
DRINTERVAL 30 # DR max time between DR buffer flushes (in sec)
DRTIMEOUT 30 # DR network timeout (in sec)
DRLOSTFOUND /usr/informix/etc/dr.lostfound # DR lost+found file path
# CDR Variables
CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-vp,additional)
CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR queue (Kbytes)CDR_LOGDELTA 30 # % of log space allowed in queue memory
CDR_NUMCONNECT 16 # Expected connections per server
CDR_NIFRETRY 300 # Connection retry (seconds)
CDR_NIFCOMPRESS 0 # Link level compression (-1 never, 0 none, 9 max)
# Backup/Restore variables
BAR_ACT_LOG /usr/informix/bar_act.log # ON-Bar Log file - not in /tmp please
BAR_DEBUG_LOG /usr/informix/bar_dbug.log
# ON-Bar Debug Log - not in /tmp please
BAR_MAX_BACKUP 0
BAR_RETRY 1
BAR_NB_XPORT_COUNT 10
BAR_XFER_BUF_SIZE 31
RESTARTABLE_RESTORE off
BAR_PROGRESS_FREQ 0
# Informix Storage Manager variables
ISM_DATA_POOL ISMData
ISM_LOG_POOL ISMLogs
# Read Ahead Variables
RA_PAGES
Could this be users waiting on locks? Occasionally,
after a certain threshold, lock waiting can begin to
cascade and get worse in an apparently step-wise
fashion. Try placing a monitor on the number of locks
to see if that all of a sudden jumps. If it does, it
may indicate that a transaction in the application may
need to be made more efficient.
--John Bejarano.
--- Simon Coyne <simon.coyne@dla.com> wrote:
> Dear Informix users
>
> The firm I work for have recently purchased a 4
> processor ES45 Alpha server to run our practice
> management (sales ledger) system. This is replacing
> a 2 processor DS25 Alpha server (which struggles at
> period end). Outline specification is
>
> 4x1.2Ghz Alpha processors
> 8Gb RAM
> 5304 4 Channel Ultra SCSI controller
> 2xMA300(I think) disk cages populated with 36Gb
> disks and configured as RAID 1+0 for both OS and
> Informix raw chunks.
> Tru 64 5.1B
> Informix 9.3.FC3
>
> The instance is about 80Gb total.
>
> Unfortunately, we are experiencing crippling
> performance problems. When we get to about 55 green
> screen users on the system, performance within
> Informix becomes terrible. The 3 CPU VPs go to 100%
> and disk I/O drops to almost 0 across every one of
> the Informix disks. On occassion, when CPU
> utilisation dips below ~85%, disk I/O goes back to
> normal levels and performance is great. At all
> times, the disk performance for O/S based operations
> is more than acceptable. Informix performance is
> almost binary; it's either working or stopped. There
> doesn't seem to be an inbetween position where
> performance starts degrading.
>
> Has anyone seen anything like this at all or have
> any helpful suggestions? (Update statistics runs
> daily.)
>
> Thanks in advance
>
>
> Simon Coyne
> Infrastructure Specialist
> DLA Piper Rudnick Gray Cary LLP
>
> ONconfig follows
>
>
#**************************************************************************
> #
> # INFORMIX SOFTWARE, INC.
> #
> # Title: onconfig.std
> # Description: Informix Dynamic Server 2000
> Configuration Parameters
> #
>
#**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name
> ROOTPATH /dev/rootdbs # Path for device> containing root dbspace
> ROOTOFFSET 500 # Offset of root> dbspace into device (Kbytes)
> ROOTSIZE 2000000 # Size of root
> dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag
> (Yes = 1, No = 0)
> MIRRORPATH # Path for device> containing mirrored root
> MIRROROFFSET 0 # Offset into> mirrored device (Kbytes)
>
> # Physical Log Configuration
>
> PHYSDBS physdbs # Location (dbspace)
> of physical log
> PHYSFILE 1000000 # Physical log file
> size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 1034 # Number of logical> log files
> LOGSIZE 1500 # Logical log size
> (Kbytes)>
> # Diagnostics
>
> MSGPATH /usr/informix/online.log # System
> message log file path
> CONSOLE /usr/informix/console.log # System
> console message path
> ALARMPROGRAM /usr/informix/etc/log_full.sh #
> Alarm program path
> TBLSPACE_STATS 1 # Maintain tblspace> statistics
>
> # System Archive Tape Device
>
> TAPEDEV /devices/tape/tape17c # Tape device
> path
> #TAPEDEV /devices/tape/tape16c # FS Backup
> device
> #TAPEDEV /dev/null # Black Hole
> #TAPEDEV /nfsdump/aristadbbackup.img
> TAPEBLK 1024 # Tape block size
> (Kbytes)
> TAPESIZE 85000000 # Maximum amount of> data to put on tape (Kbytes)
>
> # Log Archive Tape Device
>
> LTAPEDEV /devices/tape/tape16
> LTAPEBLK 1024 # Log tape block
> size (Kbytes)
> LTAPESIZE 35000000 # Max amount of data> to put on log tape (Kbytes)
>
> # Optical
>
> STAGEBLOB # Informix Dynamic
> Server 2000 staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id> corresponding to a OnLine instance
> DBSERVERNAME arista # Name of default> database server
> DBSERVERALIASES arista_tcp # List of alternate> dbservernames
> DEADLOCK_TIMEOUT 60 # Max time to wait> of lock in distributed env.
> RESIDENT 0 # Forced residency
> flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 0 # 0 for> single-processor, 1 for multi-processor
> NUMCPUVPS 2 # Number of user
> (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit> number of cpu vps to one
>
> NOAGE 0 # Process aging> # NRW changed proc affinity to second cpu to help
> system performance
> AFF_SPROC 0 # Affinity start> processor
> AFF_NPROCS 0 # Affinity number of> processors
>
> # Shared Memory Parameters
>
> LOCKS 1000000 # Maximum number of> locks
> #[was] BUFFERS 300000 # Maximum
> number of shared buffers
> #[was] BUFFERS 400000 # Maximum
> number of shared buffers
> BUFFERS 500000 # Maximum number of> shared buffers
> NUMAIOVPS 16 # Number of IO vps
> #[was] PHYSBUFF 655 # Physical> log buffer size (Kbytes)
> PHYSBUFF 864 # Physical log
> buffer size (Kbytes)
> #[was] LOGBUFF 32 # Logical log
> buffer size (Kbytes)
> LOGBUFF 38 # Logical log buffer
> size (Kbytes)> LOGSMAX 2048 # Maximum number of
> logical log files
> CLEANERS 16 # Number of buffer> cleaner processes
> SHMBASE 0x200000000 # Shared memory> base address
> #[was] SHMVIRTSIZE 155648 # initial
> virtual shared memory segment size
> #[was] SHMVIRTSIZE 188416 # initial
> virtual shared memory segment size
> SHMVIRTSIZE 753644 # initial virtual> shared memory segment size
> #[was] SHMADD 16000 # Size of new
> shared memory segments (Kbytes)
> #[was] SHMADD 18841 # Size of new
> shared memory segments (Kbytes)
> SHMADD 75366 # Size of new shared> memory segments (Kbytes)
> SHMTOTAL 0 # Total shared
> memory (Kbytes). 0=>unlimited
> CKPTINTVL 300 # Check point
> interval (in sec)
> LRUS 128 # Number of LRU> queues
> LRU_MAX_DIRTY 3 # LRU percent dirty> begin cleaning limit
> # [was] LRU_MAX_DIRTY 10 # LRU
> percent dirty begin cleaning limit
>
=== message truncated ===
This
looks more of an OS problem to me.
However, here are some Informix corrections.
1. Increase NUMCPUVPS to 3 (you have it at 2 now). If the performance becomes
better and you still wants more performance then dynamically allocate 1-2
extra cpuvps using onmode -p and monitor the performance
2. Modify MULTIPROCESSOR to 1
2. If your OS is Sun Solaris then make sure you are using KAIO.
3. Also, monitor your online.log during the slow down to see any memory leaks
4. Your NETTYPE seems to be ok, but if the performance is still slow try
giving it variant combinations
5. Run the following commands and see if you see anything out of ordinary
onstat -g seg
onstat -g ioq
onstat -g rea
onstat -g lmx
onstat -g iof
onstat -g glo
onstat -g ntu
etc., etc., etc.
6. Also, when the engine is at its knees run onstat - see if the engine mode
changes for some reason.
7. I assume you are running full duplexed not half duplexed (talk to your
sysadmin and he/she will know)
Hope these helps.
Ravi.
Simon Coyne <simon.coyne@dla.com> wrote:
Dear Informix users
The firm I work for have recently purchased a 4 processor ES45 Alpha server to
run our practice management (sales ledger) system. This is replacing a 2
processor DS25 Alpha server (which struggles at period end). Outline
specification is
4x1.2Ghz Alpha processors
8Gb RAM
5304 4 Channel Ultra SCSI controller
2xMA300(I think) disk cages populated with 36Gb disks and configured as RAID
1+0 for both OS and Informix raw chunks.
Tru 64 5.1B
Informix 9.3.FC3
The instance is about 80Gb total.
Unfortunately, we are experiencing crippling performance problems. When we get
to about 55 green screen users on the system, performance within Informix
becomes terrible. The 3 CPU VPs go to 100% and disk I/O drops to almost 0
across every one of the Informix disks. On occassion, when CPU utilisation
dips below ~85%, disk I/O goes back to normal levels and performance is great.
At all times, the disk performance for O/S based operations is more than
acceptable. Informix performance is almost binary; it's either working or
stopped. There doesn't seem to be an inbetween position where performance
starts degrading.
Has anyone seen anything like this at all or have any helpful suggestions?
(Update statistics runs daily.)
Thanks in advance
Simon Coyne
Infrastructure Specialist
DLA Piper Rudnick Gray Cary LLP
ONconfig follows
#**************************************************************************
#
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig.std
# Description: Informix Dynamic Server 2000 Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace name
ROOTPATH /dev/rootdbs # Path for device containing root dbspace
ROOTOFFSET 500 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 2000000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS physdbs # Location (dbspace) of physical log
PHYSFILE 1000000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 1034 # Number of logical log files
LOGSIZE 1500 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /usr/informix/online.log # System message log file path
CONSOLE /usr/informix/console.log # System console message path
ALARMPROGRAM /usr/informix/etc/log_full.sh # Alarm program path
TBLSPACE_STATS 1 # Maintain tblspace statistics
# System Archive Tape Device
TAPEDEV /devices/tape/tape17c # Tape device path
#TAPEDEV /devices/tape/tape16c # FS Backup device
#TAPEDEV /dev/null # Black Hole
#TAPEDEV /nfsdump/aristadbbackup.img
TAPEBLK 1024 # Tape block size (Kbytes)
TAPESIZE 85000000 # Maximum amount of data to put on tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV /devices/tape/tape16
LTAPEBLK 1024 # Log tape block size (Kbytes)
LTAPESIZE 35000000 # Max amount of data to put on log tape (Kbytes)
# Optical
STAGEBLOB # Informix Dynamic Server 2000 staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLine instance
DBSERVERNAME arista # Name of default database server
DBSERVERALIASES arista_tcp # List of alternate dbservernames
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 0 # 0 for single-processor, 1 for multi-processor
NUMCPUVPS 2 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
NOAGE 0 # Process aging# NRW changed proc affinity to second cpu to help system performance
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 1000000 # Maximum number of locks#[was] BUFFERS 300000 # Maximum number of shared buffers
#[was] BUFFERS 400000 # Maximum number of shared buffers
BUFFERS 500000 # Maximum number of shared buffers
NUMAIOVPS 16 # Number of IO vps
#[was] PHYSBUFF 655 # Physical log buffer size (Kbytes)
PHYSBUFF 864 # Physical log buffer size (Kbytes)
#[was] LOGBUFF 32 # Logical log buffer size (Kbytes)
LOGBUFF 38 # Logical log buffer size (Kbytes)LOGSMAX 2048 # Maximum number of logical log files
CLEANERS 16 # Number of buffer cleaner processes
SHMBASE 0x200000000 # Shared memory base address#[was] SHMVIRTSIZE 155648 # initial virtual shared memory segment size
#[was] SHMVIRTSIZE 188416 # initial virtual shared memory segment size
SHMVIRTSIZE 753644 # initial virtual shared memory segment size#[was] SHMADD 16000 # Size of new shared memory segments (Kbytes)
#[was] SHMADD 18841 # Size of new shared memory segments (Kbytes)
SHMADD 75366 # Size of new shared memory segments (Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 128 # Number of LRU queues
LRU_MAX_DIRTY 3 # LRU percent dirty begin cleaning limit# [was] LRU_MAX_DIRTY 10 # LRU percent dirty begin cleaning limit
# [was] LRU_MIN_DIRTY 5 # LRU percent dirty end cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water mark percentage
LTXEHWM 60 # Long transaction high water mark (exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline worker threads
ON_RECVRY_THREADS 1 # Default number of online worker
Hi Simon
It's my general opinion.
It sounds like a OS config problem. Because you upgrade your system and
most
of the times, there are certain configurations that could be different
with old system.
Ensure your configuration, mapping old system with new one. Obviously
take special
care if you change OS version.
A second issue, could be nice review again INFORMIX parameters, because
sometimes you could made migration quickly and you couldn't consider
tuning Informix parameters with new server.
And a third,
Check load (VPU,NET,IO,etc) in order to find where is the overload.
Check your LRUs,CLEANERs are working fine.
Check your Logical Logs, and so on.
Regards
-----Mensaje original-----
De: forum.subscriber@iiug.org [mailto:forum.subscriber@iiug.org] En
nombre de Simon Coyne
Enviado el: Martes, 13 de Diciembre de 2005 10:04 a.m.
Para: ids@iiug.org
Asunto: Poor performance on replacement server [6087]
Dear Informix users
The firm I work for have recently purchased a 4 processor ES45 Alpha
server to run our practice management (sales ledger) system. This is
replacing a 2 processor DS25 Alpha server (which struggles at period
end). Outline specification is
4x1.2Ghz Alpha processors
8Gb RAM
5304 4 Channel Ultra SCSI controller
2xMA300(I think) disk cages populated with 36Gb disks and configured as
RAID 1+0 for both OS and Informix raw chunks. Tru 64 5.1B Informix
9.3.FC3
The instance is about 80Gb total.
Unfortunately, we are experiencing crippling performance problems. When
we get to about 55 green screen users on the system, performance within
Informix becomes terrible. The 3 CPU VPs go to 100% and disk I/O drops
to almost 0 across every one of the Informix disks. On occassion, when
CPU utilisation dips below ~85%, disk I/O goes back to normal levels and
performance is great. At all times, the disk performance for O/S based
operations is more than acceptable. Informix performance is almost
binary; it's either working or stopped. There doesn't seem to be an
inbetween position where performance starts degrading.
Has anyone seen anything like this at all or have any helpful
suggestions? (Update statistics runs daily.)
Thanks in advance
Simon Coyne
Infrastructure Specialist
DLA Piper Rudnick Gray Cary LLP
ONconfig follows
#***********************************************************************
***
#
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig.std
# Description: Informix Dynamic Server 2000 Configuration Parameters #
#***********************************************************************
***
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace name
ROOTPATH /dev/rootdbs # Path for device containing rootdbspace
ROOTOFFSET 500 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 2000000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirroredroot
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS physdbs # Location (dbspace) of physical log
PHYSFILE 1000000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 1034 # Number of logical log files
LOGSIZE 1500 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /usr/informix/online.log # System message log file path
CONSOLE /usr/informix/console.log # System console message path
ALARMPROGRAM /usr/informix/etc/log_full.sh # Alarm program path
TBLSPACE_STATS 1 # Maintain tblspace statistics
# System Archive Tape Device
TAPEDEV /devices/tape/tape17c # Tape device path
#TAPEDEV /devices/tape/tape16c # FS Backup device
#TAPEDEV /dev/null # Black Hole
#TAPEDEV /nfsdump/aristadbbackup.img
TAPEBLK 1024 # Tape block size (Kbytes)
TAPESIZE 85000000 # Maximum amount of data to put on tape
(Kbytes)
# Log Archive Tape Device
LTAPEDEV /devices/tape/tape16
LTAPEBLK 1024 # Log tape block size (Kbytes)
LTAPESIZE 35000000 # Max amount of data to put on log tape
(Kbytes)
# Optical
STAGEBLOB # Informix Dynamic Server 2000 staging
area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLineinstance
DBSERVERNAME arista # Name of default database server
DBSERVERALIASES arista_tcp # List of alternate dbservernames
DEADLOCK_TIMEOUT 60 # Max time to wait of lock indistributed env.
RESIDENT 0 # Forced residency flag (Yes = 1, No =
0)
MULTIPROCESSOR 0 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 2 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vpsto one
NOAGE 0 # Process aging# NRW changed proc affinity to second cpu to help system performance
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 1000000 # Maximum number of locks#[was] BUFFERS 300000 # Maximum number of shared
buffers
#[was] BUFFERS 400000 # Maximum number of shared
buffers
BUFFERS 500000 # Maximum number of shared buffers
NUMAIOVPS 16 # Number of IO vps#[was] PHYSBUFF 655 # Physical log buffer size
(Kbytes)
PHYSBUFF 864 # Physical log buffer size (Kbytes)#[was] LOGBUFF 32 # Logical log buffer size
(Kbytes)
LOGBUFF 38 # Logical log buffer size (Kbytes)LOGSMAX 2048 # Maximum number of logical log files
CLEANERS 16 # Number of buffer cleaner processes
SHMBASE 0x200000000 # Shared memory base address#[was] SHMVIRTSIZE 155648 # initial virtual shared memory
segment size
#[was] SHMVIRTSIZE 188416 # initial virtual shared memory
segment size
SHMVIRTSIZE 753644 # initial virtual shared memory segmentsize
#[was] SHMADD 16000 # Size of new shared memory
segments (Kbytes)
#[was] SHMADD 18841 # Size of new shared memory
segments (Kbytes)
SHMADD 75366 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 128 # Number of LRU queues
LRU_MAX_DIRTY 3 # LRU percent dirty begin cleaning limit
# [was] LRU_MAX_DIRTY 10 # LRU percent dirty begincleaning limit
# [was] LRU_MIN_DIRTY 5 # LRU percent dirty end cleaning
limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water markpercentage
LTXEHWM 60 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore. # ON_RECVRY_THREADS: # Number of parallel worker threads during
an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline workerthreads
ON_RECVRY_THREADS 1 # Default number of online workerthre
Simon:
Hi, see below:
Art S. Kagel
> ----- Original Message -----
> From: Simon Coyne <simon.coyne@dla.com>
> At: 12/13 11:42
>
> Dear Informix users
>
> The firm I work for have recently purchased a 4 processor ES45 Alpha server
to
> run our practice management (sales ledger) system. This is replacing a 2
> processor DS25 Alpha server (which struggles at period end). Outline
> specification is
>
> 4x1.2Ghz Alpha processors
> 8Gb RAM
> 5304 4 Channel Ultra SCSI controller
> 2xMA300(I think) disk cages populated with 36Gb disks and configured as RAID
1+0
> for both OS and Informix raw chunks.
Be careful. I recommend no more than 4 spindles per controller channel. It
only takes 5-6 drives to exceed the controllers throughput capability. I
also recommend that you set up separate RAID10 structures for the OS and for
IDS.
> Tru 64 5.1B
> Informix 9.3.FC3
>
> The instance is about 80Gb total.
>
> Unfortunately, we are experiencing crippling performance problems. When we
get
> to about 55 green screen users on the system, performance within Informix
> becomes terrible. The 3 CPU VPs go to 100% and disk I/O drops to almost 0
across
Actually the ONCONFIG file below shows only 2 CPU VPs. Are those users
accessing the DB locally using shared memory connections?
> every one of the Informix disks. On occassion, when CPU utilisation dips
below
> ~85%, disk I/O goes back to normal levels and performance is great. At all
Someone suggested that lock contention is your problem. I would definitely
not rule that out. You say a 'green screen' application is using the DB?
Is it possible that the app locks records while the users is modifying the
coy on screen? This is a common DBMS application problem. Apps that access
RDBMS's for interactive users input systems should use optimistic locking
techniques (ie don't lock the row until the users finishing modifying it,
then lock the row, compare the original version to the current DB contents,
if no changes perform the update and commit, if the record has been changed
since the user fetched it, reject the changes and make the user try again.)
> times, the disk performance for O/S based operations is more than acceptable.
> Informix performance is almost binary; it's either working or stopped. There
> doesn't seem to be an inbetween position where performance starts degrading.
>
> Has anyone seen anything like this at all or have any helpful suggestions?
> (Update statistics runs daily.)
>
> Thanks in advance
>
>
> Simon Coyne
> Infrastructure Specialist
> DLA Piper Rudnick Gray Cary LLP
>
> ONconfig follows
>
> #**************************************************************************
> #
> # INFORMIX SOFTWARE, INC.
> #
> # Title: onconfig.std
> # Description: Informix Dynamic Server 2000 Configuration Parameters
> #
> #**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name
> ROOTPATH /dev/rootdbs # Path for device containing root dbspace
> ROOTOFFSET 500 # Offset of root dbspace into device (Kbytes)
> ROOTSIZE 2000000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH # Path for device containing mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS physdbs # Location (dbspace) of physical log
> PHYSFILE 1000000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 1034 # Number of logical log files
> LOGSIZE 1500 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /usr/informix/online.log # System message log file path
> CONSOLE /usr/informix/console.log # System console message path
> ALARMPROGRAM /usr/informix/etc/log_full.sh # Alarm program path
> TBLSPACE_STATS 1 # Maintain tblspace statistics>
> # System Archive Tape Device
>
> TAPEDEV /devices/tape/tape17c # Tape device path
> #TAPEDEV /devices/tape/tape16c # FS Backup device
> #TAPEDEV /dev/null # Black Hole
> #TAPEDEV /nfsdump/aristadbbackup.img
> TAPEBLK 1024 # Tape block size (Kbytes)
> TAPESIZE 85000000 # Maximum amount of data to put on tape (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /devices/tape/tape16
> LTAPEBLK 1024 # Log tape block size (Kbytes)
> LTAPESIZE 35000000 # Max amount of data to put on log tape (Kbytes)>
> # Optical
>
> STAGEBLOB # Informix Dynamic Server 2000 staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a OnLine instance
> DBSERVERNAME arista # Name of default database server
> DBSERVERALIASES arista_tcp # List of alternate dbservernames
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
> RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
I don't remember any problems using RESIDENT with tru64, so I'm going to
suggest you change this to 1 (resident segment locked in memory) or -1
(initial virtual segment locked into memory as well as the resident
segment). You have plenty of RAM, use it. This can be a huge performance
gain on a busy system.
> MULTIPROCESSOR 0 # 0 for single-processor, 1 for multi-processor
You have 4 CPUs and you set CPUVPs >1, MULTIPROCESSOR must be 1 as well to
properly implement resource control.
> NUMCPUVPS 2 # Number of user (cpu) vps
Don't be squeemish, set this to at least 4. You have 4 very fast CPUs, use
them. Processors faster than about 400MHZ (~600MHZ for Intel) can support
more than one CPU VP per physical CPU according to a round of user testing
several years ago.
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
>
> NOAGE 0 # Process aging
Check you release notes to see if this is supported on tru64. I don't know
how TRU64 works but many UNIXes, especially Solaris and HPUX, are agressive
about lowering the runtime priority of long running processes, and you
oninits are normally running for a VERY LONG time before you restart them.
Especially compared to user tasks that come and go in minutes. So, best if
NOAGE is available on your platform to set it to 1.
> # NRW changed proc affinity to second cpu to help system performance
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 1000000 # Maximum number of locks
> BUFFERS 500000 # Maximum number of shared buffers
> NUMAIOVPS 16 # Number of IO vps
If all of your chunks are RAW and tru64 supports it you are using KAIO for
chunk IO, so you typically need from 2-6 AIO VPs. Check onstat -g iov, if
there are kaio threads reported for each CPU VP, then you should look to
reduce NUMAIOPVS. On the same report, any aio vps with an io/wup value of
0.0 is completely unused and you can safely reduce NUMAIOVPS by the number
of such.
> PHYSBUFF 864 # Physical log buffer size (Kbytes)
This is a bit large. A larger PHYSUFF will delay physically writing
physical log pages to disk and will slow down check
Thanks to everyone for their
replies. First, I've been a bit of a clot and given
the wrong ONconfig file; I took the config from the old server. It is not much
different and indeed the changes are around MULTIPROCESSOR and the number of
CPU
VPs. (I've also changed the LRUMAX & MIN entries and the amount of DS_MEMORY.
However, generally, they are very similar. My intention was to get the thing
working and then tune it up properly.)
Taking peoples comments and questions
Could it be locking? Possibly. I cannot say with any certainty because I
cannot get
a response from a query to sysmaster because the CPUs are maxed out. However,
the
application (both 4GL and browser based) haven't changed at all; it is the same
programs attaching to both systems. We do experience some problems on the DS25
with
locking, but not like this.
Multiprocessor support - I've tried this both ways around with no variation in
performance.
onstat information - Unfortunately, management took the decision to back the
change
out and we are now back on the DS25. Therefore collecting current stats is
impossible.
Is it an OS problem? I'm not convinced; as I said in my original post, if IDS
is
taken out of the equation, the OS seems to be OK. Disk IO is good and CPU
usage is
very low.
Hardware Setup - The card is a 4 channel controller with 256Mb cache and the
disk
cages have 2 channels per disk tray. The first 7 disks on channel 0, tray 0 are
configured as a 4 disk stripe for the OS and a 2 disk stripe containing the
rootdbs, physdbs and logdbs. Last disk is a hot spare. Channel 1, tray 0 is
configured as a 6 disk stripe with a hot spare split into lots of 15Gb logical
volumes for the data. All these stripes are then mirrored to tray 1.
The one thing I maybe didn't make clear in my original post is that the same
configuration (both IDS and OS) runs significantly quicker on the DS25. The
DS25
only struggles to support the business during peak load; on a day to day basis
it
is fine. Currently, the ES45 isn't even coping with a 20% of day to day
loading (it
starts struggling with less than 50 users logged in as opposed to 300).
Regards
Simon Coyne
You can
also check the number of sql statements which are performing
sequential scans or just have a high cost associated with them.
Onstat -p will give you the number of seq scans performed on the system
since the system was started (or you zero'd the stats) then you could
profile the sql running through your system by collecting a snapshot of
current sql every 5 minutes (this period seems to work for us and is a
good place to start) using the sysmaster tables (there's a flag in there
which tells you if the query is perfoming a seq scan and also gives you
the estimated cost of the query). This will enable you to identify any
poor performing sql during the peak load time.
If you want to pursue this let me know if you need the sysmaster sql.
-----Original Message-----
From: forum.subscriber@iiug.org [mailto:forum.subscriber@iiug.org] On
Behalf Of John Bejarano
Sent: 13 December 2005 17:04
To: ids@iiug.org
Subject: Re: Poor performance on replacement server [6089]
Could this be users waiting on locks? Occasionally, after a certain
threshold, lock waiting can begin to cascade and get worse in an
apparently step-wise fashion. Try placing a monitor on the number of
locks to see if that all of a sudden jumps. If it does, it may indicate
that a transaction in the application may need to be made more
efficient.
--John Bejarano.
--- Simon Coyne <simon.coyne@dla.com> wrote:
> Dear Informix users
>
> The firm I work for have recently purchased a 4 processor ES45 Alpha
> server to run our practice management (sales ledger) system. This is
> replacing a 2 processor DS25 Alpha server (which struggles at period
> end). Outline specification is
>
> 4x1.2Ghz Alpha processors
> 8Gb RAM
> 5304 4 Channel Ultra SCSI controller
> 2xMA300(I think) disk cages populated with 36Gb disks and configured
> as RAID 1+0 for both OS and Informix raw chunks.
> Tru 64 5.1B
> Informix 9.3.FC3
>
> The instance is about 80Gb total.
>
> Unfortunately, we are experiencing crippling performance problems.
> When we get to about 55 green screen users on the system, performance
> within Informix becomes terrible. The 3 CPU VPs go to 100% and disk
> I/O drops to almost 0 across every one of the Informix disks. On
> occassion, when CPU utilisation dips below ~85%, disk I/O goes back to
> normal levels and performance is great. At all times, the disk
> performance for O/S based operations is more than acceptable. Informix
> performance is almost binary; it's either working or stopped. There
> doesn't seem to be an inbetween position where performance starts
> degrading.
>
> Has anyone seen anything like this at all or have any helpful
> suggestions? (Update statistics runs
> daily.)
>
> Thanks in advance
>
>
> Simon Coyne
> Infrastructure Specialist
> DLA Piper Rudnick Gray Cary LLP
>
> ONconfig follows
>
>
#***********************************************************************
***
> #
> # INFORMIX SOFTWARE, INC.
> #
> # Title: onconfig.std
> # Description: Informix Dynamic Server 2000 Configuration Parameters
> #
>
#***********************************************************************
***
>
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name
> ROOTPATH /dev/rootdbs # Path for device> containing root dbspace
> ROOTOFFSET 500 # Offset of root> dbspace into device (Kbytes)
> ROOTSIZE 2000000 # Size of root
> dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag
> (Yes = 1, No = 0)
> MIRRORPATH # Path for device> containing mirrored root
> MIRROROFFSET 0 # Offset into> mirrored device (Kbytes)
>
> # Physical Log Configuration
>
> PHYSDBS physdbs # Location (dbspace)
> of physical log
> PHYSFILE 1000000 # Physical log file
> size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 1034 # Number of logical> log files
> LOGSIZE 1500 # Logical log size
> (Kbytes)>
> # Diagnostics
>
> MSGPATH /usr/informix/online.log # System
> message log file path
> CONSOLE /usr/informix/console.log # System
> console message path
> ALARMPROGRAM /usr/informix/etc/log_full.sh #
> Alarm program path
> TBLSPACE_STATS 1 # Maintain tblspace> statistics
>
> # System Archive Tape Device
>
> TAPEDEV /devices/tape/tape17c # Tape device
> path
> #TAPEDEV /devices/tape/tape16c # FS Backup
> device
> #TAPEDEV /dev/null # Black Hole
> #TAPEDEV /nfsdump/aristadbbackup.img
> TAPEBLK 1024 # Tape block size
> (Kbytes)
> TAPESIZE 85000000 # Maximum amount of> data to put on tape (Kbytes)
>
> # Log Archive Tape Device
>
> LTAPEDEV /devices/tape/tape16
> LTAPEBLK 1024 # Log tape block
> size (Kbytes)
> LTAPESIZE 35000000 # Max amount of data> to put on log tape (Kbytes)
>
> # Optical
>
> STAGEBLOB # Informix Dynamic
> Server 2000 staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id> corresponding to a OnLine instance
> DBSERVERNAME arista # Name of default> database server
> DBSERVERALIASES arista_tcp # List of alternate> dbservernames
> DEADLOCK_TIMEOUT 60 # Max time to wait> of lock in distributed env.
> RESIDENT 0 # Forced residency
> flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 0 # 0 for> single-processor, 1 for multi-processor
> NUMCPUVPS 2 # Number of user
> (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit> number of cpu vps to one
>
> NOAGE 0 # Process aging> # NRW changed proc affinity to second cpu to help system performance
> AFF_SPROC 0 # Affinity start> processor
> AFF_NPROCS 0 # Affinity number of> processors
>
> # Shared Memory Parameters
>
> LOCKS 1000000 # Maximum number of> locks
> #[was] BUFFERS 300000 # Maximum
> number of shared buffers
> #[was] BUFFERS 400000 # Maximum
> number of shared buffers
> BUFFERS 500000 # Maximum number of> shared buffers
> NUMAIOVPS 16 # Number of IO vps
> #[was] PHYSBUFF 655 # Physical> log buffer size (Kbytes)
> PHYSBUFF 864 # Physical log
> buffer size (Kbytes)
> #[was] LOGBUFF 32 # Logical log
> buffer size (Kbytes)
> LOGBUFF 38 # Logical log buffer
> size (Kbytes)> LOGSMAX 2048 # Maximum number of
> logical log files
> CLEANERS 16 # Number of buffer> cleaner processes
> SHMBASE 0x200000000 # Shared memory> base address
> #[was] SHMVIRTSIZE 155648 # initial
> virtual shared memory segment size
> #[was] SHMVIRTSIZE 188416 # initial
> virtual shared memory segment size
> SHMVIRTSIZE 753644 # initial virtual> shared memory segment size
> #[was] SHMADD 16000 # Size of new
> shared memory segments (Kbytes)
> #[was] SHMADD 18841 # Size of new
> shared memory segments (Kbytes)
> SHMADD 75366 # Size of new shared> memory segments (Kbytes)
> SHMTOTAL 0 # Total shared
> memory (Kbytes). 0=>unlimited
> CKPTINTVL 300 # Check point
> interval (in sec)
> LRUS 128 # Number of LRU> queues
> LRU_MAX_DIRTY 3 # LRU percent dirty> begin cleaning limit
> # [was] L
Simon,
just to get the 'too obvious' out of the way, take a dbschema -ss of the
database on each server and compare. Are the locking modes of the tables the
same? It's a common error to recreate the database using dbschema or dbexport
without the -ss flag and so have the target tables default to PAGE level
locking. This will cause extensive contention versus ROW level locking and
cause the CPU VPs to spin.
Art S. Kagel
----- Original Message -----
From: simon@coynes.eclipse.co.uk
At: 12/14 5:20
Thanks to everyone for their replies. First, I've been a bit of a clot and
given
the wrong ONconfig file; I took the config from the old server. It is not much
different and indeed the changes are around MULTIPROCESSOR and the number of
CPU
VPs. (I've also changed the LRUMAX & MIN entries and the amount of DS_MEMORY.
However, generally, they are very similar. My intention was to get the thing
working and then tune it up properly.)
Taking peoples comments and questions
Could it be locking? Possibly. I cannot say with any certainty because I cannot
get
a response from a query to sysmaster because the CPUs are maxed out. However,
the
application (both 4GL and browser based) haven't changed at all; it is the same
programs attaching to both systems. We do experience some problems on the DS25
with
locking, but not like this.
Multiprocessor support - I've tried this both ways around with no variation in
performance.
onstat information - Unfortunately, management took the decision to back the
change
out and we are now back on the DS25. Therefore collecting current stats is
impossible.
Is it an OS problem? I'm not convinced; as I said in my original post, if IDS
is
taken out of the equation, the OS seems to be OK. Disk IO is good and CPU usage
is
very low.
Hardware Setup - The card is a 4 channel controller with 256Mb cache and the
disk
cages have 2 channels per disk tray. The first 7 disks on channel 0, tray 0 are
configured as a 4 disk stripe for the OS and a 2 disk stripe containing the
rootdbs, physdbs and logdbs. Last disk is a hot spare. Channel 1, tray 0 is
configured as a 6 disk stripe with a hot spare split into lots of 15Gb logical
volumes for the data. All these stripes are then mirrored to tray 1.
The one thing I maybe didn't make clear in my original post is that the same
configuration (both IDS and OS) runs significantly quicker on the DS25. The
DS25
only struggles to support the business during peak load; on a day to day basis
it
is fine. Currently, the ES45 isn't even coping with a 20% of day to day loading
(it
starts struggling with less than 50 users logged in as opposed to 300).
Regards
Simon Coyne
Hi Simon,
More comments, suggestions and thoughts :
1. It seems all rootdbs, physdbs and logdbs mapped to the same disk/partition
if I am not confused. Check and see if the DS25 is mapped the same way. (I bet
you that it is not, just kidding) I would put the rootdbs on its owb
partition, nothing else just the rootdbs and then put logdbs in its own
partition as well.
2. Check the /etc/system files in both boxes and see if the memory parameters
are set differently for the OS
3. If/when ever /etc/system file is changed it requires an OS reboot for the
changes to take effect
4. Are the MHZ on the CPUs in both boxes the same?
5. Run ipcs command during the lock up/ slow down and see if OS is having some
memory allocation issues or memory segment issues. When you run the ipcs
command you will see the segments cearly, compare them in both boxes and see
if you see adverse behavior in the new box
6. Also, see if the message portion of the memory in the old box, DS25 is
larger than the one in the new box
7. Also, I remember someone had suggested to set NOAGE, it is a good thing.
8. 9.3.FC3 check for the bugs, ask the support guys to check and give you the
bug lists for 9.3.FC3, I know in one shop I was in we upgraded to 9.3.FC4
because we didn't want to deal with the problems we were getting with FC3.
They may ask you to upgrade to 9.4 since some of the 9.3 versions are
desupported now (I think).
Hope these helps.
Ravi.
simon@coynes.eclipse.co.uk wrote:
Thanks to everyone for their replies. First, I've been a bit of a clot and
given
the wrong ONconfig file; I took the config from the old server. It is not much
different and indeed the changes are around MULTIPROCESSOR and the number of
CPU
VPs. (I've also changed the LRUMAX & MIN entries and the amount of DS_MEMORY.
However, generally, they are very similar. My intention was to get the thing
working and then tune it up properly.)
Taking peoples comments and questions
Could it be locking? Possibly. I cannot say with any certainty because I
cannot get
a response from a query to sysmaster because the CPUs are maxed out. However,
the
application (both 4GL and browser based) haven't changed at all; it is the same
programs attaching to both systems. We do experience some problems on the DS25
with
locking, but not like this.
Multiprocessor support - I've tried this both ways around with no variation in
performance.
onstat information - Unfortunately, management took the decision to back the
change
out and we are now back on the DS25. Therefore collecting current stats is
impossible.
Is it an OS problem? I'm not convinced; as I said in my original post, if IDS
is
taken out of the equation, the OS seems to be OK. Disk IO is good and CPU
usage is
very low.
Hardware Setup - The card is a 4 channel controller with 256Mb cache and the
disk
cages have 2 channels per disk tray. The first 7 disks on channel 0, tray 0 are
configured as a 4 disk stripe for the OS and a 2 disk stripe containing the
rootdbs, physdbs and logdbs. Last disk is a hot spare. Channel 1, tray 0 is
configured as a 6 disk stripe with a hot spare split into lots of 15Gb logical
volumes for the data. All these stripes are then mirrored to tray 1.
The one thing I maybe didn't make clear in my original post is that the same
configuration (both IDS and OS) runs significantly quicker on the DS25. The
DS25
only struggles to support the business during peak load; on a day to day basis
it
is fine. Currently, the ES45 isn't even coping with a 20% of day to day
loading (it
starts struggling with less than 50 users logged in as opposed to 300).
Regards
Simon Coyne
---------------------------------
Yahoo! Shopping
Find Great Deals on Holiday Gifts at Yahoo! Shopping
Dear Simon, If I remember well, SA 5304 RAID controller is Compaq's product supporting Ultra160 SCSI or WideUltra3, with maximum throughput of 160 MB/s. You did not mention the speed of your 36GB disks, but if it's 10K rpm or (desirably) 15K rpm, with 7 disks/channel it could be the limiting factor. 5304 is rather dated product, launched before 4-5 years. Also, you are perhaps overloading its engine by populating 4 channels with 7 disks each. Even if you had a controller supporting higher maximum transfer rates, I would suggest to use at least 2 controllers, if possible supporting Ultra320 SCSI and placed to different buses. That being said, I am convinced that this is not the only place for improvement, and you will get much more by considering suggestions from great people in the mailing group that had already responeded (like Art Kagel and many others). Perhaps it could be advantageous to place rootdbs, physdbs and logdbs to more than 2 disk stripe, or separate them. Sorry for my far-from-perfect English, Darko Krstic
The
database was restored from an ontape archive and subsequently had update
statistics run against it (using your "dostats" scripts).
-----Original Message-----
From: kagel@bloomberg.net [mailto:kagel@bloomberg.net]
Sent: 14 December 2005 17:32
To: simon@coynes.eclipse.co.uk
Subject: RE: Poor performance on replacement server [6097]
Simon, just to get the 'too obvious' out of the way, take a dbschema -ss of
the
database on each server and compare. Are the locking modes of the tables
the
same? It's a common error to recreate the database using dbschema or
dbexportwithout the -ss flag and so have the target tables default to PAGE level
locking. This will cause extensive contention versus ROW level locking and
cause the CPU VPs to spin.
Art
----- Original Message -----
From: simon@coynes.eclipse.co.uk
At: 12/14 5:20
Thanks to everyone for their replies. First, I've been a bit of a clot and
given
the wrong ONconfig file; I took the config from the old server. It is not
much
different and indeed the changes are around MULTIPROCESSOR and the number of
CPU
VPs. (I've also changed the LRUMAX & MIN entries and the amount of
DS_MEMORY.
However, generally, they are very similar. My intention was to get the thing
working and then tune it up properly.)
Taking peoples comments and questions
Could it be locking? Possibly. I cannot say with any certainty because I
cannot
get
a response from a query to sysmaster because the CPUs are maxed out.
However,
the
application (both 4GL and browser based) haven't changed at all; it is the
same
programs attaching to both systems. We do experience some problems on the
DS25
with
locking, but not like this.
Multiprocessor support - I've tried this both ways around with no variation
in
performance.
onstat information - Unfortunately, management took the decision to back the
change
out and we are now back on the DS25. Therefore collecting current stats is
impossible.
Is it an OS problem? I'm not convinced; as I said in my original post, if
IDS is
taken out of the equation, the OS seems to be OK. Disk IO is good and CPU
usage
is
very low.
Hardware Setup - The card is a 4 channel controller with 256Mb cache and the
disk
cages have 2 channels per disk tray. The first 7 disks on channel 0, tray 0
are
configured as a 4 disk stripe for the OS and a 2 disk stripe containing the
rootdbs, physdbs and logdbs. Last disk is a hot spare. Channel 1, tray 0 is
configured as a 6 disk stripe with a hot spare split into lots of 15Gb
logical
volumes for the data. All these stripes are then mirrored to tray 1.
The one thing I maybe didn't make clear in my original post is that the same
configuration (both IDS and OS) runs significantly quicker on the DS25. The
DS25
only struggles to support the business during peak load; on a day to day
basis
it
is fine. Currently, the ES45 isn't even coping with a 20% of day to day
loading
(it
starts struggling with less than 50 users logged in as opposed to 300).
Regards
Simon Coyne
--
No virus found in this incoming message.
Checked by AVG Free Edition.
Version: 7.1.371 / Virus Database: 267.13.13/200 - Release Date: 14/12/2005
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix