Problem with Long Checkpoints
Posted in 1999
Topics: Backup & Restore, Performance & Tuning, Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues
Hello,
I have a problem with long checkpoint durations on some of my systems.
Let me describe my environment.
HW Environment:
7 HP 9000 K460 boxes w/
HP-UX 10.20
768mb or 1 gig of ram on each box
16 to 24 - 2.1 gb disks on each box depending on the number of Informix
instances. Each instance has 10 dbspaces on only 2 mirrored pairs(4
disks in total). I'm sure this is at least part of the problem. But
I'm sure there is some tuning the can help?
DB Environment
1 to 4 instances on each box. The database is identical on each
instance, (except for the contents of course!), and is roughly 4 gb is
total size. Theses are by no stretch large databases! All storage is
mirrored using HP LVM mirroring. Each instances probably has an
average of 30 users hitting them 24x7. The application runs on the
box also, so each user is telnetting in and running the 4GL app locally
through shared memory connections.
Here is the problem. On the best performing db instance the checkpoint
duration is 2 or 3 seconds at every CKPTINTVL. On the worst
performing instance the duration is usually around 8 or 9 seconds and
sometimes up to 25 seconds. I am including all the info that I saw Art
Kagel ask of someone a few days ago. Hope this isn't tooooooo much
info.
All suggestions welcome(and very much needed!)
Thanks....Jeff S.
INFORMIX-OnLine Version 7.24.UC5 -- On-Line -- Up 4 days 11:53:14 --
192072 Kbytes
Configuration File: /informix/etc/onconfig_lff
#**************************************************************************
#
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig.std
# Description: INFORMIX-OnLine Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace nameROOTPATH /links/ROOTDBS_LFF # Path for device containing root
dbspace
ROOTOFFSET 0 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 128000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirroredroot
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS phylog # Location (dbspace) of physical log
PHYSFILE 25000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 30 # Number of logical log files
LOGSIZE 10000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /informix/online.log_lff # System message log file path
CONSOLE /dev/console # System console message path
ALARMPROGRAM /usr/local/publix/publix.infx_alarm
# System Archive Tape Device
#TAPEDEV /dev/null
TAPEDEV /dev/rmt/0m # Tape device path
#TAPEDEV /bu_lff/ontape_post_reorg.lff # Tape device path
TAPEBLK 16 # Tape block size (Kbytes)
TAPESIZE 5000000 # Maximum amount of data to put on tape
(Kbytes)
# Log Archive Tape Device
LTAPEDEV /bu_lff/log.lff # Log tape device path
#LTAPEDEV /dev/rmt/0m # Log tape device path
LTAPEBLK 16 # Log tape block size (Kbytes)
LTAPESIZE 5000000 # Max amount of data to put on log tape
(Kbytes)
# Optical
STAGEBLOB ,1 # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLineinstance
DBSERVERNAME onlineipc_lff # Name of default database server
DBSERVERALIASES onlinetcp_lff # List of alternate dbservernames
NETTYPE ipcshm,1,100,CPU # Override sqlhosts
NETTYPE soctcp,1,75,NET # Override sqlhosts
DEADLOCK_TIMEOUT 300 # Max time to wait of lock indistributed env.
RESIDENT 1 # Forced residency flag (Yes = 1, No =
0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 4 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vpsto one
NOAGE 0 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 300000 # Maximum number of locks
BUFFERS 60000 # Maximum number of shared buffers
NUMAIOVPS 20 # Number of IO vps
PHYSBUFF 260 # Physical log buffer size (Kbytes)
LOGBUFF 32 # Logical log buffer size (Kbytes)LOGSMAX 50 # Maximum number of logical log files
CLEANERS 50 # Number of buffer cleaner processes
SHMBASE 0x0 # Shared memory base address
SHMVIRTSIZE 48000 # initial virtual shared memory segmentsize
SHMADD 16384 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 650000 # Total shared memory (Kbytes).
0=>unlimited
CKPTINTVL 3000 # Check point interval (in sec)
LRUS 50 # Number of LRU queues
LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaninglimit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 40 # Long transaction high water markpercentage
LTXEHWM 50 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 48 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline workerthreads
ON_RECVRY_THREADS 1 # Default number of online workerthreads
# Data Replication Variables
# DRAUTO: 0 manual, 1 retain type, 2 reverse type
DRAUTO 0 # DR automatic switchover
DRINTERVAL 30 # DR max time between DR buffer flushes
(in sec)
DRTIMEOUT 30 # DR network timeout (in sec)DRLOSTFOUND /informix/etc/dr.lostfound # DR lost+found file path
# Backup/Restore variables
BAR_ACT_LOG /tmp/bar_act.log
BAR_MAX_BACKUP 0
BAR_RETRY 1
BAR_NB_XPORT_COUNT 10
BAR_XFER_BUF_SIZE 31
# Read Ahead Variables
RA_PAGES 32 # Number of pages to attempt to readahead
RA_THRESHOLD 30 # Number of pages left before nextg
Jeffrey Screws wrote: > > Hello, > > I have a problem with long checkpoint durations on some of my systems. > Let me describe my environment. > > HW Environment: > > 7 HP 9000 K460 boxes w/ > HP-UX 10.20 > 768mb or 1 gig of ram on each box > 16 to 24 - 2.1 gb disks on each box depending on the number of Informix > instances. Each instance has 10 dbspaces on only 2 mirrored pairs(4 > disks in total). I'm sure this is at least part of the problem. But > I'm sure there is some tuning the can help? > > DB Environment > > 1 to 4 instances on each box. The database is identical on each > instance, (except for the contents of course!), and is roughly 4 gb is > total size. Theses are by no stretch large databases! All storage is > mirrored using HP LVM mirroring. Each instances probably has an > average of 30 users hitting them 24x7. The application runs on the > box also, so each user is telnetting in and running the 4GL app locally > through shared memory connections. > > Here is the problem. On the best performing db instance the checkpoint > duration is 2 or 3 seconds at every CKPTINTVL. On the worst > performing instance the duration is usually around 8 or 9 seconds and > sometimes up to 25 seconds. I am including all the info that I saw Art > Kagel ask of someone a few days ago. Hope this isn't tooooooo much > info. Actually the instance is mostly running well. I see you could take advantage of more BUFFERS (there are a few FG writes on the -F report) and you can probably take advantage of a few more aio VPs since they are all currently performing >1 I/O per wakeup and there should be at least a few at <1 I/O per wakeup indicating not waiting for I/O requests to be serviced. You can try changing LRU_MAX/MIN_DIRTY to 1/0 from 2/1 it will help somewhat. The real problem on this server is that the disks are not keeping up with the demands placed on them. I see that from the fact that you have a HUGE RA_THRESHOLD and yet 98% of your read ahead is being used. That means to me that the engine has not finished reading an RA_PAGES request when the query completes else there would be many orphaned RA pages with so large a threshold. This is to be expected from singleton and simple mirror drive disk farms. Also 2.1GB drives were mostly 5400RPM drives with lower data density per cylinder which are considered slow compared to the latest 9-36GB drives with spindles turning 10,000RPM and uch higher data density. I do not think that a 4GB database can gain much from striping or RAID but I would strongly suggest that you replace the 2.1 GB drives with newer faster ones. That seems to be your biggest bottleneck. In addition it is likely that those were SCSI-2 or wide SCSI-2 drives at 10-20MB/S interface speeds. You may look into replacing the controllers with Ultra-SCSI2 at 40MB/S or Ultra-SCSI3 at 80MB/S since the 10000RPM drives are approaching 40-60MB/S transfer to the bus themselves. Art S. Kagel
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix