Checkpoints up to 137 seconds
Posted in 1999
Topics: Backup & Restore, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration
Hi
I'm using informix 730UC5 on a Sequent box, I thought I had everything
set just right, that was till Wednesday this week, when my checkpoints
started rocketing. Normally my checkpoints average at about 4 seconds,
now they are averaging at about 60 seconds.
My LRU MAX is set to 2 my LRU_MIN is set to 0. The have 17 LRU queues
and 8 CLEANERS. The checkpoint interval was recently changed to 600
seconds - I have have changed this back to 300 seconds. (I have
attached a copy of my onconfig file.)
There hasn't been a major increase in users the load is only slightly
higher than normal. It is the end of the month after all. But does
anyone have any ideas of what other step I could take to decrease the
checkpoint interval?
thanks Bridget
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace name
ROOTPATH /dev/vx/rdsk/proddg/psprodrdbs # Path for device containing root
dbspace
ROOTOFFSET 0 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 257040 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 1 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirroredroot
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS rootdbs # Location (dbspace) of physical log
PHYSFILE 40000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 40 # Number of logical log files
LOGSIZE 40000 # Logical log size (Kbytes)
# Diagnostics
#MSGPATH /usr/informix/online.log # System message log file path
MSGPATH /product/informix/730UC5/log/psprod.log
CONSOLE /dev/console # System console message path
# ALARMPROGRAM /product/informix/724UC1/etc/log_full.sh
ALARMPROGRAM /product/informix/730UC5/etc/no_log.sh
SYSALARMPROGRAM /product/informix/730UC5/etc/evidence.sh
# System Archive Tape Device
TAPEDEV /dev/rmt/td0c # Tape device path#TAPEDEV /dev/null # Tape device path
TAPEBLK 32 # Tape block size (Kbytes)
TAPESIZE 8000000 # Maximum amount of data to put on tape
(Kbytes)
# Log Archive Tape Device
LTAPEDEV /dev/rmt/td2c # Log tape device path
LTAPEBLK 32 # Log tape block size (Kbytes)
LTAPESIZE 5000000 # Max amount of data to put on log tape
(Kbytes)
# Optical
STAGEBLOB # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 1 # Unique id corresponding to a OnLineinstance
DBSERVERNAME psprodtcp # Name of default database server
DBSERVERALIASES psprodshm # List of alternate dbservernames#NETTYPE tlitcp,,, # Override sqlhosts nettype parameters
NETTYPE tlitcp,2,100,NET # Override sqlhosts nettype parameters
DEADLOCK_TIMEOUT 60 # Max time to wait of lock indistributed env.
RESIDENT 1 # Forced residency flag (Yes = 1, No =
0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 8 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vpsto one
NOAGE 1 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 8 # Affinity number of processors
# Shared Memory Parameters
LOCKS 3000000 # Maximum number of locks
BUFFERS 100000 # Maximum number of shared buffers
NUMAIOVPS 1 # Number of IO vps
PHYSBUFF 32 # Physical log buffer size (Kbytes)
LOGBUFF 32 # Logical log buffer size (Kbytes)LOGSMAX 40 # Maximum number of logical log files
CLEANERS 8 # Number of buffer cleaner processes
SHMBASE 0x10000000 # Shared memory base address
SHMVIRTSIZE 192000 # initial virtual shared memory segmentsize
SHMADD 16000 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 17 # Number of LRU queues
LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 0 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water markpercentage
LTXEHWM 60 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline workerthreads
ON_RECVRY_THREADS 1 # Default number of online workerthreads
# Data Replication Variables
# DRAUTO: 0 manual, 1 retain type, 2 reverse type
DRAUTO 0 # DR automatic switchover
DRINTERVAL 30 # DR max time between DR buffer flushes
(in sec)
DRTIMEOUT 30 # DR network timeout (in sec)DRLOSTFOUND /usr/informix/etc/dr.lostfound # DR lost+found file path
# Informix Storage Manager variables
ISM_DATA_POOL ISMData # If the data pool name is changed, be
sure to
# update $INFORMIXDIR/bin/onbar. Change
to
# ism_catalog -create_bootstrap -pool
# <new name>
ISM_LOG_POOL ISMLogs
# Read Ahead Variables
RA_PAGES 36 # Number of pages to attempt to readahead
RA_THRESHOLD 24 # Number of pages left before next group
# DBSPACETEMP:
# OnLine equivalent of DBTEMP for SE. This is the list of dbspaces
# that the OnLine SQL Engine will use to create temp tables etc.
# If specified it must be a colon separated list of dbspaces that exist
# when the OnLine system is brought online. If not specified, or if
# all dbspaces specified are invalid, various ad hoc queries will create
# temporary files in /tmp instead.
DBSPACETEMP psprodtemp01:psprodtemp02 # Default temp dbspaces
# DUMP*:
# The following parameters control the type of diagnostics information
which
# is preserved when an unanticipated error condition (assertion failure)
occurs
# during OnLine operations.
# For DUMPSHMEM
Bridget Reitsma wrote: > > Hi > > I'm using informix 730UC5 on a Sequent box, I thought I had everything > set just right, that was till Wednesday this week, when my checkpoints > started rocketing. Normally my checkpoints average at about 4 seconds, > now they are averaging at about 60 seconds. > My LRU MAX is set to 2 my LRU_MIN is set to 0. The have 17 LRU queues > and 8 CLEANERS. The checkpoint interval was recently changed to 600 > seconds - I have have changed this back to 300 seconds. (I have > attached a copy of my onconfig file.) It might be the increase in the checkpoint interval, but I doubt it. BTW increasing CLEANERS to 17 or increasing both LRUS and CLEANERS to 31 or 32 MAY reduce the checkpoint duration but I don't think that any of these items is to blame for such a substantial increase. More likely it is a process which is in a critical section causing the checkpoint to wait for it to release its latch before beginning. The checkpoint blocks new requests for a critical section latch and then waits for the existing latches to clear before continuing with the checkpoint. So if there are a later number of processed in crcitical sections, especially if they are waiting for each other to free resources, this can cause the checkpoint to become blocked. Just now I had a server checkpoint for 15 seconds with only 53 dirty buffers and 3 pages of physical log and 3 pages of logical log buffer to flush to disk. I run 127 LRUs and 127 CLEANERS and have 49 AIO VPs. This checkpoint should have taken less than a second. It did not because I have 21 VERY active middleware processes querying the system. > There hasn't been a major increase in users the load is only slightly > higher than normal. It is the end of the month after all. But does > anyone have any ideas of what other step I could take to decrease the > checkpoint interval? [SNIP] Art S. Kagel
In article <37AB1378.8D75857D@bloomberg.net>, kagel@bloomberg.net wrote: > Bridget Reitsma wrote: > > > > Hi > > > > I'm using informix 730UC5 on a Sequent box, I thought I had everything > > set just right, that was till Wednesday this week, when my checkpoints > > started rocketing. Normally my checkpoints average at about 4 seconds, > > now they are averaging at about 60 seconds. > > > My LRU MAX is set to 2 my LRU_MIN is set to 0. The have 17 LRU queues > > and 8 CLEANERS. The checkpoint interval was recently changed to 600 > > seconds - I have have changed this back to 300 seconds. (I have > > attached a copy of my onconfig file.) > > It might be the increase in the checkpoint interval, but I doubt it. BTW > increasing CLEANERS to 17 or increasing both LRUS and CLEANERS to 31 or 32 > MAY reduce the checkpoint duration but I don't think that any of these > items is to blame for such a substantial increase. More likely it is > a process which is in a critical section causing the checkpoint to wait > for it to release its latch before beginning. The checkpoint blocks new > requests for a critical section latch and then waits for the existing > latches to clear before continuing with the checkpoint. So if there are > a later number of processed in crcitical sections, especially if they are > waiting for each other to free resources, this can cause the checkpoint > to become blocked. > > Just now I had a server checkpoint for 15 seconds with only 53 dirty buffers and > 3 pages of physical log and 3 pages of logical log buffer to > flush to disk. I run 127 LRUs and 127 CLEANERS and have 49 AIO VPs. This > checkpoint should have taken less than a second. It did not because I > have 21 VERY active middleware processes querying the system. > > > There hasn't been a major increase in users the load is only slightly > > higher than normal. It is the end of the month after all. But does > > anyone have any ideas of what other step I could take to decrease the > > checkpoint interval? > [SNIP] > > Art S. Kagel > Try reducing your buffers until the read cache % starts to suffer or you start to do foreground writes. Ted Gajary Sent via Deja.com http://www.deja.com/ Share what you know. Learn what you don't.
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix