Long checkpoints
Posted in 1999
Doug McAllister reported 60-second checkpoints flushing ~18,000 dirty buffers during batch loads on a Solaris 2.6 IDS 7.31 box with 300,000 buffers. Suggestions included lowering LRU_MIN/LRU_MAX (to roughly 1/2 so fewer dirty pages accumulate), raising CLEANERS from 20 to 127 to match LRUS, increasing PHYSBUFF and SHMVIRTSIZE, and checking NUMCPUVPS/KAIO. Others argued the real bottleneck is disk layout: spread chunks across many spindles, use fragmentation/detached indexes, and prefer RAID 10 over RAID 5, which cut one poster's checkpoints to 1-2 seconds. Doug said he'd raise CLEANERS, but no confirmed outcome is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues
Greetings all.....
I watch this group every day and sometimes offer my help. Now it is my
turn.
I have a system that is generally a lightly loaded system but runs a
batch load a couple times a day. During these loads, checkpoints get
horrendous averaging 60 seconds to flush about 18,000 dirty buffers.
(See below) In a recent conversation with the tech-support guys they
suggested that I change my cleaners from 20 to 127 (to match my LRU
queues). OK, but during a checkpoint, all of the cleaners are idle
anyway. LRU MIN/MAX is set to 5 and 7 respectively.
This is a big system running Solaris 2.6 with 8gig memory and 8 400mhz
processors
Any suggestions?
tnx..........
Informix Dynamic Server Version 7.31.UC3 -- On-Line -- Up 1 days
02:55:42 --
776448 Kbytes
Configuration File: /apps/informix/etc/onconfig.GIM2
#**************************************************************************
#
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig.std
# Description: INFORMIX-OnLine Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace nameROOTPATH /dev/vx/rdsk/infchunk3 # Path for device containing root
dbspace
ROOTOFFSET 0 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 100000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirroredroot
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS phydbs1 # Location (dbspace) of physical log
PHYSFILE 75000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 50 # Number of logical log files
LOGSIZE 5000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /apps/informix/logs/gim2.log # System message log file
path
CONSOLE /apps/informix/logs/gim2.console #System console message
path
ALARMPROGRAM /apps/informix/etc/log_full.sh # Alarm program path
# System Archive Tape Device
TAPEDEV /dev/rmt/0ub # Tape device path
TAPEBLK 128 # Tape block size (Kbytes)
TAPESIZE 2000000 # Maximum amount of data to put on tape
(Kbytes)
# Log Archive Tape Device
LTAPEDEV /dev/rmt/0ub # Log tape device path
LTAPEBLK 16 # Log tape block size (Kbytes)
LTAPESIZE 10240 # Max amount of data to put on log tape
(Kbytes)
# Optical
STAGEBLOB ,1 # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 0 # Unique id corresponding to a OnLineinstance
DBSERVERNAME unxprod04_infx # Name of default database serverDBSERVERALIASES unxprod04_infx_shm # List of alternate dbservernames
NETTYPE tlitcp,2,50,NET # Configure poll thread(s) for nettype
NETTYPE ipcshm,3,50,CPU # Configure poll thread(s) for nettype
DEADLOCK_TIMEOUT 60 # Max time to wait of lock indistributed env.
RESIDENT 0 # Forced residency flag (Yes = 1, No =
0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 6 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vpsto one
NOAGE 1 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 500000 # Maximum number of locks
BUFFERS 300000 # Maximum number of shared buffers
NUMAIOVPS 2 # Number of IO vps
PHYSBUFF 32 # Physical log buffer size (Kbytes)
LOGBUFF 10 # Logical log buffer size (Kbytes)LOGSMAX 60 # Maximum number of logical log files
CLEANERS 20 # Number of buffer cleaner processes
SHMBASE 0xa000000 # Shared memory base address
SHMVIRTSIZE 32000 # initial virtual shared memory segmentsize
SHMADD 8192 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
CKPTINTVL 900 # Check point interval (in sec)
LRUS 127 # Number of LRU queues
LRU_MAX_DIRTY 7 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 5 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water markpercentage
LTXEHWM 60 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline workerthreads
ON_RECVRY_THREADS 1 # Default number of online workerthreads
# Data Replication Variables
# DRAUTO: 0 manual, 1 retain type, 2 reverse type
DRAUTO 0 # DR automatic switchover
DRINTERVAL 30 # DR max time between DR buffer flushes
(in sec)
DRTIMEOUT 30 # DR network timeout (in sec)DRLOSTFOUND /usr/informix/etc/dr.lostfound # DR lost+found file path
# CDR Variables
CDR_LOGBUFFERS 2048 # size of log reading buffer pool
(Kbytes)
CDR_EVALTHREADS 1,2 # evaluator threads
(per-cpu-vp,additional)
CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
queue (Kb
ytes)
# Backup/Restore variables
BAR_ACT_LOG /apps/informix/logs/bar_act.log
BAR_MAX_BACKUP 12
BAR_RETRY 1
BAR_NB_XPORT_COUNT 10
BAR_XFER_BUF_SIZE 31
BAR_DEBUG_LOG /tmp/bar_debug.log
BAR_BSALIB_PATH /usr/lib/ibsad001.so
# Read Ahead Variables
RA_PAGES 32 # Number of pages to attempt to readahead
RA_THRESHOLD 30 # Number of pages left before next group
# DBSPACETEMP:
# OnLine equivalent of DBTEMP for SE. This is the list of dbspaces
# that the OnLine SQL Engine will use to create temp tables etc.
# If specified it must be a colon separated list of dbspaces that exist
# when the OnLine system is brought online. If not spe
Doug McAllister wrote in message <38234139.F3266E75@fmr.com>...
>Greetings all.....
>
>I watch this group every day and sometimes offer my help. Now it is my
>turn.
>
>I have a system that is generally a lightly loaded system but runs a
>batch load a couple times a day. During these loads, checkpoints get
>horrendous averaging 60 seconds to flush about 18,000 dirty buffers.
>(See below) In a recent conversation with the tech-support guys they
>suggested that I change my cleaners from 20 to 127 (to match my LRU
>queues). OK, but during a checkpoint, all of the cleaners are idle
>anyway. LRU MIN/MAX is set to 5 and 7 respectively.
I would set them to 1 and 0 respectively, then when the checkpoint does
occur; hopefully, there are less modified pages in the buffers to flush.
Also, you might lower the number of BUFFERS a tad. Oops, I just checked
your BUFWAITS, nevermind on lowering the BUFFERS.
Oh, increase NUMCPUVPS to say a value between 8 - 12.
and you do have KAIO turned on, right???
David Weis
dweis@louisville-i-market.com
Voice: 502-637-7316
>
>This is a big system running Solaris 2.6 with 8gig memory and 8 400mhz
>processors
>
>Any suggestions?
>
>
>tnx..........
>
>Informix Dynamic Server Version 7.31.UC3 -- On-Line -- Up 1 days
>02:55:42 --
>776448 Kbytes
>
>Configuration File: /apps/informix/etc/onconfig.GIM2
>#**************************************************************************
>
>#
># INFORMIX SOFTWARE, INC.
>#
># Title: onconfig.std
># Description: INFORMIX-OnLine Configuration Parameters
>#
>#**************************************************************************
>
># Root Dbspace Configuration
>
>ROOTNAME rootdbs # Root dbspace name>ROOTPATH /dev/vx/rdsk/infchunk3 # Path for device containing root
>dbspace
>ROOTOFFSET 0 # Offset of root dbspace into device
>(Kbytes)
>ROOTSIZE 100000 # Size of root dbspace (Kbytes)>
># Disk Mirroring Configuration Parameters
>
>MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
>MIRRORPATH # Path for device containing mirrored>root
>MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
># Physical Log Configuration
>
>PHYSDBS phydbs1 # Location (dbspace) of physical log
>PHYSFILE 75000 # Physical log file size (Kbytes)>
># Logical Log Configuration
>
>LOGFILES 50 # Number of logical log files
>LOGSIZE 5000 # Logical log size (Kbytes)>
># Diagnostics
>
>MSGPATH /apps/informix/logs/gim2.log # System message log file
>path
>CONSOLE /apps/informix/logs/gim2.console #System console message
>path
>ALARMPROGRAM /apps/informix/etc/log_full.sh # Alarm program path
>
># System Archive Tape Device
>
>TAPEDEV /dev/rmt/0ub # Tape device path
>TAPEBLK 128 # Tape block size (Kbytes)
>TAPESIZE 2000000 # Maximum amount of data to put on tape
>(Kbytes)>
># Log Archive Tape Device
>
>LTAPEDEV /dev/rmt/0ub # Log tape device path
>LTAPEBLK 16 # Log tape block size (Kbytes)
>LTAPESIZE 10240 # Max amount of data to put on log tape
>(Kbytes)>
># Optical
>
>STAGEBLOB ,1 # INFORMIX-OnLine/Optical staging area
>
># System Configuration
>
>SERVERNUM 0 # Unique id corresponding to a OnLine>instance
>DBSERVERNAME unxprod04_infx # Name of default database server>DBSERVERALIASES unxprod04_infx_shm # List of alternate dbservernames
>NETTYPE tlitcp,2,50,NET # Configure poll thread(s) for nettype
>NETTYPE ipcshm,3,50,CPU # Configure poll thread(s) for nettype
>DEADLOCK_TIMEOUT 60 # Max time to wait of lock in>distributed env.
>RESIDENT 0 # Forced residency flag (Yes = 1, No =
>0)
>
>MULTIPROCESSOR 1 # 0 for single-processor, 1 for>multi-processor
>NUMCPUVPS 6 # Number of user (cpu) vps
>SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps>to one
>
>NOAGE 1 # Process aging
>AFF_SPROC 0 # Affinity start processor
>AFF_NPROCS 0 # Affinity number of processors>
># Shared Memory Parameters
>
>LOCKS 500000 # Maximum number of locks
>BUFFERS 300000 # Maximum number of shared buffers
>NUMAIOVPS 2 # Number of IO vps
>PHYSBUFF 32 # Physical log buffer size (Kbytes)
>LOGBUFF 10 # Logical log buffer size (Kbytes)>LOGSMAX 60 # Maximum number of logical log files
>CLEANERS 20 # Number of buffer cleaner processes
>SHMBASE 0xa000000 # Shared memory base address
>SHMVIRTSIZE 32000 # initial virtual shared memory segment>size
>SHMADD 8192 # Size of new shared memory segments
>(Kbytes)
>SHMTOTAL 0 # Total shared memory (Kbytes).
>0=>unlimited
>CKPTINTVL 900 # Check point interval (in sec)
>LRUS 127 # Number of LRU queues
>LRU_MAX_DIRTY 7 # LRU percent dirty begin cleaning limit
>LRU_MIN_DIRTY 5 # LRU percent dirty end cleaning limit
>LTXHWM 50 # Long transaction high water mark>percentage
>LTXEHWM 60 # Long transaction high water mark
>(exclusive)
>TXTIMEOUT 0x12c # Transaction timeout (in sec)
>STACKSIZE 32 # Stack size (Kbytes)>
># System Page Size
># BUFFSIZE - OnLine no longer supports this configuration parameter.
># To determine the page size used by OnLine on your platform
># see the last line of output from the command, 'onstat -b'.
>
>
># Recovery Variables
># OFF_RECVRY_THREADS:
># Number of parallel worker threads during fast recovery or an offline
>restore.
># ON_RECVRY_THREADS:
># Number of parallel worker threads during an online restore.
>
>OFF_RECVRY_THREADS 10 # Default number of offline worker>threads
>ON_RECVRY_THREADS 1 # Default number of online worker>threads
>
># Data Replication Variables
># DRAUTO: 0 manual, 1 retain type, 2 reverse type
>DRAUTO 0 # DR automatic switchover
>DRINTERVAL 30 # DR max time between DR buffer flushes
>(in sec)
>DRTIMEOUT 30 # DR network timeout (in sec)>DRLOSTFOUND /usr/informix/etc/dr.lostfound # DR lost+found file path
>
># CDR Variables
>CDR_LOGBUFFERS 2048 # size of log reading buffer pool
>(Kbytes)
>CDR_EVALTHREADS 1,2 # evaluator threads
>(per-cpu-vp,additional)
>CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
>CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
>queue (Kb
>ytes)>
># Back
The first things to do are change LRU MIN to 1, MAX to 2. Avoid 0 and 1, you'll get too much redundant writing too quickly. Let the page age to at least 1 percent before writing it out. Up CLEANERS to 127, and make sure all of the stuff you're writing out isn't all going to the same disk. If it is, none of the above will help. And it would be consistent with idle cleaners during the checkpoint. On the non checkpoint related front, I'd also up PHYSBUFF to ~1024 if you use unbuffered logging. Is this an OLTP or DSS type system? If OLTP, you probably want OPTCOMPIND = 0. With so much physical memory, I'd think about upping BUFFERS too, though this won't shorten checkpoints, and your hit ratio is fine. SHMVIRTSIZE and SHMADD look very low to me for a "large" system. Good Luck Doug McAllister wrote: > Greetings all..... > > I watch this group every day and sometimes offer my help. Now it is my > turn. >
Greg <greg@fastlane.net> wrote in message news:C2332003A8DE349E.2EBC5ECFCC04FF94.9492CC951EE5F7C9@lp.airnews.net... > The first things to do are change LRU MIN to 1, MAX to 2. Avoid 0 and 1, > you'll get too much redundant writing too quickly. Let the page age to at > least 1 percent before writing it out. Up CLEANERS to 127, and make sure all > of the stuff you're writing out isn't all going to the same disk. If it is, > none of the above will help. And it would be consistent with idle cleaners > during the checkpoint. Since the long checkpoints occur during a batch load of new rows, this may be the culprit. > > On the non checkpoint related front, I'd also up PHYSBUFF to ~1024 if you use > unbuffered logging. Is this an OLTP or DSS type system? If OLTP, you probably > want OPTCOMPIND = 0. With so much physical memory, I'd think about upping > BUFFERS too, though this won't shorten checkpoints, and your hit ratio is > fine. SHMVIRTSIZE and SHMADD look very low to me for a "large" system. > I agree, I had shmvirt set to 1000000 and the vendor of the app that runs here, gim2, said that they needed to parm set low. We had a discussion and I was trying their "suggestion". This is an OLTP system that is loaded by batch. I intend to up cleaners to 127. > Good Luck > > Doug McAllister wrote: > > > Greetings all..... > > > > I watch this group every day and sometimes offer my help. Now it is my > > turn. > > >
David Weis <dweis@louisville-i-market.com> wrote in message
news:w9PU3.33610$23.1749988@typ11.nn.bcandid.com...
> Doug McAllister wrote in message <38234139.F3266E75@fmr.com>...
> >Greetings all.....
> >
> >I watch this group every day and sometimes offer my help. Now it is my
> >turn.
> >
> >I have a system that is generally a lightly loaded system but runs a
> >batch load a couple times a day. During these loads, checkpoints get
> >horrendous averaging 60 seconds to flush about 18,000 dirty buffers.
> >(See below) In a recent conversation with the tech-support guys they
> >suggested that I change my cleaners from 20 to 127 (to match my LRU
> >queues). OK, but during a checkpoint, all of the cleaners are idle
> >anyway. LRU MIN/MAX is set to 5 and 7 respectively.
>
> I would set them to 1 and 0 respectively, then when the checkpoint does
> occur; hopefully, there are less modified pages in the buffers to flush.
> Also, you might lower the number of BUFFERS a tad. Oops, I just checked
> your BUFWAITS, nevermind on lowering the BUFFERS.
>
> Oh, increase NUMCPUVPS to say a value between 8 - 12.
Can't set numcpuvp's to greater than physical which is 8.
>
> and you do have KAIO turned on, right???
yes, Solaris defaults to KAIO.
>
> David Weis
> dweis@louisville-i-market.com
> Voice: 502-637-7316
>
>
>
>
>
>
>
> >
> >This is a big system running Solaris 2.6 with 8gig memory and 8 400mhz
> >processors
> >
> >Any suggestions?
> >
> >
> >tnx..........
> >
> >Informix Dynamic Server Version 7.31.UC3 -- On-Line -- Up 1 days
> >02:55:42 --
> >776448 Kbytes
> >
> >Configuration File: /apps/informix/etc/onconfig.GIM2
>
>#**************************************************************************
> >
> >#
> ># INFORMIX SOFTWARE, INC.
> >#
> ># Title: onconfig.std
> ># Description: INFORMIX-OnLine Configuration Parameters
> >#
>
>#**************************************************************************
> >
> ># Root Dbspace Configuration
> >
> >ROOTNAME rootdbs # Root dbspace name> >ROOTPATH /dev/vx/rdsk/infchunk3 # Path for device containing root
> >dbspace
> >ROOTOFFSET 0 # Offset of root dbspace into device
> >(Kbytes)
> >ROOTSIZE 100000 # Size of root dbspace (Kbytes)> >
> ># Disk Mirroring Configuration Parameters
> >
> >MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> >MIRRORPATH # Path for device containing mirrored> >root
> >MIRROROFFSET 0 # Offset into mirrored device (Kbytes)> >
> ># Physical Log Configuration
> >
> >PHYSDBS phydbs1 # Location (dbspace) of physical log
> >PHYSFILE 75000 # Physical log file size (Kbytes)> >
> ># Logical Log Configuration
> >
> >LOGFILES 50 # Number of logical log files
> >LOGSIZE 5000 # Logical log size (Kbytes)> >
> ># Diagnostics
> >
> >MSGPATH /apps/informix/logs/gim2.log # System message log file
> >path
> >CONSOLE /apps/informix/logs/gim2.console #System console message
> >path
> >ALARMPROGRAM /apps/informix/etc/log_full.sh # Alarm program path
> >
> ># System Archive Tape Device
> >
> >TAPEDEV /dev/rmt/0ub # Tape device path
> >TAPEBLK 128 # Tape block size (Kbytes)
> >TAPESIZE 2000000 # Maximum amount of data to put on tape
> >(Kbytes)> >
> ># Log Archive Tape Device
> >
> >LTAPEDEV /dev/rmt/0ub # Log tape device path
> >LTAPEBLK 16 # Log tape block size (Kbytes)
> >LTAPESIZE 10240 # Max amount of data to put on log tape
> >(Kbytes)> >
> ># Optical
> >
> >STAGEBLOB ,1 # INFORMIX-OnLine/Optical staging area
> >
> ># System Configuration
> >
> >SERVERNUM 0 # Unique id corresponding to a OnLine> >instance
> >DBSERVERNAME unxprod04_infx # Name of default database server> >DBSERVERALIASES unxprod04_infx_shm # List of alternate dbservernames
> >NETTYPE tlitcp,2,50,NET # Configure poll thread(s) for nettype
> >NETTYPE ipcshm,3,50,CPU # Configure poll thread(s) for nettype
> >DEADLOCK_TIMEOUT 60 # Max time to wait of lock in> >distributed env.
> >RESIDENT 0 # Forced residency flag (Yes = 1, No =
> >0)
> >
> >MULTIPROCESSOR 1 # 0 for single-processor, 1 for> >multi-processor
> >NUMCPUVPS 6 # Number of user (cpu) vps
> >SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> >to one
> >
> >NOAGE 1 # Process aging
> >AFF_SPROC 0 # Affinity start processor
> >AFF_NPROCS 0 # Affinity number of processors> >
> ># Shared Memory Parameters
> >
> >LOCKS 500000 # Maximum number of locks
> >BUFFERS 300000 # Maximum number of shared buffers
> >NUMAIOVPS 2 # Number of IO vps
> >PHYSBUFF 32 # Physical log buffer size (Kbytes)
> >LOGBUFF 10 # Logical log buffer size (Kbytes)> >LOGSMAX 60 # Maximum number of logical log files
> >CLEANERS 20 # Number of buffer cleaner processes
> >SHMBASE 0xa000000 # Shared memory base address
> >SHMVIRTSIZE 32000 # initial virtual shared memory segment> >size
> >SHMADD 8192 # Size of new shared memory segments
> >(Kbytes)
> >SHMTOTAL 0 # Total shared memory (Kbytes).
> >0=>unlimited
> >CKPTINTVL 900 # Check point interval (in sec)
> >LRUS 127 # Number of LRU queues
> >LRU_MAX_DIRTY 7 # LRU percent dirty begin cleaning limit
> >LRU_MIN_DIRTY 5 # LRU percent dirty end cleaning limit
> >LTXHWM 50 # Long transaction high water mark> >percentage
> >LTXEHWM 60 # Long transaction high water mark
> >(exclusive)
> >TXTIMEOUT 0x12c # Transaction timeout (in sec)
> >STACKSIZE 32 # Stack size (Kbytes)> >
> ># System Page Size
> ># BUFFSIZE - OnLine no longer supports this configuration parameter.
> ># To determine the page size used by OnLine on your platform
> ># see the last line of output from the command, 'onstat -b'.
> >
> >
> ># Recovery Variables
> ># OFF_RECVRY_THREADS:
> ># Number of parallel worker threads during fast recovery or an offline
> >restore.
> ># ON_RECVRY_THREADS:
> ># Number of parallel worker threads during an online restore.
> >
> >OFF_RECVRY_THREADS 10 # Default number of offline worker> >threads
> >ON_RECVRY_THREADS 1 # Default number of online worker> >threads
> >
> ># Data Replication Variables
> ># DRAUTO: 0 manual, 1 retain type, 2 reverse type
> >DRAUTO 0 # DR automatic switchover
> >DRINTERVAL
After doing all the standard LRU_MAX/LRU_MIN/BUFFERS/CLEANERS/LRUS
tuning on my HP K260 quad processor 7.31UC2, my checkpoints were still
5-10 seconds. I was really starting to wonder what the heck was going
on. Then I got an HP Model 20 disk array with all RAID 10. My
checkpoints are now 1-2 seconds max, and are less than a second most of
the time.
Moral: disk/chunk configuration really counts. Have you allocated
chunks _across_ the drives or _down_ the drives? Get as many spindles
in on the action as you can. Use fragmentation if its available. Check
out detached indexes.
hth
allen
Doug McAllister wrote:
>
> Greetings all.....
>
> I watch this group every day and sometimes offer my help. Now it is my
> turn.
>
> I have a system that is generally a lightly loaded system but runs a
> batch load a couple times a day. During these loads, checkpoints get
> horrendous averaging 60 seconds to flush about 18,000 dirty buffers.
> (See below) In a recent conversation with the tech-support guys they
> suggested that I change my cleaners from 20 to 127 (to match my LRU
> queues). OK, but during a checkpoint, all of the cleaners are idle
> anyway. LRU MIN/MAX is set to 5 and 7 respectively.
>
> This is a big system running Solaris 2.6 with 8gig memory and 8 400mhz
> processors
>
> Any suggestions?
>
> tnx..........
>
> Informix Dynamic Server Version 7.31.UC3 -- On-Line -- Up 1 days
> 02:55:42 --
> 776448 Kbytes
>
> Configuration File: /apps/informix/etc/onconfig.GIM2
> #**************************************************************************
>
> #
> # INFORMIX SOFTWARE, INC.
> #
> # Title: onconfig.std
> # Description: INFORMIX-OnLine Configuration Parameters
> #
> #**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name> ROOTPATH /dev/vx/rdsk/infchunk3 # Path for device containing root
> dbspace
> ROOTOFFSET 0 # Offset of root dbspace into device
> (Kbytes)
> ROOTSIZE 100000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH # Path for device containing mirrored> root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS phydbs1 # Location (dbspace) of physical log
> PHYSFILE 75000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 50 # Number of logical log files
> LOGSIZE 5000 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /apps/informix/logs/gim2.log # System message log file
> path
> CONSOLE /apps/informix/logs/gim2.console #System console message
> path
> ALARMPROGRAM /apps/informix/etc/log_full.sh # Alarm program path
>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt/0ub # Tape device path
> TAPEBLK 128 # Tape block size (Kbytes)
> TAPESIZE 2000000 # Maximum amount of data to put on tape
> (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /dev/rmt/0ub # Log tape device path
> LTAPEBLK 16 # Log tape block size (Kbytes)
> LTAPESIZE 10240 # Max amount of data to put on log tape
> (Kbytes)>
> # Optical
>
> STAGEBLOB ,1 # INFORMIX-OnLine/Optical staging area
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a OnLine> instance
> DBSERVERNAME unxprod04_infx # Name of default database server> DBSERVERALIASES unxprod04_infx_shm # List of alternate dbservernames
> NETTYPE tlitcp,2,50,NET # Configure poll thread(s) for nettype
> NETTYPE ipcshm,3,50,CPU # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in> distributed env.
> RESIDENT 0 # Forced residency flag (Yes = 1, No =
> 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 6 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
>
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 500000 # Maximum number of locks
> BUFFERS 300000 # Maximum number of shared buffers
> NUMAIOVPS 2 # Number of IO vps
> PHYSBUFF 32 # Physical log buffer size (Kbytes)
> LOGBUFF 10 # Logical log buffer size (Kbytes)> LOGSMAX 60 # Maximum number of logical log files
> CLEANERS 20 # Number of buffer cleaner processes
> SHMBASE 0xa000000 # Shared memory base address
> SHMVIRTSIZE 32000 # initial virtual shared memory segment> size
> SHMADD 8192 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
> 0=>unlimited
> CKPTINTVL 900 # Check point interval (in sec)
> LRUS 127 # Number of LRU queues
> LRU_MAX_DIRTY 7 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 5 # LRU percent dirty end cleaning limit
> LTXHWM 50 # Long transaction high water mark> percentage
> LTXEHWM 60 # Long transaction high water mark
> (exclusive)
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 32 # Stack size (Kbytes)>
> # System Page Size
> # BUFFSIZE - OnLine no longer supports this configuration parameter.
> # To determine the page size used by OnLine on your platform
> # see the last line of output from the command, 'onstat -b'.
>
> # Recovery Variables
> # OFF_RECVRY_THREADS:
> # Number of parallel worker threads during fast recovery or an offline
> restore.
> # ON_RECVRY_THREADS:
> # Number of parallel worker threads during an online restore.
>
> OFF_RECVRY_THREADS 10 # Default number of offline worker> threads
> ON_RECVRY_THREADS 1 # Default number of online worker> threads
>
> # Data Replication Variables
> # DRAUTO: 0 manual, 1 retain type, 2 reverse type
> DRAUTO 0 # DR automatic switchover
> DRINTERVAL 30 # DR max time between DR buffer flushes
> (in sec)
> DRTIMEOUT 30 # DR network timeout (in sec)> DRLOSTFOUND /usr/informix/etc/dr.lostfound # DR lost+found file path
>
> # CDR Variables
> CDR_LOGBUFFERS 2048 # size of log reading buffer pool
> (Kbytes)
> CDR_EVALTHRE
allenj wrote: > > After doing all the standard LRU_MAX/LRU_MIN/BUFFERS/CLEANERS/LRUS > tuning on my HP K260 quad processor 7.31UC2, my checkpoints were still > 5-10 seconds. I was really starting to wonder what the heck was going > on. Then I got an HP Model 20 disk array with all RAID 10. My > checkpoints are now 1-2 seconds max, and are less than a second most of > the time. > > Moral: disk/chunk configuration really counts. Have you allocated > chunks _across_ the drives or _down_ the drives? Get as many spindles > in on the action as you can. Use fragmentation if its available. Check > out detached indexes. [SNIP] Another happy RAID10 convert, YES! I just got word that the manager responsible for specing our systems has recently realized that I have been right all these years and many of his disk problems are RAID5 related. The sysadmins tell me they have orders to use ONLY RAID10 for all our arrays, not just Informix drives. Another hard won victory! Art S. Kagel
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix