Res: long checkpoints on informix
Posted in 2007
Hi Sakura,
I wrote a shell to monitor my checkpoints and discover what dbspace/chunk is take more time.
Copy the code below into a file on your unix , execute a chmod 700 and execute with informix... when occur a checkpoint , it will display to you the dbspaces/chunks where is written second by second...
Works with IDS 7.31 , 9.40 and 10 and Unix/Linux
Sorry my poor english...
#!/bin/sh
####
#### by Cesar Inacio Martins
#####################################################################
vTmp1=/tmp/tmp1.$$
vTmp2=/tmp/tmp2.$$
vTmp3=/tmp/tmp3.$$
vTmp4=/tmp/tmp4.$$
## get the dbspaces names...
onstat -d |awk '$1 == "Dbspaces",/active/ { match($0," [a-zA-Z0-9_\\-]*$") ; X=substr($0,RSTART,RLENGTH); print $2,X } ' 2>/dev/null | tail -n +3 | grep -v active >$vTmp2
## get the chunks
onstat -d |awk '$1 == "Chunks",/active/ { print $2,$3,$8 } ' | tail -n +3 | grep -v active >$vTmp3
echo '
BEGIN {
while ( ( getline < lArqDB ) > 0 ) {
vDB[$1]=$2
}
close(lArqDB)
while ( ( getline < lArqChk ) > 0 ) {
vChk[$1]=$3
vChkDB[$1]=$2
}
close(lArqChk)
}
/Informix.*CKPT/ || ( length($3) == 1 ) {
if ( $3 == "C" ) print $0,vDB[vChkDB[$4]],vChk[$4] ;
if ( $3 == "L" ) print $0,"LRU-"$4 ;
}
' >$vTmp4
while sleep 1
do
onstat -F | awk -v lArqDB=$vTmp2 -v lArqChk=$vTmp3 -f $vTmp4 > $vTmp1
if [ -s $vTmp1 ]; then
echo " "
date +"%d/%m/%Y %H:%M"
cat $vTmp1 | awk '{print " "$0 } '
onstat -m |grep -i "checkpoint.*seconds" | tail -n1
fidone
rm -f $vTmp1 $vTmp2 $vTmp3 $vTmp4
########################
----- Mensagem original ----
De: "Sakura.ggxx@gmail.com" <Sakura.ggxx@gmail.com>
Para: informix-list@iiug.org
Enviadas: Terça-feira, 17 de Julho de 2007 12:44:15
Assunto: Re: long checkpoints on informix
On 16 jul, 06:34, Ben Thompson <b...@nomonitorsoftspam.com> wrote:
> Sakura.g...@gmail.com wrote:
> > hi everyone
> > i recently start to have some long and constant checkpoint
> > i did lot of change but cant reduce it
> > i have Informix Dynamic Server Version 9.40.FC2
> > and here is some parameter of the onconfig
> > ---------------------------------------------------------------------------------------------------------
>
> > TBLSPACE_STATS 1 # Maintain tblspace statistics
> > MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-> > processor
> > NUMCPUVPS 2 # Number of user (cpu) vps
> > SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> > to one
> > NOAGE 1 # Process aging
> > AFF_SPROC 0 # Affinity start processor
> > AFF_NPROCS 0 # Affinity number of processors
> > LOCKS 100000 # Maximum number of locks
> > BUFFERS 200000 # Maximum number of shared buffers
> > NUMAIOVPS 1 # Number of IO vps
> > CLEANERS 127 # Number of buffer cleaner processes
> > SHMBASE 0x10A000000L # Shared memory base address
> > SHMVIRTSIZE 300000 # initial virtual shared memory> > segment size
> > SHMADD 16384 # Size of new shared memory segments
> > (Kbytes)
> > CKPTINTVL 600 # Check point interval (in sec)(10
> > minutos)
> > LRUS 127 # Number of LRU queues
> > LRU_MAX_DIRTY 2.000000 # LRU percent dirty begin cleaning> > limit
> > LRU_MIN_DIRTY 1.000000 # LRU percent dirty end cleaning limit
> > DYNAMIC_LOGS 2
> > OFF_RECVRY_THREADS 10 # Default number of offline worker> > threads
> > ON_RECVRY_THREADS 1 # Default number of online worker> > threads
> > CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-
> > vp,additional)
> > CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
> > CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
> > queue (Kbytes)
> > RA_PAGES 32 # Number of pages to attempt> > to read ahead
> > RA_THRESHOLD 30 # Number of pages left before> > next group
> > MAX_PDQPRIORITY 0 # Maximum allowed pdqpriority
> > DS_MAX_QUERIES # Maximum number of decision support> > queries
> > DS_TOTAL_MEMORY # Decision support memory (Kbytes)
> > DS_MAX_SCANS 1048576 # Maximum number of decision support> > scans
> > OPTCOMPIND 2 # To hint the optimizer
> > ---------------------------------------------------------------------------------------------------------> > Profile
> > dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> > 69174631 143507619 16152454169 99.57 10204375 98766167 29414237
> > 65.31
>
> Some comments and things to try that may be helpful, may not.
>
> You're running 9.40.FC2 so you clearly have a 64-bit system. Which
> platform? How much memory do you have? What are your processors? What is
> the disc controller and RAID level?
>
> You've got 65% cached writes which is pretty low. Your other post shows
> around 70 io operations/sec which is quite high so your system is busy.
>
> Your BUFFERS are just 200000 which is just 400Mb or 800Mb depending on
> your page size, plus you've got 300Mb of shared memory. It's not a lot
> really by today's standards and not for a 64-bit system. If you have the
> memory, have you tried increasing these values? It would get your cached
> writes up.
>
> You've got two CPU VPs and just one AIO VP. You could definitely have
> more AIO VPs and maybe you could try two CPU VPs per processor (if you
> have a multi-processor system on modern processors).
>
> You've got a lot of LRUs, presumably to keep checkpoints down, but
> mostly chunk writes which is a little confusing. Maybe others can help
> here. However performance is poor. Is your RAID set healthy?
>
> Your read-ahead values are quite high. Maybe you could reduce these to
> cut down the number of pages read in. Obviously there is a trade-off
> here with the number of disc access requests requests required so
> perhaps a little tuning?
>
> You've got a lot of writes so presumably you've got an OLTP set up?
> Perhaps try changing OPTCOMPIND to '0' (zero).
>
> Try monitoring using "onstat -F" during the checkpoint by running it
> every second for analysis later. You have only posted truncated "onstat
> -F" output.
>
> Also check locks using "onstat -k" and cross-reference with "onstat -u".
>
> Regards, Ben.
SunOS sun4u sparc SUNW,Sun-Fire-V240
Memory size: 4096 Megabytes
RAID5
i guest is a OLTP system
> However performance is poor. Is your RAID set healthy?
i guest is working right, by the way, how can i check this ?
i dont find the way to check it
> Try monitoring using "onstat -F" during the checkpoint by running it
> every second for analysis later. You have only posted truncated @