Re: Res: long checkpoints on informix
Posted in 2007
On 17 jul, 15:43, Cesar Inacio Martins
<cesar_inacio_mart...@yahoo.com.br> wrote:
> Hi Sakura,
>
> I wrote a shell to monitor my checkpoints and discover what dbspace/chunk is take more time.
>
> Copy the code below into a file on your unix , execute a chmod 700 and execute with informix... when occur a checkpoint , it will display to you the dbspaces/chunks where is written second by second...
> Works with IDS 7.31 , 9.40 and 10 and Unix/Linux
>
> Sorry my poor english...
>
> #!/bin/sh
> ####
> #### by Cesar Inacio Martins
> #####################################################################
>
> vTmp1=/tmp/tmp1.$$
> vTmp2=/tmp/tmp2.$$
> vTmp3=/tmp/tmp3.$$
> vTmp4=/tmp/tmp4.$$
>
> ## get the dbspaces names...
> onstat -d |awk '$1 == "Dbspaces",/active/ { match($0," [a-zA-Z0-9_\\-]*$") ; X=substr($0,RSTART,RLENGTH); print $2,X } ' 2>/dev/null | tail -n +3 | grep -v active >$vTmp2
> ## get the chunks
> onstat -d |awk '$1 == "Chunks",/active/ { print $2,$3,$8 } ' | tail -n +3 | grep -v active >$vTmp3>
> echo '
> BEGIN {
> while ( ( getline < lArqDB ) > 0 ) {
> vDB[$1]=$2
> }
> close(lArqDB)
> while ( ( getline < lArqChk ) > 0 ) {
> vChk[$1]=$3
> vChkDB[$1]=$2
> }
> close(lArqChk)
>
> }
>
> /Informix.*CKPT/ || ( length($3) == 1 ) {
> if ( $3 == "C" ) print $0,vDB[vChkDB[$4]],vChk[$4] ;
> if ( $3 == "L" ) print $0,"LRU-"$4 ;
> }
> ' >$vTmp4
>
> while sleep 1
> do
> onstat -F | awk -v lArqDB=$vTmp2 -v lArqChk=$vTmp3 -f $vTmp4 > $vTmp1
> if [ -s $vTmp1 ]; then
> echo " "
> date +"%d/%m/%Y %H:%M"
> cat $vTmp1 | awk '{print " "$0 } '
> onstat -m |grep -i "checkpoint.*seconds" | tail -n1
> fi> done
>
> rm -f $vTmp1 $vTmp2 $vTmp3 $vTmp4
> ########################
>
> ----- Mensagem original ----
> De: "Sakura.g...@gmail.com" <Sakura.g...@gmail.com>
> Para: informix-l...@iiug.org
> Enviadas: Terça-feira, 17 de Julho de 2007 12:44:15
> Assunto: Re: long checkpoints on informix
>
> On 16 jul, 06:34, Ben Thompson <b...@nomonitorsoftspam.com> wrote:
>
>
>
> > Sakura.g...@gmail.com wrote:
> > > hi everyone
> > > i recently start to have some long and constant checkpoint
> > > i did lot of change but cant reduce it
> > > i have Informix Dynamic Server Version 9.40.FC2
> > > and here is some parameter of the onconfig
> > > ---------------------------------------------------------------------------------------------------------
>
> > > TBLSPACE_STATS 1 # Maintain tblspace statistics
> > > MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-> > > processor
> > > NUMCPUVPS 2 # Number of user (cpu) vps
> > > SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> > > to one
> > > NOAGE 1 # Process aging
> > > AFF_SPROC 0 # Affinity start processor
> > > AFF_NPROCS 0 # Affinity number of processors
> > > LOCKS 100000 # Maximum number of locks
> > > BUFFERS 200000 # Maximum number of shared buffers
> > > NUMAIOVPS 1 # Number of IO vps
> > > CLEANERS 127 # Number of buffer cleaner processes
> > > SHMBASE 0x10A000000L # Shared memory base address
> > > SHMVIRTSIZE 300000 # initial virtual shared memory> > > segment size
> > > SHMADD 16384 # Size of new shared memory segments
> > > (Kbytes)
> > > CKPTINTVL 600 # Check point interval (in sec)(10
> > > minutos)
> > > LRUS 127 # Number of LRU queues
> > > LRU_MAX_DIRTY 2.000000 # LRU percent dirty begin cleaning> > > limit
> > > LRU_MIN_DIRTY 1.000000 # LRU percent dirty end cleaning limit
> > > DYNAMIC_LOGS 2
> > > OFF_RECVRY_THREADS 10 # Default number of offline worker> > > threads
> > > ON_RECVRY_THREADS 1 # Default number of online worker> > > threads
> > > CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-
> > > vp,additional)
> > > CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
> > > CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
> > > queue (Kbytes)
> > > RA_PAGES 32 # Number of pages to attempt> > > to read ahead
> > > RA_THRESHOLD 30 # Number of pages left before> > > next group
> > > MAX_PDQPRIORITY 0 # Maximum allowed pdqpriority
> > > DS_MAX_QUERIES # Maximum number of decision support> > > queries
> > > DS_TOTAL_MEMORY # Decision support memory (Kbytes)
> > > DS_MAX_SCANS 1048576 # Maximum number of decision support> > > scans
> > > OPTCOMPIND 2 # To hint the optimizer
> > > ---------------------------------------------------------------------------------------------------------> > > Profile
> > > dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> > > 69174631 143507619 16152454169 99.57 10204375 98766167 29414237
> > > 65.31
>
> > Some comments and things to try that may be helpful, may not.
>
> > You're running 9.40.FC2 so you clearly have a 64-bit system. Which
> > platform? How much memory do you have? What are your processors? What is
> > the disc controller and RAID level?
>
> > You've got 65% cached writes which is pretty low. Your other post shows
> > around 70 io operations/sec which is quite high so your system is busy.
>
> > Your BUFFERS are just 200000 which is just 400Mb or 800Mb depending on
> > your page size, plus you've got 300Mb of shared memory. It's not a lot
> > really by today's standards and not for a 64-bit system. If you have the
> > memory, have you tried increasing these values? It would get your cached
> > writes up.
>
> > You've got two CPU VPs and just one AIO VP. You could definitely have
> > more AIO VPs and maybe you could try two CPU VPs per processor (if you
> > have a multi-processor system on modern processors).
>
> > You've got a lot of LRUs, presumably to keep checkpoints down, but
> > mostly chunk writes which is a little confusing. Maybe others can help
> > here. However performance is poor. Is your RAID set healthy?
>
> > Your read-ahead values are quite high. Maybe you could reduce these to
> > cut down the number of pages read in. Obviously there is a trade-off
> > here with the number of disc access requests requests required so
> > perhaps a little tuning?
>
> > You've got a lot of writes so presumably you've got an OLTP set up?
> > Perhaps try changing OPTCOMPIND to '0' (zero).
>
> > Try monitoring using "onstat -F" during the checkpoint by running it
> > every second for analysis later. You have only posted truncated "onstat
> > -F" output.
>
> > Also check locks using "onstat -k" and cross-reference with "onstat -u".
>
> > Regards, Ben.
>
> SunOS sun4u sparc SUNW,Sun-Fire-V240
>