Re: checkpoints taking as much as 25 secs
Posted in 1999
Great answers from Peter, Obnoxio, and Lloyd. Yes you are seeing a bufwaits
ratio of around 40% which is horrific, we need to try to get that under 10%
with a preference for a value under 7%. Take Lloyd's or Obnoxio's
recommendations for LRUS and CLEANERS and the LRU_MIN/MAX_DIRTY parameters
but also look at adding more buffers. You have 42,000 buffers but in the
21 hours since the engine was bounced (that also resets the stats BTW) you
have used over 6,148,644 buffers (pagreads plus any newly created pages)
which means that on average you are turning over the entire buffer pool
every 10 minutes! It looks like a combination of slow disks (yes RAID5 can
be a bottleneck however your level of writes, as you say, is low so this is
not the big problem) and buffer thrashing. Also you have 384000 sequential
scans with works out to five per second, if nothing else all those sequential
scans are a big part of the buffer thrashing! You need to look at your
database design vis-a-vis how your applications are using it and add some
indexes somewhere.
Art S. Kagel
Tam McLaughlin wrote:
>
> The online log file shows that some checkpoints can take from 4 secs up to
> 25 secs
> but the average is prob 0 secs roughly every 5 mins. The long checkpoint
> times prob
> occur when some large progs are running e.g. cron jobs. I have not been on
> an IDS7
> admin course yet so not too sure what to look for. Our system is:
>
> IDS 7.30.UC2 on SCO UnixWare 7.0.1 with 3 pentium pro processors running
> RAID 5.
> We have approx 10 Gb of dbspace used and the largest table is ~ 3 Gb and our
> app
> is mostly read only.
>
> I realise raid is not the best option and we may have a disk i/o bottleneck.
> It may be that
> my question is too general and that it could take a long time to work out
> why we are getting
> these long checkpoints but if anyone has some clues where to start or what
> to check it may
> help me a great deal.
>
> Some info may help here:
>
> onstat -p : the bufwaits look high (i done an onstat -z a few days ago) but> this may be normal
>
> Informix Dynamic Server Version 7.30.UC2 -- On-Line -- Up 21:14:51 --
> 114040 Kbytes
>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 26169757 6148644 1242311687 97.89 140931 242136 5077195 97.22
>
> isamtot open start read write rewrite delete commit
> rollbk
> 162742135 3237700 9110352 121063233 2131212 225000 50794 40789
> 5439
>
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 0 0 0 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 0 37389.09 6435.06 214 508
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 4750714 2 633141750 0 0 112 29328 384018
>
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 15832893 12764 5910997 21753370 168421
>
> my onconfig file is:
>
> ROOTNAME rootdbs> ROOTPATH /dev/rdsk/rootdbs
> ROOTOFFSET 0
> ROOTSIZE 512000> MIRROR 0
> MIRRORPATH
> MIRROROFFSET 0> PHYSDBS rootdbs
> PHYSFILE 10000
> LOGFILES 21
> LOGSIZE 5000> MSGPATH /users/informix/pluto.log
> CONSOLE /users/informix/pluto_console.log
> ALARMPROGRAM /users/informix/etc/log_full.sh
> SYSALARMPROGRAM /users/informix/etc/evidence.sh
> TBLSPACE_STATS 1> STAGEBLOB
> SERVERNUM 0
> DBSERVERNAME pluto
> DBSERVERALIASES pluto_net
> NETTYPE ipcshm,1,64,CPU
> NETTYPE tlitcp,1,10,NET
> DEADLOCK_TIMEOUT 60
> RESIDENT 0
> MULTIPROCESSOR 1
> NUMCPUVPS 3
> SINGLE_CPU_VP 0
> NOAGE 0
> AFF_SPROC 0
> AFF_NPROCS 0
> LOCKS 20000
> BUFFERS 42000
> NUMAIOVPS 16
> PHYSBUFF 120
> LOGBUFF 40> LOGSMAX 60
> CLEANERS 6
> SHMBASE 0xa000000
> SHMVIRTSIZE 24000
> SHMADD 8192
> SHMTOTAL 0
> CKPTINTVL 300
> LRUS 8
> LRU_MAX_DIRTY 20
> LRU_MIN_DIRTY 15
> LTXHWM 45
> LTXEHWM 50
> TXTIMEOUT 0x12c
> STACKSIZE 32
> OFF_RECVRY_THREADS 10
> ON_RECVRY_THREADS 1
> DRAUTO 0
> DRINTERVAL 30
> DRTIMEOUT 30
> DRLOSTFOUND /users/informix/etc/dr.lostfound> CDR_LOGBUFFERS 2048
> CDR_EVALTHREADS 1,2
> CDR_DSLOCKWAIT 5
> CDR_QUEUEMEM 4096
> BAR_ACT_LOG /tmp/bar_act.log
> BAR_MAX_BACKUP 0
> BAR_RETRY 1
> BAR_NB_XPORT_COUNT 10
> BAR_XFER_BUF_SIZE 31> ISM_DATA_POOL ISMData
> ISM_LOG_POOL ISMLogs
> RA_PAGES
> RA_THRESHOLD
> DBSPACETEMP tempspace
> FILLFACTOR 90
> USEOSTIME 0
> MAX_PDQPRIORITY 100
> DS_MAX_QUERIES
> DS_TOTAL_MEMORY
> DS_MAX_SCANS 1048576
> DATASKIP off
> OPTCOMPIND 0
> ONDBSPACEDOWN 2> LBU_PRESERVE 0
> OPCACHEMAX 0
> HETERO_COMMIT 0
> OPT_GOAL -1
> DIRECTIVES 1
> RESTARTABLE_RESTORE off