performance with very active threads
Posted in 2000
An insurance-company DBA on IDS 7.30 (UnixWare, 3 CPUs, 764MB RAM) saw intermittent OLTP slowness with checkpoints lasting 20-30 seconds, apparently when long DSS/report jobs ran. Advice: too few page cleaners (8 cleaners for 63 LRUs = several passes) - set CLEANERS >= LRUS (some suggested 65/65 to dodge a power-of-2 bug), lengthen CKPTINTVL (600s-30min rather than the onconfig.std sample 300s) and enlarge rather than shrink PHYSFILE, move physical/logical logs out of rootdbs, turn off TBLSPACE_STATS, set RESIDENT 1 and NOAGE, add buffers/RA_PAGES, use multiple temp dbspaces and NO LOG temp tables, and apply PDQ to reports. The poster questioned the long checkpoint interval; no confirmed outcome is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Logging & Checkpoints
I have some intermittent performance problems with our instance that lead
users to
complain their programs are running slow.
The users mainly run OLTP type programs (we are an insurance company
processing claims etc)
and most of the time, checkpoint duration < 0 but at certain times, the
interval can be as as ,much
as 20-30 seconds. I now believe this is probably due to some DSS type
programs that run via cron and by some users. These programs can create some
very large queries (170 line sql from onstat -g ses) and look through a
number of tables to produce reports and large (20-200Mb) delimited files.
Now, there are a number of things I can tune and will do so this weekend
i.e. get away from RAID 5 and use RAID 10, place my tables in different
dbspaces, put logs in different dbspace etc etc.
However, when I run: onstat -g acr -r 1, there apart from the oninit
threads, there are a few threads
that are constantly working away. But so what ? is this good/bad ? could it
be that these are intensive programs that require more work rather than the
casual select from a user program? So what do I look
for ? what can i check to see what can be tuned. From onstat -p, my
read/write ratios are good, I believe I have plenty of buffers etc. I tried
putting setting PDQPRIORITY=40 in a script but onstat -g mgm showed no
queries active. I have about 20 5 Mb logical log files which can get filled
up over a few days.
So, I have a few programs that run for a few hours searching thru the
database. I hink most are using indexes as some of our tables have several
indexes on them. I cannot get the developers to use sqexplain. The unix box
is not used for anything else. Any ideas where I should be looking ? I take
it I should be trying to get these progs to use PDQ.
Details: 3 intel pent pro processors, 764 Mb Ram, sco unixware 7.0.1, IDS
7.30
onstat -p
Informix Dynamic Server Version 7.30.UC10X2 -- On-Line -- Up 8 days
06:10:09 -- 323080 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
202085809 5929898 1316346513 84.65 2332438 2278714 480555135 99.51
isamtot open start read write rewrite delete commit
rollbk 1916535217 24709189 69257312 1375228628 233538432 854964 21011
792707 48174
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 271449.86 33836.99 1836 4740
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
45003781 157 2881589552 2 0 997 100381 2362772
ixda-RA idx-RA da-RA RA-pgsused lchwaits
107363614 57987 58116582 165258340 1112639
onstat -g ioq
AIO I/O queues:q name/id len maxlen totalops dskread dskwrite dskcopy
adt 0 0 0 0 0 0 0
msc 0 0 2 58243 0 0 0
aio 0 0 7 29811 14892 6 0
pio 0 0 1 6756 0 6756 0
lio 0 0 1 804868 0 804868 0
gfd 3 0 160 461550 355892 105658 0
gfd 4 0 177 24085993 23955385 130608 0
gfd 5 0 177 86571328 86335846 235482 0
gfd 6 0 164 71626634 71548062 78572 0
gfd 7 0 172 8899865 8796684 103181 0
gfd 8 0 166 1056297 494293 562004 0
gfd 9 0 151 7142010 7048660 93350 0
gfd 10 0 164 901532 829320 72212 0
gfd 11 0 182 2874388 2734613 139775 0
I forgot to include my onconfig file. From a posting i read the other day, I
can see there a few things
I can do the next time I restart the engine: unset TBLSPACE_STATS and reduce
my PHYSFILE size.
PHYSDBS rootdbs
PHYSFILE 10000
LOGFILES 21
LOGSIZE 5000
TBLSPACE_STATS 1TAPEDEV /dev/rmt/ctape3
TAPEBLK 32
TAPESIZE 8000000LTAPEDEV /dev/rmt/ctape2
LTAPEBLK 16
LTAPESIZE 525000STAGEBLOB
SERVERNUM 0
DBSERVERNAME pluto
DBSERVERALIASES pluto_net
NETTYPE ipcshm,3,64,CPU
NETTYPE tlitcp,1,10,NET
DEADLOCK_TIMEOUT 60
RESIDENT 0
MULTIPROCESSOR 1
NUMCPUVPS 3
SINGLE_CPU_VP 0
NOAGE 0
AFF_SPROC 0
AFF_NPROCS 0
LOCKS 20000
BUFFERS 96000
NUMAIOVPS 8
PHYSBUFF 120
LOGBUFF 26LOGSMAX 60
CLEANERS 8
SHMBASE 0xa000000
SHMVIRTSIZE 120000
SHMADD 32000
SHMTOTAL 0
CKPTINTVL 300
LRUS 63
LRU_MAX_DIRTY 2
LRU_MIN_DIRTY 1
LTXHWM 40
LTXEHWM 45
TXTIMEOUT 0x12c
STACKSIZE 32
OFF_RECVRY_THREADS 10
ON_RECVRY_THREADS 1
DRAUTO 0
DRINTERVAL 30
DRTIMEOUT 30
CDR_LOGBUFFERS 2048
CDR_EVALTHREADS 1,2
CDR_DSLOCKWAIT 5
CDR_QUEUEMEM 4096
BAR_ACT_LOG /tmp/bar_act.log
BAR_MAX_BACKUP 0
BAR_RETRY 1
BAR_NB_XPORT_COUNT 10
BAR_XFER_BUF_SIZE 31ISM_DATA_POOL ISMData
ISM_LOG_POOL ISMLogs
RA_PAGES
RA_THRESHOLD
DBSPACETEMP tempspace
DUMPDIR /ascii
DUMPSHMEM 1
DUMPGCORE 0
DUMPCORE 0
DUMPCNT 1
FILLFACTOR 90
USEOSTIME 0
MAX_PDQPRIORITY 100
DS_MAX_QUERIES
DS_TOTAL_MEMORY
DS_MAX_SCANS 1048576
DATASKIP off
OPTCOMPIND 0
ONDBSPACEDOWN 2
LBU_PRESERVE 1OPCACHEMAX 0
HETERO_COMMIT 0
OPT_GOAL -1
DIRECTIVES 1
RESTARTABLE_RESTORE off
IO contention could be a problem, but the checkpoints at 20-30 seconds
is a definite impact on the OLTP users. You should attack this first.
First, the most likely reason for this is that there are a lot of dirty
buffers in the buffer cache so there is more work to be done. The page
cleaners have not had time to flush these pages in the background
because the CKPTINTVL is too frequent or the PHYSFILE is too small.
Increase both of these (I like a checkpoint to occur every 30 minutes or
more, let the PHYSFILE be large enough so it doesn't kick off a CP prior
to this). Current LRUS, page cleaners, and chunk settings are needed in
order to make better recommendations. This could be a problem too.
Second, who's modifying all those pages? My guess is its those big
reports. Reports are typically RO, but if you grep them for temp you'll
probably find a lot of temp tables being used. Do you have a
DBSPACETEMP defined? Are reports creating temp tables with NO LOG
option? That will help a lot.
PDQ will help these reports a lot. It will get them out of the buffer
pool and reduce contention with the OLTP users.
You're also only using 50% of available RAM for Informix. Give 100MB to
the OS and the rest to Informix.
What's your concurrency controls for oltp and reports? Run onstat -g
sql. The OLTP slowness could be attributed to waiting on locks held.
You should also upgrade to 7.31.
Ray
Tam McLaughlin wrote:
>
> I have some intermittent performance problems with our instance that lead
> users to
> complain their programs are running slow.
>
> The users mainly run OLTP type programs (we are an insurance company
> processing claims etc)
> and most of the time, checkpoint duration < 0 but at certain times, the
> interval can be as as ,much
> as 20-30 seconds. I now believe this is probably due to some DSS type
> programs that run via cron and by some users. These programs can create some
> very large queries (170 line sql from onstat -g ses) and look through a
> number of tables to produce reports and large (20-200Mb) delimited files.
>
> Now, there are a number of things I can tune and will do so this weekend
> i.e. get away from RAID 5 and use RAID 10, place my tables in different
> dbspaces, put logs in different dbspace etc etc.
>
> However, when I run: onstat -g acr -r 1, there apart from the oninit
> threads, there are a few threads
> that are constantly working away. But so what ? is this good/bad ? could it
> be that these are intensive programs that require more work rather than the
> casual select from a user program? So what do I look
> for ? what can i check to see what can be tuned. From onstat -p, my
> read/write ratios are good, I believe I have plenty of buffers etc. I tried
> putting setting PDQPRIORITY=40 in a script but onstat -g mgm showed no
> queries active. I have about 20 5 Mb logical log files which can get filled
> up over a few days.
>
> So, I have a few programs that run for a few hours searching thru the
> database. I hink most are using indexes as some of our tables have several
> indexes on them. I cannot get the developers to use sqexplain. The unix box
> is not used for anything else. Any ideas where I should be looking ? I take
> it I should be trying to get these progs to use PDQ.
>
> Details: 3 intel pent pro processors, 764 Mb Ram, sco unixware 7.0.1, IDS
> 7.30
>
> onstat -p
> Informix Dynamic Server Version 7.30.UC10X2 -- On-Line -- Up 8 days
> 06:10:09 -- 323080 Kbytes>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 202085809 5929898 1316346513 84.65 2332438 2278714 480555135 99.51
> isamtot open start read write rewrite delete commit
> rollbk 1916535217 24709189 69257312 1375228628 233538432 854964 21011
> 792707 48174
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 0 0 0 0 0 0 0
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 0 271449.86 33836.99 1836 4740
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 45003781 157 2881589552 2 0 997 100381 2362772
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 107363614 57987 58116582 165258340 1112639
>
> onstat -g ioq
> AIO I/O queues:> q name/id len maxlen totalops dskread dskwrite dskcopy
> adt 0 0 0 0 0 0 0
> msc 0 0 2 58243 0 0 0
> aio 0 0 7 29811 14892 6 0
> pio 0 0 1 6756 0 6756 0
> lio 0 0 1 804868 0 804868 0
> gfd 3 0 160 461550 355892 105658 0
> gfd 4 0 177 24085993 23955385 130608 0
> gfd 5 0 177 86571328 86335846 235482 0
> gfd 6 0 164 71626634 71548062 78572 0
> gfd 7 0 172 8899865 8796684 103181 0
> gfd 8 0 166 1056297 494293 562004 0
> gfd 9 0 151 7142010 7048660 93350 0
> gfd 10 0 164 901532 829320 72212 0
> gfd 11 0 182 2874388 2734613 139775 0
During a checkpoint you only have 8 cleaners utilized to flush 63
LRU's. That's 4 passes. Make LRUS=64 and CLEANERS=64. That's one
pass. Let the disk subsystem figure out any contention. They're pretty
good at that:) Also increase CKPTINTVL to every 30 minutes. Why have a
20-30 second interruption every 5 minutes when you can have it every 1/2
hour or every hour! Don't reduce the PHYSFILE. Make it bigger. If you
make it smaller, you will only get more frequent checkpoints and OLTP
interruptions.
Ray
Tam McLaughlin wrote:
>
> I forgot to include my onconfig file. From a posting i read the other day, I
> can see there a few things
> I can do the next time I restart the engine: unset TBLSPACE_STATS and reduce
> my PHYSFILE size.
>
> PHYSDBS rootdbs
> PHYSFILE 10000
> LOGFILES 21
> LOGSIZE 5000
> TBLSPACE_STATS 1> TAPEDEV /dev/rmt/ctape3
> TAPEBLK 32
> TAPESIZE 8000000> LTAPEDEV /dev/rmt/ctape2
> LTAPEBLK 16
> LTAPESIZE 525000> STAGEBLOB
> SERVERNUM 0
> DBSERVERNAME pluto
> DBSERVERALIASES pluto_net
> NETTYPE ipcshm,3,64,CPU
> NETTYPE tlitcp,1,10,NET
> DEADLOCK_TIMEOUT 60
> RESIDENT 0
> MULTIPROCESSOR 1
> NUMCPUVPS 3
> SINGLE_CPU_VP 0
> NOAGE 0
> AFF_SPROC 0
> AFF_NPROCS 0
> LOCKS 20000
> BUFFERS 96000
> NUMAIOVPS 8
> PHYSBUFF 120
> LOGBUFF 26> LOGSMAX 60
> CLEANERS 8
> SHMBASE 0xa000000
> SHMVIRTSIZE 120000
> SHMADD 32000
> SHMTOTAL 0
> CKPTINTVL 300
> LRUS 63
> LRU_MAX_DIRTY 2
> LRU_MIN_DIRTY 1
> LTXHWM 40
> LTXEHWM 45
> TXTIMEOUT 0x12c
> STACKSIZE 32
> OFF_RECVRY_THREADS 10
> ON_RECVRY_THREADS 1
> DRAUTO 0
> DRINTERVAL 30
> DRTIMEOUT 30
> CDR_LOGBUFFERS 2048
> CDR_EVALTHREADS 1,2
> CDR_DSLOCKWAIT 5
> CDR_QUEUEMEM 4096
> BAR_ACT_LOG /tmp/bar_act.log
> BAR_MAX_BACKUP 0
> BAR_RETRY 1
> BAR_NB_XPORT_COUNT 10
> BAR_XFER_BUF_SIZE 31> ISM_DATA_POOL ISMData
> ISM_LOG_POOL ISMLogs
> RA_PAGES
> RA_THRESHOLD
> DBSPACETEMP tempspace
> DUMPDIR /ascii
> DUMPSHMEM 1
> DUMPGCORE 0
> DUMPCORE 0
> DUMPCNT 1
> FILLFACTOR 90
> USEOSTIME 0
> MAX_PDQPRIORITY 100
> DS_MAX_QUERIES
> DS_TOTAL_MEMORY
> DS_MAX_SCANS 1048576
> DATASKIP off
> OPTCOMPIND 0
> ONDBSPACEDOWN 2
> LBU_PRESERVE 1> OPCACHEMAX 0
> HETERO_COMMIT 0
> OPT_GOAL -1
> DIRECTIVES 1
> RESTARTABLE_RESTORE off
Ray Canuel wrote in message <39B6574D.1213E610@informix.com>... > >During a checkpoint you only have 8 cleaners utilized to flush 63 >LRU's. That's 4 passes. Make LRUS=64 and CLEANERS=64. That's one >pass. Let the disk subsystem figure out any contention. They're pretty >good at that:) Also increase CKPTINTVL to every 30 minutes. Why have a >20-30 second interruption every 5 minutes when you can have it every 1/2 >hour or every hour! Don't reduce the PHYSFILE. Make it bigger. If you >make it smaller, you will only get more frequent checkpoints and OLTP >interruptions. > >Ray > > Thanks for the help. Informix always reccomends checkpoint interval every 5 mins so 1/2 hr seems a long time betweem checkpoints which also means there there is more risk of data loss if something goes wrong before the next checkpoint.
See comments below
--
Tony Flaherty
Snr. A/P, Informix DBA, HpUx Admin, Gimmi a broom!
MFS Ltd.
These comments are all mine, and do not represent advice from MFS Ltd in any
way.
Tam McLaughlin wrote in message <8p5j0q$40n$1@plutonium.btinternet.com>...
>I forgot to include my onconfig file. From a posting i read the other day,
I
>can see there a few things
>I can do the next time I restart the engine: unset TBLSPACE_STATS and
reduce
>my PHYSFILE size.
>
>PHYSDBS rootdbs
Put the physical log in its own dbspace
>PHYSFILE 10000
>LOGFILES 21
>LOGSIZE 5000
if these are also in the rootdbs then move them
>TBLSPACE_STATS 1
as you say turn these off
>TAPEDEV /dev/rmt/ctape3
>TAPEBLK 32
>TAPESIZE 8000000>LTAPEDEV /dev/rmt/ctape2
>LTAPEBLK 16
>LTAPESIZE 525000>STAGEBLOB
>SERVERNUM 0
> DBSERVERNAME pluto
> DBSERVERALIASES pluto_net
> NETTYPE ipcshm,3,64,CPU
>NETTYPE tlitcp,1,10,NET
> DEADLOCK_TIMEOUT 60
> RESIDENT 0
Set this to 1 to lock your buffers in memory
>MULTIPROCESSOR 1
>NUMCPUVPS 3
>SINGLE_CPU_VP 0
>NOAGE 0
I don't know about your platform but on mine (hpux) this is death! NOAGE 1
> AFF_SPROC 0
>AFF_NPROCS 0
>LOCKS 20000
> BUFFERS 96000
From your onstat -p in your previous post your cached reads is low.
perhaps you do need more buffers?
>NUMAIOVPS 8
Are you using KAIO, if not set this to the number of chunks
>PHYSBUFF 120
>LOGBUFF 26>LOGSMAX 60
>CLEANERS 8
This should be => LRUS
>SHMBASE 0xa000000
>SHMVIRTSIZE 120000
>SHMADD 32000
>SHMTOTAL 0
>CKPTINTVL 300
>LRUS 63
> LRU_MAX_DIRTY 2
>LRU_MIN_DIRTY 1
>LTXHWM 40
>LTXEHWM 45
>TXTIMEOUT 0x12c
>STACKSIZE 32
>OFF_RECVRY_THREADS 10
>ON_RECVRY_THREADS 1
> DRAUTO 0
>DRINTERVAL 30
>DRTIMEOUT 30> CDR_LOGBUFFERS 2048
>CDR_EVALTHREADS 1,2
>CDR_DSLOCKWAIT 5
>CDR_QUEUEMEM 4096
>BAR_ACT_LOG /tmp/bar_act.log
>BAR_MAX_BACKUP 0
>BAR_RETRY 1
>BAR_NB_XPORT_COUNT 10
>BAR_XFER_BUF_SIZE 31>ISM_DATA_POOL ISMData
>ISM_LOG_POOL ISMLogs
>RA_PAGES
>RA_THRESHOLD
Try, RA_PAGES 32, RA_THRESHOLD 24
>DBSPACETEMP tempspace
Create multiple temp dbspaces, don't create them all as temp though as temp
dbspaces are not logged.
>DUMPDIR /ascii
>DUMPSHMEM 1
>DUMPGCORE 0
>DUMPCORE 0
>DUMPCNT 1
>FILLFACTOR 90
>USEOSTIME 0
>MAX_PDQPRIORITY 100
>DS_MAX_QUERIES
>DS_TOTAL_MEMORY
>DS_MAX_SCANS 1048576
>DATASKIP off
>OPTCOMPIND 0
>ONDBSPACEDOWN 2> LBU_PRESERVE 1
>OPCACHEMAX 0
>HETERO_COMMIT 0
>OPT_GOAL -1
>DIRECTIVES 1
>RESTARTABLE_RESTORE off>
>
Hope these comments help.
Ray Canuel wrote in message <39B6574D.1213E610@informix.com>... > >During a checkpoint you only have 8 cleaners utilized to flush 63 >LRU's. That's 4 passes. Make LRUS=64 and CLEANERS=64. That's one >pass. Let the disk subsystem figure out any contention. They're pretty >good at that:) Also increase CKPTINTVL to every 30 minutes. Why have a >20-30 second interruption every 5 minutes when you can have it every 1/2 >hour or every hour! Don't reduce the PHYSFILE. Make it bigger. If you >make it smaller, you will only get more frequent checkpoints and OLTP >interruptions. > >Ray I agree with your comments on LRUS, I think in some versions there's a bug if the number of LRUS/CLEANERS is a power of 2, to be safe try 65/65 I would say a checkpoint interval of 30 minutes on an OLTP system is a no-no! to prevent users experiencing apparant hangs whilst the engine checkpoints you want them to be <= 1 second in which case little and often is the key. There have been discussions on this group about optimum physical log file size. I can't remember enough detail to comment but there is an optimum size. Search the archives for the thread(s). > -- Tony Flaherty Snr. A/P, Informix DBA, HpUx Admin, Gimmi a broom! MFS Ltd.
>
>First, the most likely reason for this is that there are a lot of dirty
>buffers in the buffer cache so there is more work to be done. The page
>cleaners have not had time to flush these pages in the background
>because the CKPTINTVL is too frequent or the PHYSFILE is too small.
>Increase both of these (I like a checkpoint to occur every 30 minutes or
>more, let the PHYSFILE be large enough so it doesn't kick off a CP prior
>to this). Current LRUS, page cleaners, and chunk settings are needed in
>order to make better recommendations. This could be a problem too.
>
>Second, who's modifying all those pages? My guess is its those big
>reports. Reports are typically RO, but if you grep them for temp you'll
>probably find a lot of temp tables being used. Do you have a
>DBSPACETEMP defined? Are reports creating temp tables with NO LOG
>option? That will help a lot.
there are some temp tables being created and I had to get a developer to
modify
a program to include the NO LOG option the other day. I do have DBSPACETEMP
defined but when i move from RAID5 to RAID10, I will be adding another
tempspace
onto a differet disk.
>
>PDQ will help these reports a lot. It will get them out of the buffer
>pool and reduce contention with the OLTP users.
>
>You're also only using 50% of available RAM for Informix. Give 100MB to
>the OS and the rest to Informix.
I chose that value on a reccomendation which I cannot remember now - some
percentage of total memory. I also think that the idea was if the read cache
%
was not increasing, then there was not much point in adding more buffers.
>
>What's your concurrency controls for oltp and reports? Run onstat -g
>sql. The OLTP slowness could be attributed to waiting on locks held.
>
>You should also upgrade to 7.31.
>
>Ray
>
>
>Tam McLaughlin wrote:
>>
>> I have some intermittent performance problems with our instance that lead
>> users to
>> complain their programs are running slow.
>>
>> The users mainly run OLTP type programs (we are an insurance company
>> processing claims etc)
>> and most of the time, checkpoint duration < 0 but at certain times, the
>> interval can be as as ,much
>> as 20-30 seconds. I now believe this is probably due to some DSS type
>> programs that run via cron and by some users. These programs can create
some
>> very large queries (170 line sql from onstat -g ses) and look through a
>> number of tables to produce reports and large (20-200Mb) delimited files.
>>
>> Now, there are a number of things I can tune and will do so this weekend
>> i.e. get away from RAID 5 and use RAID 10, place my tables in different
>> dbspaces, put logs in different dbspace etc etc.
>>
>> However, when I run: onstat -g acr -r 1, there apart from the oninit
>> threads, there are a few threads
>> that are constantly working away. But so what ? is this good/bad ? could
it
>> be that these are intensive programs that require more work rather than
the
>> casual select from a user program? So what do I look
>> for ? what can i check to see what can be tuned. From onstat -p, my
>> read/write ratios are good, I believe I have plenty of buffers etc. I
tried
>> putting setting PDQPRIORITY=40 in a script but onstat -g mgm showed no
>> queries active. I have about 20 5 Mb logical log files which can get
filled
>> up over a few days.
>>
>> So, I have a few programs that run for a few hours searching thru the
>> database. I hink most are using indexes as some of our tables have
several
>> indexes on them. I cannot get the developers to use sqexplain. The unix
box
>> is not used for anything else. Any ideas where I should be looking ? I
take
>> it I should be trying to get these progs to use PDQ.
>>
>> Details: 3 intel pent pro processors, 764 Mb Ram, sco unixware 7.0.1, IDS
>> 7.30
>>
>> onstat -p
>> Informix Dynamic Server Version 7.30.UC10X2 -- On-Line -- Up 8 days
>> 06:10:09 -- 323080 Kbytes>>
>> Profile
>> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
>> 202085809 5929898 1316346513 84.65 2332438 2278714 480555135 99.51
>> isamtot open start read write rewrite delete commit
>> rollbk 1916535217 24709189 69257312 1375228628 233538432 854964 21011
>> 792707 48174
>> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
>> 0 0 0 0 0 0 0
>> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
>> 0 0 0 271449.86 33836.99 1836 4740
>> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
>> 45003781 157 2881589552 2 0 997 100381 2362772
>> ixda-RA idx-RA da-RA RA-pgsused lchwaits
>> 107363614 57987 58116582 165258340 1112639
>>
>> onstat -g ioq
>> AIO I/O queues:>> q name/id len maxlen totalops dskread dskwrite dskcopy
>> adt 0 0 0 0 0 0 0
>> msc 0 0 2 58243 0 0 0
>> aio 0 0 7 29811 14892 6 0
>> pio 0 0 1 6756 0 6756 0
>> lio 0 0 1 804868 0 804868 0
>> gfd 3 0 160 461550 355892 105658 0
>> gfd 4 0 177 24085993 23955385 130608 0
>> gfd 5 0 177 86571328 86335846 235482 0
>> gfd 6 0 164 71626634 71548062 78572 0
>> gfd 7 0 172 8899865 8796684 103181 0
>> gfd 8 0 166 1056297 494293 562004 0
>> gfd 9 0 151 7142010 7048660 93350 0
>> gfd 10 0 164 901532 829320 72212 0
>> gfd 11 0 182 2874388 2734613 139775 0
Tam McLaughlin wrote: > > Ray Canuel wrote in message <39B6574D.1213E610@informix.com>... > > > >During a checkpoint you only have 8 cleaners utilized to flush 63 > >LRU's. That's 4 passes. Make LRUS=64 and CLEANERS=64. That's one > >pass. Let the disk subsystem figure out any contention. They're pretty > >good at that:) Also increase CKPTINTVL to every 30 minutes. Why have a > >20-30 second interruption every 5 minutes when you can have it every 1/2 > >hour or every hour! Don't reduce the PHYSFILE. Make it bigger. If you > >make it smaller, you will only get more frequent checkpoints and OLTP > >interruptions. > > > >Ray > > > > > Thanks for the help. Informix always reccomends checkpoint interval every 5 > mins so 1/2 hr > seems a long time betweem checkpoints which also means there there is more > risk of data > loss if something goes wrong before the next checkpoint. 1) Most of us use CHKPTINTVL of at least 600 secs. 2) Checkpointing more frequently will not prevent data loss it will only speed recovery by simplifying fast recovery a bit. 3) Where do you see a recommendation to checkpoint every 5 mins? That is the sample value in onconfig.std but certainly not a recommendation. Art S. Kagel
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g