long checkpoints on informix
Posted in 2007
A user on IDS 9.40.FC2 reported checkpoints regularly lasting 9-14 seconds and posted onconfig and onstat output (200000 BUFFERS, 127 LRUS, 2 CPU VPs, 1 AIO VP, ~65% write cache). Suggestions included enlarging the physical log (ruled out, as it never exceeded 20% full), and Art Kagel noted that 14-second fuzzy checkpoints with only ~450 dirty buffers point to contention between the checkpoint thread and user threads holding locks in critical sections, rather than I/O volume. Ben Thompson suggested more buffers/memory, more AIO VPs, lower read-ahead, checking RAID health, and monitoring onstat -F/-k/-u during checkpoints. No confirmed resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Server Administration, Logging & Checkpoints
hi everyone
i recently start to have some long and constant checkpoint
i did lot of change but cant reduce it
i have Informix Dynamic Server Version 9.40.FC2
and here is some parameter of the onconfig
---------------------------------------------------------------------------------------------------------
TBLSPACE_STATS 1 # Maintain tblspace statistics
MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-processor
NUMCPUVPS 2 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vpsto one
NOAGE 1 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
LOCKS 100000 # Maximum number of locks
BUFFERS 200000 # Maximum number of shared buffers
NUMAIOVPS 1 # Number of IO vps
CLEANERS 127 # Number of buffer cleaner processes
SHMBASE 0x10A000000L # Shared memory base address
SHMVIRTSIZE 300000 # initial virtual shared memorysegment size
SHMADD 16384 # Size of new shared memory segments
(Kbytes)
CKPTINTVL 600 # Check point interval (in sec)(10
minutos)
LRUS 127 # Number of LRU queues
LRU_MAX_DIRTY 2.000000 # LRU percent dirty begin cleaninglimit
LRU_MIN_DIRTY 1.000000 # LRU percent dirty end cleaning limit
DYNAMIC_LOGS 2
OFF_RECVRY_THREADS 10 # Default number of offline workerthreads
ON_RECVRY_THREADS 1 # Default number of online workerthreads
CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-
vp,additional)
CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
queue (Kbytes)
RA_PAGES 32 # Number of pages to attemptto read ahead
RA_THRESHOLD 30 # Number of pages left beforenext group
MAX_PDQPRIORITY 0 # Maximum allowed pdqpriority
DS_MAX_QUERIES # Maximum number of decision supportqueries
DS_TOTAL_MEMORY # Decision support memory (Kbytes)
DS_MAX_SCANS 1048576 # Maximum number of decision supportscans
OPTCOMPIND 2 # To hint the optimizer
---------------------------------------------------------------------------------------------------------
and onstat -p
---------------------------------------------------------------------------------------------------------
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
69174631 143507619 16152454169 99.57 10204375 98766167 29414237
65.31
isamtot open start read write rewrite delete
commit rollbk
13431042812 16518600 2227313233 7879740008 418219403 1700547 22851
7167133 69
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 210120.79 6646.62 655 2076
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress
seqscans
5606856 42 7665171312 0 0 1665 689438
16106139
ixda-RA idx-RA da-RA RA-pgsused lchwaits
10238122 3271 36204247 46437468 5764864
---------------------------------------------------------------------------------------------------------
any help plz ?
Check your Physical Log and increase the size. On Jul 13, 2007, at 11:07 AM, Sakura.ggxx@gmail.com wrote: > 69174631 143507619 16152454169 99.57 10204375 98766167 29414237 > 65.31
On 13 jul, 13:37, Christine Normile <christ...@meta-today.com> wrote: > Check your Physical Log and increase the size. > On Jul 13, 2007, at 11:07 AM, Sakura.g...@gmail.com wrote: > > > 69174631 143507619 16152454169 99.57 10204375 98766167 29414237 > > 65.31 the physical log never reach the 75% to provoke a chekpoint, so i dont think could be that the max physical log percent reached was 20 %
On 13 jul, 13:37, Christine Normile <christ...@meta-today.com> wrote: > Check your Physical Log and increase the size. > On Jul 13, 2007, at 11:07 AM, Sakura.g...@gmail.com wrote: > > > 69174631 143507619 16152454169 99.57 10204375 98766167 29414237 > > 65.31 the physical log never reach the 75% to provoke a chekpoint, so i dont think could be that the max physical log percent reached was 20 %
Please post the output from:
onstat -m
onstat -d
onstat -D
onstat -u
onstat -g iov
Sakura.ggxx@gmail.com said:
> hi everyone
> i recently start to have some long and constant checkpoint
> i did lot of change but cant reduce it
> i have Informix Dynamic Server Version 9.40.FC2
> and here is some parameter of the onconfig
> ---------------------------------------------------------------------------------------------------------
>
> TBLSPACE_STATS 1 # Maintain tblspace statistics
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-> processor
> NUMCPUVPS 2 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors
> LOCKS 100000 # Maximum number of locks
> BUFFERS 200000 # Maximum number of shared buffers
> NUMAIOVPS 1 # Number of IO vps
> CLEANERS 127 # Number of buffer cleaner processes
> SHMBASE 0x10A000000L # Shared memory base address
> SHMVIRTSIZE 300000 # initial virtual shared memory> segment size
> SHMADD 16384 # Size of new shared memory segments
> (Kbytes)
> CKPTINTVL 600 # Check point interval (in sec)(10
> minutos)
> LRUS 127 # Number of LRU queues
> LRU_MAX_DIRTY 2.000000 # LRU percent dirty begin cleaning> limit
> LRU_MIN_DIRTY 1.000000 # LRU percent dirty end cleaning limit
> DYNAMIC_LOGS 2
> OFF_RECVRY_THREADS 10 # Default number of offline worker> threads
> ON_RECVRY_THREADS 1 # Default number of online worker> threads
> CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-
> vp,additional)
> CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
> CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
> queue (Kbytes)
> RA_PAGES 32 # Number of pages to attempt> to read ahead
> RA_THRESHOLD 30 # Number of pages left before> next group
> MAX_PDQPRIORITY 0 # Maximum allowed pdqpriority
> DS_MAX_QUERIES # Maximum number of decision support> queries
> DS_TOTAL_MEMORY # Decision support memory (Kbytes)
> DS_MAX_SCANS 1048576 # Maximum number of decision support> scans
> OPTCOMPIND 2 # To hint the optimizer
> --------------------------------------------------------------------------------------------------------->
>
>
> and onstat -p
>
>
>
> ---------------------------------------------------------------------------------------------------------
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 69174631 143507619 16152454169 99.57 10204375 98766167 29414237
> 65.31
>
> isamtot open start read write rewrite delete
> commit rollbk
> 13431042812 16518600 2227313233 7879740008 418219403 1700547 22851
> 7167133 69
>
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 0 0 0 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 0 210120.79 6646.62 655 2076
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress
> seqscans
> 5606856 42 7665171312 0 0 1665 689438
> 16106139
>
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 10238122 3271 36204247 46437468 5764864
> ---------------------------------------------------------------------------------------------------------
>
>
> any help plz ?
>
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
> --
> This message has been scanned for viruses and
> dangerous content by OpenProtect(http://www.openprotect.com), and is
> believed to be clean.
>
--
Bye now,
Obnoxio
"I'm astonished anyone pays real money for this crap."
-- Cosmo
"Cluster in my trousers"
-- Guy Bowerman
--
This message has been scanned for viruses and
dangerous content by OpenProtect(http://www.openprotect.com), and is
believed to be clean.
On 13 jul, 14:05, "Obnoxio The Clown" <obno...@serendipita.com> wrote:
> Please post the output from:
>
> onstat -m
> onstat -d
> onstat -D
> onstat -u
> onstat -g iov>
> Sakura.g...@gmail.com said:
>
>
>
> > hi everyone
> > i recently start to have some long and constant checkpoint
> > i did lot of change but cant reduce it
> > i have Informix Dynamic Server Version 9.40.FC2
> > and here is some parameter of the onconfig
> > ---------------------------------------------------------------------------------------------------------
>
> > TBLSPACE_STATS 1 # Maintain tblspace statistics
> > MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-> > processor
> > NUMCPUVPS 2 # Number of user (cpu) vps
> > SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> > to one
> > NOAGE 1 # Process aging
> > AFF_SPROC 0 # Affinity start processor
> > AFF_NPROCS 0 # Affinity number of processors
> > LOCKS 100000 # Maximum number of locks
> > BUFFERS 200000 # Maximum number of shared buffers
> > NUMAIOVPS 1 # Number of IO vps
> > CLEANERS 127 # Number of buffer cleaner processes
> > SHMBASE 0x10A000000L # Shared memory base address
> > SHMVIRTSIZE 300000 # initial virtual shared memory> > segment size
> > SHMADD 16384 # Size of new shared memory segments
> > (Kbytes)
> > CKPTINTVL 600 # Check point interval (in sec)(10
> > minutos)
> > LRUS 127 # Number of LRU queues
> > LRU_MAX_DIRTY 2.000000 # LRU percent dirty begin cleaning> > limit
> > LRU_MIN_DIRTY 1.000000 # LRU percent dirty end cleaning limit
> > DYNAMIC_LOGS 2
> > OFF_RECVRY_THREADS 10 # Default number of offline worker> > threads
> > ON_RECVRY_THREADS 1 # Default number of online worker> > threads
> > CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-
> > vp,additional)
> > CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
> > CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
> > queue (Kbytes)
> > RA_PAGES 32 # Number of pages to attempt> > to read ahead
> > RA_THRESHOLD 30 # Number of pages left before> > next group
> > MAX_PDQPRIORITY 0 # Maximum allowed pdqpriority
> > DS_MAX_QUERIES # Maximum number of decision support> > queries
> > DS_TOTAL_MEMORY # Decision support memory (Kbytes)
> > DS_MAX_SCANS 1048576 # Maximum number of decision support> > scans
> > OPTCOMPIND 2 # To hint the optimizer
> > --------------------------------------------------------------------------------------------------------->
> > and onstat -p
>
> > ---------------------------------------------------------------------------------------------------------
> > Profile
> > dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> > 69174631 143507619 16152454169 99.57 10204375 98766167 29414237
> > 65.31
>
> > isamtot open start read write rewrite delete
> > commit rollbk
> > 13431042812 16518600 2227313233 7879740008 418219403 1700547 22851
> > 7167133 69
>
> > gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> > 0 0 0 0 0 0 0
>
> > ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> > 0 0 0 210120.79 6646.62 655 2076
>
> > bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress
> > seqscans
> > 5606856 42 7665171312 0 0 1665 689438
> > 16106139
>
> > ixda-RA idx-RA da-RA RA-pgsused lchwaits
> > 10238122 3271 36204247 46437468 5764864
> > ---------------------------------------------------------------------------------------------------------
>
> > any help plz ?
>
> > _______________________________________________
> > Informix-list mailing list
> > Informix-l...@iiug.org
> >http://www.iiug.org/mailman/listinfo/informix-list
>
> > --
> > This message has been scanned for viruses and
> > dangerous content by OpenProtect(http://www.openprotect.com), and is
> > believed to be clean.
>
> --
> Bye now,
> Obnoxio
>
> "I'm astonished anyone pays real money for this crap."
> -- Cosmo
>
> "Cluster in my trousers"
> -- Guy Bowerman
>
> --
> This message has been scanned for viruses and
> dangerous content by OpenProtect(http://www.openprotect.com), and is
> believed to be clean.
onstat -m
---------------------------------------------------------------------------------------------------------10:32:30 Maximum server connections 282
10:42:41 Fuzzy Checkpoint Completed: duration was 10 seconds, 176buffers not flushed,
timestamp: 56716651.
10:42:41 Checkpoint loguniq 31535, logpos 0x249a6ec, timestamp:56716651
10:42:41 Maximum server connections 282
10:52:54 Fuzzy Checkpoint Completed: duration was 14 seconds, 295buffers not flushed,
timestamp: 58155048.
10:52:54 Checkpoint loguniq 31535, logpos 0x314b0e0, timestamp:58155048
10:52:54 Maximum server connections 282
11:03:04 Fuzzy Checkpoint Completed: duration was 10 seconds, 373buffers not flushed,
timestamp: 62108820.
11:03:04 Checkpoint loguniq 31535, logpos 0x3a431f4, timestamp:62108820
11:03:04 Maximum server connections 282
11:13:15 Fuzzy Checkpoint Completed: duration was 11 seconds, 411buffers not flushed,
timestamp: 76781814.
11:13:15 Checkpoint loguniq 31535, logpos 0x4a103bc, timestamp:76781816
11:13:15 Maximum server connections 282
11:23:27 Fuzzy Checkpoint Completed: duration was 9 seconds, 471buffers not flushed,
timestamp: 79611828.
11:23:27 Checkpoint loguniq 31535, logpos 0x55bf68c, timestamp:79611828
---------------------------------------------------------------------------------------------------------
onstat -d
---------------------------------------------------------------------------------------------------------Dbspaces
address number flags fchunk nchunks flags
owner name
125981e58 1 0x1 1 2 N
informix rootdbs
126f9cd58 2 0x1 2 1 N
informix tempdbs
1272f7028 3 0x1 3 1 N
informix dbsfbs01
3 active, 2047 maximum
Chunks
address chunk/dbs offset size free
bpages flags pathname
125982028 1 1 0 70000
67330 PO-- /dev/chunk00fbs
126f9c890 2 2 0 1024000
997065 PO-- /dev/chunk01fbs
126f9ca28 3 3 0 1024000
108760 PO-- /dev/chunk02fbs
126f9cbc0 4 1 0 1000000
770797 PO-- /dev/chunk03fbs
-----------------------
Sakura.ggxx@gmail.com said:
> if u need anything else just tell me
onstat -F
onstat -R
--
Bye now,
Obnoxio
"I'm astonished anyone pays real money for this crap."
-- Cosmo
"Cluster in my trousers"
-- Guy Bowerman
--
This message has been scanned for viruses and
dangerous content by OpenProtect(http://www.openprotect.com), and is
believed to be clean.
On 13 jul, 14:49, "Obnoxio The Clown" <obno...@serendipita.com> wrote:
> Sakura.g...@gmail.com said:
>
> > if u need anything else just tell me
>
> onstat -F
> onstat -R>
> --
> Bye now,
> Obnoxio
>
> "I'm astonished anyone pays real money for this crap."
> -- Cosmo
>
> "Cluster in my trousers"
> -- Guy Bowerman
>
> --
> This message has been scanned for viruses and
> dangerous content by OpenProtect(http://www.openprotect.com), and is
> believed to be clean.
onstat -F
-----------------------------------------------------------------------------------------Fg Writes LRU Writes Chunk Writes
0 55128 571421
-----------------------------------------------------------------------------------------
onstat -R
-----------------------------------------------------------------------------------------
127 buffer LRU queue pairs priority levels
# f/m pair total % of length LOW HIGH
0 f 1575 98.3% 1548 385 1163
1 m 1.7% 27 15 12
2 f 1574 98.9% 1556 393 1163
3 m 1.1% 18 9 9
4 F 1574 98.5% 1551 388 1163
5 m 1.5% 23 11 12
6 f 1574 98.5% 1550 387 1163
7 m 1.5% 24 16 8
8 f 1576 98.5% 1552 389 1163
9 m 1.5% 24 13 11
10 f 1575 98.9% 1558 395 1163
11 m 1.1% 17 14 3
12 f 1576 98.2% 1548 385 1163
13 m 1.8% 28 18 10
14 f 1575 99.2% 1562 399 1163
15 m 0.8% 13 7 6
16 f 1575 99.0% 1559 396 1163
17 m 1.0% 16 6 10
18 f 1575 98.2% 1547 384 1163
19 m 1.8% 28 13 15
20 f 1575 98.7% 1554 391 1163
21 m 1.3% 21 12 9
22 f 1574 98.6% 1552 389 1163
23 m 1.4% 22 10 12
24 f 1574 98.6% 1552 389 1163
25 m 1.4% 22 10 12
26 f 1575 98.6% 1553 390 1163
27 m 1.4% 22 16 6
28 f 1574 98.3% 1548 385 1163
29 m 1.7% 26 17 9
30 f 1575 98.3% 1549 386 1163
31 m 1.7% 26 15 11
32 f 1574 98.3% 1547 384 1163
33 m 1.7% 27 10 17
34 f 1575 98.4% 1550 387 1163
35 m 1.6% 25 16 9
36 f 1575 98.9% 1558 395 1163
37 m 1.1% 17 9 8
38 f 1574 98.7% 1553 390 1163
39 m 1.3% 21 8 13
40 f 1576 99.0% 1561 398 1163
41 m 1.0% 15 8 7
42 f 1575 98.7% 1554 391 1163
43 m 1.3% 21 10 11
44 f 1575 98.3% 1549 386 1163
45 m 1.7% 26 16 10
46 f 1575 98.5% 1551 388 1163
47 m 1.5% 24 16 8
48 f 1575 98.0% 1544 381 1163
49 m 2.0% 31 18 13
50 f 1575 98.7% 1554 391 1163
51 m 1.3% 21 12 9
52 f 1574 98.7% 1554 391 1163
53 m 1.3% 20 8 12
54 f 1575 98.5% 1552 389 1163
55 m 1.5% 23 13 10
56 f 1575 98.4% 1550 387 1163
57 m 1.6% 25 17 8
58 f 1574 98.5% 1550 387 1163
59 m 1.5% 24 15 9
60 f 1575 98.3% 1549 386 1163
61 m 1.7% 26 18 8
62 f 1575 98.7% 1554 391 1163
63 m 1.3% 21 11 10
64 f 1575 98.6% 1553 390 1163
65 m 1.4% 22 12 10
66 f 1575 98.6% 1553 390 1163
67 m 1.4% 22 13 9
68 f 1575 98.9% 1557 394 1163
69 m 1.1% 18 14 4
70 f 1575 98.8% 1556 393 1163
71 m 1.2% 19 12 7
72 f 1576 98.9% 1558 395 1163
73 m 1.1% 18 11 7
74 f 1575 99.0% 1560 397 1163
75 m 1.0% 15 5 10
76 f 1575 98.9% 1558 395 1163
77 m 1.1% 17 9 8
78 f 1574 98.7% 1553 390 1163
79 m 1.3% 21 12 9
80 f 1575 98.3% 1549 386 1163
81 m 1.7% 26 15 11
82 f 1574 98.8% 1555 392 1163
83 m 1.2% 19 9 10
84 f 1575 98.3% 1548 385 1163
85 m 1.7% 27 14 13
86 f 1574 98.5% 1550 387 1163
87 m 1.5% 24 9 15
88 f 1575 98.7% 1554 391 1163
89 m 1.3% 21 16 5
90 f 1575 98.2% 1546 383 1163
91 m 1.8% 29 15 14
92 f 1574 98.4% 1549 386 1163
93 m 1.6% 25 11 14
94 f 1575 98.5% 1552 389 1163
95 m 1.5% 23 15 8
96 f 1575 98.7% 1555 392 1163
97 m 1.3% 20 7 13
98 f 1575 98.3% 1549 386 1163
99 m 1.7% 26 14 12
100 f 1574 98.5% 1551 388 1163
101 m 1.5% 23 10 13
102 f 1575 98.5% 1552 389 1163
103 m 1.5% 23 16 7
104 f 1574 98.2% 1545 382 1163
105 m 1.8% 29 20 9
106 f 1574 98.4% 1549 386 1163@@NL
On Jul 13, 12:07 pm, Sakura.g...@gmail.com wrote: > hi everyone > i recently start to have some long and constant checkpoint > i did lot of change but cant reduce it > i have Informix Dynamic Server Version 9.40.FC2 > and here is some parameter of the onconfig <SNIP> You should NOT be seeing 14 second Fuzzy checkpoints with only ~450 dirty buffers per checkpoint. I might suspect a RAID5 disk farm or singleton drives, but you say this started recently. There must be contention between the checkpoint thread and user threads in critical sections delaying the beginning of the checkpoint. Look for applications that lock records and hold locks while presenting data for interactive users to update. Art S. Kagel
On 13 jul, 15:44, "Art S. Kagel" <art.ka...@gmail.com> wrote: > On Jul 13, 12:07 pm, Sakura.g...@gmail.com wrote:> hi everyone > > i recently start to have some long and constant checkpoint > > i did lot of change but cant reduce it > > i have Informix Dynamic Server Version 9.40.FC2 > > and here is some parameter of the onconfig > > <SNIP> > > You should NOT be seeing 14 second Fuzzy checkpoints with only ~450 > dirty buffers per checkpoint. I might suspect a RAID5 disk farm or > singleton drives, but you say this started recently. There must be > contention between the checkpoint thread and user threads in critical > sections delaying the beginning of the checkpoint. Look for > applications that lock records and hold locks while presenting data > for interactive users to update. > > Art S. Kagel you are saying: > There must be contention between the checkpoint thread and user threads in critical how can i check this contention, there is some script or commando to check that??
Sakura.ggxx@gmail.com wrote:
> hi everyone
> i recently start to have some long and constant checkpoint
> i did lot of change but cant reduce it
> i have Informix Dynamic Server Version 9.40.FC2
> and here is some parameter of the onconfig
> ---------------------------------------------------------------------------------------------------------
>
> TBLSPACE_STATS 1 # Maintain tblspace statistics
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-> processor
> NUMCPUVPS 2 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> to one
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors
> LOCKS 100000 # Maximum number of locks
> BUFFERS 200000 # Maximum number of shared buffers
> NUMAIOVPS 1 # Number of IO vps
> CLEANERS 127 # Number of buffer cleaner processes
> SHMBASE 0x10A000000L # Shared memory base address
> SHMVIRTSIZE 300000 # initial virtual shared memory> segment size
> SHMADD 16384 # Size of new shared memory segments
> (Kbytes)
> CKPTINTVL 600 # Check point interval (in sec)(10
> minutos)
> LRUS 127 # Number of LRU queues
> LRU_MAX_DIRTY 2.000000 # LRU percent dirty begin cleaning> limit
> LRU_MIN_DIRTY 1.000000 # LRU percent dirty end cleaning limit
> DYNAMIC_LOGS 2
> OFF_RECVRY_THREADS 10 # Default number of offline worker> threads
> ON_RECVRY_THREADS 1 # Default number of online worker> threads
> CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-
> vp,additional)
> CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
> CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
> queue (Kbytes)
> RA_PAGES 32 # Number of pages to attempt> to read ahead
> RA_THRESHOLD 30 # Number of pages left before> next group
> MAX_PDQPRIORITY 0 # Maximum allowed pdqpriority
> DS_MAX_QUERIES # Maximum number of decision support> queries
> DS_TOTAL_MEMORY # Decision support memory (Kbytes)
> DS_MAX_SCANS 1048576 # Maximum number of decision support> scans
> OPTCOMPIND 2 # To hint the optimizer
> ---------------------------------------------------------------------------------------------------------
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 69174631 143507619 16152454169 99.57 10204375 98766167 29414237
> 65.31
Some comments and things to try that may be helpful, may not.
You're running 9.40.FC2 so you clearly have a 64-bit system. Which
platform? How much memory do you have? What are your processors? What is
the disc controller and RAID level?
You've got 65% cached writes which is pretty low. Your other post shows
around 70 io operations/sec which is quite high so your system is busy.
Your BUFFERS are just 200000 which is just 400Mb or 800Mb depending on
your page size, plus you've got 300Mb of shared memory. It's not a lot
really by today's standards and not for a 64-bit system. If you have the
memory, have you tried increasing these values? It would get your cached
writes up.
You've got two CPU VPs and just one AIO VP. You could definitely have
more AIO VPs and maybe you could try two CPU VPs per processor (if you
have a multi-processor system on modern processors).
You've got a lot of LRUs, presumably to keep checkpoints down, but
mostly chunk writes which is a little confusing. Maybe others can help
here. However performance is poor. Is your RAID set healthy?
Your read-ahead values are quite high. Maybe you could reduce these to
cut down the number of pages read in. Obviously there is a trade-off
here with the number of disc access requests requests required so
perhaps a little tuning?
You've got a lot of writes so presumably you've got an OLTP set up?
Perhaps try changing OPTCOMPIND to '0' (zero).
Try monitoring using "onstat -F" during the checkpoint by running it
every second for analysis later. You have only posted truncated "onstat
-F" output.
Also check locks using "onstat -k" and cross-reference with "onstat -u".
Regards, Ben.