Re[2]: Long checkpoints (LRU_xxx_DIRTY/LRUS)->NETTYPE values
Posted in 1999
Topics: Performance & Tuning, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration
On Thu, 25 Mar 1999 11:25:10 -0800, MOPSOMER@raychem.com wrote:
> Here is some additional information regarding these long checkpoints
> we have in our productional SAP R/3 environment (Informix 7.24.UC3).
[snip]
> Yesterday, I had to extent a couple of dbspaces, which apparently
> forces a checkpoint. So this forced several checkpoints within a
> couple of minutes. Each of these took more than 15 seconds ! At one
> of these long checkpoints, there were even less than 1000 dirty pages.
Given your LRU params below (1 and 0 for MAX & MIN DIRTY), your
checkpoints should be running much faster. I'd guess one of 2 things
is happening: either the i/o throughput on your machine is terribly
slow for some reason; or your checkpoints are being blocked due to
very high user activity (remember: users already in critical sections
of code when the checkpoint is requested must first finish their work
before the checkpoint is allowed to proceed). You might try running
onstat - -r5 >>filename to capture the status of the system at 5
second intervals. If you see that a checkpoint, when it is requested,
is being blocked, the second of the 2 things is happening. If not, it
must be the first (slow i/o).
> I don't know if my NETTYPE values can cause problems. I use CPUVPS
> for both protocols (SHM and TLI). I use 1 poll thread for the shared
> memory connections and 5 poll threads for the used TLI connections.
> I know that the manuals say that only one protocol can run his poll
> threads on CPU vps, but there are no error messages in the message log
> at startup of the Informix instance and also the onstat -g ath output
> indicates that poll threads of both protocols are running on CPU VPs.
1 poll thread of any type can be run on each cpuvp. Since you have 16
cpuvps, you could run up to 16 poll threads on a cpuvp, either shm,
tli or a combination of both. If most of your work is coming over via
TLI, don't bother increasing shm poll threads; if anything, increase
tli poll threads.
Dave
Here is some additional information regarding these long checkpoints
we have in our productional SAP R/3 environment (Informix 7.24.UC3).
Our physical log is 250 MB, which is quite big (but I want to avoid
physical log overflows by all means). The value for LOGBUFF is 128.
PHYSBUFF is 4096. We are using SAP, so we must use unbuffered logging.
The disk containing rootdbs and physdbs contains nothing else. The
disks containing our logical log files also contain nothing else.
As said in my previous mail, NUMCPUVPS is 16 (we have 18 processors).
The database server is almost exclusively used for that Informix
instance. We currently have 8 application servers (SAP instances).
Yesterday, I had to extent a couple of dbspaces, which apparently
forces a checkpoint. So this forced several checkpoints within a
couple of minutes. Each of these took more than 15 seconds ! At one
of these long checkpoints, there were even less than 1000 dirty pages.
I don't know if my NETTYPE values can cause problems. I use CPUVPS
for both protocols (SHM and TLI). I use 1 poll thread for the shared
memory connections and 5 poll threads for the used TLI connections.
I know that the manuals say that only one protocol can run his poll
threads on CPU vps, but there are no error messages in the message log
at startup of the Informix instance and also the onstat -g ath output
indicates that poll threads of both protocols are running on CPU VPs.
I also read the recent interesting TechNote about "Performance and
Tuning" from Art Kagel. He indicates in that article that it is
recommended to use CPUVPs for shared memory connections and NETVPs for
TLI connections. But in my case, 95% of the work is coming over TLI
connections, so I am afraid that using NETVPs for TLI connections
could hurt performance. Art also indicated that you should have as
many poll threads for shared memory connections as you have CPUVPs.
There is almost no work generated by the few shared memory connections
I have, so I don't know if I indeed should increase the number of poll
threads for shared memory connections ? Any ideas on NETTYPE values ?
MY NETTYPE ONCONFIG SETTINGS :
NETTYPE ipcshm,1,60,CPU
NETTYPE tlitcp,5,60,CPU
Output from running command : onstat -g ath | egrep "(tli|pol|list)"
7 590f2710 0 2 running 1cpu sm_poll
8 591067f8 0 2 yield forever 3cpu tlitcppoll
9 59106e30 0 2 yield forever 4cpu tlitcppoll
10 59107498 0 2 running 5cpu tlitcppoll
11 59107b00 0 2 running 6cpu tlitcppoll
12 5911a7a8 0 2 yield forever 7cpu tlitcppoll
13 5911ade8 0 2 yield forever 1cpu sm_listen
15 5918b088 0 3 yield forever 1cpu tlitcplst
Below you can also find the output from onstat -g glo after almost 4
days uptime. At the end of this mail, I also added the ONCONFIG file.
Thank you in advance,
Mario Opsomer
Output from running onstat -g glo (after more than 3 days uptime) :
-------------------------------------------------------------------
INFORMIX-OnLine Version 7.24.UC3 On-Line Up 3 days 23:48:41 2979040 Kb
MT global info:
sessions threads vps lngspins
290 381 21 27332
sched calls thread switches yield 0 yield n yield forever
total: 1241527105 2614504089 3347825973 18543159 569461645
per sec: 173405 11258 164391 157 2357
Virtual processor summary:
class vps usercpu syscpu total
cpu 16 950350.34 237212.28 1187562.62
aio 1 1.75 0.28 2.03
lio 1 0.08 0.08 0.16
pio 1 0.17 0.06 0.23
adm 1 0.68 4.22 4.90
msc 1 4.12 1.57 5.69
total 21 950357.14 237218.49 1187575.63
Individual virtual processors:
vp pid class usercpu syscpu total
1 10371 cpu 82649.04 10900.23 93549.27
2 10399 adm 0.68 4.22 4.90
3 10533 cpu 72277.10 43882.34 116159.44
4 10534 cpu 69965.15 34856.66 104821.81
5 10535 cpu 68178.97 32219.61 100398.58
6 10536 cpu 71699.63 30683.54 102383.17
7 10537 cpu 63981.07 31965.36 95946.43
8 10538 cpu 69132.96 7014.14 76147.10
9 10539 cpu 64988.33 6640.57 71628.90
10 10543 cpu 61200.16 6331.60 67531.76
11 10548 cpu 56559.13 5831.54 62390.67
12 10558 cpu 52217.77 5540.37 57758.14
13 10568 cpu 55063.34 4839.85 59903.19
14 10576 cpu 45065.26 4560.10 49625.36
15 10585 cpu 42429.83 4232.98 46662.81
16 10593 cpu 39331.77 3986.80 43318.57
17 10596 cpu 35610.83 3726.59 39337.42
18 10599 lio 0.08 0.08 0.16
19 10633 pio 0.17 0.06 0.23
20 10647 aio 1.75 0.28 2.03
21 10648 msc 4.12 1.57 5.69
tot 950357.14 237218.49 1187575.63
My ONCONFIG file :
------------------
AFF_NPROCS 0
AFF_SPROC 0
ALARMPROGRAM /informix/PR1/etc/alarm_raychem.sh
BAR_ACT_LOG /tmp/bar_act.log
BAR_MAX_BACKUP 0
BAR_NB_XPORT_COUNT 10
BAR_RETRY 1
BAR_XFER_BUF_SIZE 31
BUFFERS 512000
CDR_DSLOCKWAIT 5
CDR_EVALTHREADS 1,2CDR_LOGBUFFERS 2048
CDR_QUEUEMEM 4096
CKPTINTVL 1200
CLEANERS 52
CONSOLE /informix/PR1/console.sp00m.pr1.log
DATASKIP off
DBSERVERALIASES sp00mpr1tcp
DBSERVERNAME sp00mpr1shmDBSPACETEMP tempdbs1,tempdbs2,tempdbs3,tempdbs4,tempdbs5
DD_HASHMAX 20
DD_HASHSIZE 511
DEADLOCK_TIMEOUT 60
DRAUTO 0
DRINTERVAL 30
DRLOSTFOUND /informix/PR1/etc/dr.lostfound
DRTIMEOUT 30
DS_MAX_QUERIES 8
DS_MAX_SCANS 1048576
DS_TOTAL_MEMORY 10240
DUMPCNT 1
DUMPCORE 0
DUMPDIR /archusrdata
DUMPGCORE 0
DUMPSHMEM 1
FILLFACTOR 90
HETERO_COMMIT 0LBU_PRESERVE 1
LOCKS 4000000
LOGBUFF 128
LOGFILES 832
LOGSIZE 10000LOGSMAX 1280
LRUS 75
LRU_MAX_DIRTY 1
LRU_MIN_DIRTY 0
LTAPEBLK 256
LTAPEDEV /dev/3590/4stcb
LTAPESIZE 1950000
LTXEHWM 80
LTXHWM 5
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g