Re: Mapping cpuvp's to IBM virtual processors ....
Posted in 2007
Topics: Storage & Space Management, Server Administration
On Thu, 2007-02-22 at 17:27 -0500, ART KAGEL, BLOOMBERG/ 731 LEXIN
wrote:
> For soctcp use 1 or 2 NET VPs (a net VP should handle over 200 connections
> efficiently but I like to keep that below ~150 myself).
I added 2.
>
> #LRUs controls contention between users. In order to read new data into a
> buffer or to write/rewrite a buffer, the user much latch the LRU in which the
> buffer is listed in order to move the buffer to the front of the LRU's queue
> or
> to move the buffer from the LRU's clean queue to its dirty queue. More LRUS
> mean less contention for LRU resources. Some buffer contention is cause by
> insufficient buffers, but most is from LRU contention. If you have more users
> than LRUs then more LRUs will help.
I doubled the LRU's from 8 to 16.
>
> I also notice that you have only one AIO VP configured, even if you are using
> only RAW devices you should have 2-3 configured. Add to that 1 to 1.5 per
> COOKED chunk.
All space is raw, but I changed from 1 to 3 anyway.
>
> What else do we need to see to help? Post the following:
>
> onstat -p (again - with the time since reset)
> onstat -d
> onstat -F
> onstat -R
> onstat -g glo
> onstat -g iov
> onstat -P (at least the summaries at the end, the partnum zero line, and any> partnum lines that have a significant % of the buffer cache allocated to
them)
> onstat -g iof>
> A full copy of the ONCONFIG file.
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Dave Thacker <ids@iiug.org>
> At: 2/22 17:12:14
>
> On Thu, 2007-02-22 at 16:35 -0500, ART KAGEL, BLOOMBERG/ 731 LEXIN
> wrote:
> > Your assumption is partially correct. You don't NEED all 6 CPU's allocated
> to
> > CPU VPs, but it's a good thing. The IDS scheduler will allocate sessions to
> > CPU
> > VPs as needed so it's possible that one or a few logical CPUS will be
harder
> > hit than the others. That said there are tunables that can affect that. For
> > example if you are using shared memory connections to the instance you
> should
> > have ipcshm poll threads running in all 6 CPU VPs or the load will not be
> > balanced as well as it could be and responsiveness will suffer. I don't see
> > any
> > NETTYPE entries in the partial ONCONFIG you posted so you're probably
> running
> > with the default of one poll thread.
>
> NETTYPE is at default. Workload is 90% TCP/10$ SHM. What type of
> threads should I add for soctcp?
> >
> > Also note that your instance is poorly tuned. BR is over 20 and values > 7
> > indicate a slow server needing more LRUs and more CLEANERS (and possibly
> more
> > BUFFERS). Your BTR (estimated since you didn't post any BUFFERS or
> BUFFERPOOL> > parameters) looks very high indicating that you may need more buffers.
I suspected as much. Last night increased buffers from 12000 to 100000
Today's stats look much healthier.
I'm still not seeing any LRU writes, and I'm not sure why. The output
you asked for is posted below. Thank You for helping with this!
All stats were run 305 minutes after the last stats reset (onstat -z)
===========onstat -p===========
IBM Informix Dynamic Server Version 10.00.FC6 -- On-Line -- Up
13:26:02 -- 1075472 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits
%cached
2919 3085 184065849 100.00 34530 55350 364888
90.54
isamtot open start read write rewrite delete
commit rollbk
157746043 1243273 12306759 107186718 41753 69057 35664
53002 0
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free
gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 1802.65 1566.00 47 94
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits
===========onstat -d===========
IBM Informix Dynamic Server Version 10.00.FC6 -- On-Line -- Up
13:26:02 -- 1075472 Kbytes
Dbspaces
address number flags fchunk nchunks pgsize flags
owner name
7000000201b5e78 1 0x60001 1 1 4096 N B
informix uswdbs
7000000201b6710 2 0x68001 2 1 4096 N SB
informix replsbdbs
7000000201b68a8 3 0x60001 3 1 4096 N B
informix repltxdbs
7000000201b6a40 4 0x60001 4 1 4096 N B
informix logdbs
7000000201b6bd8 5 0x60001 5 1 4096 N B
informix physlogdbs
7000000201b6d70 6 0x42001 6 1 4096 N TB
informix tmpdbs1
700000021146028 7 0x42001 7 1 4096 N TB
informix tmpdbs2
7000000211461c0 8 0x42001 8 1 4096 N TB
informix tmpdbs3
700000021146358 9 0x42001 9 1 4096 N TB
informix tmpdbs4
7000000211464f0 10 0x60001 10 1 4096 N B
informix uswdata1
10 active, 2047 maximum
Chunks
address chunk/dbs offset size free bpages
flags pathname
7000000201b6028 1 1 0 20000 11885
PO-B /usr/informix/links/uswdbs
7000000203a8910 2 2 20000 50000 46623 46624
POSB /usr/informix/links/uswdbs
Metadata 3323 2144 3323
7000000203a8ab0 3 3 70000 50000 49907
PO-B /usr/informix/links/uswdbs
7000000203a8c50 4 4 120000 25000 24947
PO-B /usr/informix/links/uswdbs
7000000203a8df0 5 5 145000 25000 24947
PO-B /usr/informix/links/uswdbs
700000020391c50 6 6 170000 25000 24947
PO-B /usr/informix/links/uswdbs
700000020391df0 7 7 195000 25000 24947
PO-B /usr/informix/links/uswdbs
7000000201b6230 8 8 220000 25000 24947
PO-B /usr/informix/links/uswdbs
7000000201b63d0 9 9 245000 25000 24947
PO-B /usr/informix/links/uswdbs
7000000201b6570 10 10 270000 500000 294141
PO-B /usr/informix/links/uswdbs
10 active, 32766 maximum
compress seqscans
751 20 34654452 0 0 29 3340
110132
ixda-RA idx-RA da-RA RA-pgsused lchwaits
287 248 0 535 20459
===========onstat -F===========
IBM Informix Dynamic Server Version 10.00.FC6 -- On-Line -- Up
13:26:02 -- 1075472 Kbytes
Fg Writes LRU Writes Chunk Writes
0 0 30832
address flusher state data
700000020358760 0 I 0 = 0X0
700000020358e98 1 I 0 = 0X0
7000000203595d0 2 I 0 = 0X0
700000020359d08 3 I 0 = 0X0
70000002035a440 4 I 0 = 0X0
70000002035ab78 5 I 0 = 0X0
states: Exit Idle Chunk Lru
===========onstat -R===========
IBM Informix Dynamic Server Version 10.00.FC6 -- On-Line -- Up
13:26:02 -- 1075472 Kbytes
Buffer pool page size: 4096
16 buffer LRU queue pairs priority levels
# f/m pair total % of length LOW HIGH
0 f 6250 99.1% 6195 5076 1119
1 m 0.9% 55 0 55
2 f 6250 99.2% 6199 5033 1166
3 m 0.8% 51 0 51
4 f 6250 99.1% 6194 5093 1101
5 m 0.9% 56 0 56
6 f 6250 98.9% 6184 5070 1114
7 m 1.1% 66 0 66
8 f 6250 99.0% 6190 5075 1115
9 m 1.0% 60 0 60
10 f 6250 98.9% 6183 5061 1122
11 m 1.1% 67 1 66
12 f 6250 99.1% 6193 5094 1099
13 m 0.9% 57 0 57
14 f 6250 99.1% 6195 5082 1113
15 m 0.9% 55 0 55
16 F 6250 99.3% 6204 5095 1109
17 m 0.7% 46 0 46
18 f 6250 99.0% 6187 5079 1108
19 m 1.0% 63 0 63
20 f 6250 98.8% 6178 5085 1093
21 m 1.2% 72 0 72
22 f 6250 99.2% 6202 5094 1108
23 m 0.8% 48 0 48
24 f 6250 99.2% 6197 5108 1089
25 m 0.8% 53 1 52
26 f 6250 99.3% 6205 5106 1099
27 m 0.7% 45 1 44
28 f 6250 99.2% 6198 5074 1124
29 m 0.8% 52 0 52
30 f 6250 98.8% 6178 5034 1144
31 m 1.2% 72 0 72
918 dirty,
Looks like you're cruising now! Nearly 100% reads from cache, >90% writes are
to cache, BR=0.20 BTR=0.72, RAU=100%. Good CPU VP balance the 1st 4 are taking
about 75% of the load and it's spread fairly evenly so responsiveness should be
noticably better now. Your CLEANERS are still 6 and should be >= LRUS so that
should be 16 for best checkpoint and cache flush performance.
You are still seeing almost all CHUNK writes because of the LRU_MAX/MIN_DIRTY
settings on the buffer pool of 20/10. If your checkpoints are not exceedingly
long (I suspect they are not) and external tasks are not contending for IO
bandwidth at checkpoint times, I'd leave it. If checkpoint duration is/becomes
a problem drop those down as low as 2/1 or even 1/0.5 and you'll see LRU writes
balancing out at about 70% of total.
Art S. Kagel
----- Original Message -----
From: Dave Thacker <ids@iiug.org>
At: 2/23 16:20:11
On Thu, 2007-02-22 at 17:27 -0500, ART KAGEL, BLOOMBERG/ 731 LEXIN
wrote:
> For soctcp use 1 or 2 NET VPs (a net VP should handle over 200 connections
> efficiently but I like to keep that below ~150 myself).
I added 2.
>
> #LRUs controls contention between users. In order to read new data into a
> buffer or to write/rewrite a buffer, the user much latch the LRU in which
the
> buffer is listed in order to move the buffer to the front of the LRU's queue
> or
> to move the buffer from the LRU's clean queue to its dirty queue. More LRUS
> mean less contention for LRU resources. Some buffer contention is cause by
> insufficient buffers, but most is from LRU contention. If you have more
users
> than LRUs then more LRUs will help.
I doubled the LRU's from 8 to 16.
>
> I also notice that you have only one AIO VP configured, even if you are
using
> only RAW devices you should have 2-3 configured. Add to that 1 to 1.5 per
> COOKED chunk.
All space is raw, but I changed from 1 to 3 anyway.
>
> What else do we need to see to help? Post the following:
>
> onstat -p (again - with the time since reset)
> onstat -d
> onstat -F
> onstat -R
> onstat -g glo
> onstat -g iov
> onstat -P (at least the summaries at the end, the partnum zero line, and any> partnum lines that have a significant % of the buffer cache allocated to
them)
> onstat -g iof>
> A full copy of the ONCONFIG file.
>
> Art S. Kagel
>
> ----- Original Message -----
> From: Dave Thacker <ids@iiug.org>
> At: 2/22 17:12:14
>
> On Thu, 2007-02-22 at 16:35 -0500, ART KAGEL, BLOOMBERG/ 731 LEXIN
> wrote:
> > Your assumption is partially correct. You don't NEED all 6 CPU's allocated
> to
> > CPU VPs, but it's a good thing. The IDS scheduler will allocate sessions
to
> > CPU
> > VPs as needed so it's possible that one or a few logical CPUS will be
harder
> > hit than the others. That said there are tunables that can affect that.
For
> > example if you are using shared memory connections to the instance you
> should
> > have ipcshm poll threads running in all 6 CPU VPs or the load will not be
> > balanced as well as it could be and responsiveness will suffer. I don't
see
> > any
> > NETTYPE entries in the partial ONCONFIG you posted so you're probably
> running
> > with the default of one poll thread.
>
> NETTYPE is at default. Workload is 90% TCP/10$ SHM. What type of
> threads should I add for soctcp?
> >
> > Also note that your instance is poorly tuned. BR is over 20 and values > 7
> > indicate a slow server needing more LRUs and more CLEANERS (and possibly
> more
> > BUFFERS). Your BTR (estimated since you didn't post any BUFFERS or
> BUFFERPOOL> > parameters) looks very high indicating that you may need more buffers.
I suspected as much. Last night increased buffers from 12000 to 100000
Today's stats look much healthier.
I'm still not seeing any LRU writes, and I'm not sure why. The output
you asked for is posted below. Thank You for helping with this!
All stats were run 305 minutes after the last stats reset (onstat -z)
===========onstat -p===========
IBM Informix Dynamic Server Version 10.00.FC6 -- On-Line -- Up
13:26:02 -- 1075472 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits
%cached
2919 3085 184065849 100.00 34530 55350 364888
90.54
isamtot open start read write rewrite delete
commit rollbk
157746043 1243273 12306759 107186718 41753 69057 35664
53002 0
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free
gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 1802.65 1566.00 47 94
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits
===========onstat -d===========
IBM Informix Dynamic Server Version 10.00.FC6 -- On-Line -- Up
13:26:02 -- 1075472 Kbytes
Dbspaces
address number flags fchunk nchunks pgsize flags
owner name
7000000201b5e78 1 0x60001 1 1 4096 N B
informix uswdbs
7000000201b6710 2 0x68001 2 1 4096 N SB
informix replsbdbs
7000000201b68a8 3 0x60001 3 1 4096 N B
informix repltxdbs
7000000201b6a40 4 0x60001 4 1 4096 N B
informix logdbs
7000000201b6bd8 5 0x60001 5 1 4096 N B
informix physlogdbs
7000000201b6d70 6 0x42001 6 1 4096 N TB
informix tmpdbs1
700000021146028 7 0x42001 7 1 4096 N TB
informix tmpdbs2
7000000211461c0 8 0x42001 8 1 4096 N TB
informix tmpdbs3
700000021146358 9 0x42001 9 1 4096 N TB
informix tmpdbs4
7000000211464f0 10 0x60001 10 1 4096 N B
informix uswdata1
10 active, 2047 maximum
Chunks
address chunk/dbs offset size free bpages
flags pathname
7000000201b6028 1 1 0 20000 11885
PO-B /usr/informix/links/uswdbs
7000000203a8910 2 2 20000 50000 46623 46624
POSB /usr/informix/links/uswdbs
Metadata 3323 2144 3323
7000000203a8ab0 3 3 70000 50000 49907
PO-B /usr/informix/links/uswdbs
7000000203a8c50 4 4 120000 25000 24947
PO-B /usr/informix/links/uswdbs
7000000203a8df0 5 5 145000 25000 24947
PO-B /usr/informix/links/uswdbs
700000020391c50 6 6 170000 25000 24947
PO-B /usr/informix/links/uswdbs
700000020391df0 7 7 195000 25000 24947
PO-B /usr/informix/links/uswdbs
7000000201b6230 8 8 220000 25000 24947
PO-B /usr/informix/links/uswdbs
7000000201b63d0 9 9 245000 25000 24947
PO-B /usr/informix/links/uswdbs
7000000201b6570 10 10 270000 500000 294141
PO-B /usr/informix/links/uswdbs
10 active, 32766 maximum
compress seqscans
751 20 34654452 0 0 29 3340
110132
ixda-RA idx-RA da-RA RA-pgsused lchwaits
287 248 0 535 20459
===========onstat -F===========
IBM Informix Dynamic Server Version 10.00.FC6 -- On-Line -- Up
13:26:02 -- 1075472 Kbytes
Fg Writes LRU Writes Chunk Writes
0 0 30832
address flusher state data
700000020358760 0 I 0 = 0X0
700000020358e98 1 I 0 = 0X0
7000000203595d0 2 I 0 = 0X0
700000020359d08 3 I 0 = 0X0
70000002035a440 4 I 0 = 0X0
70000002035ab78 5 I 0 = 0X0
states: Exit Idle Chunk Lru
===========onstat -R===========
IBM Informix Dynamic Server Version 10.00.FC6 -- On-Line -- Up
13:26:02 -- 1075472 Kbytes
Buffer pool page size: 4096
16 buffer LRU