Mutex logs on logical log flip
Posted in 2016
On IDS 11.50 (Solaris T5, 230 CPU VPs, ~4,500 connections), queries stalled 10-15s a few seconds after each logical log filled, with hundreds of threads waiting on nsf.lock. Respondents explained nsf has nothing to do with NFS: it's contention on the shared file-descriptor table between CPU VPs, typically triggered by bursts of new connections; NUMFDSERVERS helps little and the real fix came in 11.70. Advice centred on cutting CPU VPs (count physical cores, not hardware threads, ~4 VPs/core), avoiding thread/priority inversion, NOAGE, and tuning NETTYPE poll threads. No confirmed resolution is recorded; an IBM PMR was open.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Triggers, Constraints & Referential Integrity, Logging & Checkpoints, Platform-Specific Issues
IDS 11.50 on Solaris We've only just started noticing this odd occurrence. I t only just have started happening or it may have been happening for ages but there's more focus upon it now. Our customer's monitoring was detecting occasional periods of slow queries, perhaps persisting for 10 or 15s. Well drilling down these do seem to occur. They coincide with a short but intense period whereby 100s or even 1,000s of threads are waiting on an nsf.lock. These always occur a few seconds after each logical log fills. We use ALARMPROGRAM to call Netbackup to back up the recently-filled log so assume that this is the trigger, but we're really struggling to see the connection. Any ideas? IBM PMR 70953,001,866 refers. cheers Neil
Are the log backups being written to NFS drives? Are any server chunks on NFS drives? Art On Feb 28, 2016 8:29 PM, "NEIL TRUBY" <neil.truby@ardenta.com> wrote: > IDS 11.50 on Solaris > > We've only just started noticing this odd occurrence. I t only just have > started happening or it may have been happening for ages but there's more > focus upon it now. > > Our customer's monitoring was detecting occasional periods of slow queries, > perhaps persisting for 10 or 15s. > > Well drilling down these do seem to occur. They coincide with a short but > intense period whereby 100s or even 1,000s of threads are waiting on an > nsf.lock. > > These always occur a few seconds after each logical log fills. We use > ALARMPROGRAM to call Netbackup to back up the recently-filled log so assume > that this is the trigger, but we're really struggling to see the > connection. > > Any ideas? IBM PMR 70953,001,866 refers. > > cheers > Neil > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --047d7bfea0b62e8967052cdd2376
Log backups go to some sort of SAN-attached data domain device. Chunks on NFS!?! I'm not a bloody amateur! ;-). They're raw devices, fronted by Veritas Storage Foundation on an EMC VNX-2 SAN. cheers N
Consumate professional as always Neil. Just trying to get the dumb stuff out of the way. Art On Feb 28, 2016 8:55 PM, "NEIL TRUBY" <neil.truby@ardenta.com> wrote: > Log backups go to some sort of SAN-attached data domain device. > > Chunks on NFS!?! I'm not a bloody amateur! ;-). They're raw devices, > fronted > by Veritas Storage Foundation on an EMC VNX-2 SAN. > > cheers > N > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --001a113ff31865a552052cdd4cf3
"that" nsf has nothing to do with NSF. It's related to the sharing of file descriptors among different CPU VPs.... It's also a classical issue on large instances.... Several improvements were made but I think they were put on 11.70+ only. I'm a bit puzzled by the relation with the log change. Typical situations ate related to high CPU usage and most frequent with a large number of new connections. There were cases of customers with ten's or even hundreds per second. I'd say the PMR is your best options. Support colleagues will find several cases.... eventually too many to be really useful. Newer versions would help. I also wonder if Solaris may be an increased "risk" factor... Your NETTYPE settings coul be interesting to see. And please check the number of new connections per second.... I belìeve I've also seen some interestin theory from Art regarding Sparc CPUs an their multi-threading.... Regards On Feb 28, 2016 23:58, "Art Kagel" <art.kagel@gmail.com> wrote: > Consumate professional as always Neil. Just trying to get the dumb stuff > out of the way. > > Art > On Feb 28, 2016 8:55 PM, "NEIL TRUBY" <neil.truby@ardenta.com> wrote: > > > Log backups go to some sort of SAN-attached data domain device. > > > > Chunks on NFS!?! I'm not a bloody amateur! ;-). They're raw devices, > > fronted > > by Veritas Storage Foundation on an EMC VNX-2 SAN. > > > > cheers > > N > > > > > > > > > > ******************************************************************************* > > Forum Note: Use "Reply" to post a response in the discussion forum. > > > > > > --001a113ff31865a552052cdd4cf3 > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --089e01176343698942052cde1500
blockquote, div.yahoo_quoted { margin-left: 0 !important; border-left:1px #715FFA solid !important; padding-left:1ex !important; background-color:white !important; } You also might check on the total number of VPs vrs physical CPUs. Also check on the actual usage of each of the CPU VPs. Sent from Yahoo Mail for iPad On Sunday, February 28, 2016, 7:58 PM, Fernando Nunes <domusonline@gmail.com> wrote: "that" nsf has nothing to do with NSF. It's related to the sharing of file descriptors among different CPU VPs.... It's also a classical issue on large instances.... Several improvements were made but I think they were put on 11.70+ only. I'm a bit puzzled by the relation with the log change. Typical situations ate related to high CPU usage and most frequent with a large number of new connections. There were cases of customers with ten's or even hundreds per second. I'd say the PMR is your best options. Support colleagues will find several cases.... eventually too many to be really useful. Newer versions would help. I also wonder if Solaris may be an increased "risk" factor... Your NETTYPE settings coul be interesting to see. And please check the number of new connections per second.... I belìeve I've also seen some interestin theory from Art regarding Sparc CPUs an their multi-threading.... Regards On Feb 28, 2016 23:58, "Art Kagel" <art.kagel@gmail.com> wrote: > Consumate professional as always Neil. Just trying to get the dumb stuff > out of the way. > > Art > On Feb 28, 2016 8:55 PM, "NEIL TRUBY" <neil.truby@ardenta.com> wrote: > > > Log backups go to some sort of SAN-attached data domain device. > > > > Chunks on NFS!?! I'm not a bloody amateur! ;-). They're raw devices, > > fronted > > by Veritas Storage Foundation on an EMC VNX-2 SAN. > > > > cheers > > N > > > > > > > > > > ******************************************************************************* > > Forum Note: Use "Reply" to post a response in the discussion forum. > > > > > > --001a113ff31865a552052cdd4cf3 > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --089e01176343698942052cde1500 ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
What I meant was "that nsf has nothing to do with NFS".
The first improvement tried was to increase the number of file descriptor
servers.... there's a parameter for that:
NUMFDSERVERS
But my understanding is that it was not very useful as more FDSERVERS could
cause even greater contetnion. Later the table where the file descriptors
are stored was split and that could effectively improve the situation. But
this later change was done only on 11.70. I believe NUMFDSERVERS was
introduced in some later fixpack of 11.50. A query on sysmaster:syscfgtab
may eliminate any doubts.
On Mon, Feb 29, 2016 at 12:54 AM, Fernando Nunes <domusonline@gmail.com>
wrote:
> "that" nsf has nothing to do with NSF. It's related to the sharing of file
> descriptors among different CPU VPs.... It's also a classical issue on
> large instances....
> Several improvements were made but I think they were put on 11.70+ only.
>
> I'm a bit puzzled by the relation with the log change. Typical situations
> ate related to high CPU usage and most frequent with a large number of new
> connections. There were cases of customers with ten's or even hundreds per
> second.
>
> I'd say the PMR is your best options. Support colleagues will find several
> cases.... eventually too many to be really useful.
>
> Newer versions would help. I also wonder if Solaris may be an increased
> "risk" factor...
> Your NETTYPE settings coul be interesting to see. And please check the
> number of new connections per second.... I belìeve I've also seen some
> interestin theory from Art regarding Sparc CPUs an their
> multi-threading....
>
> Regards
> On Feb 28, 2016 23:58, "Art Kagel" <art.kagel@gmail.com> wrote:
>
> > Consumate professional as always Neil. Just trying to get the dumb stuff
> > out of the way.
> >
> > Art
> > On Feb 28, 2016 8:55 PM, "NEIL TRUBY" <neil.truby@ardenta.com> wrote:
> >
> > > Log backups go to some sort of SAN-attached data domain device.
> > >
> > > Chunks on NFS!?! I'm not a bloody amateur! ;-). They're raw devices,
> > > fronted
> > > by Veritas Storage Foundation on an EMC VNX-2 SAN.
> > >
> > > cheers
> > > N
> > >
> > >
> > >
> > >
> >
> >
>
>
*******************************************************************************
> > > Forum Note: Use "Reply" to post a response in the discussion forum.
> > >
> > >
> >
> > --001a113ff31865a552052cdd4cf3
> >
> >
> >
> >
>
>
*******************************************************************************
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
> --089e01176343698942052cde1500
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--089e01184b0cd5f6f4052ce59cb6
Thanks. We already knew from Marco Greco, ex of IBM, that the handling of this
in 11.70 is vastly imrpoved. This is the last large customer resisting that
move.
VPCLASS CPU,NUM=230 # VPCLASS CPU - CPU #T5 Migration
NETTYPE tlitcp,11,800,NET
... are the settings. The server has 32 physical cores each with 8 threads.
232 of these are assigned to this "Container" (ie VM).
Typically we have 4,500 connections and IBM has suggested reducing the number
of poll threads to 250 per tlitcp connection something like:
NETTYPE tlitcp,36,250,NET
VPCLASS cpu,num=230,noage
(NOAGE needs a Solaris change to work)
Do you really have 230 physical CPUs? If not you could be running into process
inversion which is making it difficult for the connection file descriptor to
be received by all of the CPUVPs
Madison Pruet
Retired and Loving it
On Monday, February 29, 2016 11:36 AM, NEIL TRUBY <neil.truby@ardenta.com>
wrote:
Thanks. We already knew from Marco Greco, ex of IBM, that the handling of this
in 11.70 is vastly imrpoved. This is the last large customer resisting that
move.
VPCLASS CPU,NUM=230 # VPCLASS CPU - CPU #T5 Migration
NETTYPE tlitcp,11,800,NET
.... are the settings. The server has 32 physical cores each with 8 threads.
232 of these are assigned to this "Container" (ie VM).
Typically we have 4,500 connections and IBM has suggested reducing the number
of poll threads to 250 per tlitcp connection something like:
NETTYPE tlitcp,36,250,NET
VPCLASS cpu,num=230,noage
(NOAGE needs a Solaris change to work)
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
>> Do you really have 230 physical CPUs? If not you could be running into process inversion which is making it difficult for the connection file descriptor to be received by all of the CPUVPs ... No, nothing like it. There are 8 cpus each with 32 "threads", seen as 256 "cpus" by the OS. We went through this about 5 years ago with the help of Spokey Wheeler (whatever happened to him?) using a paper written by Simon "Cosmo" David (whatever happened to him?) who argued that you had to be really careful with this type of multi-threading as Informix assumes (you'll know this far better than I so feel free to correct me) that each "cpu" (as it sees it) is an independent processing unit whereas in reality it isn't: some program on the cpu itself is "time-slicing" between its cores. It seems to me that this must be true of any multi-core CPU but presumably architecture like the Solaris T-5 is particularly susceptible? Bottom line is, unless we set cpu vps close to the number of threads we can't push the server very hard at all.
I did take a look at the PMR... The number of CPU VPs was a bit of a
shock...
I've seen your other email.... I have some reservations regarding the fact
that T5 has a lot of threads but not as much "internal pieces" so satisfy a
real database process in each of those "threads". But you say you can't
push it very hard without that... did you look at "onstat -g eff"? I'd be
curious to see that, but i think there were some issues with the
calculation on 11.50...
Also, I have the feeling Art shared some thoughs about this in the past....?
Regarding the number of connections... I do understand the
recommendation... But you're on Solaris... and a few years ago after some
deep investigation of some errors I've found some indicators that on
Solaris, and related to the fact it uses TLI instead of sockets, this value
is used to define the number of pending TCP requests on a specific port
(something along these lines).
But I suppose this wouldn't be related to your problem.... I can drop a
note to Markus if I find this relevant.
The only drawback of an high value is a bit more memory consumption that in
an instance like this is most probably irrelevant.
And I don't understand the rational behind the 36 poll threads... if each
can handle 200-300 connections...15 should be a good value...
On a previous situation I also got the feeling that too many poll threads
wouldn't help...
One very important thing.... do you have lots of new connections per second
or not? This is typically THE trigger for the nsf.lock problem.
Also.... do you have a large ready queue when this happens?
Something I've seen for myself in the past was that occasionally the
session holding the mutex got "stuck" in the ready queue if it get's too
big....
The CPU VP configuration may have impact on this. I'm afraid your "oninit"
processes may be strugling for real CPU resources.
As a final note, many years ago, a customer hitting a similar problem got a
fix that basically changed the number of spins a thread would do before
putting itself on hold waiting for a mutex. The argument was that in
certain cases would be better to "waste" a few more CP cycles instead of
going into the ready queue and struggle to "emerge" again.
That would be a possible path... but it's such a sensitive area that I
would not consider it until everything else was completely cleared.
Still regarding the CPU and the "pushing the server hard".... be careful on
how you measure this... it's better to measure the work done, and not the
idle CPU....
Finally, Spokey AFAIK left IBM and is wokring on his own... it's alive in
facebook at least.
Cosmo is still with us. AFAIK still with Informix. You may mention the
previous history to Makus. Maybe he can shced some more light on this.
On Mon, Feb 29, 2016 at 5:35 PM, NEIL TRUBY <neil.truby@ardenta.com> wrote:
> Thanks. We already knew from Marco Greco, ex of IBM, that the handling of
> this
> in 11.70 is vastly imrpoved. This is the last large customer resisting that
> move.
>
> VPCLASS CPU,NUM=230 # VPCLASS CPU - CPU #T5 Migration
> NETTYPE tlitcp,11,800,NET>
> .... are the settings. The server has 32 physical cores each with 8
> threads.
> 232 of these are assigned to this "Container" (ie VM).
>
> Typically we have 4,500 connections and IBM has suggested reducing the
> number
> of poll threads to 250 per tlitcp connection something like:
>
> NETTYPE tlitcp,36,250,NET
> VPCLASS cpu,num=230,noage>
> (NOAGE needs a Solaris change to work)
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--001a113f3bd4a535c2052cedc450
Hardware processor threads do help, but to what degree depends on the program. They rely on memory fetch stalls to swap in another thread to execute on the core. Database in generally have removed bad memory stalls long before threads were invented. While there will always be stalls, Java programs and webservers will have many more stalls and can take advantage of more processor threads. Due to the database's efficient code, the OS will recognize that threads are not memory stalling enough and forcibly swap a processor threads, if this happens to a cpu vp that is bad, if it happens to a cpu vp holding a mutex this is horrible. If you have many processor threads that are being starved the chance of the OS forcibly swapping out a VP holding a much need mutex is very high. This creates the inversion my esteemed colleague was alluding to Hope this helps, John Sent from my iPhone On Feb 29, 2016, at 10:10 AM, NEIL TRUBY <neil.truby@ardenta.com> wrote: >>> Do you really have 230 physical CPUs? If not you could be running into > process inversion which is making it difficult for the connection file > descriptor to be received by all of the CPUVPs ... > > No, nothing like it. There are 8 cpus each with 32 "threads", seen as 256 > "cpus" by the OS. > > We went through this about 5 years ago with the help of Spokey Wheeler > (whatever happened to him?) using a paper written by Simon "Cosmo" David > (whatever happened to him?) who argued that you had to be really careful with > this type of multi-threading as Informix assumes (you'll know this far better > than I so feel free to correct me) that each "cpu" (as it sees it) is an > independent processing unit whereas in reality it isn't: some program on the > cpu itself is "time-slicing" between its cores. > > It seems to me that this must be true of any multi-core CPU but presumably > architecture like the Solaris T-5 is particularly susceptible? > > Bottom line is, unless we set cpu vps close to the number of threads we can't > push the server very hard at all. > > > ***************************************************************************= **** > Forum Note: Use "Reply" to post a response in the discussion forum. >
Neil:
Ignore the "threads" and just count the number of physical cores. You can
run up to 4 cpu vps per core on T5 chips. The problems with database
performance on the T* chipsets were not completelt solved until the T6 & T7
processors, but tbe T5s are far better than the T1-4 processors were.
Art
Art
On Feb 29, 2016 2:37 PM, "Fernando Nunes" <domusonline@gmail.com> wrote:
> I did take a look at the PMR... The number of CPU VPs was a bit of a
> shock...
> I've seen your other email.... I have some reservations regarding the fact
> that T5 has a lot of threads but not as much "internal pieces" so satisfy a
> real database process in each of those "threads". But you say you can't
> push it very hard without that... did you look at "onstat -g eff"? I'd be
> curious to see that, but i think there were some issues with the
> calculation on 11.50...
>
> Also, I have the feeling Art shared some thoughs about this in the
> past....?
>
> Regarding the number of connections... I do understand the
> recommendation... But you're on Solaris... and a few years ago after some
> deep investigation of some errors I've found some indicators that on
> Solaris, and related to the fact it uses TLI instead of sockets, this value
> is used to define the number of pending TCP requests on a specific port
> (something along these lines).
> But I suppose this wouldn't be related to your problem.... I can drop a
> note to Markus if I find this relevant.
> The only drawback of an high value is a bit more memory consumption that in
> an instance like this is most probably irrelevant.
>
> And I don't understand the rational behind the 36 poll threads... if each
> can handle 200-300 connections...15 should be a good value...
> On a previous situation I also got the feeling that too many poll threads
> wouldn't help...
>
> One very important thing.... do you have lots of new connections per second
> or not? This is typically THE trigger for the nsf.lock problem.
> Also.... do you have a large ready queue when this happens?
> Something I've seen for myself in the past was that occasionally the
> session holding the mutex got "stuck" in the ready queue if it get's too
> big....
> The CPU VP configuration may have impact on this. I'm afraid your "oninit"
> processes may be strugling for real CPU resources.
>
> As a final note, many years ago, a customer hitting a similar problem got a
> fix that basically changed the number of spins a thread would do before
> putting itself on hold waiting for a mutex. The argument was that in
> certain cases would be better to "waste" a few more CP cycles instead of
> going into the ready queue and struggle to "emerge" again.
> That would be a possible path... but it's such a sensitive area that I
> would not consider it until everything else was completely cleared.
>
> Still regarding the CPU and the "pushing the server hard".... be careful on
> how you measure this... it's better to measure the work done, and not the
> idle CPU....
>
> Finally, Spokey AFAIK left IBM and is wokring on his own... it's alive in
> facebook at least.
> Cosmo is still with us. AFAIK still with Informix. You may mention the
> previous history to Makus. Maybe he can shced some more light on this.
>
> On Mon, Feb 29, 2016 at 5:35 PM, NEIL TRUBY <neil.truby@ardenta.com>
> wrote:
>
> > Thanks. We already knew from Marco Greco, ex of IBM, that the handling of
> > this
> > in 11.70 is vastly imrpoved. This is the last large customer resisting
> that
> > move.
> >
> > VPCLASS CPU,NUM=230 # VPCLASS CPU - CPU #T5 Migration
> > NETTYPE tlitcp,11,800,NET> >
> > .... are the settings. The server has 32 physical cores each with 8
> > threads.
> > 232 of these are assigned to this "Container" (ie VM).
> >
> > Typically we have 4,500 connections and IBM has suggested reducing the
> > number
> > of poll threads to 250 per tlitcp connection something like:
> >
> > NETTYPE tlitcp,36,250,NET
> > VPCLASS cpu,num=230,noage> >
> > (NOAGE needs a Solaris change to work)
> >
> >
> >
> >
>
>
*******************************************************************************
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...
>
> --001a113f3bd4a535c2052cedc450
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--047d7bea43fcabc3f5052cefc26b
Perhaps this output (23h of a fairly quiet day) does suggest an excess of
vcpus, as you and Fernando suggest?:
#> onstat -g glo
IBM Informix Dynamic Server Version 11.50.FC8W2XG -- On-Line -- Up 4 days
21:36:09 -- 181565440 Kbytes
MT global info:
sessions threads vps lngspins
4172 4571 249 52381
sched calls thread switches yield 0 yield n yield forever
total: 2339548457 1510104260 866927191 24553758 574356289
per sec: 4773 1256 3529 112 56
Virtual processor summary:
class vps usercpu syscpu total
cpu 230 355973.20 39224.01 395197.21
aio 3 16.41 31.44 47.85
tli 11 11229.75 15119.84 26349.59
lio 1 0.67 1.59 2.26
pio 1 0.68 1.65 2.33
adm 1 64.13 48.12 112.25
msc 1 9.67 6.41 16.08
fifo 1 0.61 1.84 2.45
total 249 367295.12 54434.90 421730.02
Individual virtual processors:
vp pid class usercpu syscpu total Thread Eff
1 14260 cpu 32113.33 4296.34 36409.67 56821.20 64%
2 22966 adm 64.13 48.12 112.25 0.00 0%
3 23000 cpu 35526.07 4025.03 39551.10 62900.86 62%
4 23028 cpu 35104.15 3771.14 38875.29 59487.19 65%
5 23052 cpu 33151.37 3554.63 36706.00 55526.56 66%
6 23085 cpu 31003.00 3328.35 34331.35 50600.00 67%
7 23157 cpu 28009.59 3195.80 31205.39 46715.58 66%
8 23234 cpu 25239.36 2768.71 28008.07 40782.84 68%
9 23317 cpu 21835.58 2364.25 24199.83 34944.57 69%
10 23385 cpu 18587.20 2008.86 20596.06 29000.17 71%
11 23462 cpu 16066.57 1807.27 17873.84 25486.23 70%
12 23562 cpu 13576.85 1410.67 14987.52 20876.59 71%
13 23675 cpu 10522.30 1015.38 11537.68 15688.10 73%
14 23794 cpu 9751.94 947.65 10699.59 14483.11 73%
15 23906 cpu 6791.94 611.73 7403.67 10044.38 73%
16 24008 cpu 5208.83 478.94 5687.77 7658.48 74%
17 24080 cpu 5492.45 490.71 5983.16 7967.34 75%
18 24144 cpu 4265.32 345.48 4610.80 6005.20 76%
19 24196 cpu 3323.42 280.62 3604.04 4722.64 76%
20 24277 cpu 4135.04 357.02 4492.06 5760.58 77%
21 24360 cpu 3199.34 277.11 3476.45 4502.18 77%
22 24415 cpu 2477.45 195.08 2672.53 3380.14 79%
23 24470 cpu 1908.09 159.20 2067.29 2658.57 77%
24 24538 cpu 1107.25 95.54 1202.79 1591.77 75%
25 24614 cpu 699.67 55.40 755.07 1027.14 73%
26 24702 cpu 667.87 54.20 722.07 978.27 73%
27 24801 cpu 235.00 26.78 261.78 421.51 62%
28 24890 cpu 285.63 27.23 312.86 473.44 66%
29 24967 cpu 126.96 16.09 143.05 268.64 53%
30 25044 cpu 110.45 15.14 125.59 244.31 51%
31 25099 cpu 96.55 13.71 110.26 210.78 52%
32 25182 cpu 235.46 25.99 261.45 390.56 66%
33 25189 cpu 66.15 10.92 77.07 169.14 45%
34 25301 cpu 46.38 10.26 56.64 145.72 38%
35 25387 cpu 27.15 9.46 36.61 116.87 31%
36 25473 cpu 26.38 8.76 35.14 115.57 30%
37 25516 cpu 35.33 7.40 42.73 121.14 35%
38 25563 cpu 16.84 7.70 24.54 99.38 24%
39 25608 cpu 21.77 6.70 28.47 99.24 28%
40 25672 cpu 14.79 7.09 21.88 87.49 25%
41 25732 cpu 14.62 6.73 21.35 84.90 25%
42 25782 cpu 14.46 6.55 21.01 82.03 25%
43 25832 cpu 11.90 5.48 17.38 72.58 23%
44 25890 cpu 12.71 6.21 18.92 75.08 25%
45 25969 cpu 7.82 5.08 12.90 63.38 20%
46 26039 cpu 10.59 5.11 15.70 66.63 23%
47 26096 cpu 9.48 5.05 14.53 64.41 22%
48 26185 cpu 7.78 5.28 13.06 63.30 20%
49 26258 cpu 10.85 4.55 15.40 61.39 25%
50 26348 cpu 8.31 4.83 13.14 59.89 21%
51 26443 cpu 6.31 4.54 10.85 56.19 19%
52 26548 cpu 10.37 4.26 14.63 60.40 24%
53 26669 cpu 5.19 5.15 10.34 55.62 18%
54 26829 cpu 6.07 5.19 11.26 53.53 21%
55 26934 cpu 5.37 4.27 9.64 50.03 19%
56 27033 cpu 6.72 4.60 11.32 50.67 22%
57 27130 cpu 4.41 4.26 8.67 47.35 18%
58 27211 cpu 4.39 4.05 8.44 47.58 17%
59 27260 cpu 4.81 5.10 9.91 48.72 20%
60 27295 cpu 4.21 4.89 9.10 45.21 20%
61 27339 cpu 4.73 5.10 9.83 44.40 22%
62 27359 cpu 3.89 4.71 8.60 45.01 19%
63 27377 cpu 3.66 4.77 8.43 43.03 19%
64 27431 cpu 3.62 4.56 8.18 42.17 19%
65 27494 cpu 4.66 4.76 9.42 42.29 22%
66 27531 cpu 4.70 4.10 8.80 41.61 21%
67 27565 cpu 3.70 4.24 7.94 42.07 18%
68 27614 cpu 3.01 4.11 7.12 37.99 18%
69 27698 cpu 2.88 4.62 7.50 38.40 19%
70 27783 cpu 3.12 4.66 7.78 36.49 21%
71 27861 cpu 2.73 4.13 6.86 36.36 18%
72 27945 cpu 2.78 4.40 7.18 36.00 19%
73 28048 cpu 2.52 4.72 7.24 33.93 21%
74 28173 cpu 2.80 4.56 7.36 35.19 20%
75 28288 cpu 2.19 4.41 6.60 33.51 19%
76 28377 cpu 2.75 3.72 6.47 33.09 19%
77 28476 cpu 2.61 4.32 6.93 33.81 20%
78 28577 cpu 2.38 4.55 6.93 30.98 22%
79 28668 cpu 2.74 3.65 6.39 31.86 20%
80 28728 cpu 2.34 4.50 6.84 32.74 20%
81 28811 cpu 2.44 4.08 6.52 32.43 20%
82 28886 cpu 9.02 4.77 13.79 36.95 37%
83 28941 cpu 3.15 4.88 8.03 30.98 25%
84 28991 cpu 3.60 4.37 7.97 30.36 26%
85 29043 cpu 2.66 3.94 6.60 30.82 21%
86 29091 cpu 2.32 4.44 6.76 30.40 22%
87 29138 cpu 2.35 4.44 6.79 27.67 24%
88 29196 cpu 2.32 4.62 6.94 28.04 24%
89 29272 cpu 2.37 4.87 7.24 28.13 25%
90 29366 cpu 2.33 4.80 7.13 29.82 23%
91 29444 cpu 4.67 4.08 8.75 27.19 32%
92 29540 cpu 1.93 4.62 6.55 26.03 25%
93 29664 cpu 1.89 3.82 5.71 25.90 22%
94 29793 cpu 1.94 4.70 6.64 26.48 25%
95 29900 cpu 1.93 4.19 6.12 25.34 24%
96 29952 cpu 2.50 4.46 6.96 24.70 28%
97 29998 cpu 5.13 5.26 10.39 29.67 35%
98 60 cpu 2.08 4.56 6.64 24.55 27%
99 127 cpu 2.16 5.16 7.32 25.44 28%
100 188 cpu 3.45 4.27 7.72 23.46 32%
101 235 cpu 1.74 4.29 6.03 21.64 27%
102 288 cpu 2.01 4.75 6.76 23.57 28%
103 354 cpu 1.93 3.80 5.73 22.98 24%
104 460 cpu 2.87 3.86 6.73 22.98 29%
105 585 cpu 1.80 3.69 5.49 20.87 26%
106 696 cpu 1.84 3.80 5.64 21.22 26%
107 768 cpu 1.61 3.61 5.22 19.94 26%
108 816 cpu 1.54 3.61 5.15 20.26 25%
109 875 cpu 2.14 3.71 5.85 21.22 27%
110 936 cpu 2.44 4.04 6.48 23.57 27%
111 1028 cpu 2.35 3.81 6.16 21.31 28%
112 1115 cpu 1.78 4.29 6.07 21.06 28%
113 1234 cpu 2.07 4.05 6.12 22.35 27%
114 1357 cpu 1.85 3.97 5.82 20.57 28%
115 1462 cpu 1.87 3.25 5.12 20.33 25%
116 1551 cpu 2.30 3.82 6.12 20.88 29%
117 1665 cpu 2.07 4.17 6.24 19.43 32%
118 1798 cpu 1.79 4.56 6.35 19.53 32%
119 1896 cpu 2.14 3.82 5.96 18.24 32%
120 2005 cpu 1.74 3.80 5.54 18.97 29%
121 2097 cpu 1.95 3.99 5.94 19.58 30%
122 2182 cpu 1.46 3.68 5.14 17.78 28%
123 2249 cpu 1.73 3.81 5.54 18.24 30%
124 2287 cpu 1.83 4.54 6.37 18.91 33%
125 2305 cpu 2.07 4.50 6.57 17.81 36%
126 2324 cpu 1.93 4.12 6.05 16.83 35%
127 2360 cpu 1.97 3.99 5.96 19.30 30%
128 2379 cpu 1.77 4.57 6.34 16.50 38%
129 2443 cpu 2.20 4.38 6.58 16.87 38%
130 2543 cpu 1.96 4.35 6.31 18.12 34%
131 2605 cpu 1.81 3.72 5.53 16.41 33%
132 2679 cpu 1.87 4.08 5.95 17.14 34%
133 2778 cpu 1.86 4.29 6.15 16.39 37%
134 2907 cpu 1.76 3.92 5.68 16.29 34%
135 3007 cpu 1.65 3.47 5.12 15.10 33%
136 3075 cpu 1.67 3.70 5.37 15.10 35%
137 3155 cpu 1.82 3.81 5.63 15.25 36%
138 3246 cpu 1.64 3.74 5.38 15.90 33%
139 3316 cpu 2.17 3.72 5.89 17.34 33%
140 3403 cpu 2.22 4.23 6.45 14.59 44%
141 3509 cpu 1.82 4.39 6.21 17.55 35%
142 3625 cpu 1.96 3.70 5.66 13.97 40%
143 3730 cpu 1.67 3.72 5.39 15.09 35%
144 3833 cpu 1.84 3.49 5.33 15.04 35%
145 3927 cpu 1.97 4.30 6.27 13.23 47%
146 4000 cpu 1.80 3.89 5.69 14.87 38%
147 4069 cpu 1.88 3.76 5.64 13.09 43%
148 4129 cpu 1.63 3.64 5.27 12.97 40%
149 4186 cpu 1.85 4.07 5.92 13.94 42%
150 4233 cpu 2.72 3.47 6.19 13.29 46%
151 4290 cpu 1.95 3.43 5.38 15.17 35%
152 4361 cpu 1.81 3.11 4.92 15.48 31%
153 4476 cpu 1.48 3.62 5.10 10.96
blockquote, div.yahoo_quoted { margin-left: 0 !important; border-left:1px
#715FFA solid !important; padding-left:1ex !important; background-color:white
!important; } That would be my analysis...
Sent from Yahoo Mail for iPad
On Monday, February 29, 2016, 6:46 PM, NEIL TRUBY <neil.truby@ardenta.com>
wrote:
Perhaps this output (23h of a fairly quiet day) does suggest an excess of
vcpus, as you and Fernando suggest?:
#> onstat -g glo
IBM Informix Dynamic Server Version 11.50.FC8W2XG -- On-Line -- Up 4 days
21:36:09 -- 181565440 Kbytes
MT global info:
sessions threads vps lngspins
4172 4571 249 52381
sched calls thread switches yield 0 yield n yield forever
total: 2339548457 1510104260 866927191 24553758 574356289
per sec: 4773 1256 3529 112 56
Virtual processor summary:
class vps usercpu syscpu total
cpu 230 355973.20 39224.01 395197.21
aio 3 16.41 31.44 47.85
tli 11 11229.75 15119.84 26349.59
lio 1 0.67 1.59 2.26
pio 1 0.68 1.65 2.33
adm 1 64.13 48.12 112.25
msc 1 9.67 6.41 16.08
fifo 1 0.61 1.84 2.45
total 249 367295.12 54434.90 421730.02
Individual virtual processors:
vp pid class usercpu syscpu total Thread Eff
1 14260 cpu 32113.33 4296.34 36409.67 56821.20 64%
2 22966 adm 64.13 48.12 112.25 0.00 0%
3 23000 cpu 35526.07 4025.03 39551.10 62900.86 62%
4 23028 cpu 35104.15 3771.14 38875.29 59487.19 65%
5 23052 cpu 33151.37 3554.63 36706.00 55526.56 66%
6 23085 cpu 31003.00 3328.35 34331.35 50600.00 67%
7 23157 cpu 28009.59 3195.80 31205.39 46715.58 66%
8 23234 cpu 25239.36 2768.71 28008.07 40782.84 68%
9 23317 cpu 21835.58 2364.25 24199.83 34944.57 69%
10 23385 cpu 18587.20 2008.86 20596.06 29000.17 71%
11 23462 cpu 16066.57 1807.27 17873.84 25486.23 70%
12 23562 cpu 13576.85 1410.67 14987.52 20876.59 71%
13 23675 cpu 10522.30 1015.38 11537.68 15688.10 73%
14 23794 cpu 9751.94 947.65 10699.59 14483.11 73%
15 23906 cpu 6791.94 611.73 7403.67 10044.38 73%
16 24008 cpu 5208.83 478.94 5687.77 7658.48 74%
17 24080 cpu 5492.45 490.71 5983.16 7967.34 75%
18 24144 cpu 4265.32 345.48 4610.80 6005.20 76%
19 24196 cpu 3323.42 280.62 3604.04 4722.64 76%
20 24277 cpu 4135.04 357.02 4492.06 5760.58 77%
21 24360 cpu 3199.34 277.11 3476.45 4502.18 77%
22 24415 cpu 2477.45 195.08 2672.53 3380.14 79%
23 24470 cpu 1908.09 159.20 2067.29 2658.57 77%
24 24538 cpu 1107.25 95.54 1202.79 1591.77 75%
25 24614 cpu 699.67 55.40 755.07 1027.14 73%
26 24702 cpu 667.87 54.20 722.07 978.27 73%
27 24801 cpu 235.00 26.78 261.78 421.51 62%
28 24890 cpu 285.63 27.23 312.86 473.44 66%
29 24967 cpu 126.96 16.09 143.05 268.64 53%
30 25044 cpu 110.45 15.14 125.59 244.31 51%
31 25099 cpu 96.55 13.71 110.26 210.78 52%
32 25182 cpu 235.46 25.99 261.45 390.56 66%
33 25189 cpu 66.15 10.92 77.07 169.14 45%
34 25301 cpu 46.38 10.26 56.64 145.72 38%
35 25387 cpu 27.15 9.46 36.61 116.87 31%
36 25473 cpu 26.38 8.76 35.14 115.57 30%
37 25516 cpu 35.33 7.40 42.73 121.14 35%
38 25563 cpu 16.84 7.70 24.54 99.38 24%
39 25608 cpu 21.77 6.70 28.47 99.24 28%
40 25672 cpu 14.79 7.09 21.88 87.49 25%
41 25732 cpu 14.62 6.73 21.35 84.90 25%
42 25782 cpu 14.46 6.55 21.01 82.03 25%
43 25832 cpu 11.90 5.48 17.38 72.58 23%
44 25890 cpu 12.71 6.21 18.92 75.08 25%
45 25969 cpu 7.82 5.08 12.90 63.38 20%
46 26039 cpu 10.59 5.11 15.70 66.63 23%
47 26096 cpu 9.48 5.05 14.53 64.41 22%
48 26185 cpu 7.78 5.28 13.06 63.30 20%
49 26258 cpu 10.85 4.55 15.40 61.39 25%
50 26348 cpu 8.31 4.83 13.14 59.89 21%
51 26443 cpu 6.31 4.54 10.85 56.19 19%
52 26548 cpu 10.37 4.26 14.63 60.40 24%
53 26669 cpu 5.19 5.15 10.34 55.62 18%
54 26829 cpu 6.07 5.19 11.26 53.53 21%
55 26934 cpu 5.37 4.27 9.64 50.03 19%
56 27033 cpu 6.72 4.60 11.32 50.67 22%
57 27130 cpu 4.41 4.26 8.67 47.35 18%
58 27211 cpu 4.39 4.05 8.44 47.58 17%
59 27260 cpu 4.81 5.10 9.91 48.72 20%
60 27295 cpu 4.21 4.89 9.10 45.21 20%
61 27339 cpu 4.73 5.10 9.83 44.40 22%
62 27359 cpu 3.89 4.71 8.60 45.01 19%
63 27377 cpu 3.66 4.77 8.43 43.03 19%
64 27431 cpu 3.62 4.56 8.18 42.17 19%
65 27494 cpu 4.66 4.76 9.42 42.29 22%
66 27531 cpu 4.70 4.10 8.80 41.61 21%
67 27565 cpu 3.70 4.24 7.94 42.07 18%
68 27614 cpu 3.01 4.11 7.12 37.99 18%
69 27698 cpu 2.88 4.62 7.50 38.40 19%
70 27783 cpu 3.12 4.66 7.78 36.49 21%
71 27861 cpu 2.73 4.13 6.86 36.36 18%
72 27945 cpu 2.78 4.40 7.18 36.00 19%
73 28048 cpu 2.52 4.72 7.24 33.93 21%
74 28173 cpu 2.80 4.56 7.36 35.19 20%
75 28288 cpu 2.19 4.41 6.60 33.51 19%
76 28377 cpu 2.75 3.72 6.47 33.09 19%
77 28476 cpu 2.61 4.32 6.93 33.81 20%
78 28577 cpu 2.38 4.55 6.93 30.98 22%
79 28668 cpu 2.74 3.65 6.39 31.86 20%
80 28728 cpu 2.34 4.50 6.84 32.74 20%
81 28811 cpu 2.44 4.08 6.52 32.43 20%
82 28886 cpu 9.02 4.77 13.79 36.95 37%
83 28941 cpu 3.15 4.88 8.03 30.98 25%
84 28991 cpu 3.60 4.37 7.97 30.36 26%
85 29043 cpu 2.66 3.94 6.60 30.82 21%
86 29091 cpu 2.32 4.44 6.76 30.40 22%
87 29138 cpu 2.35 4.44 6.79 27.67 24%
88 29196 cpu 2.32 4.62 6.94 28.04 24%
89 29272 cpu 2.37 4.87 7.24 28.13 25%
90 29366 cpu 2.33 4.80 7.13 29.82 23%
91 29444 cpu 4.67 4.08 8.75 27.19 32%
92 29540 cpu 1.93 4.62 6.55 26.03 25%
93 29664 cpu 1.89 3.82 5.71 25.90 22%
94 29793 cpu 1.94 4.70 6.64 26.48 25%
95 29900 cpu 1.93 4.19 6.12 25.34 24%
96 29952 cpu 2.50 4.46 6.96 24.70 28%
97 29998 cpu 5.13 5.26 10.39 29.67 35%
98 60 cpu 2.08 4.56 6.64 24.55 27%
99 127 cpu 2.16 5.16 7.32 25.44 28%
100 188 cpu 3.45 4.27 7.72 23.46 32%
101 235 cpu 1.74 4.29 6.03 21.64 27%
102 288 cpu 2.01 4.75 6.76 23.57 28%
103 354 cpu 1.93 3.80 5.73 22.98 24%
104 460 cpu 2.87 3.86 6.73 22.98 29%
105 585 cpu 1.80 3.69 5.49 20.87 26%
106 696 cpu 1.84 3.80 5.64 21.22 26%
107 768 cpu 1.61 3.61 5.22 19.94 26%
108 816 cpu 1.54 3.61 5.15 20.26 25%
109 875 cpu 2.14 3.71 5.85 21.22 27%
110 936 cpu 2.44 4.04 6.48 23.57 27%
111 1028 cpu 2.35 3.81 6.16 21.31 28%
112 1115 cpu 1.78 4.29 6.07 21.06 28%
113 1234 cpu 2.07 4.05 6.12 22.35 27%
114 1357 cpu 1.85 3.97 5.82 20.57 28%
115 1462 cpu 1.87 3.25 5.12 20.33 25%
116 1551 cpu 2.30 3.82 6.12 20.88 29%
117 1665 cpu 2.07 4.17 6.24 19.43 32%
118 1798 cpu 1.79 4.56 6.35 19.53 32%
119 1896 cpu 2.14 3.82 5.96 18.24 32%
120 2005 cpu 1.74 3.80 5.54 18.97 29%
121 2097 cpu 1.95 3.99 5.94 19.58 30%
122 2182 cpu 1.46 3.68 5.14 17.78 28%
123 2249 cpu 1.73 3.81 5.54 18.24 30%
124 2287 cpu 1.83 4.54 6.37 18.91 33%
125 2305 cpu 2.07 4.50 6.57 17.81 36%
126 2324 cpu 1.93 4.12 6.05 16.83 35%
127 2360 cpu 1.97 3.99 5.96 19.30 30%
128 2379 cpu 1.77 4.57 6.34 16.50 38%
129 2443 cpu 2.20 4.38 6.58 16.87 38%
130 2543 cpu 1.96 4.35 6.31 18.12 34%
131 2605 cpu 1.81 3.72 5.53 16.41 33%
132 2679 cpu 1.87 4.08 5.95 17.14 34%
133 2778 cpu 1.86 4.29 6.15 16.39 37%
134 2907 cpu 1.76 3.92 5.68 16.29 34%
135 3007 cpu 1.65 3.47 5.12 15.10 33%
136 3075 cpu 1.67 3.70 5.37 15.10 35%
137 3155 cpu 1.82 3.81 5.63 15.25 36%
138 3246 cpu 1.64 3.74 5.38 15.90 33%
139 3316 cpu 2.17 3.72 5.89 17.34 33%
140 3403 cpu 2.22 4.23 6.45 14.59 44%
141 3509 cpu 1.82 4.39 6.21 17.55 35%
142 3625 cpu 1.96 3.70 5.66 13.97 40%
143 3730 cpu 1.67 3.72 5.39 15.09 35%
144 3833 cpu 1.84 3.49 5.33 15.04 35%
145 3927 cpu 1.97 4.30
Hi Neil, what I used to do in the past when on SOL11, before T5 bacame available: - compare sum of total of the cpuvps to the sum of total of non-cpuvps This shows tli using 6.7% of sum of total and has to be considered when it comes to parallel execution. You should investigate why syscpu > usercpu in all tli data, but this is low prio on your ToDO list IMHO. - grep out the cpuvps lines, then sort descending using total as key. You wrote ' .... 23h of a fairly quiet day ' and you can see, that the top 20 cpuvps do more than 95% of all work or in other words only 20 cpuvps use more than 1% of sum of total of the cpuvps. Therefore either you only have 'a few' situations when workload requires more than 20 cpuvps, OR you see a bottleneck in receiving requests for database action. I used end-to-end monitoring transactions to see if the expected average duration of these TXes is normal and recorded the values once every 30 seconds. Finally it all comes down to single thread performance. In an environment optimized for leight weight processes the database usually does not get the love it needs. Seen from the OS, INFORMIX is a bunch of batch jobs consuming much of the CPU time when it comes to run-ready-queue sorting or priority aging calculations in the OS. This is a common problem for all type of legacy software nowadays. And noage does help indeed, but also binding will help. Both will not elimiate the problems, though. You can have proof for this when you are able to use a similar setup for a simple test: Start a number of dummy batch jobs, which do consume CPU resource, and cause comparable small I/O load. - first start as many test jobs as you have cores. - then double it - then double it again and record the wall clock duration time to see what actually can run in parallel You will see very interesting data, as it does NOT scale at all. dic_k now toying with V12.10 to show, that even the RESTful usage of our database is way faster than all other backends, when you know about tuning.
Im still here, you cheeky bastard. > On 29 Feb 2016, at 18:09, NEIL TRUBY <neil.truby@ardenta.com> wrote: > >>> Do you really have 230 physical CPUs? If not you could be running into > process inversion which is making it difficult for the connection file > descriptor to be received by all of the CPUVPs ... > > No, nothing like it. There are 8 cpus each with 32 "threads", seen as 256 > "cpus" by the OS. > > We went through this about 5 years ago with the help of Spokey Wheeler > (whatever happened to him?) using a paper written by Simon "Cosmo" David > (whatever happened to him?) who argued that you had to be really careful with > this type of multi-threading as Informix assumes (you'll know this far better > than I so feel free to correct me) that each "cpu" (as it sees it) is an > independent processing unit whereas in reality it isn't: some program on the > cpu itself is "time-slicing" between its cores. > > It seems to me that this must be true of any multi-core CPU but presumably > architecture like the Solaris T-5 is particularly susceptible? > > Bottom line is, unless we set cpu vps close to the number of threads we can't > push the server very hard at all. > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. >
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g