Many threads, some cores, not many sockets
Posted in 2016
This is a continuation of my earlier post "Mutex logs on logical log flip", prompted by musings from Madison, Art and Fernando amongst others. IBM Informix Dynamic Server Version 11.50.FC8W2XG on Solaris. Problem was that we're having sporadic slowdowns for nsf.lock mutexes (which we know 11.50 is prone to), seemingly being load-related, and being triggered whenever NetBackup pops up to back up a logical log. The physical server (a Sun/Oracle T5-2) has 2 sockets (let's call them 0 and 1); each with 16 cores (0 to 15 on each socket); each core has 8 threads presenting 256 threads in all (0-255). (Yes I know I'm mixing conventions but this is how Solaris does it!). We have 3 "zones", with threads allocated thus: Threads 0-7 ie Socket 0 Core 0 OLTP DB zone Threads 8-15 ie Socket 0 Core 1 archive DB zone Threads 16-239 ie Socket 1 Core 2-7 ; Socket 1 Core 0-5 OLTP DB zone Threads 240-255 ie Socket 1 Cores 6-7 global zone Our problems are in the OLTP zone, which can be seen to be occupying all 8 threads on 29 of the 32 cores; the archive db occupies all the threads on 1 core and the global zone all the threads on the remaining 2 cores. It's been very difficult to get any info on the extent to which each thread is a separate processing unit (flagged up as an important consideration by Cosmo and Spokey). As was pointed out our current config is showing many of our 229 cpu vps little-utilised. Perhaps a better strategy would be to allocate to the OLTP zone only (say) 4 threads from each core - so 116, and set cpu vps to 113, with the objective of removing competition between threads on the same core at the Solaris level ... I guess as well as inviting general comment my question is "Who *does* know all this stuff?". Presumably someone at Oracle rather than IBM? But (haven't tried for a few years) the Sun engineers were reluctant to admit that all 256 threads couldn't be exploited at once.