Configuring IDS for multi-threaded processors
Posted in 2009
Question: how many CPU VPs to configure for IDS 10.0 on a Sun M9000 (16 quad-core, dual-thread SPARC VII+ chips, 128 memory-threads visible to Solaris) — follow the usual 'cores minus one' rule, or stay closer to the physical chip count? Replies reported that scaling is roughly linear only up to the number of sockets (16), improves little up to ~64, and degrades beyond that, pointing at memory/bus throughput limits rather than Informix itself. Suggestions included benchmarking at 16/32/48/64/96/127 VPs, binding CPU/AIO/NET VPs across sockets and cores, checking onstat -g glo and -g spi for latch contention, and involving IBM labs. No definitive setting or resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Platform-Specific Issues, Versions, Editions & End-of-Life
Has anyone views and experience on setting cpu vps for multi-threaded chips? One customer is running on IDS 10.0FC9 on Solaris 10 on a Sun server with 16 quad-core, dual-thread, SPARC VII+ cpus. Solaris "sees" this as 128 cpus. I suppose common orthodoxy might set cpu vps to one fewer than the 128. On the other hand, in conversations with IBMers whose opinions I respect enormously, the suggestion is that, on multi-threaded chips, cpu vps should be set cautiously, perhaps close to the number of physcial cpus. Any thoughts? Or, experiences from real life? thx
Neil Truby schrieb: > Has anyone views and experience on setting cpu vps for multi-threaded > chips? > One customer is running on IDS 10.0FC9 on Solaris 10 on a Sun server > with 16 quad-core, dual-thread, SPARC VII+ cpus. > Solaris "sees" this as 128 cpus. > > I suppose common orthodoxy might set cpu vps to one fewer than the 128. > On the other hand, in conversations with IBMers whose opinions I respect > enormously, the suggestion is that, on multi-threaded chips, cpu vps > should be set cautiously, perhaps close to the number of physcial cpus. > > Any thoughts? Or, experiences from real life? > > thx Hi Neil, what is the name og the machine? How much memory does it have, and how is the memory set up in the harware of the motherboards? One thing, which did hit 2 customers of mine, both on M9000, is that memory thruput in the shared memory, where the bufferpools are, is such, that reducing BUFFERS and doing more I/O, but assisting this I/O with DASD SSD (CMOS, that is *real* SSDs, not flash mem junk SSDs) was approx 4 times faster at 12000 IOPS per SSD and using 4 of them. In other words: I've seen a sum of mem thruput of only 9GB/sec for *4* threads one one CPU, which is per CPU-thread only 10% more than my laptop does, on which I write this. A test is some work, but clarifies it. Cache a table, which is fragmented 128 fold and scan it not unsing any index. You must monitor LRU contention, and subtract that. I had to find out several times, that incore parallelism only scaled linear until the number of sockets is reached, so in your case this is 16. [ going from 15 fragments to 16 will add another amount like the initial amount was of scanned records/sec - or almost as many as the not fragmented table is scanned ] Next scaling goes up to 64, but expect an additional 0.6 of the initial amount, when going from 16 fragments to 17. Then is gets even worse after 64 - like adding one fragment gives you only 0.20 * initial scan rate additional performance (in terms of scanned recs/second) instead of initial recs/sec if you add fragment 65. Example: Assuming you have 1 CPUvp per frag plus 1 unfragmeted: 1 seq scan is 1000 recs/sec total sum of thruput 2 frags: 2000 3 frags: 3000 ..... 16 frags: 16000 this will not decrease until 16 or almost not decrease 17 frags: 16600 18 frags: 17200 ........ 64 frags: 44800 expect the additional 600/sec to decrease somewhat 65 frags: 45000 this goes down from +200 with every added fragment. But it depens heavily on the physical setup of your motherboards. The same behaviour can hit you when it comes to thruput of PCI (the way to your disks) ..... This depens on the bus layout, which is not documented in detail :( I was not able to saturate I/O on a big player's I/O subsystem connected via trunced infiniband (4 times 8Gbit), multpathing and only one M5xxx host. Needed 3 hosts to get to saturation point in the sunsystem box. And no, it was not EMC. Because of all this on the big and very expensive boxes hear me say: If you need performance you must learn to spend less money. And a corrolarium of this is: expect 40% addition database performance, if you spend 30% less bucks. dic_k -- Richard Kofler SOLID STATE EDV Dienstleistungen GmbH Vienna/Austria/Europe
"Richard Kofler" <richard.kofler@chello.at> wrote in message news:36527$4a2a7165$547019ee$25794@news.chello.at... > Neil Truby schrieb: >> Has anyone views and experience on setting cpu vps for multi-threaded >> chips? >> One customer is running on IDS 10.0FC9 on Solaris 10 on a Sun server with >> 16 quad-core, dual-thread, SPARC VII+ cpus. >> Solaris "sees" this as 128 cpus. >> >> I suppose common orthodoxy might set cpu vps to one fewer than the 128. >> On the other hand, in conversations with IBMers whose opinions I respect >> enormously, the suggestion is that, on multi-threaded chips, cpu vps >> should be set cautiously, perhaps close to the number of physcial cpus. >> >> Any thoughts? Or, experiences from real life? >> >> thx > > Hi Neil, > > what is the name og the machine? > How much memory does it have, and how is the memory set up > in the harware of the motherboards? It's an M9000 with 128g memory. /usr/sbin/prtdiag -v | more System Configuration: Sun Microsystems sun4u Sun SPARC Enterprise M9000 Server System clock frequency: 960 MHz Memory size: 131072 Megabytes ==================================== CPUs ==================================== CPU CPU Run L2$ CPU CPU LSB Chip ID MHz MB Impl. Mask --- ---- ---------------------------------------- ---- --- ----- ---- 00 0 0, 1, 2, 3, 4, 5, 6, 7 2520 6.0 7 145 00 1 8, 9, 10, 11, 12, 13, 14, 15 2520 6.0 7 145 00 2 16, 17, 18, 19, 20, 21, 22, 23 2520 6.0 7 145 00 3 24, 25, 26, 27, 28, 29, 30, 31 2520 6.0 7 145 01 0 32, 33, 34, 35, 36, 37, 38, 39 2520 6.0 7 145 01 1 40, 41, 42, 43, 44, 45, 46, 47 2520 6.0 7 145 01 2 48, 49, 50, 51, 52, 53, 54, 55 2520 6.0 7 145 01 3 56, 57, 58, 59, 60, 61, 62, 63 2520 6.0 7 145 02 0 64, 65, 66, 67, 68, 69, 70, 71 2520 6.0 7 145 02 1 72, 73, 74, 75, 76, 77, 78, 79 2520 6.0 7 145 02 2 80, 81, 82, 83, 84, 85, 86, 87 2520 6.0 7 145 02 3 88, 89, 90, 91, 92, 93, 94, 95 2520 6.0 7 145 03 0 96, 97, 98, 99, 100, 101, 102, 103 2520 6.0 7 145 03 1 104, 105, 106, 107, 108, 109, 110, 111 2520 6.0 7 145 03 2 112, 113, 114, 115, 116, 117, 118, 119 2520 6.0 7 145 03 3 120, 121, 122, 123, 124, 125, 126, 127 2520 6.0 7 145 ============================ Memory Configuration ============================ Memory Available Memory DIMM # of Mirror Interleave LSB Group Size Status Size DIMMs Mode Factor --- ------ ------------------ ------- ------ ----- ------- ---------- 00 A 32768MB okay 2048MB 16 no 8-way 01 A 32768MB okay 2048MB 16 no 8-way 02 A 32768MB okay 2048MB 16 no 8-way 03 A 32768MB okay 2048MB 16 no 8-way
> From: neil.truby@ardenta.com > Subject: Re: Configuring IDS for multi-threaded processors > Date: Sat, 6 Jun 2009 17:48:32 +0100 > To: informix-list@iiug.org > > > "Richard Kofler" <richard.kofler@chello.at> wrote in message > news:36527$4a2a7165$547019ee$25794@news.chello.at... > > Neil Truby schrieb: > >> Has anyone views and experience on setting cpu vps for multi-threaded > >> chips? > >> One customer is running on IDS 10.0FC9 on Solaris 10 on a Sun server with > >> 16 quad-core, dual-thread, SPARC VII+ cpus. > >> Solaris "sees" this as 128 cpus. > >> > >> I suppose common orthodoxy might set cpu vps to one fewer than the 128. > >> On the other hand, in conversations with IBMers whose opinions I respect > >> enormously, the suggestion is that, on multi-threaded chips, cpu vps > >> should be set cautiously, perhaps close to the number of physcial cpus. > >> > >> Any thoughts? Or, experiences from real life? > >> > >> thx > > > > Hi Neil, > > > > what is the name og the machine? > > How much memory does it have, and how is the memory set up > > in the harware of the motherboards? > > It's an M9000 with 128g memory. > /usr/sbin/prtdiag -v | more > System Configuration: Sun Microsystems sun4u Sun SPARC Enterprise M9000 > Server > System clock frequency: 960 MHz > Memory size: 131072 Megabytes > What's the rule of thumb? N-1 where N is the number of CPUs? If that's true, then you'd want to do 63vcpus and tune it to that configuration. I'm not sure by what is meant that Sun sees each core as 2 virtual cpus. Do they say that there's two process stacks per core? Are they saying that their cores have 2 different sets of paths to external hardware not part of the core, like memory? (I think the latter and not the former) A processor stack at the core level would allow for core affinity, or at least a preference for a process to stay on the core on which it was started but would have the option to move if necessary. Also how much parallelism exists within Informix? One of the 'gripes' of the IT community is that most of the software written today can no take advantage of the parallelism now offered by the higher number of cores in a SMP machine. (That the hardware has outpaced the software.) If you have the time, do the following: Test 32 vcpus, 48 vcpus, 64 vcpus, and then 96 vcpus. Tune for each set up and run some tests. I would suspect that the performance to be near linear to a point and then drop due to a combination of software and hardware issues. It would be interesting to see if the scaling is near linear through 96 and then drops off around 127/128 vcpus. (You did say rule of thumb was one less...) My guess is that the Sparc VII (Jupiter) series, 4 cores 8 threads, will top out at 63/64 vcpus. (Since the cores have 2 threads, I think you could safely go to 64.) Again, I base this on IDS, the OS and the hardware limitations. I have my reasons for thinking this and I think that it would be possible to modify IDS to run more efficiently on Sun, however, I doubt that this would happen unless IBM was also planning on going down the Rock path. Note: With Oracle's purchase of Sun, you can better believe that there will be a massive rewrite of their architecture to perform better on Rock. How much of a rewrite will depend on what Oracle looks like to start with. If successful, Oracle will have a distinct advantage over anyone. But hey! What do I know? ;-) I mean if a silly git like Clive can tell me I'm full of shit, then what I say doesn't matter. ;-) Of course if Clive is wrong, then you'd better think about what I said, just take it with a grain of salt. There are two messages here. 1) A suggestion to test the scalability of Sun and IDS. 2) IBM's SWG management has a new serious competitor in Oracle/Sun. Post acquisitions, if Larry does what I think he's going to do, IBM's going to be hurting. _________________________________________________________________ Lauren found her dream laptop. Find the PC that’s right for you. http://www.microsoft.com/windows/choosepc/?ocid=ftp_val_wl_290
"Ian Michael Gumby" <im_gumby@hotmail.com> wrote in message news:mailman.6.1244378383.4791.informix-list@iiug.org... >> What's the rule of thumb? N-1 where N is the number of CPUs? Depends upon your defintion of "cpu", doesn't it? >>If that's true, then you'd want to do 63vcpus and tune it to that >>configuration. >> I'm not sure by what is meant that Sun sees each core as 2 virtual cpus. I didn't say "virtual cpus". Using psrinfo you see each of the 127 threads listed, that is what I meant. cheers N
> From: neil.truby@ardenta.com > Subject: Re: Configuring IDS for multi-threaded processors > Date: Sun, 7 Jun 2009 14:02:10 +0100 > To: informix-list@iiug.org > > > "Ian Michael Gumby" <im_gumby@hotmail.com> wrote in message > news:mailman.6.1244378383.4791.informix-list@iiug.org... > > >> What's the rule of thumb? N-1 where N is the number of CPUs? > > Depends upon your defintion of "cpu", doesn't it? > > >>If that's true, then you'd want to do 63vcpus and tune it to that > >>configuration. > >> I'm not sure by what is meant that Sun sees each core as 2 virtual cpus. > > I didn't say "virtual cpus". Using psrinfo you see each of the 127 threads > listed, that is what I meant. > Well what do you call it when there are 64 cores and 128 threads. Sun's lit says each core has 2 threads. I guess the real question is if you tune based on thread or based on core count. I seem to recall some criticism against Sun's design, but I don't know if it was fair or FUD from a competitor. This is why I was suggesting that you start your configurations with 32, 48, 64, 96, 127 vcpus and do some basic tuning. It would be very interesting to see if there is near linear scalability. Note that the other poster who responded, Richard K, indicated that there were some memory constraints. " I had to find out several times, that incore parallelism only scaled linear until the number of sockets is reached, so in your case this is 16. [ going from 15 fragments to 16 will add another amount like the initial amount was of scanned records/sec - or almost as many as the not fragmented table is scanned ] Next scaling goes up to 64, but expect an additional 0.6 of the initial amount, when going from 16 fragments to 17. Then is gets even worse after 64 - like adding one fragment gives you only 0.20 * initial scan rate additional performance (in terms of scanned recs/second) instead of initial " This could be related to some of the FUD comments that were made. (So maybe starting with 16 vcpus maybe the best starting point? ) The key is that if you have the machine and the time, even if you do only basic tuning, I would hope that you see a roughly linear line. That would be the interesting thing. Does this make sense? Just forget the clown. Unless he wants to talk about tuning, he's just being in a pissy mood. -G PS. I'm assuming that Richard was talking about Fragmenting a table. So does it make sense to create 64+ vcpus even if your tables are going to be limited to 16 frags? I mean if your system is under a heavy db load, could you have multiple people hitting multiple tables and utilize the additional vcpus? (Or your instance is supporting multiple databases?) Not that I could afford a Sun box with that much memory or cpus. :-( (My wife keeps spending me in to the poor house... :-( ) _________________________________________________________________ Lauren found her dream laptop. Find the PC that’s right for you. http://www.microsoft.com/windows/choosepc/?ocid=ftp_val_wl_290
Ian Michael Gumby schrieb: > > >> From: neil.truby@ardenta.com >> Subject: Re: Configuring IDS for multi-threaded processors >> Date: Sun, 7 Jun 2009 14:02:10 +0100 >> To: informix-list@iiug.org >> >> >> "Ian Michael Gumby" <im_gumby@hotmail.com> wrote in message >> news:mailman.6.1244378383.4791.informix-list@iiug.org... >> >>>> What's the rule of thumb? N-1 where N is the number of CPUs? >> Depends upon your defintion of "cpu", doesn't it? >> >>>> If that's true, then you'd want to do 63vcpus and tune it to that >>>> configuration. >>>> I'm not sure by what is meant that Sun sees each core as 2 virtual cpus. >> I didn't say "virtual cpus". Using psrinfo you see each of the 127 threads >> listed, that is what I meant. >> > > Well what do you call it when there are 64 cores and 128 threads. Sun's lit says each core has 2 threads. > > I guess the real question is if you tune based on thread or based on core count. I seem to recall some criticism against Sun's design, but I don't know if it was fair or FUD from a competitor. > > This is why I was suggesting that you start your configurations with 32, 48, 64, 96, 127 vcpus and do some basic tuning. It would be very interesting to see if there is near linear scalability. > > Note that the other poster who responded, Richard K, indicated that there were some memory constraints. > " > I had to find out several times, that incore parallelism only > scaled linear until the number of sockets is reached, so in your > case this is 16. [ going from 15 fragments to 16 will add another > amount like the initial amount was of scanned records/sec - > or almost as many as the not fragmented table is scanned ] > Next scaling goes up to 64, but expect an additional 0.6 of the > initial amount, when going from 16 fragments to 17. > Then is gets even worse after 64 - like adding one fragment gives you only > 0.20 * initial scan rate additional performance > (in terms of scanned recs/second) instead of initial > > " > > This could be related to some of the FUD comments that were made. > > (So maybe starting with 16 vcpus maybe the best starting point? ) > > The key is that if you have the machine and the time, even if you do only basic tuning, > I would hope that you see a roughly linear line. That would be the interesting thing. > > Does this make sense? > Just forget the clown. Unless he wants to talk about tuning, he's just being in a pissy mood. > > -G > > PS. > I'm assuming that Richard was talking about Fragmenting a table. So does it make sense to > create 64+ vcpus even if your tables are going to be limited to 16 frags? I mean if your > system is under a heavy db load, could you have multiple people hitting multiple tables and > utilize the additional vcpus? (Or your instance is supporting multiple databases?) > > Not that I could afford a Sun box with that much memory or cpus. :-( (My wife keeps spending > me in to the poor house... :-( ) > > > _________________________________________________________________ > Lauren found her dream laptop. Find the PC that's right for you. > http://www.microsoft.com/windows/choosepc/?ocid=ftp_val_wl_290 Hi people, here is another test result of what I did, and it is much easier to explain in my bad English, than the more complicated testing with precached fragmented tables. We had a database batch job and we were able to precache almost all data needed for this batch. I/O was well under 10 IOPS. When it run alone it run 1 hour. Whe we started 2 (for testing only, as they produced the same result, using the same input data filters), both finished after 1 hour. This was also so if we started 16: total duration was 1 hour 01 minutes. BUT: when we started 17, then 16 were finished after 1 hour, and one after 01:40. When starting 16 + 8 -> 16 finished in 01:00, 8 finished in 01:40 plus a bit. ... and so on.... When starting more than 64 things got worse: 16 finished in 01:05, 48 fished in 01:55, and the rest in 02:10 or slower ^^All this basically w/o any I/O, so CPU bound. One problem we faced was nomadic CPU usage, so binding CPUvps, but also the NETvps, to 'their' CPU made a difference of more than 10% or 6 minutes shorter duration for the batches per duration-hour. After this customer was lost and went to Oracle, the features of the M9000 remained the same, basically. Switching off multithreading made it better on ORA 10g, but not good - only better. HTH dic_k -- Richard Kofler SOLID STATE EDV Dienstleistungen GmbH Vienna/Austria/Europe
On 7 June, 14:02, "Neil Truby" <neil.tr...@ardenta.com> wrote:
> "Ian Michael Gumby" <im_gu...@hotmail.com> wrote in messagenews:mailman.6.1244378383.4791.informix-list@iiug.org...
>
> >> What's the rule of thumb? N-1 where N is the number of CPUs?
>
> Depends upon your defintion of "cpu", doesn't it?
>
> >>If that's true, then you'd want to do 63vcpus and tune it to that
> >>configuration.
> >> I'm not sure by what is meant that Sun sees each core as 2 virtual cpus.
>
> I didn't say "virtual cpus". Using psrinfo you see each of the 127 threads
> listed, that is what I meant.
>
> cheers
> N
Are you using KAIO or AIO vps?
What are the NETTYPE settings?
What does onstat -g glo give after say a week (or several days, i.e.
one cycle of processing activity including jobs that do not run daily)
of activity with 127 cpu vps?
Also what are the top 30 entries what the output of onstat -g spi is
sorted by num waits?
I would pin each CPU vp in turn to the first cpu from each on the 16
sockets, then pin the next 16 in core2 on sockets 16-1 (assume cpus
are busiest in vp number order), the next the 16 cpu vps on cpu 3 in
sockets 8-16,1-7 etc. Try to balance cpu activity across sockets. Once
the first 64 cpu vps are bound then bind
the remaining 64 cpus to the other cores each time binding the next
cpu vp to the lighest loaded socket.
Also bind AIO vps (in socket order 16-1) on core3. Finally netvps on
sockets 8-16,1-7 on core 4.
Basically you want to spread the cpu <-> external bus (memory,
network, disk controlller) activity first across sockets, then across
cores and finally threads
within cores.
You probably need to know how (in order of importance)
1. How the cpu to memory interconnect works. I assume above that this
that each socket has equal access to memory.
2. How network card interrupts map to cpus. If network cards
interrupts are bound to partiucular cores/sockets then it may be best
to not bind NET vps
to those sockets/cores. If network interrupts bind to a given core
rather than socket it may be best to bind the corresponding NET vp to
another core
on the same socket.
On thing you may find is that certain latches/mutexs with in the
instance run hot as they are not tuned for such a large number of CPU
VPs, making
the latches more fine grained make removed bottlenecks with the number
of CPU VPs.
I would get someone from the IBM Developement Labs involved as this is
probably the instance with the largest number of CPU VPS in existance.
The IBM engineers should be keen to make Informix scale to 100+ cpus
and would enjoy the task of tuning this instance.
> From: david@smooth1.co.uk > > On thing you may find is that certain latches/mutexs with in the > instance run hot as they are not tuned for such a large number of CPU > VPs, making > the latches more fine grained make removed bottlenecks with the number > of CPU VPs. > > > I would get someone from the IBM Developement Labs involved as this is > probably the instance with the largest number of CPU VPS in existance. > > The IBM engineers should be keen to make Informix scale to 100+ cpus > and would enjoy the task of tuning this instance. > David, I'd take networking out of the picture and focus on the engine running tests via a shared memory connection. Not that I don't think that there *will* be issues with networking, but that I think we need some form of isolation to force focus on fewer moving parts. I do agree that there are some issues with context switching you're going to have issues, but context switching for 128 vcpus should not be that difficult. Of course having said that... I imagine that when they did certain design decisions, they didn't consider context switching for anything remotely close to 128 vcpus. So is it safe to say that for massively large databases that going the shared nothing distributed route like... dare I say it? ... XPS? 8^0 :-P ... is the way to go? (Each node can be SMPed). The other reason I would considered a clustered shared nothing is that using OTC components, you can build it cheaper than massive SMP boxes. (But I could be wrong... Which is more, a 16 socket Sun box, or a blade center with 8 dual socket blades? (Assuming Sun's sparc VII+ chips are available for blades... I'd wager that IBM is aware of this and working out a solution for it. Of course I don't know if they'll do a DB2 version first, or an IDS version. My gut tells me that in the world of IBM, DB2 is still queen and IDS is that evil stepchild who needs to sweep the floors and do the dishes... ;-) (OK, so I fractured two fairy tales... but you get the idea!) -G _________________________________________________________________ Windows Live™ SkyDrive™: Get 25 GB of free online storage. http://windowslive.com/online/skydrive?ocid=TXT_TAGLM_WL_SD_25GB_062009
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g