Re: Configuring IDS for multi-threaded processors
Posted in 2009
Topics: Performance & Tuning, Clustering, Grid & MACH11
Hi Art,
Can you tell me if this behave occur with processor Niagra 2 (UltraSPARC T2) ?
Where I can get documentation about this kind of information? threads optimization on Sun Processors (T1 and T2)...
--- Em seg, 8/6/09, Art Kagel <art.kagel@gmail.com> escreveu:
De: Art Kagel <art.kagel@gmail.com>
Assunto: Re: Configuring IDS for multi-threaded processors
Para: "Ian Michael Gumby" <im_gumby@hotmail.com>
Cc: informix-list@iiug.org, david@smooth1.co.uk
Data: Segunda-feira, 8 de Junho de 2009, 8:02
David,
You cannot take advantage of the virtual threads in the Niagra processors. They were optimized for POSIX threads and perform poorly for single -OS LEVEL- threaded applications such as the oninit processes in IDS. To take best advantage of the Niagra processors with IDS you have to disable the hardware virtual threads and just count on the physical cores when calculating the number of CPU VPs to use. This advice from my own testing with a client, several posts on the Internet, and the experience of some IBM Informix people. Obviously YMMV applies and one should test things oneself, but that's my advice. Note that the individual physical cores of even the fastest Niagra processor are much slower than the cores in the latest Intel, AMD, and PowerPC processors with the obvious results in cross testing.
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves.
On Sun, Jun 7, 2009 at 8:50 PM, Ian Michael Gumby <im_gumby@hotmail.com> wrote:
> From: david@smooth1.co.uk
>
> On thing you may find is that certain latches/mutexs with in the
> instance run hot as they are not tuned for such a large number of CPU
> VPs, making
> the latches more fine grained make removed bottlenecks with the number
> of CPU VPs.
>
>
> I would get someone from the IBM Developement Labs involved as this is
> probably the instance with the largest number of CPU VPS in existance.
>
> The IBM engineers should be keen to make Informix scale to 100+ cpus
> and would enjoy the task of tuning this instance.
>
David,
I'd take networking out of the picture and focus on the engine running tests via a shared memory connection.
Not that I don't think that there *will* be issues with networking, but that I think we need some form of isolation to force focus on fewer moving parts.
I do agree that there are some issues with context switching you're going to have issues, but context switching for 128 vcpus should not be that difficult. Of course having said that... I imagine that when they did certain design decisions, they didn't consider context switching for anything remotely close to 128 vcpus.
So is it safe to say that for massively large databases that going the shared nothing distributed route like... dare I say it? ... XPS? 8^0 :-P ... is the way to go? (Each node can be SMPed). The other reason I would considered a clustered shared nothing is that using OTC components, you can build it cheaper than massive SMP boxes. (But I could be wrong... Which is more, a 16 socket Sun box, or a blade center with 8 dual socket blades? (Assuming Sun's sparc VII+ chips are available for blades...
I'd wager that IBM is aware of this and working out a solution for it. Of course I don't know if they'll do a DB2 version first, or an IDS version. My gut tells me that in the world of IBM, DB2 is still queen and IDS is that evil stepchild who needs to sweep the floors and do the dishes... ;-) (OK, so I fractured two fairy tales... but you get the idea!)
-G
Windows Live™ SkyDrive™: Get 25 GB of free online storage. Get it on your BlackBerry or iPhone.
_______________________________________________
Informix-list mailing list
Informix-list@iiug.org
http://www.iiug.org/mailman/listinfo/informix-list
-----Anexo incorporado-----
_______________________________________________
Informix-list mailing list
Informix-list@iiug.org
http://www.iiug.org/mailman/listinfo/informix-list
Veja quais são os assuntos do momento no Yahoo! +Buscados
http://br.maisbuscados.yahoo.com
Cesar Inacio Martins schrieb:
> Hi Art,
>
> Can you tell me if this behave occur with processor Niagra 2 (UltraSPARC T2) ?
> Where I can get documentation about this kind of information? threads optimization on Sun Processors (T1 and T2)...
>
>
> --- Em seg, 8/6/09, Art Kagel <art.kagel@gmail.com> escreveu:
>
> De: Art Kagel <art.kagel@gmail.com>
> Assunto: Re: Configuring IDS for multi-threaded processors
> Para: "Ian Michael Gumby" <im_gumby@hotmail.com>
> Cc: informix-list@iiug.org, david@smooth1.co.uk
> Data: Segunda-feira, 8 de Junho de 2009, 8:02
>
> David,
>
> You cannot take advantage of the virtual threads in the Niagra processors. They were optimized for POSIX threads and perform poorly for single -OS LEVEL- threaded applications such as the oninit processes in IDS. To take best advantage of the Niagra processors with IDS you have to disable the hardware virtual threads and just count on the physical cores when calculating the number of CPU VPs to use. This advice from my own testing with a client, several posts on the Internet, and the experience of some IBM Informix people. Obviously YMMV applies and one should test things oneself, but that's my advice. Note that the individual physical cores of even the fastest Niagra processor are much slower than the cores in the latest Intel, AMD, and PowerPC processors with the obvious results in cross testing.
>
>
>
> Art
>
> Art S. Kagel
> Oninit (www.oninit.com)
> IIUG Board of Directors (art@iiug.org)
>
> Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves.
>
>
>
>
>
>
> On Sun, Jun 7, 2009 at 8:50 PM, Ian Michael Gumby <im_gumby@hotmail.com> wrote:
>
>
>
>
>
>
>
>
>
>> From: david@smooth1.co.uk
>
>> On thing you may find is that certain latches/mutexs with in the
>> instance run hot as they are not tuned for such a large number of CPU
>
>
>> VPs, making
>> the latches more fine grained make removed bottlenecks with the number
>> of CPU VPs.
>>
>>
>> I would get someone from the IBM Developement Labs involved as this is
>> probably the instance with the largest number of CPU VPS in existance.
>
>
>> The IBM engineers should be keen to make Informix scale to 100+ cpus
>> and would enjoy the task of tuning this instance.
>>
>
> David,
>
> I'd take networking out of the picture and focus on the engine running tests via a shared memory connection.
>
>
> Not that I don't think that there *will* be issues with networking, but that I think we need some form of isolation to force focus on fewer moving parts.
>
> I do agree that there are some issues with context switching you're going to have issues, but context switching for 128 vcpus should not be that difficult. Of course having said that... I imagine that when they did certain design decisions, they didn't consider context switching for anything remotely close to 128 vcpus.
>
>
>
> So is it safe to say that for massively large databases that going the shared nothing distributed route like... dare I say it? ... XPS? 8^0 :-P ... is the way to go? (Each node can be SMPed). The other reason I would considered a clustered shared nothing is that using OTC components, you can build it cheaper than massive SMP boxes. (But I could be wrong... Which is more, a 16 socket Sun box, or a blade center with 8 dual socket blades? (Assuming Sun's sparc VII+ chips are available for blades...
>
>
>
> I'd wager that IBM is aware of this and working out a solution for it. Of course I don't know if they'll do a DB2 version first, or an IDS version. My gut tells me that in the world of IBM, DB2 is still queen and IDS is that evil stepchild who needs to sweep the floors and do the dishes... ;-) (OK, so I fractured two fairy tales... but you get the idea!)
>
>
>
> -G
>
>
> Windows Live™ SkyDrive™: Get 25 GB of free online storage. Get it on your BlackBerry or iPhone.
>
>
>
> _______________________________________________
>
> Informix-list mailing list
>
> Informix-list@iiug.org
>
> http://www.iiug.org/mailman/listinfo/informix-list
>
>
>
>
>
> -----Anexo incorporado-----
>
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
>
>
> Veja quais são os assuntos do momento no Yahoo! +Buscados
> http://br.maisbuscados.yahoo.com
Hi Cesar,
Look at page 54 in this one
http://www.inout.ch/files/pdf/download/inout_oracle_performance_on_solaris-server_oct-2008.pdf
Read about Strands in context of >>sun fujitsu solaris M9000<<
using the search engine of your choice.
Read all of, even it seems to be boring of
http://wikis.sun.com/download/attachments/57517586/820-6882.pdf?version=1
Then read page 24 a second time. Especially the second paragraph.
Look for more of this kind.
Then read http://www.iiug.org/iiug08/D05_Sun_Harnessing_The_Power_Of_Solaris_10_For_IDS_11.pdf
which is still at the IIUG.
IMHO this really should
- either stay there at IIUG, but must be commented
- or better simply delete it
as it is IMHO pure {vapour|paper}-ware
Fact is, THIS type of hyperthreading has nothing to do with
what Intel uses the term for since long time (2000?).
Strands is the contrary of what one does need when running
CPUvps, which
- are the contrary of light weight threads
- are CPU intensive
- run best when never have to yield and seldom lose
the CPU which they are scheduled on
- seem to be very CPU intensive batch jobs from the OS
scheduler perspective
- run best if there is no penalty for comsuming many
CPU seconds
Also fact is, that for apps running like 500+ leight weight threads
per process, strands ARE a very GOOD solution. But then, when I was
young, lwp were invented because exec (-vp) was so slow and had so
much overhead.....
So I was never convinced, that it is better to construct hardware
to compensate for ugly design patterns of the software.
But now, whom do I tell this? To the Java folks, who gladly
use the absolutely rotten 8000++ of 20K classes of J2EE?
AND call this best practise?
All they have for me is *roll eyes* and the occasional hint
where the next graveyard is, or a muttered 'shut up, old man'
So if you want to know, what best practise is:
Take exactly this machine, let it perform well under the
expectations evoked by marketing hype & lies, then upp your
consultance fee, bring in ORA (Sun) consultants as well and
leave 1 month before customer runs out of budget ;)
The braindead are there to make us rich, aren't they?
dic_k
P.S. consider it as a lucky incident, that I left out the rants about
I/O (i.e.PCI) performance or network performance of this type of
hardware.
But you have to take my lil sentence:
If you want performance , you must learn to spend less money.
--
Richard