Weird one
Posted in 2008
Art Kagel benchmarked parallel dbload runs on IDS 10.00.UC8 under Linux and saw periodic IO-wait spikes. Assuming extent additions were the cause, he pre-created the table with one large extent — but load times got worse (16–64 seconds slower) instead of better. Madison Pruet suggested possible causes: direct I/O/KAIO priority inversion with many CPU VPs but few AIO VPs, swap thrashing, or too few page cleaners/badly tuned LRU and checkpoint triggers. The thread drifts into OS thread/stack questions and records no confirmed cause or fix.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Platform-Specific Issues
<Cross posted to CDI & IDS Forum>
OK, an oddity for you. I'm doing some performance and tuning testing using
multiple copies of dbload.
I'll tell you what I'm dealing with. Anyone with any idea as to why the
engine is behaving as it is please pipe in.
Platform: Linux 2.6.9-67.ELsmp, x86-32 with 2 single core 3.4GHZ processors,
single disk spindle for IDS another for the OS
IDS Vers: 10.00.UC8
OK, so I'm getting reasonable timings running various numbers of load
clients against the engine with 1, 2, 4, 6, & 8 CPU VPs truncating the table
between runs.
I notice in TOP that at some point in the run IO Wait goes way up and CPU
consumption way down for 10-30 seconds then back to 'normal'. I figure,
"Oh! Extends are being added." So after completing the set of timing runs
I change the script to drop and recreate the table with an extent size
sufficient for the loaded data in the initial load instead of the truncate
and run the tests again expecting to see that the runtimes have decrease by
10-20 seconds per run. NO! the runtimes INCREASED by 16 to 64 seconds!
Yes, I checked that there was indeed still a single extent the size of the
initial extent size allocation after each load.
Any ideas as to why this should be?
--
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do those
opinions reflect those of other individuals affiliated with any entity with
which I am affiliated nor those of the entities themselves.
> Any ideas as to why this should be? Informix is crap ?
Art Kagel wrote:
> <Cross posted to CDI & IDS Forum>
>
> OK, an oddity for you. I'm doing some performance and tuning testing
> using multiple copies of dbload.
>
> I'll tell you what I'm dealing with. Anyone with any idea as to why the
> engine is behaving as it is please pipe in.
>
> Platform: Linux 2.6.9-67.ELsmp, x86-32 with 2 single core 3.4GHZ
> processors, single disk spindle for IDS another for the OS
> IDS Vers: 10.00.UC8
If you are running with directIO or KAIO, I'd be inclined to think that
you might have run into some form of inversion type of problem,
considering so many CPUVPS with so few read CPUVPS and considering that
you are seeing IOWait time. You might also want to check to see if
you're thrashing the swap file.
>
> OK, so I'm getting reasonable timings running various numbers of load
> clients against the engine with 1, 2, 4, 6, & 8 CPU VPs truncating the
> table between runs.
>
> I notice in TOP that at some point in the run IO Wait goes way up and
> CPU consumption way down for 10-30 seconds then back to 'normal'. I
> figure, "Oh! Extends are being added." So after completing the set of
> timing runs I change the script to drop and recreate the table with an
> extent size sufficient for the loaded data in the initial load instead
> of the truncate and run the tests again expecting to see that the
> runtimes have decrease by 10-20 seconds per run. NO! the runtimes
> INCREASED by 16 to 64 seconds!
>
> Yes, I checked that there was indeed still a single extent the size of
> the initial extent size allocation after each load.
>
> Any ideas as to why this should be?
>
> --
> Art S. Kagel
> Oninit (www.oninit.com <http://www.oninit.com>)
> IIUG Board of Directors (art@iiug.org <mailto:art@iiug.org>)
>
> Disclaimer: Please keep in mind that my own opinions are my own opinions
> and do not reflect on my employer, Oninit, the IIUG, nor any other
> organization with which I am associated either explicitly or implicitly.
> Neither do those opinions reflect those of other individuals affiliated
> with any entity with which I am affiliated nor those of the entities
> themselves.
>
Mark Townsend wrote: > >> Any ideas as to why this should be? > > Informix is crap ? Come on Mark. You can do better than that. You're surely aware of the meaning of seeing IOWait time. ;-)
Art Kagel wrote:
> <Cross posted to CDI & IDS Forum>
>
> OK, an oddity for you. I'm doing some performance and tuning testing
> using multiple copies of dbload.
>
> I'll tell you what I'm dealing with. Anyone with any idea as to why the
> engine is behaving as it is please pipe in.
>
> Platform: Linux 2.6.9-67.ELsmp, x86-32 with 2 single core 3.4GHZ
> processors, single disk spindle for IDS another for the OS
> IDS Vers: 10.00.UC8
>
> OK, so I'm getting reasonable timings running various numbers of load
> clients against the engine with 1, 2, 4, 6, & 8 CPU VPs truncating the
> table between runs.
>
> I notice in TOP that at some point in the run IO Wait goes way up and
> CPU consumption way down for 10-30 seconds then back to 'normal'. I
> figure, "Oh! Extends are being added." So after completing the set of
> timing runs I change the script to drop and recreate the table with an
> extent size sufficient for the loaded data in the initial load instead
> of the truncate and run the tests again expecting to see that the
> runtimes have decrease by 10-20 seconds per run. NO! the runtimes
> INCREASED by 16 to 64 seconds!
Also sounds to me like you might need additional page cleaners and maybe
adjust the triggers for firing them. IDS 11 might have non-blocking
checkpoint so LRU MIN/MAX can be pretty high. But that's not the case
with IDS10.
>
> Yes, I checked that there was indeed still a single extent the size of
> the initial extent size allocation after each load.
>
> Any ideas as to why this should be?
>
> --
> Art S. Kagel
> Oninit (www.oninit.com <http://www.oninit.com>)
> IIUG Board of Directors (art@iiug.org <mailto:art@iiug.org>)
>
> Disclaimer: Please keep in mind that my own opinions are my own opinions
> and do not reflect on my employer, Oninit, the IIUG, nor any other
> organization with which I am associated either explicitly or implicitly.
> Neither do those opinions reflect those of other individuals affiliated
> with any entity with which I am affiliated nor those of the entities
> themselves.
>
Just out of curiosity, and this may not be at issue...
On a multi-core /smp Linux box, does each CPU have its own stack or does it share the same stack?
If they share the same stack, how does one track to see if their are thread contention issues.
Meaning if Thread A from the same process is on core 1, and Thread B from the same process is on core 2, what sort of contention occurs?
If each CPU core has its own stack, is there load balancing among the stacks and would there be process migration between stacks?
Ok, this is an OS question and not an Informix question per se.
It would be interesting to see how solaris would handle this too.
> From: mpruet1@verizon.net
> Subject: Re: Weird one
> Date: Thu, 25 Sep 2008 05:07:28 +0000
> To: informix-list@iiug.org
>
> Art Kagel wrote:
> > <Cross posted to CDI & IDS Forum>
> >
> > OK, an oddity for you. I'm doing some performance and tuning testing
> > using multiple copies of dbload.
> >
> > I'll tell you what I'm dealing with. Anyone with any idea as to why the
> > engine is behaving as it is please pipe in.
> >
> > Platform: Linux 2.6.9-67.ELsmp, x86-32 with 2 single core 3.4GHZ
> > processors, single disk spindle for IDS another for the OS
> > IDS Vers: 10.00.UC8
> >
> > OK, so I'm getting reasonable timings running various numbers of load
> > clients against the engine with 1, 2, 4, 6, & 8 CPU VPs truncating the
> > table between runs.
> >
> > I notice in TOP that at some point in the run IO Wait goes way up and
> > CPU consumption way down for 10-30 seconds then back to 'normal'. I
> > figure, "Oh! Extends are being added." So after completing the set of
> > timing runs I change the script to drop and recreate the table with an
> > extent size sufficient for the loaded data in the initial load instead
> > of the truncate and run the tests again expecting to see that the
> > runtimes have decrease by 10-20 seconds per run. NO! the runtimes
> > INCREASED by 16 to 64 seconds!
>
> Also sounds to me like you might need additional page cleaners and maybe
> adjust the triggers for firing them. IDS 11 might have non-blocking
> checkpoint so LRU MIN/MAX can be pretty high. But that's not the case
> with IDS10.
>
> >
> > Yes, I checked that there was indeed still a single extent the size of
> > the initial extent size allocation after each load.
> >
> > Any ideas as to why this should be?
> >
> > --
> > Art S. Kagel
> > Oninit (www.oninit.com <http://www.oninit.com>)
> > IIUG Board of Directors (art@iiug.org <mailto:art@iiug.org>)
> >
> > Disclaimer: Please keep in mind that my own opinions are my own opinions
> > and do not reflect on my employer, Oninit, the IIUG, nor any other
> > organization with which I am associated either explicitly or implicitly.
> > Neither do those opinions reflect those of other individuals affiliated
> > with any entity with which I am affiliated nor those of the entities
> > themselves.
> >
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
_________________________________________________________________
Get more out of the Web. Learn 10 hidden secrets of Windows Live.
http://windowslive.com/connect/post/jamiethomson.spaces.live.com-Blog-cns!550F681DAD532637!5295.entry?ocid=TXT_TAGLM_WL_domore_092008
Read the Solaris Internals book by Jack Mauro ???
Ian Michael Gumby wrote:
> Just out of curiosity, and this may not be at issue...
>
> On a multi-core /smp Linux box, does each CPU have its own stack or does
> it share the same stack?
>
> If they share the same stack, how does one track to see if their are
> thread contention issues.
> Meaning if Thread A from the same process is on core 1, and Thread B
> from the same process is on core 2, what sort of contention occurs?
>
> If each CPU core has its own stack, is there load balancing among the
> stacks and would there be process migration between stacks?
>
> Ok, this is an OS question and not an Informix question per se.
>
> It would be interesting to see how solaris would handle this too.
>
>
> > From: mpruet1@verizon.net
> > Subject: Re: Weird one
> > Date: Thu, 25 Sep 2008 05:07:28 +0000
> > To: informix-list@iiug.org
> >
> > Art Kagel wrote:
> > > <Cross posted to CDI & IDS Forum>
> > >
> > > OK, an oddity for you. I'm doing some performance and tuning testing
> > > using multiple copies of dbload.
> > >
> > > I'll tell you what I'm dealing with. Anyone with any idea as to why
> the
> > > engine is behaving as it is please pipe in.
> > >
> > > Platform: Linux 2.6.9-67.ELsmp, x86-32 with 2 single core 3.4GHZ
> > > processors, single disk spindle for IDS another for the OS
> > > IDS Vers: 10.00.UC8
> > >
> > > OK, so I'm getting reasonable timings running various numbers of load
> > > clients against the engine with 1, 2, 4, 6, & 8 CPU VPs truncating the
> > > table between runs.
> > >
> > > I notice in TOP that at some point in the run IO Wait goes way up and
> > > CPU consumption way down for 10-30 seconds then back to 'normal'. I
> > > figure, "Oh! Extends are being added." So after completing the set of
> > > timing runs I change the script to drop and recreate the table with an
> > > extent size sufficient for the loaded data in the initial load instead
> > > of the truncate and run the tests again expecting to see that the
> > > runtimes have decrease by 10-20 seconds per run. NO! the runtimes
> > > INCREASED by 16 to 64 seconds!
> >
> > Also sounds to me like you might need additional page cleaners and maybe
> > adjust the triggers for firing them. IDS 11 might have non-blocking
> > checkpoint so LRU MIN/MAX can be pretty high. But that's not the case
> > with IDS10.
> >
> > >
> > > Yes, I checked that there was indeed still a single extent the size of
> > > the initial extent size allocation after each load.
> > >
> > > Any ideas as to why this should be?
> > >
> > > --
> > > Art S. Kagel
> > > Oninit (www.oninit.com <http://www.oninit.com>)
> > > IIUG Board of Directors (art@iiug.org <mailto:art@iiug.org>)
> > >
> > > Disclaimer: Please keep in mind that my own opinions are my own
> opinions
> > > and do not reflect on my employer, Oninit, the IIUG, nor any other
> > > organization with which I am associated either explicitly or
> implicitly.
> > > Neither do those opinions reflect those of other individuals
> affiliated
> > > with any entity with which I am affiliated nor those of the entities
> > > themselves.
> > >
> > _______________________________________________
> > Informix-list mailing list
> > Informix-list@iiug.org
> > http://www.iiug.org/mailman/listinfo/informix-list
>
> ------------------------------------------------------------------------
> Get more out of the Web. Learn 10 hidden secrets of Windows Live. Learn
> Now
> <http://windowslive.com/connect/post/jamiethomson.spaces.live.com-Blog-cns!550F681DAD532637!5295.entry?ocid=TXT_TAGLM_WL_getmore_092008>