Network tuning advice needed
Answered: green (solid confidence) — Fernando Nunes' suggestion to set the OPTOFC environment variable (optimize open/fetch/close network round-trips) is tried by the OP and confirmed to cut runtime roughly in half (65min to 35min), directly resolving the reported new-server-vs-old-server slowdown.
Advisory only.
Posted in 2009
Tom moved his app off the database server (AIX 5.3 to a dedicated AIX 6.1 box, both IDS 11.5) and one program ran over 3x slower, with sessions stuck in 'cond wait netnorm', suggesting a network bottleneck that IBM's PMR review found nothing wrong with. Suggestions included checking SQLIDEBUG traces, bonded NICs and query plans, but running the program locally (onipcshm) restored the old timing, confirming network round-trip latency from many tiny queries. Setting OPTOFC=1 cut the run from 65 to 35 minutes; further gains were expected from rewriting the app to use joins or cached lookups instead of per-row queries.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Server Administration, Platform-Specific Issues
Greetings All,
I am looking for some advice on a new (to me) configuration. Up until
now, we have always had our engine running on the same server as our
application. That has now changed. We now have a setup with the
application on one server and the engine on a dedicated database
server. The old server is running AIX 5.3 and the new server is
running AIX 6.1. Both servers are running IDS 11.5.
Here is the problem: A particular program we run takes slightly more
than 3 times longer to run on the new server than it does on the old
server. Everything else runs faster on the new server. Using onstat -g
ses ses_id and onstat -g act we can see that the program is constantly
shifting into cond wait netnorm on the new server and we can catch a
cond wait sm_read on the old server. This leads us to believe that our
bottleneck is with the network. The two servers are connected to a
1Gig switch using two NIC's on each server configured as
EtherChannels. I opened a PMR with IBM, sent them all of the requested
network data, and they see no issues.
Here is my question: Are there any tunables in IDS 11.5 that I can
monitor and tune? The only significant change in my onconfig files
between the old and new servers is here:
Old server:
VPCLASS cpu,num=8,noage
VPCLASS net,num=1,noage
VP_MEMORY_CACHE_KB 1000
New server:
VPCLASS cpu,num=5,noage
VPCLASS net,num=4,noage
VP_MEMORY_CACHE_KB 10240
Thanks,
Tom
On Oct 9, 10:59 am, Tom <tlow...@munis.com> wrote:
> Greetings All,
>
> I am looking for some advice on a new (to me) configuration. Up until
> now, we have always had our engine running on the same server as our
> application. That has now changed. We now have a setup with the
> application on one server and the engine on a dedicated database
> server. The old server is running AIX 5.3 and the new server is
> running AIX 6.1. Both servers are running IDS 11.5.
>
> Here is the problem: A particular program we run takes slightly more
> than 3 times longer to run on the new server than it does on the old
> server. Everything else runs faster on the new server. Using onstat -g
> ses ses_id and onstat -g act we can see that the program is constantly
> shifting into cond wait netnorm on the new server and we can catch a
> cond wait sm_read on the old server. This leads us to believe that our
> bottleneck is with the network. The two servers are connected to a
> 1Gig switch using two NIC's on each server configured as
> EtherChannels. I opened a PMR with IBM, sent them all of the requested
> network data, and they see no issues.
>
> Here is my question: Are there any tunables in IDS 11.5 that I can
> monitor and tune? The only significant change in my onconfig files
> between the old and new servers is here:
>
> Old server:
> VPCLASS cpu,num=8,noage
> VPCLASS net,num=1,noage
> VP_MEMORY_CACHE_KB 1000>
> New server:
> VPCLASS cpu,num=5,noage
> VPCLASS net,num=4,noage
> VP_MEMORY_CACHE_KB 10240>
> Thanks,
>
> Tom
I know in your post you said both the old server (aix 5.3) and the new
(aix 6.1) are running 11.5, but are you sure this is a network issue?
Could something have happened to the database when it was migrated
from the old to new machine that is causing some other performance
problem? Like a query plan issue or something that is what's actually
causing the slow down? One possibility would be to turn on SQLIDEBUG
for the client app that you are having problems with and ideally if
you could get the output for both new and old for comparison
purposes. That might be helpful in trying to determine if there is 1,
or a set of SQL that seems to be getting significant performance
decreases for deeper investigation that could be non-network related.
Jacques Renaut
IBM Informix Advanced Support
APD team
> From: jprenaut@yahoo.com > > I know in your post you said both the old server (aix 5.3) and the new > (aix 6.1) are running 11.5, but are you sure this is a network issue? > Could something have happened to the database when it was migrated > from the old to new machine that is causing some other performance > problem? Like a query plan issue or something that is what's actually > causing the slow down? One possibility would be to turn on SQLIDEBUG > for the client app that you are having problems with and ideally if > you could get the output for both new and old for comparison > purposes. That might be helpful in trying to determine if there is 1, > or a set of SQL that seems to be getting significant performance > decreases for deeper investigation that could be non-network related. > > Jacques Renaut > IBM Informix Advanced Support > APD team While this is true, why not eliminate the hardware issue before looking at IDS? Tom's running two nics bonded so that they appear as one. What happens if he shuts down one nic and uses one for his network connection? In a separate e-mail his disks are on fiber channel and are not FoE so they aren't an issue. The other issue is that the OS level changed. No offense to IBM, but IDS and DB2 are not the only products that are prone to bugs. It could be AIX and hardware issues. Of course baring that, I'd then say that if you're right, what about an "UPDATE STATISTICS" command? ;-) I'd also recommend seeing what happens if the app runs local to the machine. If the issue is in the database, then he'd see some negative performance that would remove the network connection. If he doesn't see anything, then one has to ask what else is going on with the network. If there's no real traffic, then I'd go back to the machine and its hardware. The point is, and I know Clive doesn't grok this... before jumping to conclusions, see if you can remove potential problem areas one at a time. ;-) -G _________________________________________________________________ Hotmail: Trusted email with Microsoft’s powerful SPAM protection. http://clk.atdmt.com/GBL/go/177141664/direct/01/
On Oct 9, 1:46 pm, Ian Michael Gumby <im_gu...@hotmail.com> wrote: > > From: jpren...@yahoo.com > > > I know in your post you said both the old server (aix 5.3) and the new > > (aix 6.1) are running 11.5, but are you sure this is a network issue? > > Could something have happened to the database when it was migrated > > from the old to new machine that is causing some other performance > > problem? Like a query plan issue or something that is what's actually > > causing the slow down? One possibility would be to turn on SQLIDEBUG > > for the client app that you are having problems with and ideally if > > you could get the output for both new and old for comparison > > purposes. That might be helpful in trying to determine if there is 1, > > or a set of SQL that seems to be getting significant performance > > decreases for deeper investigation that could be non-network related. > > > Jacques Renaut > > IBM Informix Advanced Support > > APD team > > While this is true, why not eliminate the hardware issue before looking at IDS? > > Tom's running two nics bonded so that they appear as one. > > What happens if he shuts down one nic and uses one for his network connection? > In a separate e-mail his disks are on fiber channel and are not FoE so they aren't an issue. > > The other issue is that the OS level changed. No offense to IBM, but IDS and DB2 are not the only products that are prone to bugs. It could be AIX and hardware issues. > > Of course baring that, I'd then say that if you're right, what about an "UPDATE STATISTICS" command? ;-) > > I'd also recommend seeing what happens if the app runs local to the machine. If the issue is in the database, then he'd see some negative performance that would remove the network connection. > > If he doesn't see anything, then one has to ask what else is going on with the network. If there's no real traffic, then I'd go back to the machine and its hardware. > > The point is, and I know Clive doesn't grok this... before jumping to conclusions, see if you can remove potential problem areas one at a time. ;-) > > -G > > _________________________________________________________________ > Hotmail: Trusted email with Microsoft’s powerful SPAM protection.http://clk.atdmt.com/GBL/go/177141664/direct/01/- Hide quoted text - > > - Show quoted text - I am configuring the new server to run the software locally, as it is on the old server. It will be ready for a Monday run. I am expecting that to cure the problem. But, I have been wrong before. We are following Jacques advice and creating SQLIDEBUG files on the old server and the new server now. So, we'll have that info ready for Monday as well. Even if running the program locally fixes the issue, I'd like to see this problem fixed. And, with the SQLIDEBUG files we should have a shot at doing that. Jacques - will I need to open a trouble call with IBM Informix to get the files analyzed? Thanks, Tom
> I am configuring the new server to run the software locally, as it is > on the old server. It will be ready for a Monday run. I am expecting > that to cure the problem. But, I have been wrong before. > > We are following Jacques advice and creating SQLIDEBUG files on the > old server and the new server now. So, we'll have that info ready for > Monday as well. Even if running the program locally fixes the issue, > I'd like to see this problem fixed. And, with the SQLIDEBUG files we > should have a shot at doing that. > > Jacques - will I need to open a trouble call with IBM Informix to get > the files analyzed? > > Thanks, > > Tom I think you will. The SQLIDEBUG output needs to get parsed by a program, sqliprint, and in I know in some versions of the product we included it, but in others it was yanked out, so I have a hard time keeping that straight. Even if you did have sqliprint, if this is a problem, eventually you'd have to open a call, so you might as well get that out of the way. Jacques Renaut IBM Informix Advanced Support ADP Team
Please check the other answer, but looking at your description it reminds me
of a typical situation I see once in a while...
I would bet your program makes a very large amount of very quick queries.
I've seen that scenario in typical ETL processes with lookup queries.
For each value it reads from somewhere the process queries a table in the
remote hosts database. This query is virtually instantaneous, but the
process suffers from latency. Most of the times this is impossible to solve
unless you take another approach to the problem.
One thing that may help in the informix side is setting up the OPTOFC
variable. It means "optimize open/fetch/close"... basically it tries to
reduce the number of messages exchanged between the client and server for
situations where the programs reuse cursors. Check the manual...
Another option sometimes is to rethink the process... let me give you an
example that will also uncover another option...
Sometime ago a developer asked me to "optimize" a lookup query he was doing
in an ETL job. After looking at it was a simple index (primay key) access.
He's process was doing million of times the same query. It would take
several hours to complete (don't recall the exact time but I think it was
around 10H).
I told him that it was not possible to optimize that query. But gave him two
options:
1- do a prior sort to his input file
2- Give me the file so that I could do the lookup inside the database with a
join
Option 1 cut his time in half... basically, by accessing the table in an
ordered way (sorted by the primary key), consecutive accesses took advantage
of the most recently read index pages.
Option 2 made the process run in about 20m... Another option would be to
pull the complete lookup table into the process engine...
So I would say that besides OPTOFC you'll probably have to rethink the
process, assmuming it's making a lot of very simple and efficient queries...
Regards.
On Fri, Oct 9, 2009 at 4:59 PM, Tom <tlowrie@munis.com> wrote:
> Greetings All,
>
> I am looking for some advice on a new (to me) configuration. Up until
> now, we have always had our engine running on the same server as our
> application. That has now changed. We now have a setup with the
> application on one server and the engine on a dedicated database
> server. The old server is running AIX 5.3 and the new server is
> running AIX 6.1. Both servers are running IDS 11.5.
>
> Here is the problem: A particular program we run takes slightly more
> than 3 times longer to run on the new server than it does on the old
> server. Everything else runs faster on the new server. Using onstat -g
> ses ses_id and onstat -g act we can see that the program is constantly
> shifting into cond wait netnorm on the new server and we can catch a
> cond wait sm_read on the old server. This leads us to believe that our
> bottleneck is with the network. The two servers are connected to a
> 1Gig switch using two NIC's on each server configured as
> EtherChannels. I opened a PMR with IBM, sent them all of the requested
> network data, and they see no issues.
>
> Here is my question: Are there any tunables in IDS 11.5 that I can
> monitor and tune? The only significant change in my onconfig files
> between the old and new servers is here:
>
> Old server:
> VPCLASS cpu,num=8,noage
> VPCLASS net,num=1,noage
> VP_MEMORY_CACHE_KB 1000>
> New server:
> VPCLASS cpu,num=5,noage
> VPCLASS net,num=4,noage
> VP_MEMORY_CACHE_KB 10240>
> Thanks,
>
> Tom
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
On Oct 9, 5:31 pm, Fernando Nunes <domusonl...@gmail.com> wrote:
> Please check the other answer, but looking at your description it reminds me
> of a typical situation I see once in a while...
> I would bet your program makes a very large amount of very quick queries.
> I've seen that scenario in typical ETL processes with lookup queries.
> For each value it reads from somewhere the process queries a table in the
> remote hosts database. This query is virtually instantaneous, but the
> process suffers from latency. Most of the times this is impossible to solve
> unless you take another approach to the problem.
>
> One thing that may help in the informix side is setting up the OPTOFC
> variable. It means "optimize open/fetch/close"... basically it tries to
> reduce the number of messages exchanged between the client and server for
> situations where the programs reuse cursors. Check the manual...
>
> Another option sometimes is to rethink the process... let me give you an
> example that will also uncover another option...
>
> Sometime ago a developer asked me to "optimize" a lookup query he was doing
> in an ETL job. After looking at it was a simple index (primay key) access.
> He's process was doing million of times the same query. It would take
> several hours to complete (don't recall the exact time but I think it was
> around 10H).
> I told him that it was not possible to optimize that query. But gave him two
> options:
> 1- do a prior sort to his input file
> 2- Give me the file so that I could do the lookup inside the database with a
> join
>
> Option 1 cut his time in half... basically, by accessing the table in an
> ordered way (sorted by the primary key), consecutive accesses took advantage
> of the most recently read index pages.
> Option 2 made the process run in about 20m... Another option would be to
> pull the complete lookup table into the process engine...
>
> So I would say that besides OPTOFC you'll probably have to rethink the
> process, assmuming it's making a lot of very simple and efficient queries...
>
> Regards.
>
>
>
>
>
> On Fri, Oct 9, 2009 at 4:59 PM, Tom <tlow...@munis.com> wrote:
> > Greetings All,
>
> > I am looking for some advice on a new (to me) configuration. Up until
> > now, we have always had our engine running on the same server as our
> > application. That has now changed. We now have a setup with the
> > application on one server and the engine on a dedicated database
> > server. The old server is running AIX 5.3 and the new server is
> > running AIX 6.1. Both servers are running IDS 11.5.
>
> > Here is the problem: A particular program we run takes slightly more
> > than 3 times longer to run on the new server than it does on the old
> > server. Everything else runs faster on the new server. Using onstat -g
> > ses ses_id and onstat -g act we can see that the program is constantly
> > shifting into cond wait netnorm on the new server and we can catch a
> > cond wait sm_read on the old server. This leads us to believe that our
> > bottleneck is with the network. The two servers are connected to a
> > 1Gig switch using two NIC's on each server configured as
> > EtherChannels. I opened a PMR with IBM, sent them all of the requested
> > network data, and they see no issues.
>
> > Here is my question: Are there any tunables in IDS 11.5 that I can
> > monitor and tune? The only significant change in my onconfig files
> > between the old and new servers is here:
>
> > Old server:
> > VPCLASS cpu,num=8,noage
> > VPCLASS net,num=1,noage
> > VP_MEMORY_CACHE_KB 1000>
> > New server:
> > VPCLASS cpu,num=5,noage
> > VPCLASS net,num=4,noage
> > VP_MEMORY_CACHE_KB 10240>
> > Thanks,
>
> > Tom
> > _______________________________________________
> > Informix-list mailing list
> > Informix-l...@iiug.org
> >http://www.iiug.org/mailman/listinfo/informix-list
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...- Hide quoted text -
>
> - Show quoted text -
We ran the program locally and it completed in about 20 minutes,
approximately the same time it takes on the old system.
So, by eliminating the network we now have fairly equal execution
times. Which tells me that by going from a locally run program using
onipcshm to a remotely run program using onsoctcp we suffer one heck
of a performance hit. Does this sound normal to anyone else out there?
Again, this is our first usage of a dedicated "back-end" database
server. If this is normal, so be it. If not, I'd like to know.
I will next set the OPTOFC environment variable as suggested by
Fernando and see what that does for the original run.
I'll postthe results.
Tom
As was pointed out, this is common if the application is not written with
network latency costs in mind. If, for example, as Fernando noted, your
application does thousands of small lookups on the same table, it would be
FAR cheaper to make the lookup as a join in the original query if possible
or, even better, to suck the entire lookup table (or at least relevant parts
of it) into a local data structure that can be queried quickly. For
example, SELECT the codes and resulting data ORDER BY the code into an array
of structures that you can then quickly bsearch() or into a hash table (or
the C++/Java) equivalent list type) and perform the lookups within the
application's code space. Vastly more efficient.
If you are planning to attend this year's IIUG Conference in Overland Park,
I have proposed a presentation "Best Practices for IDS Developers". This is
one topic I will spend time on if the presentation is selected by the
Conference Committee. You and at least some of your developers should plan
to attend.
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
On Mon, Oct 12, 2009 at 12:18 PM, Tom <tlowrie@munis.com> wrote:
> On Oct 9, 5:31 pm, Fernando Nunes <domusonl...@gmail.com> wrote:
> > Please check the other answer, but looking at your description it reminds
> me
> > of a typical situation I see once in a while...
> > I would bet your program makes a very large amount of very quick queries.
> > I've seen that scenario in typical ETL processes with lookup queries.
> > For each value it reads from somewhere the process queries a table in the
> > remote hosts database. This query is virtually instantaneous, but the
> > process suffers from latency. Most of the times this is impossible to
> solve
> > unless you take another approach to the problem.
> >
> > One thing that may help in the informix side is setting up the OPTOFC
> > variable. It means "optimize open/fetch/close"... basically it tries to
> > reduce the number of messages exchanged between the client and server for
> > situations where the programs reuse cursors. Check the manual...
> >
> > Another option sometimes is to rethink the process... let me give you an
> > example that will also uncover another option...
> >
> > Sometime ago a developer asked me to "optimize" a lookup query he was
> doing
> > in an ETL job. After looking at it was a simple index (primay key)
> access.
> > He's process was doing million of times the same query. It would take
> > several hours to complete (don't recall the exact time but I think it was
> > around 10H).
> > I told him that it was not possible to optimize that query. But gave him
> two
> > options:
> > 1- do a prior sort to his input file
> > 2- Give me the file so that I could do the lookup inside the database
> with a
> > join
> >
> > Option 1 cut his time in half... basically, by accessing the table in an
> > ordered way (sorted by the primary key), consecutive accesses took
> advantage
> > of the most recently read index pages.
> > Option 2 made the process run in about 20m... Another option would be to
> > pull the complete lookup table into the process engine...
> >
> > So I would say that besides OPTOFC you'll probably have to rethink the
> > process, assmuming it's making a lot of very simple and efficient
> queries...
> >
> > Regards.
> >
> >
> >
> >
> >
> > On Fri, Oct 9, 2009 at 4:59 PM, Tom <tlow...@munis.com> wrote:
> > > Greetings All,
> >
> > > I am looking for some advice on a new (to me) configuration. Up until
> > > now, we have always had our engine running on the same server as our
> > > application. That has now changed. We now have a setup with the
> > > application on one server and the engine on a dedicated database
> > > server. The old server is running AIX 5.3 and the new server is
> > > running AIX 6.1. Both servers are running IDS 11.5.
> >
> > > Here is the problem: A particular program we run takes slightly more
> > > than 3 times longer to run on the new server than it does on the old
> > > server. Everything else runs faster on the new server. Using onstat -g
> > > ses ses_id and onstat -g act we can see that the program is constantly
> > > shifting into cond wait netnorm on the new server and we can catch a
> > > cond wait sm_read on the old server. This leads us to believe that our
> > > bottleneck is with the network. The two servers are connected to a
> > > 1Gig switch using two NIC's on each server configured as
> > > EtherChannels. I opened a PMR with IBM, sent them all of the requested
> > > network data, and they see no issues.
> >
> > > Here is my question: Are there any tunables in IDS 11.5 that I can
> > > monitor and tune? The only significant change in my onconfig files
> > > between the old and new servers is here:
> >
> > > Old server:
> > > VPCLASS cpu,num=8,noage
> > > VPCLASS net,num=1,noage
> > > VP_MEMORY_CACHE_KB 1000> >
> > > New server:
> > > VPCLASS cpu,num=5,noage
> > > VPCLASS net,num=4,noage
> > > VP_MEMORY_CACHE_KB 10240> >
> > > Thanks,
> >
> > > Tom
> > > _______________________________________________
> > > Informix-list mailing list
> > > Informix-l...@iiug.org
> > >http://www.iiug.org/mailman/listinfo/informix-list
> >
> > --
> > Fernando Nunes
> > Portugal
> >
> > http://informix-technology.blogspot.com
> > My email works... but I don't check it frequently...- Hide quoted text -
> >
> > - Show quoted text -
>
> We ran the program locally and it completed in about 20 minutes,
> approximately the same time it takes on the old system.
>
> So, by eliminating the network we now have fairly equal execution
> times. Which tells me that by going from a locally run program using
> onipcshm to a remotely run program using onsoctcp we suffer one heck
> of a performance hit. Does this sound normal to anyone else out there?
> Again, this is our first usage of a dedicated "back-end" database
> server. If this is normal, so be it. If not, I'd like to know.
>
> I will next set the OPTOFC environment variable as suggested by
> Fernando and see what that does for the original run.
>
> I'll postthe results.
>
> Tom
>
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
Tom wrote:
> On Oct 9, 5:31 pm, Fernando Nunes <domusonl...@gmail.com> wrote:
>> Please check the other answer, but looking at your description it reminds me
>> of a typical situation I see once in a while...
>> I would bet your program makes a very large amount of very quick queries.
>> I've seen that scenario in typical ETL processes with lookup queries.
>> For each value it reads from somewhere the process queries a table in the
>> remote hosts database. This query is virtually instantaneous, but the
>> process suffers from latency. Most of the times this is impossible to solve
>> unless you take another approach to the problem.
>>
>> One thing that may help in the informix side is setting up the OPTOFC
>> variable. It means "optimize open/fetch/close"... basically it tries to
>> reduce the number of messages exchanged between the client and server for
>> situations where the programs reuse cursors. Check the manual...
>>
>> Another option sometimes is to rethink the process... let me give you an
>> example that will also uncover another option...
>>
>> Sometime ago a developer asked me to "optimize" a lookup query he was doing
>> in an ETL job. After looking at it was a simple index (primay key) access.
>> He's process was doing million of times the same query. It would take
>> several hours to complete (don't recall the exact time but I think it was
>> around 10H).
>> I told him that it was not possible to optimize that query. But gave him two
>> options:
>> 1- do a prior sort to his input file
>> 2- Give me the file so that I could do the lookup inside the database with a
>> join
>>
>> Option 1 cut his time in half... basically, by accessing the table in an
>> ordered way (sorted by the primary key), consecutive accesses took advantage
>> of the most recently read index pages.
>> Option 2 made the process run in about 20m... Another option would be to
>> pull the complete lookup table into the process engine...
>>
>> So I would say that besides OPTOFC you'll probably have to rethink the
>> process, assmuming it's making a lot of very simple and efficient queries...
>>
>> Regards.
>>
>>
>>
>>
>>
>> On Fri, Oct 9, 2009 at 4:59 PM, Tom <tlow...@munis.com> wrote:
>>> Greetings All,
>>> I am looking for some advice on a new (to me) configuration. Up until
>>> now, we have always had our engine running on the same server as our
>>> application. That has now changed. We now have a setup with the
>>> application on one server and the engine on a dedicated database
>>> server. The old server is running AIX 5.3 and the new server is
>>> running AIX 6.1. Both servers are running IDS 11.5.
>>> Here is the problem: A particular program we run takes slightly more
>>> than 3 times longer to run on the new server than it does on the old
>>> server. Everything else runs faster on the new server. Using onstat -g
>>> ses ses_id and onstat -g act we can see that the program is constantly
>>> shifting into cond wait netnorm on the new server and we can catch a
>>> cond wait sm_read on the old server. This leads us to believe that our
>>> bottleneck is with the network. The two servers are connected to a
>>> 1Gig switch using two NIC's on each server configured as
>>> EtherChannels. I opened a PMR with IBM, sent them all of the requested
>>> network data, and they see no issues.
>>> Here is my question: Are there any tunables in IDS 11.5 that I can
>>> monitor and tune? The only significant change in my onconfig files
>>> between the old and new servers is here:
>>> Old server:
>>> VPCLASS cpu,num=8,noage
>>> VPCLASS net,num=1,noage
>>> VP_MEMORY_CACHE_KB 1000>>> New server:
>>> VPCLASS cpu,num=5,noage
>>> VPCLASS net,num=4,noage
>>> VP_MEMORY_CACHE_KB 10240
>>> Thanks,>>> Tom
>>> _______________________________________________
>>> Informix-list mailing list
>>> Informix-l...@iiug.org
>>> http://www.iiug.org/mailman/listinfo/informix-list
>> --
>> Fernando Nunes
>> Portugal
>>
>> http://informix-technology.blogspot.com
>> My email works... but I don't check it frequently...- Hide quoted text -
>>
>> - Show quoted text -
>
> We ran the program locally and it completed in about 20 minutes,
> approximately the same time it takes on the old system.
>
> So, by eliminating the network we now have fairly equal execution
> times. Which tells me that by going from a locally run program using
> onipcshm to a remotely run program using onsoctcp we suffer one heck
> of a performance hit. Does this sound normal to anyone else out there?
> Again, this is our first usage of a dedicated "back-end" database
> server. If this is normal, so be it. If not, I'd like to know.
>
> I will next set the OPTOFC environment variable as suggested by
> Fernando and see what that does for the original run.
>
> I'll postthe results.
>
> Tom
>
You could (just for interest I suppose) try connecting locally over soctcp rather than ipcshm.
On Oct 12, 12:52 pm, theBP <th...@Usenet-News.Net> wrote:
> Tom wrote:
> > On Oct 9, 5:31 pm, Fernando Nunes <domusonl...@gmail.com> wrote:
> >> Please check the other answer, but looking at your description it reminds me
> >> of a typical situation I see once in a while...
> >> I would bet your program makes a very large amount of very quick queries.
> >> I've seen that scenario in typical ETL processes with lookup queries.
> >> For each value it reads from somewhere the process queries a table in the
> >> remote hosts database. This query is virtually instantaneous, but the
> >> process suffers from latency. Most of the times this is impossible to solve
> >> unless you take another approach to the problem.
>
> >> One thing that may help in the informix side is setting up the OPTOFC
> >> variable. It means "optimize open/fetch/close"... basically it tries to
> >> reduce the number of messages exchanged between the client and server for
> >> situations where the programs reuse cursors. Check the manual...
>
> >> Another option sometimes is to rethink the process... let me give you an
> >> example that will also uncover another option...
>
> >> Sometime ago a developer asked me to "optimize" a lookup query he was doing
> >> in an ETL job. After looking at it was a simple index (primay key) access.
> >> He's process was doing million of times the same query. It would take
> >> several hours to complete (don't recall the exact time but I think it was
> >> around 10H).
> >> I told him that it was not possible to optimize that query. But gave him two
> >> options:
> >> 1- do a prior sort to his input file
> >> 2- Give me the file so that I could do the lookup inside the database with a
> >> join
>
> >> Option 1 cut his time in half... basically, by accessing the table in an
> >> ordered way (sorted by the primary key), consecutive accesses took advantage
> >> of the most recently read index pages.
> >> Option 2 made the process run in about 20m... Another option would be to
> >> pull the complete lookup table into the process engine...
>
> >> So I would say that besides OPTOFC you'll probably have to rethink the
> >> process, assmuming it's making a lot of very simple and efficient queries...
>
> >> Regards.
>
> >> On Fri, Oct 9, 2009 at 4:59 PM, Tom <tlow...@munis.com> wrote:
> >>> Greetings All,
> >>> I am looking for some advice on a new (to me) configuration. Up until
> >>> now, we have always had our engine running on the same server as our
> >>> application. That has now changed. We now have a setup with the
> >>> application on one server and the engine on a dedicated database
> >>> server. The old server is running AIX 5.3 and the new server is
> >>> running AIX 6.1. Both servers are running IDS 11.5.
> >>> Here is the problem: A particular program we run takes slightly more
> >>> than 3 times longer to run on the new server than it does on the old
> >>> server. Everything else runs faster on the new server. Using onstat -g
> >>> ses ses_id and onstat -g act we can see that the program is constantly
> >>> shifting into cond wait netnorm on the new server and we can catch a
> >>> cond wait sm_read on the old server. This leads us to believe that our
> >>> bottleneck is with the network. The two servers are connected to a
> >>> 1Gig switch using two NIC's on each server configured as
> >>> EtherChannels. I opened a PMR with IBM, sent them all of the requested
> >>> network data, and they see no issues.
> >>> Here is my question: Are there any tunables in IDS 11.5 that I can
> >>> monitor and tune? The only significant change in my onconfig files
> >>> between the old and new servers is here:
> >>> Old server:
> >>> VPCLASS cpu,num=8,noage
> >>> VPCLASS net,num=1,noage
> >>> VP_MEMORY_CACHE_KB 1000> >>> New server:
> >>> VPCLASS cpu,num=5,noage
> >>> VPCLASS net,num=4,noage
> >>> VP_MEMORY_CACHE_KB 10240
> >>> Thanks,> >>> Tom
> >>> _______________________________________________
> >>> Informix-list mailing list
> >>> Informix-l...@iiug.org
> >>>http://www.iiug.org/mailman/listinfo/informix-list
> >> --
> >> Fernando Nunes
> >> Portugal
>
> >>http://informix-technology.blogspot.com
> >> My email works... but I don't check it frequently...- Hide quoted text -
>
> >> - Show quoted text -
>
> > We ran the program locally and it completed in about 20 minutes,
> > approximately the same time it takes on the old system.
>
> > So, by eliminating the network we now have fairly equal execution
> > times. Which tells me that by going from a locally run program using
> > onipcshm to a remotely run program using onsoctcp we suffer one heck
> > of a performance hit. Does this sound normal to anyone else out there?
> > Again, this is our first usage of a dedicated "back-end" database
> > server. If this is normal, so be it. If not, I'd like to know.
>
> > I will next set the OPTOFC environment variable as suggested by
> > Fernando and see what that does for the original run.
>
> > I'll postthe results.
>
> > Tom
>
> You could (just for interest I suppose) try connecting locally over soctcp rather than ipcshm.- Hide quoted text -
>
> - Show quoted text -
We did one last set of runs yesterday with OPTOFC=0 and OPTOFC=1 and
what we got was a significant improvement. With OPTOFC not set the
process took 65 minutes to run. With OPTOFC=1 the process took 35
minutes. Pretty close to a 50% performance gain by setting one
variable. Thanks for the tip Fernando.
If time permits, we will try using onsoctcp locally and see if that
makes a difference. If we do, I will post the results. I will be
taking the code-change advice to the folks who write the software
package.
Thanks again to all who took the time to help out. I appreciate it
greatly.
Tom
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g