Slow communication performance of Informix 11.50
Posted in 2011
Scott saw distributed queries between Informix 11.50 servers on AIX 6.1 crawl at ~4 rows/sec versus ~6000 rows/sec on the old 9.40 setup, even between instances on the same host; truss showed the delay sitting between network _esend/_erecv calls. Suggestions included avoiding per-row reconnects, checking hosts-before-DNS resolution order, comparing query plans, and testing a shared-memory connection. It turned out to be DNS: AIX was still querying Active Directory DNS despite the changed order. Using iptrace (readable via Wireshark) exposed heavy DNS chatter; forcing local /etc/hosts resolution and recycling the node restored fast performance.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Server Administration, Platform-Specific Issues
OS: AIX 6.1 Informix 11.50 FC8 When issuing queries from one Informix 11.50 database to another (9.40 or 11.50), we are experiencing very slow performance. It does not matter if both Informix servers are on the same AIX host or if they are communicating over the network on separate AIX hosts . We issue a simple job on the Informix 11.50 server to reach out to another Informix server (9.40 or 11.50), verify a row exists in a database, fetch that row and bring it back to the initial Informix server 11.50 server and insert it into a database there. In our current Informix 9.40 environment, we are seeing acceptable performance times of about 6000 rows/second. With our new Informix 11.50 environment, we are seeing about 4 rows/second. Yes, it is taking 0.25 second for each row. We have done extensive, detailed investigation (i.e. timestamps before/after each insert) and it appears to be during the phase in which the two nodes communicate. Running the jobs separately (select & insert) on each node, prove to run very fast. The network and DNS appear fine. We have pushed 100 mb files to/from various AIX nodes in our Informix environment and they are all ranging from 8 12 mb/ seconds in both directions. Not too bad. I have a call open with IBM support at the current time, but wondered if anything jumps out to folks who have built 11.50 environments. IBM has verified our onconfig looks good. Any suggestions would be appreciated. Thanks, Scott
SCOTT YOHONN WROTE: ================================================================================ OS: AIX 6.1 Informix 11.50 FC8 When issuing queries from one Informix 11.50 database to another (9.40 or 11.50), we are experiencing very slow performance. It does not matter if both Informix servers are on the same AIX host or if they are communicating over the network on separate AIX hosts . We issue a simple job on the Informix 11.50 server to reach out to another Informix server (9.40 or 11.50), verify a row exists in a database, fetch that row and bring it back to the initial Informix server 11.50 server and insert it into a database there. In our current Informix 9.40 environment, we are seeing acceptable performance times of about 6000 rows/second. With our new Informix 11.50 environment, we are seeing about 4 rows/second. Yes, it is taking 0.25 second for each row. We have done extensive, detailed investigation (i.e. timestamps before/after each insert) and it appears to be during the phase in which the two nodes communicate. Running the jobs separately (select & insert) on each node, prove to run very fast. The network and DNS appear fine. We have pushed 100 mb files to/from various AIX nodes in our Informix environment and they are all ranging from 8 12 mb/ seconds in both directions. Not too bad. I have a call open with IBM support at the current time, but wondered if anything jumps out to folks who have built 11.50 environments. IBM has verified our onconfig looks good. Any suggestions would be appreciated. Thanks, Scott ================================================================================ RESPONSE: Does your simple job initiate a new connection to the remote database for every row it processes? If your re-establishing connection again and again, it might be the difference between a local host reference and a call to a DNS server. Double check that your hosts files have your other servers listed and that the OS is configured to use hosts before DNS.
Can you explain the situation in terms of SQL? When you write "server to reach out to another Informix server [...] verify if a row exists, fetch that row and bring it back to the initial Informix ...." what exaclty are you doing? A distributed query? More details please.... If it's a distributed query did you check the query plan? Did you compare it between the "good" and the "bad" environment? Regards. On Mon, Feb 7, 2011 at 7:42 PM, SCOTT YOHONN <syohonn@gmail.com> wrote: > OS: AIX 6.1 > Informix 11.50 FC8 > > When issuing queries from one Informix 11.50 database to another (9.40 or > 11.50), we are experiencing very slow performance. It does not matter if > both > Informix servers are on the same AIX host or if they are communicating over > the network on separate AIX hosts . We issue a simple job on the Informix > 11.50 server to reach out to another Informix server (9.40 or 11.50), > verify a > row exists in a database, fetch that row and bring it back to the initial > Informix server 11.50 server and insert it into a database there. In our > current Informix 9.40 environment, we are seeing acceptable performance > times > of about 6000 rows/second. With our new Informix 11.50 environment, we are > seeing about 4 rows/second. Yes, it is taking 0.25 second for each row. > > We have done extensive, detailed investigation (i.e. timestamps > before/after > each insert) and it appears to be during the phase in which the two nodes > communicate. Running the jobs separately (select & insert) on each node, > prove > to run very fast. The network and DNS appear fine. We have pushed 100 mb > files > to/from various AIX nodes in our Informix environment and they are all > ranging > from 8 12 mb/ seconds in both directions. Not too bad. I have a call open > with IBM support at the current time, but wondered if anything jumps out to > folks who have built 11.50 environments. IBM has verified our onconfig > looks > good. Any suggestions would be appreciated. > > Thanks, > Scott > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently... --0015174c46265d29c8049bbac8b9
This simple SQL job has been running in our current 9.40 production environment for years. It moves data at about 6000/rows a second so I assumed it was opening one connection and moving the data then closing the connection. The open/close connection theory is something we have investigated. By default, our DNS resolution is first done with Active Directory, then local to /etc/hosts. Yesterday, we switched AIX to resolve DNS first to /etc/hosts, then to AD - no change in performance. Also yesterday, we started running the UNIX truss utility during a test job to try to capture exactly where the delay is occurring. We do see the delay occur between network call of _esend and _erecv - whatever that is. It is definitely pointing to network communications, but it's tough to nail down because simple tests done done to transfer files to/from the nodes involved yield decent network speeds. I have not written off a DNS issue though. We have seem other issues in our environment with some devices (i.e. Windows 2003 Servers) that communicate DNS wit IP version 4 & 6 and network devices that are listening for IP version 4 only and it caused a major delay. Stay tuned...
The SQL I am talking about has been running in our 9.40 production environment for years moving about 6000 rows/second. We run the identical SQL code in the 11.50 environment. It is looking more like a network slow down somewhere.
Do you have any evidences that the new version is following the correct (or usual) query plans? On Tue, Feb 8, 2011 at 2:00 PM, SCOTT YOHONN <syohonn@gmail.com> wrote: > The SQL I am talking about has been running in our 9.40 production > environment > for years moving about 6000 rows/second. We run the identical SQL code in > the > 11.50 environment. It is looking more like a network slow down somewhere. > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently... --0015174c12d4a0ac02049bc723cb
Just to make sure, IF you connect via shared memory (in the database machine), the performance issue is the same? Celso Cabral Coimbra Administrador de Banco de Dados ClearTech Ltda "Trust at the heart of Communications" Tel. (11) 3576-4509 -----Mensagem original----- De: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] Em nome de Fernando Nunes Enviada em: terça-feira, 8 de fevereiro de 2011 13:38 Para: ids@iiug.org Assunto: Re: Slow communication performance of Informix 11. [22720] Do you have any evidences that the new version is following the correct (or usual) query plans? On Tue, Feb 8, 2011 at 2:00 PM, SCOTT YOHONN <syohonn@gmail.com> wrote: > The SQL I am talking about has been running in our 9.40 production > environment > for years moving about 6000 rows/second. We run the identical SQL code in > the > 11.50 environment. It is looking more like a network slow down somewhere. > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently... --0015174c12d4a0ac02049bc723cb ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
It did turn out to be a DNS issue. Although we initially thought that we forced the AIX node to use local /etc/hosts, we ran a handy AIX utility called iptrace (i.e. iptrace -a) and found that the node was still using Active Directory DNS. Once we finally were able to make the node resolve DNS locally, the response time per row was drastically faster, by a large magnitude! The iptrace output yielded great information. We could see large amounts of communication among the AIX node to various DNS servers. We recycled the AIX node and it must of cleaned up various network processes because it resolve to one or two primary DNS servers and the speed was fast going through AD DNS servers again. We are going to pick up tomorrow on another AIX/Informix node exhibiting the same issues to pinpoint exactly where it is. It is somewhere in the AIX DNS configuration. The AIX iptrace utility proved invaluable. It really helped show all of the thrashing going on with the AIX node and the DNS servers. A good tool to have in the toolbox. You can also download a freeware item called wireshark to read the iptrace output more clearly. Thanks for everyone's input! Scott