TCP Connections Hang Intermittently while SHM are fine?
Posted in 2004
Topics: Networking & sqlhosts Configuration, Platform-Specific Issues, Versions, Editions & End-of-Life
Hello all,
IDS 7.31.FD3 Solaris 9.
SCO 5.07 I-SQL Version 7.20.UD2
Clients connect from SCO to the Solaris DB server. I've noticed that
often clients will hang trying to establish a new connection. This
seems to happen as many as 10 times in an hour, or as few as none.
The 'hang' never lasts longer than 30 seconds. On average, the hanging
seems to last about 5-15 secs. I've tried modifying my NETTYPE
settings increasing the number of polling threads, but this seems to
not have helped. My NETTYPE currently is:
NETTYPE tlitcp,5,500,NET
I've setup test scripts and have found that clients with already
established connections, or clients opening new 'ipcshm' connections
are fine. Just clients connecting using 'tlitcp' are affected. Also,
the hanging happens whether you are local or remote to the DB server.
Now I know DNS is a possibility, but the timing of the hanging is so
inconsistent, and I have reviewed my '/etc/hosts' and made sure to use
only IP addresses in 'sqlhosts'. How could I go about ruling DNS out?
When and how would IDS make DNS lookups?
Thanks so much for your time. I am quite desperate and any advice
would be appreciated immensely.
Douglas Howser
Douglas Howser wrote:
> Hello all,
>
> IDS 7.31.FD3 Solaris 9.
> SCO 5.07 I-SQL Version 7.20.UD2
>
> Clients connect from SCO to the Solaris DB server. I've noticed that
> often clients will hang trying to establish a new connection. This
> seems to happen as many as 10 times in an hour, or as few as none.
>
> The 'hang' never lasts longer than 30 seconds. On average, the hanging
> seems to last about 5-15 secs. I've tried modifying my NETTYPE
> settings increasing the number of polling threads, but this seems to
> not have helped. My NETTYPE currently is:
>
> NETTYPE tlitcp,5,500,NET>
> I've setup test scripts and have found that clients with already
> established connections, or clients opening new 'ipcshm' connections
> are fine. Just clients connecting using 'tlitcp' are affected. Also,
> the hanging happens whether you are local or remote to the DB server.
>
> Now I know DNS is a possibility, but the timing of the hanging is so
> inconsistent, and I have reviewed my '/etc/hosts' and made sure to use
> only IP addresses in 'sqlhosts'. How could I go about ruling DNS out?
> When and how would IDS make DNS lookups?
>
> Thanks so much for your time. I am quite desperate and any advice
> would be appreciated immensely.
>
> Douglas Howser
Well reverse DNS can cause this:
When you connect to a service, the server tries to reverse solve your
clinet IP to a name. Depending on server configuration it can go through
a DNS server or through /etc/hosts file. When going through DNS you can
have a momentary hanging. But I don't think this would happen to local
clients... This would also be noticeable in telnet/ssh connections or
other services... It's a guess, but maybe it's worth a bit of investigation.
Regards.
Fernando,
Thanks for the reply. I've checked and double-checked DNS and all
seems well.
And yes, even locally the problem exists.
If I am local to the DB instance, and try executing query with
dbaccess through tlitcp I get an intermittent hang, maybe a handful of
times in an hour, never lasting longer than 10-20 secs. But if I try
to run dbaccess query through ipcshm, all is well, never hangs.
Please help!!!
What could make tlitcp hang besides DNS? Perhaps some parameter of DB
is misconfigured and is causing problems for the poll threads of type
NET? Is this possible? Normally don't exceed 100 concurrent users, so
am pretty confident it's not too many users.
Is it possible that a particular user connection might be sdoing
something bad, spinning in a loop or something, and causing problems
for other sessions trying to connect?
Douglas Howser
Fernando Nunes <spam@domus.online.pt> wrote in message news:<2ts465F22f2c8U1@uni-berlin.de>...
> Douglas Howser wrote:
>
> > Hello all,
> >
> > IDS 7.31.FD3 Solaris 9.
> > SCO 5.07 I-SQL Version 7.20.UD2
> >
> > Clients connect from SCO to the Solaris DB server. I've noticed that
> > often clients will hang trying to establish a new connection. This
> > seems to happen as many as 10 times in an hour, or as few as none.
> >
> > The 'hang' never lasts longer than 30 seconds. On average, the hanging
> > seems to last about 5-15 secs. I've tried modifying my NETTYPE
> > settings increasing the number of polling threads, but this seems to
> > not have helped. My NETTYPE currently is:
> >
> > NETTYPE tlitcp,5,500,NET> >
> > I've setup test scripts and have found that clients with already
> > established connections, or clients opening new 'ipcshm' connections
> > are fine. Just clients connecting using 'tlitcp' are affected. Also,
> > the hanging happens whether you are local or remote to the DB server.
> >
> > Now I know DNS is a possibility, but the timing of the hanging is so
> > inconsistent, and I have reviewed my '/etc/hosts' and made sure to use
> > only IP addresses in 'sqlhosts'. How could I go about ruling DNS out?
> > When and how would IDS make DNS lookups?
> >
> > Thanks so much for your time. I am quite desperate and any advice
> > would be appreciated immensely.
> >
> > Douglas Howser
>
>
> Well reverse DNS can cause this:
>
> When you connect to a service, the server tries to reverse solve your
> clinet IP to a name. Depending on server configuration it can go through
> a DNS server or through /etc/hosts file. When going through DNS you can
> have a momentary hanging. But I don't think this would happen to local
> clients... This would also be noticeable in telnet/ssh connections or
> other services... It's a guess, but maybe it's worth a bit of investigation.
>
> Regards.
Hello all, Just wanted to post a follow up to this. We made two changes last night and the problem now appears to be remedied. The two changes we made were to clean up our '/etc/hosts' file and also 'sqlhosts'. In our '/etc/hosts' file, there were multiple entries for the DB server. For instance: 1.1.1.1 db1.test.com 1.1.1.1 anothernamefordb1.test.com ... So we changed this to be all one line, like: 1.1.1.1 db1.test.com anothernamefordb1.test.com Second change we made which I believe is most likely the cause is that in our 'sqlhosts' file, we hade an entry for 'tliimc'. Now this apparently is for a product called MaxConnect which AFAIK we never used. When I did a search on CDI for 'tliimc' I came across one single post with someone complaining about hanging 'tlitcp' connections. So this could have been the cause. We removed the line in 'sqlhosts' referring to the MaxConnect protocol. So one of these changes was responsible. I'd be suprised if DNS was the culprit because things seemed to be resovling fine on the machine and I've never known '/etc/hosts' to be real picky. Thanks for your help Fernando. Regards, Douglas Howser