Slow connect to database?
Posted in 2007
On HP-UX 11i with IDS 9.30, a user reported intermittent multi-minute hangs when connecting via dbaccess/isql/4gl: the Unix process appeared but nothing showed in onstat -u/-g sql, and it affected both soctcp and shm connections, with no OS or NETTYPE issues. Suggestions were to trace system calls on the dbaccess and MSC VP processes with tusc/truss and compare good vs bad periods, and check DNS. DNS response times were tested and looked fine; Fernando Nunes explained that connection threads call gethostbyaddr() and can stall on reverse DNS lookups (he had hit a Tru64 bug), recommending truss/lsof on the MSC VP, adding client IPs to /etc/hosts, and (on v10) LISTEN_TIMEOUT/MAX_INCOMPLETE_CONNECTIONS. The poster bounced the instance, which seemed to help temporarily; no confirmed root cause or resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Connectivity: ESQL/C, 4GL & Embedded SQL, Server Administration, Platform-Specific Issues, Versions, Editions & End-of-Life
HP-UX 11i, IDS 9.30HC5
Can anyone shed some light on an intermittent problem we're getting ?
Every now and then we get a connection issue (using dbaccess, isql or
4gl) where we attempt a connection (e.g. 'dbaccess <database>') and
the shell prompt hangs for ages before the main screen shows; in that
time ps shows the process but onstats don't. Nothing shows up in
onstat -u, onstat -g sql etc so to all intents and purposes theprogram hasn't connected to the instance although the unix process
shows up. The delay can be up to several minutes so the users will
start screaming before long; but it's so intermittent that we can't
predict it happening.
glance, top etc show no UNIX issues and there's nothing in the
sysadmin events log; it happens using both soctcp and shm connection,
and we're well inside our NETTYPE limits.
Help!
Confused of West Kent
you could grab a utility called tusc and use that to see if you can
identify a system call
which may cause this..
try it on a dbaccess which 'hangs' and oninit which is the first one
eq
onstat -g glo...Individual virtual processors:
vp pid class usercpu syscpu total
1 3877 cpu 1.40 0.37 1.77
----^^^^^^^^
maybe you get lucky. What does HP have to say about this????
Superboer.
On 11 jun, 13:48, mal...@btinternet.com wrote:
> HP-UX 11i, IDS 9.30HC5
> Can anyone shed some light on an intermittent problem we're getting ?
> Every now and then we get a connection issue (using dbaccess, isql or
> 4gl) where we attempt a connection (e.g. 'dbaccess <database>') and
> the shell prompt hangs for ages before the main screen shows; in that
> time ps shows the process but onstats don't. Nothing shows up in
> onstat -u, onstat -g sql etc so to all intents and purposes the> program hasn't connected to the instance although the unix process
> shows up. The delay can be up to several minutes so the users will
> start screaming before long; but it's so intermittent that we can't
> predict it happening.
> glance, top etc show no UNIX issues and there's nothing in the
> sysadmin events log; it happens using both soctcp and shm connection,
> and we're well inside our NETTYPE limits.
>
> Help!
>
> Confused of West Kent
Superboer wrote:
> you could grab a utility called tusc and use that to see if you can
> identify a system call
> which may cause this..
> try it on a dbaccess which 'hangs' and oninit which is the first one
> eq
>
> onstat -g glo...> Individual virtual processors:
> vp pid class usercpu syscpu total
> 1 3877 cpu 1.40 0.37 1.77
> ----^^^^^^^^
>
> maybe you get lucky. What does HP have to say about this????
>
>
> Superboer.
We don't have truss/tusc and no-one here has experience of using it;
we'll get it and load it up but would be flying a little blind.
We're going to try an instance bounce first as there are 250 irate
users and this is always a good appeaser!
If it doesn't help we'll call HP up.
See my previous posts for why we haven't got an Informix support
contract (sorry Neil, do they listen to me? Do they F***).
More anon.....
malc_p@btinternet.com wrote:
> HP-UX 11i, IDS 9.30HC5
> Can anyone shed some light on an intermittent problem we're getting ?
> Every now and then we get a connection issue (using dbaccess, isql or
> 4gl) where we attempt a connection (e.g. 'dbaccess <database>') and
> the shell prompt hangs for ages before the main screen shows; in that
> time ps shows the process but onstats don't. Nothing shows up in
> onstat -u, onstat -g sql etc so to all intents and purposes the> program hasn't connected to the instance although the unix process
> shows up. The delay can be up to several minutes so the users will
> start screaming before long; but it's so intermittent that we can't
> predict it happening.
Have you tried testing DNS response times during the 'hang' period.
I would be surprised if it affected shm connections but......
--
Clive
On 11 Jun, 13:54, Clive Eisen <c...@serendipita.com> wrote: > Have you tried testing DNS response times during the 'hang' period. > > I would be surprised if it affected shm connections but...... > Yes thanks Clive our Unix team did this first up - no issues :-( NB the instance restart seems to have worked but we're waiting for the 250 usual connections to run up to see if it's clear now. > -- > Clive
malc_p@btinternet.com wrote:
> On 11 Jun, 13:54, Clive Eisen <c...@serendipita.com> wrote:
>
>> Have you tried testing DNS response times during the 'hang' period.
>>
>>
>>
>>
>>
>>
>>
>>
>>
>>
>> I would be surprised if it affected shm connections but......
>>
> Yes thanks Clive our Unix team did this first up - no issues :-(
>
> NB the instance restart seems to have worked but we're waiting for
> the 250 usual connections to run up to see if it's clear now.
>
>> -- Clive
>
>
Well... long post... This has been one of my favorite "hobbies" :)
I've had a long history of problems like this, and if you really want to
understand what's going on you really must get truss or a similar tool and
understand how it works.
Let me describe what I had face in the past:
Sometimes I got the same symptoms you mention. Everybody already connected was
working well, but no new connections could be established. As others mentioned
this can be DNS related, but it's only one of the causes. In my case, every
time the customer had some DNS problems they had problems in Informix.
You must understand what happens in order to study it. Everytime an attempt to
connect is made, Informix has a thread that runs on the MSC VP (you can get
it's PID from onstat -g glo). This thread calls gethostbyaddr(), an OS
function. This function usually checks the /etc/hosts file for the IP. If it
finds it, the connection proceeds and that's it. If not, it tries to talk to
the DNS in order to make a reverse DNS lookup.
By the way, there is a thread (listener thread) that "talks" to the client...
If the client does not answers in a specified time the connection is dropped.
But be aware of two issues:
1- In versions before V10, there was only one thread. In V10 you can define more.
2- The time it waits was pre-defined and could not be configured in V7 (and v9
IIRC). In v10 you can specify it
So, v10 gives you tools to avoid DoS attacks (either voluntary or caused by
some application misconfiguration).
Back to DNS... Every DNS request has a relatively short amount of time before
it gets dropped. Informix does not implement any kind of asynchronous
processing when calling gethostbyaddr(). This should be fine... BUT... my
luck... I've hit a bug in Tru64 gethostbyaddr()... A very weird one. When you
make a call to gethostbyaddr() the process opens a UDP socket to receive the
answer from the DNS server. In my situation, due to bad network configuration,
the range of available ports was overlapping some ports used by some of the
customer applications. Specifically there were some ports that were being
constantly hit by traffic generated by a middleware application. It used some
kind of broadcast to find services in a range of servers. When the UDP socket
opened by gethostbyaddr() had the address of one of this ports the function
never returned. It kept receiving the application data, and while it ignored it
(as it didn't contain valid DNS answers), it never returned.
Since this was version 7 the thread never stopped and it never stepped to
another connection attempt. Every connection attempt was being queued. If we
stoped the network traffic, then all the pending connections would be
established in a few moments.
So, my advice to you:
1- grab the MSC VP PID from onstat -g glo
2- Monitor it with trace or similar when everything is working. If your
connection rate is low make some yourself. You'll be able to identify a pattern
of system calls.
3- When you begin to have problems do this again and compare with what you got
in step 2
Be aware that truss/trace has several options. You can even get the IP address
of the DNS when it tries to contact it... this can be important (maybe your
primary is donwn?...)
You could/should also use lsof on the same PID to see what are the open file
descriptors (this includes socktets)
You have to be a firm... this can take a few minutes during which your users
will suffer...
Some things you can try to ease your troubles:
1- If using v10, check ONCONFIG parameters LISTEN_TIMEOUT and
MAX_INCOMPLETE_CONNECTIONS (can be changes dynamically with onmode -wm)
2- Put your clients IP address in /etc/hosts (SysAdmins don't usually like
this. I agree with them, but for test/investigation purposes it can be useful)
Apparently you don't have support... If it's a DNS related problem, I dare to
say it would be of much assistance... When I hit this I search internally for
similar situations and many of them were useless because there were no
conclusions. I opened a case on IBM and HP. My colleagues were helpful, but
there was not much they could do. HP found it hard to believe that it was a
problem on their side, but once they did, they provided a fix in a reasonable time.
By the way, in cases where the engines appear to hang there is a useful script
in IBM's support site. It's called tbsupport.ksh and it takes a few snapshots
of the engine state and nicely packs it. You should use it in this kind of
scenarios and send the result to the support team. Nevertheless, if this is a
DNS situation or similar, the data collected by this script won't be much of
much use. Anybody wanting it can google for it. It's easy to find. I've done
some personal changes to it, but I haven't made it available yet.
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...