What is DNS Linkage with Informix TCP service/port
Posted in 2012
Topics: Platform-Specific Issues
Dear All, Yesterday, we have faced an issue where our DB instance was unable to connect our application using tlitcp port. We found out that it has happened after we change the DNS at our system level. During this problem, I tried connecting database instance from the secondary server using the tcp service, but it was not functioning as well. The same database instance started working fine after restart. So it seems that there exists some linkage between database tcp port and DNS, can someone help me in making the logical connection between them and how can i avoid this issue in future? In the online.log, everything was normal except this line: -------------------------------------------------------------------------- listener-thread: err = -25571: oserr = 0: errstr = : Cannot create a user thread. The error -25571 has been encountered 62 times since it was last printed to the log. -------------------------------------------------------------------------- My Environment: IDS: IDS 11.50 FC8W3 System: Sunfire T2000 Sparc OS: Solaris 10 Thanks in advance!
On Wed, Jun 20, 2012 at 2:03 PM, OMER KHAN <oskhan@i2cinc.com> wrote: > Dear All, > > Yesterday, we have faced an issue where our DB instance was unable to > connect > our application using tlitcp port. > We found out that it has happened after we change the DNS at our system > level. > You'd have to restart Informix. Sorry for that. Please open a PMR and request a feature request (actually it exists already). > > During this problem, I tried connecting database instance from the > secondary > server using the tcp service, but it was not functioning as well. > > The same database instance started working fine after restart. > Yes, as expected. > > So it seems that there exists some linkage between database tcp port and > DNS, > can someone help me in making the logical connection between them and how > can > i avoid this issue in future? > You can't avoid this if you change the DNS. Unless you configure your system not to use DNS. What you can do is to plan for it with the database restart. > In the online.log, everything was normal except this line: > -------------------------------------------------------------------------- > listener-thread: err = -25571: oserr = 0: errstr = : Cannot create a > user thread. > The error -25571 has been encountered 62 times since it was last printed to > the log. > -------------------------------------------------------------------------- > My Environment: > IDS: IDS 11.50 FC8W3 > System: Sunfire T2000 Sparc > OS: Solaris 10 > > Thanks in advance! > I covered this topic in: http://informix-technology.blogspot.com/2012/01/dns-impact-on-informix-impacto-d o-dns.html A bit extensive (that's a problem with my writing), but it covers this issue in detail. Regards. -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently... --485b393aab73b7814b04c2e73012
Thanks fernando for the answers. Actually, I have further discussed the case with my system admin and here are some of the additional information which you might need: Actually, we uses the external DNS servers (/etc/resolve) and there was an internet outage due to which DB servers were unable to connect to the DNS servers. And that internet outage recovered after 45 minutes and DB servers were able to connect to DNS. But after above two incidents, our DB instance was not able to make new tcp connections? why? We also checked that we were able to ssh our DB servers successfully. Although, you made it clear in your earlier answer that if there is a change in the DNS then we have to restart Informix instance, but in our case DNS servers were temporary unavailability (we did not make any changes in /etc/resolv.conf & /etc/nsswitch.conf) and after recovery of DNS, do we still require restart? Additionally, as we are running multiple instance but we faced this issue on only one instance? And we only restarted one of the instances to fix the issue? Why our other instances were not impacted? Any idea? Look forward for your response. Thanks.
On Wed, Jun 20, 2012 at 5:13 PM, OMER KHAN <oskhan@i2cinc.com> wrote:
> Thanks fernando for the answers.
>
> Actually, I have further discussed the case with my system admin and here
> are
> some of the additional information which you might need:
>
> Actually, we uses the external DNS servers (/etc/resolve) and there was an
> internet outage due to which DB servers were unable to connect to the DNS
> servers.
> And that internet outage recovered after 45 minutes and DB servers were
> able
> to connect to DNS.
>
> But after above two incidents, our DB instance was not able to make new tcp
> connections? why?
>
Impossible to tell... Some debugging would need to take place while the
problem was happening.
More below...
> We also checked that we were able to ssh our DB servers successfully.
>
ssh may work differently....
> Although, you made it clear in your earlier answer that if there is a
> change
> in the DNS then we have to restart Informix instance, but in our case DNS
> servers were temporary unavailability (we did not make any changes in
> /etc/resolv.conf & /etc/nsswitch.conf) and after recovery of DNS, do we
> still
> require restart?
>
No. Which is contrary to your experience...
Additionally, as we are running multiple instance but we faced this issue on
> only one instance? And we only restarted one of the instances to fix the
> issue?
>
> Why our other instances were not impacted? Any idea?
>
Without very specific debugging I can only speculate. A few years ago I had
a situation like you described and I was fortunate (and stuburn enough) to
find out what happened. But this was a very strange and difficult to happen
elsewhere situation:
1- the MSC VP makes the DNS requests by calling gethostbyaddr(). This works
with UDP protocol and it opens a socket with a port number belonging to a
range that is configurable
2- The customer was using a third party tool (tibco) that uses UDP
broadcasts to find out services. These broadcasts are sent to a specific
port(s)
3- Due to misconfiguration, the port used by tibco was inside the interval
of ports configured to be used by client processes when they request a UDB
socket
4- the OS function gethostbyaddr() sent a request to a DNS server that was
down. But it opened the port ro receive the answers. When this port matched
the tibco port, the function would just sit there reading the broadcast
data (not DNS traffic). Since the traffic never stop arriving the function
would never return (the function does not have a timeout, so it kept read()
garbage)
5- Since they used only 1 MSCVP they could not do anymore connections
6- I was able to confirm all this, by stopping the network traffic on that
port. After a couple of seconds, the "hanged" instance returned to live!
So, I'm not saying that this happened to you. I'd say it's not likely. What
I'm saying is that this is a complex matter, that most customers (and even
IBM) don't usually pay enough attention to. To give you an answer I'd need
a lot of data collecting while the issue was happening.
To be a bit more positive, what can I suggest if this happens again:
1- Try to launch another MSCVP (onmode -p +1 msc) when you're sure the DNS
are ok
2- Make sure that in your /etc/nsswitch.conf you configure the system to
access the /etc/host file first. This could allow you to populate the file
with client IP addresses if you sense problems
3- You have strange errors in the online.log. I suspect these has to do
with the listener requests queuing up. This can cause some issues
(specially on Solaris if using TLI). You could try to restart the
listeners...
> Look forward for your response. Thanks.
>
Not very helpful, but I think it's impossible to answer without data
collected during the issue. This should include strace/truss outputs from
the MSCVP process.
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--20cf300faa77fc8ea504c2ea1827
We have seen this behavior a few times and the network endpoint the
listener thread is monitoring can have issues in the operating system. =
The
informix listener thread is not getting errors, but no new connections =
are
happening. To overcomes this problem without shutting down the server=
utilizes informix ability to dynamically restart the tcp listener threa=
ds.
This will not effect currently connected threads, but allows a new netw=
ork
endpoint to be created and in general solves the problem without having=
to
bounce the Informix database server.
onmode -P [start|stop|restart] <servername> dynamic listen thread contr=ol
Hope this helps,
John F. Miller III
STSM, Embedability Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
ids-bounces@iiug.org wrote on 06/20/2012 09:13:17 AM:
> From: "OMER KHAN" <oskhan@i2cinc.com>
> To: ids@iiug.org,
> Date: 06/20/2012 09:17 AM
> Subject: Re: What is DNS Linkage with Informix TCP service/ [27414]
> Sent by: ids-bounces@iiug.org
>
> Thanks fernando for the answers.
>
> Actually, I have further discussed the case with my system admin andh=
ere
are
> some of the additional information which you might need:
>
> Actually, we uses the external DNS servers (/etc/resolve) and there w=
as
an
> internet outage due to which DB servers were unable to connect to the=
DNS
> servers.
> And that internet outage recovered after 45 minutes and DB servers we=
re
able
> to connect to DNS.
>
> But after above two incidents, our DB instance was not able to make n=
ew
tcp
> connections? why?
>
> We also checked that we were able to ssh our DB servers successfully.=
>
> Although, you made it clear in your earlier answer that if there is a=
change
> in the DNS then we have to restart Informix instance, but in our case=
DNS
> servers were temporary unavailability (we did not make any changes in=
> /etc/resolv.conf & /etc/nsswitch.conf) and after recovery of DNS, dow=
e
still
> require restart?
>
> Additionally, as we are running multiple instance but we faced this i=
ssue
on
> only one instance? And we only restarted one of the instances to fix =
the
> issue?
>
> Why our other instances were not impacted? Any idea?
>
> Look forward for your response. Thanks.
>
>
>
***********************************************************************=
********
> Forum Note: Use "Reply" to post a response in the discussion forum.=
>=
Related threads
- Transaction not available
- Error -25571 errstr = informix : Cannot creta a us
- Bug? Role Separation ... informix can't access...
- NETTYPE maximum support to user connections ?