IDS and keepalive
Posted in 2016
On AIX 7.1 with IDS 11.50.FC8, Mark saw growing client timeouts plus repeated "listener-thread: err = -27001 ... Read error occurred during connection attempt" messages in the online log, and suspected a network issue. Suggestions: test with a simple ESQL/C client locally vs. remotely to isolate network/app server, raise LISTEN_TIMEOUT from 10 to ~60, check NETTYPE/MAX_INCOMPLETE_CONNECTIONS, netstat and name/reverse-lookup and multi-homed routing issues (also blamed for a new -956 "not trusted" error after adding k=1 keepalive and restarting). One reply said -27001 typically just reflects an incomplete handshake (e.g. a port check/telnet). No confirmed resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Networking & sqlhosts Configuration, Platform-Specific Issues
AIX 7.1; IDS 11.50.FC8 Seeing ever-growing number of timeouts from a client (at their end) ... the message in engine msg log: 09/06/16 09:40:06 listener-thread: err = -27001: oserr = 0: errstr = : Read error occurred during connection attempt. I still suspect it's network-related. We've turned keep alive for the OS down to 10 minutes from the default 2 hours - still seeing issues. sqlhosts has always had a keepalive option (column 5; k=1) but haven't had to use it much over the years. I would consider adding an additional listener, but the test engine is "light" - has around 100 concurrent only (prod is 2000+). Any stories/experiences from anyone? Thanks - Mark Scranton The Mark Scranton Group mark@markscranton.com
Do you see the problem if the client is on the same machine as the server but still using a TCP connection? Madison Pruet Retired and Loving it On Tuesday, September 6, 2016 12:16 PM, MARK SCRANTON <mark@markscranton.com> wrote: AIX 7.1; IDS 11.50.FC8 Seeing ever-growing number of timeouts from a client (at their end) ... the message in engine msg log: 09/06/16 09:40:06 listener-thread: err = -27001: oserr = 0: errstr = : Read error occurred during connection attempt. I still suspect it's network-related. We've turned keep alive for the OS down to 10 minutes from the default 2 hours - still seeing issues. sqlhosts has always had a keepalive option (column 5; k=1) but haven't had to use it much over the years. I would consider adding an additional listener, but the test engine is "light" - has around 100 concurrent only (prod is 2000+). Any stories/experiences from anyone? Thanks - Mark Scranton The Mark Scranton Group mark@markscranton.com ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Hard to test since the client that is having issues comes from app server/odbc. Here's something interesting ... changed sqlhosts connection to: engine_soc onsoctcp box.here.com 1521 k=1 And now any remote connections to this entry is failing due to being untrusted? You've all seen this error: 09/06/16 14:12:23 listener-thread: err = -956: oserr = 0: errstr = user@remote_box: Client host or user user@remote_box is not trusted by the server. Why would the k=1 (options column of sqlhosts) cause this? Thanks in advance - Mark Scranton The Mark Scranton Group mark@markscranton.com
Marc,
Version and platform are important, specially for this cases.
If you have a connection time issue, forget about keepalive. It shouldn't
cause any harm, but it wont' solve anything. It's useful to keep already
established connections open.
If your version has these parameters, what are they set to?
LISTEN_TIMEOUT
MAX_INCOMPLETE_CONNECTIONS
What are your NETTYPE settings and how many connections per second do you
have (from what you say I assume not many...)
netstat -i shows any errors?
I was saying that Solaris may have some differences, but from your other
post it doesn't seem to be Solaris...
These errors are probably network related as you say.... no errors on the
application side?
If you reduce INFORMIXCONTIME, maybe they'll show up ;)
Regards.
On Tue, Sep 6, 2016 at 7:16 PM, MARK SCRANTON <mark@markscranton.com> wrote:
> AIX 7.1; IDS 11.50.FC8
>
> Seeing ever-growing number of timeouts from a client (at their end) ... the
> message in engine msg log:
>
> 09/06/16 09:40:06 listener-thread: err = -27001: oserr = 0: errstr = : Read
> error occurred during connection attempt.
>
> I still suspect it's network-related. We've turned keep alive for the OS
> down
> to 10 minutes from the default 2 hours - still seeing issues. sqlhosts has
> always had a keepalive option (column 5; k=1) but haven't had to use it
> much
> over the years. I would consider adding an additional listener, but the
> test
> engine is "light" - has around 100 concurrent only (prod is 2000+).
>
> Any stories/experiences from anyone?
>
> Thanks -
> Mark Scranton
> The Mark Scranton Group
> mark@markscranton.com
>
>
> ************************************************************
> *******************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--001a11c1492878c352053bdcd63e
Not sure if you were replying to something? If yes, I missed it. What did you change exactly? In the previous post you seem to imply that "k=1" was already there. Did you change the address/hostname? Is that error, exactly what is prnted in the online.log (I assumed you masked the names, but was that all you did?) The message doesn't include something between square brackets "[....]"? Is it AIX? Again, the version and platform are important. In any case, as I wrote previously, I'd be very surprised if "k=1" made any difference. Regards. On Tue, Sep 6, 2016 at 10:21 PM, MARK SCRANTON <mark@markscranton.com> wrote: > Hard to test since the client that is having issues comes from app > server/odbc. > > Here's something interesting ... changed sqlhosts connection to: > > engine_soc onsoctcp box.here.com 1521 k=1 > > And now any remote connections to this entry is failing due to being > untrusted? You've all seen this error: > > 09/06/16 14:12:23 listener-thread: err = -956: oserr = 0: errstr = > user@remote_box: Client host or user user@remote_box is not trusted by the > server. > > Why would the k=1 (options column of sqlhosts) cause this? > > Thanks in advance - > Mark Scranton > The Mark Scranton Group > mark@markscranton.com > > > ************************************************************ > ******************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently... --001a1143e3868bab7d053bdce3f2
Thanks for the responses Fernando ...
AIX 7.1, IDS 11.50.FC8
LISTEN_TIMEOUT 10
MAX_INCOMPLETE_CONNECTIONS 1024
The issue: a client connecting via TCP is having long sessions (up to 9
minutes) that eventually timeout. They see the msg at their end and I see it
in my message log:
listener-thread: err = -27001: oserr = 0: errstr = : Read error occurred
during connection attempt.
(That is the full message of the original issue - when the client times out
and the engine notices)
Without a lot of confidence (based on lack of depth of knowledge), I added to
sqlhosts col5 was simply k=1 for the engine entries. (2, both pointed to the
same port for DBSERVERNAME and DBSERVERALIASES - setup long before me).
What are your NETTYPE settings and how many connections per second do you have
(from what you say I assume not many...)
NETTYPE soctcp,1,200,CPU # Configure poll thread(s) for nettype
netstat -i shows any errors? No netstat errors on this box.
These errors are probably network related as you say.... no errors on the

application side? If you reduce INFORMIXCONTIME, maybe they'll show up
;)
The 2nd issue is the adding of k=1 to sqlhosts. It DID seem to cause issues
with trust. The message log on the target is below:
09/06/16 15:00:59 listener-thread: err = -956: oserr = 0: errstr =
user@remote_box: Client host or user user@remote_box is not trusted by the
server.
I still think there is a connectivity/trust issue in general. We use
hosts.equiv the normal way from box to box.
Thanks again -
Mark
Maybe, but you can determine if it is a problem with the network or with the server by writing a rather simple esql/c program that could be run from both the client machine and from the server machine... A rather simple esql/c program would be something like. -----------------#include <stdlib.h>#include <stdio.h> intmain(){ $int counter; $database sysmaster; if (sqlda.sqlcode < 0) { fprintf(stderr,"database connection failed with %d\\ ", sqlda.sqlcode); return 1; } $select count(*) into :counter from systables; if (sqlda.sqlcode < 0) { fprintf(stderr,"select failed with %d\\ ", sqlda.sqlcode); return 1; } fprintf("Connect successful\\ ");}---------------------------------------------------------------- -----------If you can successfully execute this program from the server machine and from the client machine, then you know that it is neither a problem with the network or with the database. So you would then need to focus on the configuration of the app server. Madison Pruet Retired and Loving it On Tuesday, September 6, 2016 3:21 PM, MARK SCRANTON <mark@markscranton.com> wrote: Hard to test since the client that is having issues comes from app server/odbc. Here's something interesting ... changed sqlhosts connection to: engine_soc onsoctcp box.here.com 1521 k=1 And now any remote connections to this entry is failing due to being untrusted? You've all seen this error: 09/06/16 14:12:23 listener-thread: err = -956: oserr = 0: errstr = user@remote_box: Client host or user user@remote_box is not trusted by the server. Why would the k=1 (options column of sqlhosts) cause this? Thanks in advance - Mark Scranton The Mark Scranton Group mark@markscranton.com ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Ok....
1- Increase the LISTEN_TIMEOUT value to around 60
2- Being AIX... and considering you also have the trust issue... (and I
assume you restarted the engine or at least the listeners, and that may
have triggered the second issue):
I suspect because I was told by a colleague of such a situation, that your
AIX is multi-homed and there are very nsty routing issues "inside" the box.
Not sure how to debug this, but try to run a traceroute from the client, or
from inside the AIX box with target as the hostname you have in SQLHOSTS.
Algo simulate a reverse lookup on that hostname and check if it returns the
expected IP address (maybe it's returning 127.0.0.1 or something like that)
Keep us posted... It's an interesting issue.
My bet is that the reverse lookup changed. And by restarting the engine you
made Informix "suffer" from that. Usually it's a big annoyance that
Informix has no supported way to refresh the resolver low level cache. But
that lack of functionality was probably what was letting it work.
Regards.
Regards.
On Tue, Sep 6, 2016 at 11:17 PM, MARK SCRANTON <mark@markscranton.com>
wrote:
> Thanks for the responses Fernando ...
>
> AIX 7.1, IDS 11.50.FC8
>
> LISTEN_TIMEOUT 10
> MAX_INCOMPLETE_CONNECTIONS 1024>
> The issue: a client connecting via TCP is having long sessions (up to 9
> minutes) that eventually timeout. They see the msg at their end and I see
> it
> in my message log:
>
> listener-thread: err = -27001: oserr = 0: errstr = : Read error occurred
> during connection attempt.
>
> (That is the full message of the original issue - when the client times out
> and the engine notices)
>
> Without a lot of confidence (based on lack of depth of knowledge), I added
> to
> sqlhosts col5 was simply k=1 for the engine entries. (2, both pointed to
> the
> same port for DBSERVERNAME and DBSERVERALIASES - setup long before me).
>
> What are your NETTYPE settings and how many connections per second do you
> have
> (from what you say I assume not many...)
> NETTYPE soctcp,1,200,CPU # Configure poll thread(s) for nettype> 
netstat -i shows any errors? No netstat errors on this box.
> These errors are probably network related as you say.... no errors on the
> 
application side? If you reduce INFORMIXCONTIME, maybe they'll show
> up
> ;)
>
> The 2nd issue is the adding of k=1 to sqlhosts. It DID seem to cause issues
> with trust. The message log on the target is below:
> 09/06/16 15:00:59 listener-thread: err = -956: oserr = 0: errstr =
> user@remote_box: Client host or user user@remote_box is not trusted by the
> server.
> I still think there is a connectivity/trust issue in general. We use
> hosts.equiv the normal way from box to box.
>
> Thanks again -
> Mark
>
>
> ************************************************************
> *******************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--001a1134f858447d4a053bdd85db
Some points in my reply are messy....
box.here.com must resolve to the proper IP address or your clients wouldn't
be able to reach it (but check)
reverse lookup of your clients IP address maybe messed up, or if due to
internal routing Informix is getting the wrong IP address from the client
(but the message doesn't suggest it).
A glimpse of your /etc/hosts and trust files could help, but I understand
it can be tricky to share it...
Regards.
On Tue, Sep 6, 2016 at 11:31 PM, Fernando Nunes <domusonline@gmail.com>
wrote:
> Ok....
> 1- Increase the LISTEN_TIMEOUT value to around 60
> 2- Being AIX... and considering you also have the trust issue... (and I
> assume you restarted the engine or at least the listeners, and that may
> have triggered the second issue):
> I suspect because I was told by a colleague of such a situation, that your
> AIX is multi-homed and there are very nsty routing issues "inside" the box.
> Not sure how to debug this, but try to run a traceroute from the client,
> or from inside the AIX box with target as the hostname you have in
> SQLHOSTS. Algo simulate a reverse lookup on that hostname and check if it
> returns the expected IP address (maybe it's returning 127.0.0.1 or
> something like that)
>
> Keep us posted... It's an interesting issue.
> My bet is that the reverse lookup changed. And by restarting the engine
> you made Informix "suffer" from that. Usually it's a big annoyance that
> Informix has no supported way to refresh the resolver low level cache. But
> that lack of functionality was probably what was letting it work.
> Regards.
> Regards.
>
> On Tue, Sep 6, 2016 at 11:17 PM, MARK SCRANTON <mark@markscranton.com>
> wrote:
>
>> Thanks for the responses Fernando ...
>>
>> AIX 7.1, IDS 11.50.FC8
>>
>> LISTEN_TIMEOUT 10
>> MAX_INCOMPLETE_CONNECTIONS 1024>>
>> The issue: a client connecting via TCP is having long sessions (up to 9
>> minutes) that eventually timeout. They see the msg at their end and I see
>> it
>> in my message log:
>>
>> listener-thread: err = -27001: oserr = 0: errstr = : Read error occurred
>> during connection attempt.
>>
>> (That is the full message of the original issue - when the client times
>> out
>> and the engine notices)
>>
>> Without a lot of confidence (based on lack of depth of knowledge), I
>> added to
>> sqlhosts col5 was simply k=1 for the engine entries. (2, both pointed to
>> the
>> same port for DBSERVERNAME and DBSERVERALIASES - setup long before me).
>>
>> What are your NETTYPE settings and how many connections per second do you
>> have
>> (from what you say I assume not many...)
>> NETTYPE soctcp,1,200,CPU # Configure poll thread(s) for nettype>> 
netstat -i shows any errors? No netstat errors on this box.
>> These errors are probably network related as you say.... no errors on the
>> 
application side? If you reduce INFORMIXCONTIME, maybe they'll
>> show up
>> ;)
>>
>> The 2nd issue is the adding of k=1 to sqlhosts. It DID seem to cause
>> issues
>> with trust. The message log on the target is below:
>> 09/06/16 15:00:59 listener-thread: err = -956: oserr = 0: errstr =
>> user@remote_box: Client host or user user@remote_box is not trusted by
>> the
>> server.
>> I still think there is a connectivity/trust issue in general. We use
>> hosts.equiv the normal way from box to box.
>>
>> Thanks again -
>> Mark
>>
>>
>> ************************************************************
>> *******************
>> Forum Note: Use "Reply" to post a response in the discussion forum.
>>
>>
>
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--94eb2c05fb205c051c053bdda0fd
it would be due to incomplete hand shake. Any process is just checking the informix engine engine live status and not making the proper connection. You may reproduce that issue by just telnet the port of your instance and will get the same message in online message log.
Related threads
- System Or Internal Error InterruptedIOException
- Informix ODBC Error: Read error occured during con
- error 27001