Constant online log errors on all servers
Posted in 2006
Topics: Error Codes & Troubleshooting, Networking & sqlhosts Configuration, Platform-Specific Issues, Versions, Editions & End-of-Life
This is a repeat of a problem I posted a couple of months back. Searching
for inspiration really. IBM Informix Tech Support says it's a network error
(PMR 04123,019,866 refers). RedHat tech support say it's a database error.
The network administrators have no idea.
I tend to agree with IBM: it probably isn't an Informix problem per se, but
if anyone has some ideas on how to persuade IDS to give us more info, that
would be a start:
IDS 10.0FC5. We have 3 RHEL AS 4 servers on this subnet. At exactly the
same time a couple of weeks ago they all started reporting the error below
every 3 minutes or so.
These messages appear on all database servers, all day, every day. The only
way to get rid of them is to disable the database listening on this specific
tcp/ip port. Changing the port number fixes it. We did not make any
changes, eg service name, prior to the problem starting, although we arev
not responsible for the network.
07:12:46 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:15:31 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:18:16 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:20:56 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:23:33 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:26:13 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:28:54 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
Any way to get more info? IBM Informix support said:
System error 104 means "Connection reset by peer" , apparently the client
dropped the connection before it was completely set up. In other words the
end user connections are timing out before the engine is able to process the
request for a connection. Past cases have pointed to an issue with
authentication of users as the cause for the 104 error. IDS uses the MSC vp
to authenticate users, by making standard api calls to the Operating System.
While we did not capture a stack trace from the MSC VP, past cases have
pointing out that when this error is returned, the MSC vp is usually waiting
on IO. So, the call to the OS is timing out due to some kind of OS issue
with authentication of the users. There are many potential methods via which
linux can authenticate users we really need more clarification on exactly
how this is done in order to guess as to what might be going wrong.
Would suggest:
i) Check Operating System logs for further errors.
ii) Logging a call with Linux vendor, to check for any OS issues with
authentication"
Since the Linux vendor says it's a database problem, we're pretty stuck, so
any inspiration would be welcome.
Thanks and regards
Neil
Neil Truby wrote:
> This is a repeat of a problem I posted a couple of months back. Searching
> for inspiration really. IBM Informix Tech Support says it's a network error
> (PMR 04123,019,866 refers). RedHat tech support say it's a database error.
> The network administrators have no idea.
>
> I tend to agree with IBM: it probably isn't an Informix problem per se, but
> if anyone has some ideas on how to persuade IDS to give us more info, that
> would be a start:
>
> IDS 10.0FC5. We have 3 RHEL AS 4 servers on this subnet. At exactly the
> same time a couple of weeks ago they all started reporting the error below
> every 3 minutes or so.
>
> These messages appear on all database servers, all day, every day. The only
> way to get rid of them is to disable the database listening on this specific
> tcp/ip port. Changing the port number fixes it. We did not make any
> changes, eg service name, prior to the problem starting, although we arev
> not responsible for the network.
> 07:12:46 listener-thread: err = -25580: oserr = 104: errstr = : System> error occurred in network function.
> System error = 104.
> 07:15:31 listener-thread: err = -25580: oserr = 104: errstr = : System> error occurred in network function.
> System error = 104.
> 07:18:16 listener-thread: err = -25580: oserr = 104: errstr = : System> error occurred in network function.
> System error = 104.
> 07:20:56 listener-thread: err = -25580: oserr = 104: errstr = : System> error occurred in network function.
> System error = 104.
> 07:23:33 listener-thread: err = -25580: oserr = 104: errstr = : System> error occurred in network function.
> System error = 104.
> 07:26:13 listener-thread: err = -25580: oserr = 104: errstr = : System> error occurred in network function.
> System error = 104.
> 07:28:54 listener-thread: err = -25580: oserr = 104: errstr = : System> error occurred in network function.
> System error = 104.
>
> Any way to get more info? IBM Informix support said:
>
> System error 104 means "Connection reset by peer" , apparently the client
> dropped the connection before it was completely set up. In other words the
> end user connections are timing out before the engine is able to process the
> request for a connection. Past cases have pointed to an issue with
> authentication of users as the cause for the 104 error. IDS uses the MSC vp
> to authenticate users, by making standard api calls to the Operating System.
> While we did not capture a stack trace from the MSC VP, past cases have
> pointing out that when this error is returned, the MSC vp is usually waiting
> on IO. So, the call to the OS is timing out due to some kind of OS issue
> with authentication of the users. There are many potential methods via which
> linux can authenticate users we really need more clarification on exactly
> how this is done in order to guess as to what might be going wrong.
> Would suggest:
> i) Check Operating System logs for further errors.
> ii) Logging a call with Linux vendor, to check for any OS issues with
> authentication"
>
> Since the Linux vendor says it's a database problem, we're pretty stuck, so
> any inspiration would be welcome.
>
> Thanks and regards
> Neil
>
>
>
Being Linux it's probably easy to use several tools that are available.
Besides the suggestion(s) already made, I'd suggest the use of tcpdump to check what kind of traffic is arriving at the Informix port and also where does it come from.
Following the "inspiration" line you've mentioned, I can share a recent situation I had:
The DB apparently got stucked. All dbaccesses got hang. After a lot of digging with lsof, tcpdump, onstat -g stk, truss/strace etc. I've found out that a call to a gethostbyaddr() OS function was not returning. In version 7 this can hang all subsequently connections attempts.
The cause for this was a terrible coincidence that work hand in hand with a weird way of working of a middleware application.
This application broadcasts continuously UDP packets to several ports. When the gethostbyaddr() got a client port in the range of the other application broadcast ports, the OS function got stuck reading these packets. And since they don't stop arriving it didn't stop reading (and discarding) them.
Why do I give this example? Because I understand perfectly the "middle man" position where you're sitting... And I don't have strong enough words that I want to use here to qualify the position taken by the OS guys (the customer team) and the other application guys. From the IBM support colleagues the first answer I got was similar: It usually is an OS problem. I finally managed to get all the evidences needed, and currently the OS (Tru64) provider has acknowledged the problem and is working on a fix.
I find it amazing what we're seeing today at enterprise level support services. We say to customers that one of the main reasons to still buy software is the support provided. But when we see the kind of support customers get, I'm surprised that they still pay for it. Don't take me wrong: This is not about IBM support, or RH support. The environment where I work has lots of different vendors and products, and this feeling can be more or less generalized. MS and Oracle most of the times ask: Are you on the lastest version? Do you have the latest fix pack? No?! Then install it and then if the problem persists contact us again... SAP does the same. And what really astonish me is the fact that customers accept this most of the times...
Finally, this is not a critic to the support guys. I've been working with impressive people, but I feel the organizations and the way they work should be carefully reviewed. The ratio of closed cases per person per time might hide a lot of customer dissatisfaction.
So... keep digging and keep pressuring. The errors didn't start appearing from nowhere... some thing or condition must have changed...
Regards and good luck.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...