RE: Constant online log errors on all servers
Posted in 2006
Fernando,
Great points.
Wasn't it just announced that oracle will offer support for redhat?
That's bound to be a hoot!
Norma Jean Sebastian
ERP Support Administration
GIS- Enterprise Technical Services
-----Original Message-----
From: informix-list-bounces@iiug.org
[mailto:informix-list-bounces@iiug.org] On Behalf Of Fernando Nunes
Sent: Monday, November 27, 2006 7:00 PM
To: informix-list@iiug.org
Subject: Re: Constant online log errors on all servers
Neil Truby wrote:
> This is a repeat of a problem I posted a couple of months back.
Searching
> for inspiration really. IBM Informix Tech Support says it's a network
error
> (PMR 04123,019,866 refers). RedHat tech support say it's a database
error.
> The network administrators have no idea.
>
> I tend to agree with IBM: it probably isn't an Informix problem per
se, but
> if anyone has some ideas on how to persuade IDS to give us more info,
that
> would be a start:
>
> IDS 10.0FC5. We have 3 RHEL AS 4 servers on this subnet. At exactly
the
> same time a couple of weeks ago they all started reporting the error
below
> every 3 minutes or so.
>
> These messages appear on all database servers, all day, every day.
The only
> way to get rid of them is to disable the database listening on this
specific
> tcp/ip port. Changing the port number fixes it. We did not make any
> changes, eg service name, prior to the problem starting, although we
arev
> not responsible for the network.
> 07:12:46 listener-thread: err = -25580: oserr = 104: errstr = :System
> error occurred in network function.
> System error = 104.
> 07:15:31 listener-thread: err = -25580: oserr = 104: errstr = :System
> error occurred in network function.
> System error = 104.
> 07:18:16 listener-thread: err = -25580: oserr = 104: errstr = :System
> error occurred in network function.
> System error = 104.
> 07:20:56 listener-thread: err = -25580: oserr = 104: errstr = :System
> error occurred in network function.
> System error = 104.
> 07:23:33 listener-thread: err = -25580: oserr = 104: errstr = :System
> error occurred in network function.
> System error = 104.
> 07:26:13 listener-thread: err = -25580: oserr = 104: errstr = :System
> error occurred in network function.
> System error = 104.
> 07:28:54 listener-thread: err = -25580: oserr = 104: errstr = :System
> error occurred in network function.
> System error = 104.
>
> Any way to get more info? IBM Informix support said:
>
> System error 104 means "Connection reset by peer" , apparently the
client
> dropped the connection before it was completely set up. In other words
the
> end user connections are timing out before the engine is able to
process the
> request for a connection. Past cases have pointed to an issue with
> authentication of users as the cause for the 104 error. IDS uses the
MSC vp
> to authenticate users, by making standard api calls to the Operating
System.
> While we did not capture a stack trace from the MSC VP, past cases
have
> pointing out that when this error is returned, the MSC vp is usually
waiting
> on IO. So, the call to the OS is timing out due to some kind of OS
issue
> with authentication of the users. There are many potential methods via
which
> linux can authenticate users we really need more clarification on
exactly
> how this is done in order to guess as to what might be going wrong.
> Would suggest:
> i) Check Operating System logs for further errors.
> ii) Logging a call with Linux vendor, to check for any OS issues with
> authentication"
>
> Since the Linux vendor says it's a database problem, we're pretty
stuck, so
> any inspiration would be welcome.
>
> Thanks and regards
> Neil
>
>
>
Being Linux it's probably easy to use several tools that are available.
Besides the suggestion(s) already made, I'd suggest the use of tcpdump
to check what kind of traffic is arriving at the Informix port and also
where does it come from.
Following the "inspiration" line you've mentioned, I can share a recent
situation I had:
The DB apparently got stucked. All dbaccesses got hang. After a lot of
digging with lsof, tcpdump, onstat -g stk, truss/strace etc. I've found
out that a call to a gethostbyaddr() OS function was not returning. In
version 7 this can hang all subsequently connections attempts.
The cause for this was a terrible coincidence that work hand in hand
with a weird way of working of a middleware application.
This application broadcasts continuously UDP packets to several ports.
When the gethostbyaddr() got a client port in the range of the other
application broadcast ports, the OS function got stuck reading these
packets. And since they don't stop arriving it didn't stop reading (and
discarding) them.
Why do I give this example? Because I understand perfectly the "middle
man" position where you're sitting... And I don't have strong enough
words that I want to use here to qualify the position taken by the OS
guys (the customer team) and the other application guys. From the IBM
support colleagues the first answer I got was similar: It usually is an
OS problem. I finally managed to get all the evidences needed, and
currently the OS (Tru64) provider has acknowledged the problem and is
working on a fix.
I find it amazing what we're seeing today at enterprise level support
services. We say to customers that one of the main reasons to still buy
software is the support provided. But when we see the kind of support
customers get, I'm surprised that they still pay for it. Don't take me
wrong: This is not about IBM support, or RH support. The environment
where I work has lots of different vendors and products, and this
feeling can be more or less generalized. MS and Oracle most of the times
ask: Are you on the lastest version? Do you have the latest fix pack?
No?! Then install it and then if the problem persists contact us
again... SAP does the same. And what really astonish me is the fact that
customers accept this most of the times...
Finally, this is not a critic to the support guys. I've been working
with impressive people, but I feel the organizations and the way they
work should be carefully reviewed. The ratio of closed cases per person
per time might hide a lot of customer dissatisfaction.
So... keep digging and keep pressuring. The errors didn't start
appearing from nowhere... some thing or condition must have changed...
Regards and good luck.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
_______________________________________________
Informix-list mailing list
Informix-list@iiug.org
http://www.iiug.org/mailman/listinfo/informix-list
============================================================
The information contained in this message may be privileged
and confidential and protected from disclosure. If the reader
of this message is not the intended recipient, or an employee
or agent responsible for delivering this message to the