RE: Constant online log errors on all servers
Posted in 2006
Topics: Error Codes & Troubleshooting, Server Administration, Networking & sqlhosts Configuration, Platform-Specific Issues, Versions, Editions & End-of-Life
Poor you.... DB guy stuck in the middle. Oh the fun of "defending the
database" ;)
Look into this free download product to help you sniff around and watch
the traffic:
http://www.ethereal.com/
what does your sqlhost file look like, assume these errors are all
coming via network access versus shared mem?
What are you NETTYPE settings in ONCONFIG?
Norma Jean Sebastian
ERP Support Administration
GIS- Enterprise Technical Services
-----Original Message-----
From: informix-list-bounces@iiug.org
[mailto:informix-list-bounces@iiug.org] On Behalf Of Neil Truby
Sent: Monday, November 27, 2006 4:38 PM
To: informix-list@iiug.org
Subject: Constant online log errors on all servers
This is a repeat of a problem I posted a couple of months back.
Searching
for inspiration really. IBM Informix Tech Support says it's a network
error
(PMR 04123,019,866 refers). RedHat tech support say it's a database
error.
The network administrators have no idea.
I tend to agree with IBM: it probably isn't an Informix problem per se,
but
if anyone has some ideas on how to persuade IDS to give us more info,
that
would be a start:
IDS 10.0FC5. We have 3 RHEL AS 4 servers on this subnet. At exactly the
same time a couple of weeks ago they all started reporting the error
below
every 3 minutes or so.
These messages appear on all database servers, all day, every day. The
only
way to get rid of them is to disable the database listening on this
specific
tcp/ip port. Changing the port number fixes it. We did not make any
changes, eg service name, prior to the problem starting, although we
arev
not responsible for the network.
07:12:46 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:15:31 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:18:16 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:20:56 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:23:33 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:26:13 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
07:28:54 listener-thread: err = -25580: oserr = 104: errstr = : Systemerror occurred in network function.
System error = 104.
Any way to get more info? IBM Informix support said:
System error 104 means "Connection reset by peer" , apparently the
client
dropped the connection before it was completely set up. In other words
the
end user connections are timing out before the engine is able to process
the
request for a connection. Past cases have pointed to an issue with
authentication of users as the cause for the 104 error. IDS uses the MSC
vp
to authenticate users, by making standard api calls to the Operating
System.
While we did not capture a stack trace from the MSC VP, past cases have
pointing out that when this error is returned, the MSC vp is usually
waiting
on IO. So, the call to the OS is timing out due to some kind of OS issue
with authentication of the users. There are many potential methods via
which
linux can authenticate users we really need more clarification on
exactly
how this is done in order to guess as to what might be going wrong.
Would suggest:
i) Check Operating System logs for further errors.
ii) Logging a call with Linux vendor, to check for any OS issues with
authentication"
Since the Linux vendor says it's a database problem, we're pretty stuck,
so
any inspiration would be welcome.
Thanks and regards
Neil
_______________________________________________
Informix-list mailing list
Informix-list@iiug.org
http://www.iiug.org/mailman/listinfo/informix-list
============================================================
The information contained in this message may be privileged
and confidential and protected from disclosure. If the reader
of this message is not the intended recipient, or an employee
or agent responsible for delivering this message to the
intended recipient, you are hereby notified that any reproduction,
dissemination or distribution of this communication is strictly
prohibited. If you have received this communication in error,
please notify us immediately by replying to the message and
deleting it from your computer. Thank you. Tellabs
============================================================
"Sebastian, Norma J." <NormaJean.Sebastian@tellabs.com> wrote in message
news:mailman.380.1164669679.29126.informix-list@iiug.org...
>> Poor you.... DB guy stuck in the middle.
Yes, poor me. And I get paid a pittance :-(
>> what does your sqlhost file look like, assume these errors are all
coming via network access versus shared mem?
What are you NETTYPE settings in ONCONFIG?
Yes, connections will be coming via tcp/ip. There are 3 servers, live, the
one with the live HDR secondary, and a test one all affected (and the
problem started at exactly the same time on exactly the same day on all 3).
Here's the test one:
ifx_shm_11 onipcshm xxxtestdb01 dummy_11
appname_test_tcp onsoctcp xxxtestdb01 appname_test_tcp
appname_test_tcp_crap onsoctcp xxxtestdb01 appname_test_tcp_crap
(appname and the xxxs are substituted for something that might identify the
customer's name). The _crap connection listens on a different port number:
when it's the only tcp/ip connection enabled we don't get the problem.
NETTYPE ipcshm,1,50,CPU # Configure poll thread(s)for nettype
NETTYPE soctcp,2,200,NET # Configure poll thread(s)for nettype
rgds
Neil