HDR: ASF Echo-Thread Server error
Posted in 2011
Topics: High Availability & Replication, Storage & Space Management, Platform-Specific Issues
Hello All,
Primary - IDS 11.50.FC8W2 on HP-UX 11.23, 12 cores, IDS memory - 48G
Secondry - IDS 11.50.FC8W2 on HP-UX 11.23, 04 cores, IDS memory - 40G
2 SDS nodes(updatable)
DRAUTO 0
DRINTERVAL 15
DRTIMEOUT 90
Things were working fine until two days back we started getting following
messages in online.log:
AFAIK, No config/network/OS/appln changes made.
13:02:13 DR: Sending log 173640, size 150000 pages, 93.40 percent used
13:09:42 DR: Send error
13:09:42 ASF Echo-Thread Server: asfcode = -25580: oserr = 4: errstr = :
System error occurred in network function.
System error = 4.
13:09:42 DR_ERR set to -3
13:09:42 DR: Failure recovery error (14)
13:09:42 DR: Turned off on primary server
13:09:42 DR: Cannot connect to secondary server
13:09:52 DR: Primary server connected
13:09:53 DR: Secondary server needs failure recovery
Checking with the network team and the bandwidth utilization looks consistence
and fine.
Some chunks on Secondry(including root chunk) are in PI- status BUT all good
on Primary.
c0000007fa4ef1d0 1 1 0 250000 202365 PI-B- /dev/prod/rootdbs
c00000081725ad70 7 16 0 3750000 3124947 PI-B- /dev/prod1/dd01_data
During the error I see that the log applying is almost stoped on secondry:
Following o/p interver ~2 sec
c0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
$ onstat -l |grep U---C-Lc0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
$ onstat -l |grep U---C-Lc0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
$ onstat -l |grep U---C-Lc0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
$ onstat -l |grep U---C-Lc0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
$ onstat -l |grep U---C-Lc0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
$ onstat -l |grep U---C-L
Does these needs to be checked for consistency? using oncheck on HDR an option?
On Pri: "netstat -s |grep drop" (pasting only with count > 0):
11064159 connections closed (including 31881 drops)
16169 embryonic connections dropped
48 connections dropped by rexmit timeout
25 connections dropped by keepalive
6130 connect requests dropped due to no listener
1812106 ICMP messages dropped
On Sec:
330443 connections closed (including 62512 drops)
34457 embryonic connections dropped
107 connections dropped by rexmit timeout
72 connections dropped by keepalive
1094 connect requests dropped due to no listener
heavy loading in night time(we always do) for last two days the logs on HDR
are falling back as compared to Primary and we have to stop loading until both
are sync with current log.
Any pointers or suggestions what all areas I need to investigate to deal with
this errors?
Regards,
Vikas
Your secondary has 1/3 the CPU processing capability of the primary. That
could be part of it. Your secondary server should be at least as powerful
as the primary, not much less capable.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Fri, Aug 26, 2011 at 4:13 AM, VIKAS HIVARKAR
<vikas.hivarkar@gmail.com>wrote:
> Hello All,
>
> Primary - IDS 11.50.FC8W2 on HP-UX 11.23, 12 cores, IDS memory - 48G
> Secondry - IDS 11.50.FC8W2 on HP-UX 11.23, 04 cores, IDS memory - 40G
> 2 SDS nodes(updatable)
>
> DRAUTO 0
> DRINTERVAL 15
> DRTIMEOUT 90>
> Things were working fine until two days back we started getting following
> messages in online.log:
> AFAIK, No config/network/OS/appln changes made.
>
> 13:02:13 DR: Sending log 173640, size 150000 pages, 93.40 percent used
> 13:09:42 DR: Send error
> 13:09:42 ASF Echo-Thread Server: asfcode = -25580: oserr = 4: errstr = :
> System error occurred in network function.
> System error = 4.
> 13:09:42 DR_ERR set to -3
> 13:09:42 DR: Failure recovery error (14)
> 13:09:42 DR: Turned off on primary server
> 13:09:42 DR: Cannot connect to secondary server
> 13:09:52 DR: Primary server connected
> 13:09:53 DR: Secondary server needs failure recovery
>
> Checking with the network team and the bandwidth utilization looks
> consistence
> and fine.
>
> Some chunks on Secondry(including root chunk) are in PI- status BUT all
> good
> on Primary.
> c0000007fa4ef1d0 1 1 0 250000 202365 PI-B- /dev/prod/rootdbs
> c00000081725ad70 7 16 0 3750000 3124947 PI-B- /dev/prod1/dd01_data
>
> During the error I see that the log applying is almost stoped on secondry:
> Following o/p interver ~2 sec
>
> c0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
> $ onstat -l |grep U---C-L> c0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
> $ onstat -l |grep U---C-L> c0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
> $ onstat -l |grep U---C-L> c0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
> $ onstat -l |grep U---C-L> c0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
> $ onstat -l |grep U---C-L> c0000008172326c0 94 U---C-L 173641 3:1650053 150000 135680 90.45
> $ onstat -l |grep U---C-L>
> Does these needs to be checked for consistency? using oncheck on HDR an
> option?
>
> On Pri: "netstat -s |grep drop" (pasting only with count > 0):
> 11064159 connections closed (including 31881 drops)
> 16169 embryonic connections dropped
> 48 connections dropped by rexmit timeout
> 25 connections dropped by keepalive
> 6130 connect requests dropped due to no listener
> 1812106 ICMP messages dropped
>
> On Sec:
> 330443 connections closed (including 62512 drops)
> 34457 embryonic connections dropped
> 107 connections dropped by rexmit timeout
> 72 connections dropped by keepalive
> 1094 connect requests dropped due to no listener
>
> heavy loading in night time(we always do) for last two days the logs on HDR
> are falling back as compared to Primary and we have to stop loading until
> both
> are sync with current log.
>
> Any pointers or suggestions what all areas I need to investigate to deal
> with
> this errors?
>
> Regards,
> Vikas
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--bcaec52994bbeefcdf04ab65c39d