ER crashing ...
Posted in 2011
A site running IDS 10.00.UC9 reported their Enterprise Replication server crashing several times a day with an assert failure "Bad allocation size (0) for pool 'CDR'" in mt_shm_malloc_segid, with a stack through the DSI resolver thread, plus CDR NIF receive failures and lost CDR connections. IBM support blamed garbage data arriving on the ER port (network). Replies suggested changing the port or looking for port scans by security tools; Madison Pruet doubted a network cause, recalling a similar zero-length-allocation bug in 11.50 and asking whether a replicate had recently been dropped (it hadn't). No resolution or APAR is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Error Codes & Troubleshooting
10.00.UC9
our ER server is crashing several times a day since yesterday with the
following Assert Failure:
14:34:17 Assert Failed: Condition Failed (Bad allocation size (0) for pool 'CDR
' (0x715f0020)), In (mt_shm_malloc_segid)
14:34:17 Who: Session(8180, root@EastServer, 0, 73e27440)
Thread(11568, CDRD_60, 70272a60, 7)
File: mtshpool.c Line: 3218
14:34:17 Action: Please notify IBM Informix Technical Support.
14:34:17 Stack for thread: 11568 CDRD_60
base: 0x73e7f000
len: 69632
pc: 0x10070940
tos: 0x73e8e978
state: running
vp: 7
0x10070ad4 (oninit)afstack
0x10072344 (oninit)afhandler
0x1006c10c (oninit)mem_asfail
0x1006eb5c (oninit)mt_shm_malloc_segid
0x1006ec80 (oninit)mt_shm_malloc
0x10474f4c (oninit)cdr_mt_shm_malloc
0x108ccfbc (oninit)rmStreamHdrRead
0x108d3348 (oninit)rmiCreateDSCollisionBM
0x108d26a0 (oninit)rmiGetNextPendingTxn
0x108d9708 (oninit)rmNextMsg
0x108cac68 (oninit)dsiResolverThread
0x1046b820 (oninit)cdrTrampolineThread
0x10abbf98 (oninit)startup
tech has told us we need to "fix" our network as we are getting garbage data
placed on the port where ER listens. The engine will crash when the following
is detected on the other server in the ER pair :
14:39:10 CDR NIF receive failed asfcode: -25582 oserr 0
14:39:10 CDR connection to server lost, id 11, name <east_serv>
Reason: connection lost
However we feel that the engine should be able to handle this in a more
gracefull manner
anyone else thoughs ?
Mark, Sorry to sound critical of your words but you have to accept the product
for what it is and deal with it. If crap is thrown at the port you are using
how about using a different port or find the offending process and stop it
sending crap to the port? As you are using an unsupported version you are not
going to get much help. We are using 10.00.FC10 at the moment and we do not
get network problems at all and it works like a dream. I realise changing the
port is a pain in the bum but its not that hard. What port number are you
using by the way? Are you sure its unique? Regards Andy G.
> To: ids@iiug.org
> From: mark.jalkiewicz@verizon.net
> Subject: ER crashing ... [24810]
> Date: Wed, 31 Aug 2011 17:08:22 -0400
>
> 10.00.UC9
>
> our ER server is crashing several times a day since yesterday with the
> following Assert Failure:
>
> 14:34:17 Assert Failed: Condition Failed (Bad allocation size (0) for pool
> 'CDR
> ' (0x715f0020)), In (mt_shm_malloc_segid)
> 14:34:17 Who: Session(8180, root@EastServer, 0, 73e27440)
>
> Thread(11568, CDRD_60, 70272a60, 7)
>
> File: mtshpool.c Line: 3218
> 14:34:17 Action: Please notify IBM Informix Technical Support.
> 14:34:17 Stack for thread: 11568 CDRD_60
>
> base: 0x73e7f000
> len: 69632
>
> pc: 0x10070940
> tos: 0x73e8e978
> state: running
>
> vp: 7
>
> 0x10070ad4 (oninit)afstack
> 0x10072344 (oninit)afhandler
> 0x1006c10c (oninit)mem_asfail
> 0x1006eb5c (oninit)mt_shm_malloc_segid
> 0x1006ec80 (oninit)mt_shm_malloc
> 0x10474f4c (oninit)cdr_mt_shm_malloc
> 0x108ccfbc (oninit)rmStreamHdrRead
> 0x108d3348 (oninit)rmiCreateDSCollisionBM
> 0x108d26a0 (oninit)rmiGetNextPendingTxn
> 0x108d9708 (oninit)rmNextMsg
> 0x108cac68 (oninit)dsiResolverThread
> 0x1046b820 (oninit)cdrTrampolineThread
> 0x10abbf98 (oninit)startup
>
> tech has told us we need to "fix" our network as we are getting garbage data
> placed on the port where ER listens. The engine will crash when the following
> is detected on the other server in the ER pair :
>
> 14:39:10 CDR NIF receive failed asfcode: -25582 oserr 0
>
> 14:39:10 CDR connection to server lost, id 11, name <east_serv>
> Reason: connection lost
>
> However we feel that the engine should be able to handle this in a more
> gracefull manner
>
> anyone else thoughs ?
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
We have experienced similar network errors when our Info Security team port
scans a machine. Could this be the case in your environment?
________________________________
From: ids-bounces@iiug.org [ids-bounces@iiug.org] On Behalf Of Andrew Grantham
[agrantha@hotmail.com]
Sent: Wednesday, August 31, 2011 4:22 PM
To: ids@iiug.org
Subject: RE: ER crashing ... [24811]
Mark, Sorry to sound critical of your words but you have to accept the product
for what it is and deal with it. If crap is thrown at the port you are using
how about using a different port or find the offending process and stop it
sending crap to the port? As you are using an unsupported version you are not
going to get much help. We are using 10.00.FC10 at the moment and we do not
get network problems at all and it works like a dream. I realise changing the
port is a pain in the bum but its not that hard. What port number are you
using by the way? Are you sure its unique? Regards Andy G.
> To: ids@iiug.org
> From: mark.jalkiewicz@verizon.net
> Subject: ER crashing ... [24810]
> Date: Wed, 31 Aug 2011 17:08:22 -0400
>
> 10.00.UC9
>
> our ER server is crashing several times a day since yesterday with the
> following Assert Failure:
>
> 14:34:17 Assert Failed: Condition Failed (Bad allocation size (0) for pool
> 'CDR
> ' (0x715f0020)), In (mt_shm_malloc_segid)
> 14:34:17 Who: Session(8180, root@EastServer, 0, 73e27440)
>
> Thread(11568, CDRD_60, 70272a60, 7)
>
> File: mtshpool.c Line: 3218
> 14:34:17 Action: Please notify IBM Informix Technical Support.
> 14:34:17 Stack for thread: 11568 CDRD_60
>
> base: 0x73e7f000
> len: 69632
>
> pc: 0x10070940
> tos: 0x73e8e978
> state: running
>
> vp: 7
>
> 0x10070ad4 (oninit)afstack
> 0x10072344 (oninit)afhandler
> 0x1006c10c (oninit)mem_asfail
> 0x1006eb5c (oninit)mt_shm_malloc_segid
> 0x1006ec80 (oninit)mt_shm_malloc
> 0x10474f4c (oninit)cdr_mt_shm_malloc
> 0x108ccfbc (oninit)rmStreamHdrRead
> 0x108d3348 (oninit)rmiCreateDSCollisionBM
> 0x108d26a0 (oninit)rmiGetNextPendingTxn
> 0x108d9708 (oninit)rmNextMsg
> 0x108cac68 (oninit)dsiResolverThread
> 0x1046b820 (oninit)cdrTrampolineThread
> 0x10abbf98 (oninit)startup
>
> tech has told us we need to "fix" our network as we are getting garbage data
> placed on the port where ER listens. The engine will crash when the
following
> is detected on the other server in the ER pair :
>
> 14:39:10 CDR NIF receive failed asfcode: -25582 oserr 0
>
> 14:39:10 CDR connection to server lost, id 11, name <east_serv>
> Reason: connection lost
>
> However we feel that the engine should be able to handle this in a more
> gracefull manner
>
> anyone else thoughs ?
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Mark,
Somehow a request is being made for a zero length memory allocation. I=
'm
fairly certain that there was a bug in 11.50 which was similar to this
condition, but I can't remember any similar problem in version 10.
I doubt that it has anything to do with network issues,
By any chance did you drop a replicate close to when the problem first
appeared?
M.P.
=
From: "MARK JALKIEWICZ" <mark.jalkiewicz@verizon.net> =
=
To: ids@iiug.org =
=
Date: 08/31/2011 04:09 PM =
=
Subject: ER crashing ... [24810] =
=
Sent by: ids-bounces@iiug.org =
=
10.00.UC9
our ER server is crashing several times a day since yesterday with the
following Assert Failure:
14:34:17 Assert Failed: Condition Failed (Bad allocation size (0) for p=
ool
'CDR
' (0x715f0020)), In (mt_shm_malloc_segid)
14:34:17 Who: Session(8180, root@EastServer, 0, 73e27440)
Thread(11568, CDRD_60, 70272a60, 7)
File: mtshpool.c Line: 3218
14:34:17 Action: Please notify IBM Informix Technical Support.
14:34:17 Stack for thread: 11568 CDRD_60
base: 0x73e7f000
len: 69632
pc: 0x10070940
tos: 0x73e8e978
state: running
vp: 7
0x10070ad4 (oninit)afstack
0x10072344 (oninit)afhandler
0x1006c10c (oninit)mem_asfail
0x1006eb5c (oninit)mt_shm_malloc_segid
0x1006ec80 (oninit)mt_shm_malloc
0x10474f4c (oninit)cdr_mt_shm_malloc
0x108ccfbc (oninit)rmStreamHdrRead
0x108d3348 (oninit)rmiCreateDSCollisionBM
0x108d26a0 (oninit)rmiGetNextPendingTxn
0x108d9708 (oninit)rmNextMsg
0x108cac68 (oninit)dsiResolverThread
0x1046b820 (oninit)cdrTrampolineThread
0x10abbf98 (oninit)startup
tech has told us we need to "fix" our network as we are getting garbage=
data
placed on the port where ER listens. The engine will crash when the
following
is detected on the other server in the ER pair :
14:39:10 CDR NIF receive failed asfcode: -25582 oserr 0
14:39:10 CDR connection to server lost, id 11, name <east_serv>
Reason: connection lost
However we feel that the engine should be able to handle this in a more=
gracefull manner
anyone else thoughs ?
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
Madison, No replicate drops /adds etc.. going on, its just bau. No application changes made recently. If you happen to come up the apar for this defect pass it along. Thanks, Mark
Madison, No replicate deletes /create etc.. going on, its just bau. No application changes made recently. If you happen to come up with the apar for the defect you mention pass it along. Thanks, Mark
Related threads
- System Or Internal Error InterruptedIOException
- 25582 error on high volume of short live trans
- ASF Echo-Thread Server: asfcode = -25582 oserr = 4