ER crash
Posted in 2007
Frank reported IDS 10.00.UC5 on AIX 5.3 panicking (assert in mtex.c, later shared-memory pool corruption in CDRGeval threads) when deleting/updating rows under an update-anywhere ER replicate defined with --conflict=always --ats --ris, where rows existed on only one server. IBM support initially denied an IDS bug; on the list Madison Pruet and Wee Tung couldn't reproduce it (UC6 worked) and asked for a repro, onconfig and AF/stack details, while Nilesh Ozarkar noted the errno 28 messages pointed to /tmp running out of space for the shmem dump and ATS/RIS spool files. The thread ends with Frank posting a second reproduction and stack trace; no confirmed fix or root cause is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Error Codes & Troubleshooting
HI, Folks, AIX5.3, IDS UC5. the scenario is, (1) define a pair of update-anywhere ER replication server between two separate boxes, (2) define a replicate, some thing like , cdr define replicate --conflict=always --ats --ris abc_test_repl \\\\ "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\ "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test" (3) create some rows in abc_test at on server(site), but NOT in another server. (4) delete or update the rows from the server where the rows exist, then, PANIC. 17:45:09 Maximum server connections 91 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df 17:46:07 Assert Failed: No Exception Handler 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00) Thread(418964, CDRGeval2, b0302574, 7) File: mtex.c Line: 472 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem 17:46:07 Action: Please notify IBM Informix Technical Support. 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno = 28 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exception Handler. 17:46:44 Fatal error in ADM VP at mt.c:13055 17:46:44 Unexpected virtual processor termination, pid = 377262, exit = 0x100 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.pl 5 6 "Internal Subsystem failure: 'MT'" "Unexpected virtual processor termination, pid = 377262, exit = 0x100 " ' 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0 17:46:44 PANIC: Attempting to bring system down I reported it to IBM TS, but they think this is NOT IDS problem. Just wonder, any folks can reproduce the same problem. Thank you, Frank
After crash, if you restart the ER, it showed, 18:04:17 CDR NIF listening on asf://osdint1_grp 18:04:17 DDR Log Snooping - Snooping started in log 1588 18:04:18 CDR: Re-connected to server, id 3, name <ncdcprd_grp> 18:04:19 CDR CDRGeval3: error writing to spool file /tmp/RepErr/ris.osdint1_grp.0.Geval3.070123_18:04:18.1 18:04:19 CDR CDRGeval3: error writing to spool file /tmp/RepErr/ris.osdint1_grp.0.Geval3.070123_18:04:18.1 18:04:19 CDR CDRGeval3: error writing to spool file /tmp/RepErr/ris.osdint1_grp.0.Geval3.070123_18:04:18.1 18:04:19 Assert Warning: Memory free block header corruption detected in mt_shm_malloc_segid 5 18:04:19 IBM Informix Dynamic Server Version 10.00.UC5 18:04:19 Who: Session(34, informix@erika, 0, b0312118) Thread(68, CDRGeval3, b02eec8c, 4) File: mtshpool.c Line: 3422 18:04:19 Results: Pool repaired 18:04:19 Action: Please notify IBM Informix Technical Support. 18:04:19 stack trace for pid 376908 written to /tmp/af.42c4e23 18:04:19 See Also: /tmp/af.42c4e23 18:04:20 Memory free block header corruption detected in mt_shm_malloc_segid 5 18:04:32 Checkpoint Completed: duration was 0 seconds. 18:04:32 Checkpoint loguniq 1588, logpos 0x16b4018, timestamp: 0x7d519d7b Frank On 1/23/07, FRANK <yunyaoqu@gmail.com> wrote: > > > HI, Folks, > > AIX5.3, IDS UC5. > > the scenario is, > > (1) define a pair of update-anywhere ER replication server between two > separate boxes, > (2) define a replicate, some thing like , > > cdr define replicate --conflict=always --ats --ris abc_test_repl \\\\ > "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\ > "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test" > > (3) create some rows in abc_test at on server(site), but NOT in another > server. > > (4) delete or update the rows from the server where the rows exist, > > then, PANIC. > > 17:45:09 Maximum server connections 91 > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 Assert Failed: No Exception Handler > 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5 > 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00) > > Thread(418964, CDRGeval2, b0302574, 7) > > File: mtex.c Line: 472 > 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem > 17:46:07 Action: Please notify IBM Informix Technical Support. > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0 > 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno = 28 > 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exception > Handler. > 17:46:44 Fatal error in ADM VP at mt.c:13055 > 17:46:44 Unexpected virtual processor termination, pid = 377262, exit = > 0x100 > > 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.pl 5 6 > "Internal Subsystem failure: 'MT'" "Unexpected virtual processor > termination, pid = 377262, exit = 0x100 > " ' > 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0 > 17:46:44 PANIC: Attempting to bring system down > > I reported it to IBM TS, but they think this is NOT IDS problem. > > Just wonder, any folks can reproduce the same problem. > > Thank you, > Frank > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > >
What version of the server are you running? = "FRANK" = <yunyaoqu@gmail.c = om> = To Sent by: ids@iiug.org = ids-bounces@iiug. = cc org = Subj= ect ER crash [8281] = 01/23/2007 11:52 = AM = = = Please respond to = ids@iiug.org = = = HI, Folks, AIX5.3, IDS UC5. the scenario is, (1) define a pair of update-anywhere ER replication server between two separate boxes, (2) define a replicate, some thing like , cdr define replicate --conflict=3Dalways --ats --ris abc_test_repl \\\\ "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\ "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test" (3) create some rows in abc_test at on server(site), but NOT in another= server. (4) delete or update the rows from the server where the rows exist, then, PANIC. 17:45:09 Maximum server connections 91 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df 17:46:07 Assert Failed: No Exception Handler 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00) Thread(418964, CDRGeval2, b0302574, 7) File: mtex.c Line: 472 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem 17:46:07 Action: Please notify IBM Informix Technical Support. 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno =3D 28 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exception Handler. 17:46:44 Fatal error in ADM VP at mt.c:13055 17:46:44 Unexpected virtual processor termination, pid =3D 377262, exit= =3D 0x100 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.pl = 5 6 "Internal Subsystem failure: 'MT'" "Unexpected virtual processor termination, pid =3D 377262, exit =3D 0x100 " ' 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0 17:46:44 PANIC: Attempting to bring system down I reported it to IBM TS, but they think this is NOT IDS problem. Just wonder, any folks can reproduce the same problem. Thank you, Frank ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
Madison, Do you mean? AIX5.3, IDS10 UC5., or some thing else? Thanks, Frank On 1/23/07, Madison Pruet <mpruet@us.ibm.com> wrote: > > > What version of the server are you running? > > = > > "FRANK" = > > <yunyaoqu@gmail.c = > > om> = > To > > Sent by: ids@iiug.org = > > ids-bounces@iiug. = > cc > > org = > > Subj= > ect > > ER crash [8281] = > > 01/23/2007 11:52 = > > AM = > > = > > = > > Please respond to = > > ids@iiug.org = > > = > > = > > HI, Folks, > > AIX5.3, IDS UC5. > > the scenario is, > > (1) define a pair of update-anywhere ER replication server between two > separate boxes, > (2) define a replicate, some thing like , > > cdr define replicate --conflict=3Dalways --ats --ris abc_test_repl \\\\ > "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\ > "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test" > > (3) create some rows in abc_test at on server(site), but NOT in another= > > server. > > (4) delete or update the rows from the server where the rows exist, > > then, PANIC. > > 17:45:09 Maximum server connections 91 > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 Assert Failed: No Exception Handler > 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5 > 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00) > > Thread(418964, CDRGeval2, b0302574, 7) > > File: mtex.c Line: 472 > 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem > 17:46:07 Action: Please notify IBM Informix Technical Support. > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0 > 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno =3D 28 > 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exception > Handler. > 17:46:44 Fatal error in ADM VP at mt.c:13055 > 17:46:44 Unexpected virtual processor termination, pid =3D 377262, exit= > =3D > 0x100 > > 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.pl = > 5 6 > "Internal Subsystem failure: 'MT'" "Unexpected virtual processor > termination, pid =3D 377262, exit =3D 0x100 > " ' > 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0 > 17:46:44 PANIC: Attempting to bring system down > > I reported it to IBM TS, but they think this is NOT IDS problem. > > Just wonder, any folks can reproduce the same problem. > > Thank you, > Frank > > ***********************************************************************= > ******** > > Forum Note: Use "Reply" to post a response in the discussion forum. > > = > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > >
Silly me. IDS10.00.UC5 Do you have a repro? If so, then let's treat it as a bug? Thanks. M.P. = "FRANK" = <yunyaoqu@gmail.c = om> = To Sent by: ids@iiug.org = ids-bounces@iiug. = cc org = Subj= ect ER crash [8281] = 01/23/2007 11:52 = AM = = = Please respond to = ids@iiug.org = = = HI, Folks, AIX5.3, IDS UC5. the scenario is, (1) define a pair of update-anywhere ER replication server between two separate boxes, (2) define a replicate, some thing like , cdr define replicate --conflict=3Dalways --ats --ris abc_test_repl \\\\ "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\ "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test" (3) create some rows in abc_test at on server(site), but NOT in another= server. (4) delete or update the rows from the server where the rows exist, then, PANIC. 17:45:09 Maximum server connections 91 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df 17:46:07 Assert Failed: No Exception Handler 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00) Thread(418964, CDRGeval2, b0302574, 7) File: mtex.c Line: 472 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem 17:46:07 Action: Please notify IBM Informix Technical Support. 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno =3D 28 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exception Handler. 17:46:44 Fatal error in ADM VP at mt.c:13055 17:46:44 Unexpected virtual processor termination, pid =3D 377262, exit= =3D 0x100 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.pl = 5 6 "Internal Subsystem failure: 'MT'" "Unexpected virtual processor termination, pid =3D 377262, exit =3D 0x100 " ' 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0 17:46:44 PANIC: Attempting to bring system down I reported it to IBM TS, but they think this is NOT IDS problem. Just wonder, any folks can reproduce the same problem. Thank you, Frank ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
Can you post following info from AF file /tmp/af.687c49df
onstat -g ses 418266
onstat -g stk 418964
onstat -g ath
Regards,
Nilesh
"FRANK" <yunyaoqu@gmail.com>
Sent by: ids-bounces@iiug.org
01/23/2007 11:52 AM
Please respond to
ids@iiug.org
To
ids@iiug.org
cc
Subject
ER crash [8281]
HI, Folks,
AIX5.3, IDS UC5.
the scenario is,
(1) define a pair of update-anywhere ER replication server between two
separate boxes,
(2) define a replicate, some thing like ,
cdr define replicate --conflict=always --ats --ris abc_test_repl \\\\
"noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\
"noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test"
(3) create some rows in abc_test at on server(site), but NOT in another
server.
(4) delete or update the rows from the server where the rows exist,
then, PANIC.
17:45:09 Maximum server connections 91
17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df
17:46:07 Assert Failed: No Exception Handler
17:46:07 IBM Informix Dynamic Server Version 10.00.UC5
17:46:07 Who: Session(418266, informix@erika, 0, b0319d00)
Thread(418964, CDRGeval2, b0302574, 7)
File: mtex.c Line: 472
17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem
17:46:07 Action: Please notify IBM Informix Technical Support.
17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df
17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0
17:46:43 Error writing '/tmp/shmem.687c49df.0' errno = 28
17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exception
Handler.
17:46:44 Fatal error in ADM VP at mt.c:13055
17:46:44 Unexpected virtual processor termination, pid = 377262, exit =
0x100
17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.pl 5 6
"Internal Subsystem failure: 'MT'" "Unexpected virtual processor
termination, pid = 377262, exit = 0x100
" '
17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0
17:46:44 PANIC: Attempting to bring system down
I reported it to IBM TS, but they think this is NOT IDS problem.
Just wonder, any folks can reproduce the same problem.
Thank you,
Frank
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Madison, I have this from IBM, PMR 50003,487,000 Thanks, Frank On 1/23/07, Madison Pruet <mpruet@us.ibm.com> wrote: > > > Silly me. IDS10.00.UC5 > > Do you have a repro? If so, then let's treat it as a bug? > > Thanks. > > M.P. > > = > > "FRANK" = > > <yunyaoqu@gmail.c = > > om> = > To > > Sent by: ids@iiug.org = > > ids-bounces@iiug. = > cc > > org = > > Subj= > ect > > ER crash [8281] = > > 01/23/2007 11:52 = > > AM = > > = > > = > > Please respond to = > > ids@iiug.org = > > = > > = > > HI, Folks, > > AIX5.3, IDS UC5. > > the scenario is, > > (1) define a pair of update-anywhere ER replication server between two > separate boxes, > (2) define a replicate, some thing like , > > cdr define replicate --conflict=3Dalways --ats --ris abc_test_repl \\\\ > "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\ > "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test" > > (3) create some rows in abc_test at on server(site), but NOT in another= > > server. > > (4) delete or update the rows from the server where the rows exist, > > then, PANIC. > > 17:45:09 Maximum server connections 91 > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 Assert Failed: No Exception Handler > 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5 > 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00) > > Thread(418964, CDRGeval2, b0302574, 7) > > File: mtex.c Line: 472 > 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem > 17:46:07 Action: Please notify IBM Informix Technical Support. > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0 > 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno =3D 28 > 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exception > Handler. > 17:46:44 Fatal error in ADM VP at mt.c:13055 > 17:46:44 Unexpected virtual processor termination, pid =3D 377262, exit= > =3D > 0x100 > > 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.pl = > 5 6 > "Internal Subsystem failure: 'MT'" "Unexpected virtual processor > termination, pid =3D 377262, exit =3D 0x100 > " ' > 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0 > 17:46:44 PANIC: Attempting to bring system down > > I reported it to IBM TS, but they think this is NOT IDS problem. > > Just wonder, any folks can reproduce the same problem. > > Thank you, > Frank > > ***********************************************************************= > ******** > > Forum Note: Use "Reply" to post a response in the discussion forum. > > = > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > >
Have you check if the system has the required patches installed? The re= pro you describe are a basic operation and I tried recreate on 10.00.UC6 bu= t it doesn't repro. Please check your system's patches. Sen = "FRANK" = <yunyaoqu@gmail.c = om> = To Sent by: ids@iiug.org = ids-bounces@iiug. = cc org = Subj= ect Re: ER crash [8288] = 01/23/2007 12:45 = PM = = = Please respond to = ids@iiug.org = = = Madison, I have this from IBM, PMR 50003,487,000 Thanks, Frank On 1/23/07, Madison Pruet <mpruet@us.ibm.com> wrote: > > > Silly me. IDS10.00.UC5 > > Do you have a repro? If so, then let's treat it as a bug? > > Thanks. > > M.P. > > =3D > > "FRANK" =3D > > <yunyaoqu@gmail.c =3D > > om> =3D > To > > Sent by: ids@iiug.org =3D > > ids-bounces@iiug. =3D > cc > > org =3D > > Subj=3D > ect > > ER crash [8281] =3D > > 01/23/2007 11:52 =3D > > AM =3D > > =3D > > =3D > > Please respond to =3D > > ids@iiug.org =3D > > =3D > > =3D > > HI, Folks, > > AIX5.3, IDS UC5. > > the scenario is, > > (1) define a pair of update-anywhere ER replication server between tw= o > separate boxes, > (2) define a replicate, some thing like , > > cdr define replicate --conflict=3D3Dalways --ats --ris abc_test_repl = \\\\ > "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\ > "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test" > > (3) create some rows in abc_test at on server(site), but NOT in anoth= er=3D > > server. > > (4) delete or update the rows from the server where the rows exist, > > then, PANIC. > > 17:45:09 Maximum server connections 91 > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 Assert Failed: No Exception Handler > 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5 > 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00) > > Thread(418964, CDRGeval2, b0302574, 7) > > File: mtex.c Line: 472 > 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem > 17:46:07 Action: Please notify IBM Informix Technical Support. > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0 > 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno =3D3D 28 > 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exceptio= n > Handler. > 17:46:44 Fatal error in ADM VP at mt.c:13055 > 17:46:44 Unexpected virtual processor termination, pid =3D3D 377262, = exit=3D > =3D3D > 0x100 > > 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.p= l =3D > 5 6 > "Internal Subsystem failure: 'MT'" "Unexpected virtual processor > termination, pid =3D3D 377262, exit =3D3D 0x100 > " ' > 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0 > 17:46:44 PANIC: Attempting to bring system down > > I reported it to IBM TS, but they think this is NOT IDS problem. > > Just wonder, any folks can reproduce the same problem. > > Thank you, > Frank > > *********************************************************************= **=3D > ******** > > Forum Note: Use "Reply" to post a response in the discussion forum. > > =3D > > > > ***********************************************************************= ******** > Forum Note: Use "Reply" to post a response in the discussion forum. > > ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
Frank,
What you are describing in the PMR is a fairly generic test - and is
somthing which we test freqently.
Because of that, I suspect that there are some unique things on your sy=
stem
which is tending to lead into the problem.
Could you include all of the steps leading to the failure. onconfig?, =
was
this a new instance (oninit -iyv???) spooling Smart Blob? etc..
M.P.
=
"FRANK" =
<yunyaoqu@gmail.c =
om> =
To
Sent by: ids@iiug.org =
ids-bounces@iiug. =
cc
org =
Subj=
ect
Re: ER crash [8288] =
01/23/2007 12:45 =
PM =
=
=
Please respond to =
ids@iiug.org =
=
=
Madison,
I have this from IBM,
PMR 50003,487,000
Thanks,
Frank
On 1/23/07, Madison Pruet <mpruet@us.ibm.com> wrote:
>
>
> Silly me. IDS10.00.UC5
>
> Do you have a repro? If so, then let's treat it as a bug?
>
> Thanks.
>
> M.P.
>
> =3D
>
> "FRANK" =3D
>
> <yunyaoqu@gmail.c =3D
>
> om> =3D
> To
>
> Sent by: ids@iiug.org =3D
>
> ids-bounces@iiug. =3D
> cc
>
> org =3D
>
> Subj=3D
> ect
>
> ER crash [8281] =3D
>
> 01/23/2007 11:52 =3D
>
> AM =3D
>
> =3D
>
> =3D
>
> Please respond to =3D
>
> ids@iiug.org =3D
>
> =3D
>
> =3D
>
> HI, Folks,
>
> AIX5.3, IDS UC5.
>
> the scenario is,
>
> (1) define a pair of update-anywhere ER replication server between tw=
o
> separate boxes,
> (2) define a replicate, some thing like ,
>
> cdr define replicate --conflict=3D3Dalways --ats --ris abc_test_repl =
\\\\
> "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\
> "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test"
>
> (3) create some rows in abc_test at on server(site), but NOT in anoth=
er=3D
>
> server.
>
> (4) delete or update the rows from the server where the rows exist,
>
> then, PANIC.
>
> 17:45:09 Maximum server connections 91
> 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df
> 17:46:07 Assert Failed: No Exception Handler
> 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5
> 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00)
>
> Thread(418964, CDRGeval2, b0302574, 7)
>
> File: mtex.c Line: 472
> 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem
> 17:46:07 Action: Please notify IBM Informix Technical Support.
> 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df
> 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0
> 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno =3D3D 28
> 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exceptio=
n
> Handler.
> 17:46:44 Fatal error in ADM VP at mt.c:13055
> 17:46:44 Unexpected virtual processor termination, pid =3D3D 377262, =
exit=3D
> =3D3D
> 0x100
>
> 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.p=
l =3D
> 5 6
> "Internal Subsystem failure: 'MT'" "Unexpected virtual processor
> termination, pid =3D3D 377262, exit =3D3D 0x100
> " '
> 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0
> 17:46:44 PANIC: Attempting to bring system down
>
> I reported it to IBM TS, but they think this is NOT IDS problem.
>
> Just wonder, any folks can reproduce the same problem.
>
> Thank you,
> Frank
>
> *********************************************************************=
**=3D
> ********
>
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
> =3D
>
>
>
>
***********************************************************************=
********
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
Error 28 means, ENOSPC - No space left on device I suspect you are running out of /tmp space, that's why these messages in server log 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno = 28 18:04:19 CDR CDRGeval3: error writing to spool file /tmp/RepErr/ris.osdint1_grp.0.Geval3.070123_18:04:18.1 18:04:19 CDR CDRGeval3: error writing to spool file /tmp/RepErr/ris.osdint1_grp.0.Geval3.070123_18:04:18.1 Regards, Nilesh. "FRANK" <yunyaoqu@gmail.com> Sent by: ids-bounces@iiug.org 01/23/2007 11:57 AM Please respond to ids@iiug.org To ids@iiug.org cc Subject Re: ER crash [8282] After crash, if you restart the ER, it showed, 18:04:17 CDR NIF listening on asf://osdint1_grp 18:04:17 DDR Log Snooping - Snooping started in log 1588 18:04:18 CDR: Re-connected to server, id 3, name <ncdcprd_grp> 18:04:19 CDR CDRGeval3: error writing to spool file /tmp/RepErr/ris.osdint1_grp.0.Geval3.070123_18:04:18.1 18:04:19 CDR CDRGeval3: error writing to spool file /tmp/RepErr/ris.osdint1_grp.0.Geval3.070123_18:04:18.1 18:04:19 CDR CDRGeval3: error writing to spool file /tmp/RepErr/ris.osdint1_grp.0.Geval3.070123_18:04:18.1 18:04:19 Assert Warning: Memory free block header corruption detected in mt_shm_malloc_segid 5 18:04:19 IBM Informix Dynamic Server Version 10.00.UC5 18:04:19 Who: Session(34, informix@erika, 0, b0312118) Thread(68, CDRGeval3, b02eec8c, 4) File: mtshpool.c Line: 3422 18:04:19 Results: Pool repaired 18:04:19 Action: Please notify IBM Informix Technical Support. 18:04:19 stack trace for pid 376908 written to /tmp/af.42c4e23 18:04:19 See Also: /tmp/af.42c4e23 18:04:20 Memory free block header corruption detected in mt_shm_malloc_segid 5 18:04:32 Checkpoint Completed: duration was 0 seconds. 18:04:32 Checkpoint loguniq 1588, logpos 0x16b4018, timestamp: 0x7d519d7b Frank On 1/23/07, FRANK <yunyaoqu@gmail.com> wrote: > > > HI, Folks, > > AIX5.3, IDS UC5. > > the scenario is, > > (1) define a pair of update-anywhere ER replication server between two > separate boxes, > (2) define a replicate, some thing like , > > cdr define replicate --conflict=always --ats --ris abc_test_repl \\\\ > "noaa@osdint1_grp:informix.abc_test" "select * from abc_test" \\\\ > "noaa@ncdcprd_grp:informix.abc_test" "select * from abc_test" > > (3) create some rows in abc_test at on server(site), but NOT in another > server. > > (4) delete or update the rows from the server where the rows exist, > > then, PANIC. > > 17:45:09 Maximum server connections 91 > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 Assert Failed: No Exception Handler > 17:46:07 IBM Informix Dynamic Server Version 10.00.UC5 > 17:46:07 Who: Session(418266, informix@erika, 0, b0319d00) > > Thread(418964, CDRGeval2, b0302574, 7) > > File: mtex.c Line: 472 > 17:46:07 Results: Exception Caught. Type: MT_EX_OS, Context: mem > 17:46:07 Action: Please notify IBM Informix Technical Support. > 17:46:07 stack trace for pid 377262 written to /tmp/af.687c49df > 17:46:07 See Also: /tmp/af.687c49df, shmem.687c49df.0 > 17:46:43 Error writing '/tmp/shmem.687c49df.0' errno = 28 > 17:46:43 mtex.c, line 472, thread 418964, proc id 377262, No Exception > Handler. > 17:46:44 Fatal error in ADM VP at mt.c:13055 > 17:46:44 Unexpected virtual processor termination, pid = 377262, exit = > 0x100 > > 17:46:44 invoke_alarm(): /bin/sh -c '/dbbackup/logbackup/EventAlarm.pl 5 6 > "Internal Subsystem failure: 'MT'" "Unexpected virtual processor > termination, pid = 377262, exit = 0x100 > " ' > 17:46:44 invoke_alarm(): mt_exec failed, status 65280, errno 0 > 17:46:44 PANIC: Attempting to bring system down > > I reported it to IBM TS, but they think this is NOT IDS problem. > > Just wonder, any folks can reproduce the same problem. > > Thank you, > Frank > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Madison,
I replayed the similar steps between another pair of servers. Basically,
similar results.
All instances were created a couple of months ago, simple data type, no new
in onconfig,...
We have been using ER for years with no issues. Sounds new conflict rule
"always" has a bug.
Thanks,
Frank
=== define server and replicates,
cdr define server --init osdint2_grp
cdr modify server -R /tmp/RepErr -A /tmp/RepErr osdint2_grp
cdr define server --sync=osdint2_grp --init osddev2_grp
cdr define replicate --conflict=always --ats --ris abc_test_repl \\\\
"noaa@osdint2_grp:informix.abc_test" "select * from abc_test" \\\\
"noaa@osddev2_grp:informix.abc_test" "select * from abc_test"
cdr start repl abc_test_repl
cdr list repl abc_test_repl
(1) table
{ TABLE "informix".abc_test row size = 30 number of columns = 3 index size =
19 }
create table "informix".abc_test
(
id serial not null ,
name char(10) not null ,
dt datetime year to second
default current year to second,
primary key (id,name) constraint "informix".abc_test_pk
) with crcols extent size 32 next size 32 lock mode row;
revoke all on "informix".abc_test from "public";
(2) original data in each table on two sites,
----------------------- noaa@osdint2 ----------- Press CTRL-W for Help
--------
id name dt
1401 dfd 2006-10-12 16:49:26
1411 dfd 2006-10-12 16:49:27
1431 dfd 2006-10-12 16:50:07
1492 fgf 2006-10-12 18:41:42
1502 fgf 2006-10-12 18:41:43
1512 fgf 2006-10-12 18:41:43
1441 23 2006-10-12 18:42:25
1451 23 2006-10-12 18:42:26
1461 23 2006-10-12 18:42:26
----------------------- noaa@osddev2 ----------- Press CTRL-W for Help
--------
id name dt
1 start time 2006-07-10 16:20:47
2 start resu 2006-07-17 16:07:12
3 end resume 2006-07-17 16:54:06
(A) do SQL from osdint2
delete from acb_test where 1=1;
Message Log File: /dbbackup/online_log/online_int2.log
21:40:24 Checkpoint loguniq 619, logpos 0x85b6130, timestamp: 0x50685aa5
21:40:24 Maximum server connections 7
21:41:47 ER checkpoint started
21:41:48 Checkpoint Completed: duration was 0 seconds.
21:41:48 Checkpoint loguniq 619, logpos 0x85b7468, timestamp: 0x50685ad7
21:41:48 Maximum server connections 7
21:41:48 ER checkpoint logid:619 logpos:0x85b6130
21:41:48 DDR Log Snooping - Snooping started in log 619
21:45:21 Assert Warning: Memory free block header corruption detected in
mt_shm_malloc_segid
5
21:45:21 IBM Informix Dynamic Server Version 10.00.UC5
21:45:21 Who: Session(5048, informix@dora, 0, b0313f90)
Thread(5139, CDRGeval6, b02f3334, 8)
File: mtshpool.c Line: 3422
21:45:21 Results: Pool repaired
21:45:21 Action: Please notify IBM Informix Technical Support.
21:45:21 stack trace for pid 311344 written to /tmp/af.17fb81f0
21:45:21 See Also: /tmp/af.17fb81f0
21:45:22 Memory free block header corruption detected in
mt_shm_malloc_segid 5
(B)do SQL on two sites
inserts ( begin work with or without replication) on two sites ,
...
(the two tables is still not synchronized.)
delete from abc_test where 1=1 on osddev2.
IT COMES!
(B1) message on my dbacces screen,
SQL:SENDER IS NULL NO MAIL WILL BE SENTtput Choose Save Info Drop Exit
Run the current SQL statements. *** WARNING: IBM Informix Dynamic
Server is no longer r
unning. ***
----------------------- noaa@osddev2 ----------- Press CTRL-W for Help
--------
*** WARNING: IBM Informix Dynamic Server is no longer running.
***
id name dt
SENDER IS NULL
NO MAIL WILL BE SENT
(B2) online log message,
21:51:52 Maximum server connections 3
21:56:52 Fuzzy Checkpoint Completed: duration was 0 seconds, 4 buffers not
flushed.
21:56:52 Checkpoint loguniq 23, logpos 0x2cb6080, timestamp: 0xb3030e3
21:56:52 Maximum server connections 3
21:58:31 stack trace for pid 192732 written to /tmp/af.4448507
21:58:31 Assert Failed: No Exception Handler
21:58:31 IBM Informix Dynamic Server Version 10.00.UC5
21:58:31 Who: Session(56, informix@dolly, 0, 50243b40)
Thread(92, CDRGeval2, 50221ad4, 1)
File: mtex.c Line: 472
21:58:31 Results: Exception Caught. Type: MT_EX_OS, Context: mem
21:58:31 Action: Please notify IBM Informix Technical Support.
21:58:31 stack trace for pid 192732 written to /tmp/af.4448507
21:58:31 See Also: /tmp/af.4448507, shmem.4448507.0
21:58:35 mtex.c, line 472, thread 92, proc id 192732, No Exception Handler.
21:58:40 The Master Daemon Died
21:58:40 PANIC: Attempting to bring system down
informix@dolly $
==============================
informix@dora $ more /tmp/af.17fb81f0
21:45:21 Found during mt_shm_malloc_segid 5
21:45:21 Pool 'CDRGeval6' (0xb1bdb020)
21:45:21 Bad free block 0xb234502c
blk-64
blk-64
b2344fec: 00000000 00000000 00000000 00000000 ........ ........
b2344ffc *
blk+64
blk+64
b234502c: 00000000 00000000 00000000 00000000 ........ ........
b234503c *
21:45:21 Bad free block removed from pool
21:45:21 Attempting to clean up all block list...
21:45:21 Pool 'CDRGeval6' (0xb1bdb020)
21:45:21 Bad block header 0xb2345018
blk-64
blk-64
b2344fd8: 00000000 00000000 00000000 00000000 ........ ........
b2344fe8 *
21:45:21 Bad block removed from pool
21:45:21
21:45:21 IBM Informix Dynamic Server Version 10.00.UC5 Software Serial
Number AAA#B000000
21:45:21 Assert Warning: Memory free block header corruption detected in
mt_shm_malloc_segid
5
21:45:21 Who: Session(5048, informix@dora, 0, b0313f90)
Thread(5139, CDRGeval6, b02f3334, 8)
File: mtshpool.c Line: 3422
21:45:21 Results: Pool repaired
21:45:21 Action: Please notify IBM Informix Technical Support.
21:45:21 Stack for thread: 5139 CDRGeval6
base: 0xb1be7000
len: 69632
pc: 0x1006d880
tos: 0xb1bf6a68
state: running
vp: 8
0x1006da14 (oninit)afstack (0x20006770, 0x4c030, 0xb1bf6ba4, 0xf0b2, 0x0,
0xa47, 0x0, 0x0)
0x1006f334 (oninit)afhandler(0x1, 0xb1bf7128, 0x20004f94, 0xb0748ef0, 0x401,
0x1, 0x20004ed0,
0xd
5e)
0x10067b30 (oninit)recover_pool_bad_free(0xb1bf74d3, 0x200a1358, 0x200a0fb0,
0x8, 0xb1bf7998,
0xb
1bf728c, 0x25, 0x800000)
0x1006af6c (oninit)mt_shm_malloc_segid(0x0, 0x0, 0x0, 0x0, 0x0, 0x0,
0xb1bf74fa, 0x7ffffff3)
0x1006bb78 (oninit)mt_shm_malloc(0xb1bf74ee, 0x200a1090, 0x200a1074,
0x200a1444, 0xffffffff, 0
x25
00, 0x2f200074, 0x8000)
0x107f3bb4 (oninit)dsiSpoolRow(0x5f55532e, 0x38313900, 0x0, 0x0, 0x0, 0x0,
0xb1bf7958, 0x0)
0x107f501c (oninit)dsiSpool(0x55532e38, 0x31390000, 0x0, 0x0, 0x0, 0x0, 0x0,
0x0)
0x10875544 (oninit)dsiSpoolLocalDeleteRow(0x202f88a8, 0x202f88c4,
0xb1bf7c48, 0x202f8764, 0x20
20e
7cc, 0x1, 0x202f88c8, 0xb000e8a0)
0x10875d0c (oninit)dsProcessLocalDelete(0x0, 0x2020e7a8, 0xb1bf7c08, 0x2,
0x1073a974, 0xb1bf90
68,
0x3e8, 0xb1bdb918)
0x107a4930 (oninit)grouperNotifyDS(0x10000, 0xb1b8f6b8,
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g