Re: IDS 7.31.UC5 on AIX 4.3. crashes
Posted in 2000
Topics: Storage & Space Management, Error Codes & Troubleshooting, Connectivity: ESQL/C, 4GL & Embedded SQL, Server Administration, Security, Permissions & Auditing, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
I wouldn't swear to it, but I think the allocation of additional shared
memory segments by the engine is playing a role here. Try upping your
initial shared memory segment size (SHMVIRTSIZE) so Informix will
almost NEVER have to add more segments. Do this by adding up all of
the additional segments allocated from engine startup through a typical
cycle that includes some of your largest queries and/or busiest times
in terms of user sessions. It looks like you currently are adding new
segments 8MB at a time, so multiply the number of segments added by 8MB
and add this (converted to KB) to your current SHMVIRTSIZE. Consider
also increasing SHMADD so when the engine does need more memory, it
grabs bigger chunks (reducing the total number of additional segments
needed).
HTH,
Irwin Goldstein
Objective Software Systems, Inc.
http://www.objectsoft.com
In article <20000509155914.22714.00004372@ng-cg1.aol.com>,
risudi@aol.com (Risudi) wrote:
> System HAS RDS(4gl) running with it.
>
> System appears to crash periodically (see logs at the bottom of this
message)
> According to the investigating tech the following symptoms are
occurring.
> The main log message is:
> 14:35:08 The Master Daemon Died
> 14:35:08 PANIC: Attempting to bring system down>
> The system will restart without errors.
>
> AIX will occasionally capture some error logs(see below) that
basically who
> fglgo crashing.
>
> The 4gl Error log will show a Network Receive Error, however we are
not sure
> this error is due to IDS shutting down, the 4gl crashing as a result
>
> Informix Support Suggested the following:
> 1. Someone, or some process kills the oninit for the engine. (I
consider
> this to be highly unlikely in our scenario)
> 2. Someone, or some process removes the shared memory segment.
(Could it
> be possible that memory allocated by the processing of the 4gl is in
some
> sort of conflict?)
>
> Or is it possible that IDS attempts to get more Shared Memory and
crashes if it
> cannot. (note occassionally in the Online log the following is
listed, but it
> appears to work successfully)
> 16:12:17 dynamically allocated new shared memory segment (size
8388608)>
> The tech on site, mentioned that this seems to occur during process
that do
> large sorts/counts. Maybe a temp/sort space issue?
>
> any suggestions you have would be appreciated
>
> Rick
> risudi@aol.com
>
> AIX LOG
>
> ----------------------------------------------------------------------
----
> LABEL: CORE_DUMP
> IDENTIFIER: C60BB505
>
> Date/Time: Tue Apr 25 16:04:53
> Sequence Number: 882611
> Machine Id: 000201404C00
> Node Id: Rapidparts
> Class: S
> Type: PERM
> Resource Name: SYSPROC
>
> Description
> SOFTWARE PROGRAM ABNORMALLY TERMINATED
>
> Probable Causes
> SOFTWARE PROGRAM
>
> User Causes
> USER GENERATED SIGNAL
>
> Recommended Actions
> CORRECT THEN RETRY
>
> Failure Causes
> SOFTWARE PROGRAM
>
> Recommended Actions
> RERUN THE APPLICATION PROGRAM
> IF PROBLEM PERSISTS THEN DO THE FOLLOWING
> CONTACT APPROPRIATE SERVICE REPRESENTATIVE
>
> Detail Data
> SIGNAL NUMBER
> 11
> USER'S PROCESS ID:
> 40770
> FILE SYSTEM SERIAL NUMBER
> 16
> INODE NUMBER
> 90112
> PROGRAM NAME
> fglgo
> ADDITIONAL INFORMATION
> _sqremove DC
> _sqremove 24
> _iqprepar 250
> _iqnprep 124
> doprepare 3C
> prepare DC
> runop DD4
> runprog 64
> main 278
> ??
>
> Symptom Data
> REPORTABLE
> 1
> INTERNAL ERROR
> 0
> SYMPTOM CODE
> PCSS/SPI2 FLDS/fglgo SIG/11 FLDS/_sqremove VALU/dc FLDS/_iqprepar
> ----------------------------------------------------------------------
-----
> LABEL: CORE_DUMP
> IDENTIFIER: C60BB505
>
> Date/Time: Tue Apr 25 14:40:46
> Sequence Number: 882610
> Machine Id: 000201404C00
> Node Id: Rapidparts
> Class: S
> Type: PERM
> Resource Name: SYSPROC
>
> Description
> SOFTWARE PROGRAM ABNORMALLY TERMINATED
>
> Probable Causes
> SOFTWARE PROGRAM
>
> User Causes
> USER GENERATED SIGNAL
>
> Recommended Actions
> CORRECT THEN RETRY
>
> Failure Causes
> SOFTWARE PROGRAM
>
> Recommended Actions
> RERUN THE APPLICATION PROGRAM
> IF PROBLEM PERSISTS THEN DO THE FOLLOWING
> CONTACT APPROPRIATE SERVICE REPRESENTATIVE
>
> Detail Data
> SIGNAL NUMBER
> 11
> USER'S PROCESS ID:
> 37870
> FILE SYSTEM SERIAL NUMBER
> 16
> INODE NUMBER
> 169985
> PROGRAM NAME
> fglgo
> ADDITIONAL INFORMATION
> _sqfrcmem 74
> _sqreleas 70
> _sqreleas 70
>
> Cut-snip from the Informix Log
> 14:27:48 Checkpoint Completed: duration was 0 seconds.
> 14:32:48 Checkpoint Completed: duration was 0 seconds.
> 14:35:08 The Master Daemon Died
> 14:35:08 PANIC: Attempting to bring system down>
> Tue Apr 25 14:35:50 2000
>
> 14:35:50 Event alarms enabled. ALARMPROG
= '/infxprog/etc/log_full.sh'
> 14:35:54 DR: DRAUTO is 0 (Off)
> 14:35:54 AIX MP latch code enabled
> 14:35:54 Requested shared memory segment size rounded from 588KB to592KB
> 14:35:54 Informix Dynamic Server Version 7.31.UC5 Software SerialNumber
> AAC#J850321
> 14:35:55
> 14:35:55
> 14:35:56 (5) connection rejected - no calls allowed for sqlexec
> 14:35:56 listener-thread: err = -27002: oserr = 0: errstr = : Noconnections
> are allowed in Dynamic Server quiescent mode.
>
> 14:35:56 Informix Dynamic Server Initialized -- Shared MemoryInitialized.
> 14:35:56 Physical Recovery Started.
> 14:35:56 Physical Recovery Complete: 52 Pages Restored.
> 14:35:56 Logical Recovery Started.
> 14:35:59 Logical Recovery Complete.> 0 Committed, 0 Rolled Back, 0 Open, 0 Bad Locks
>
> 14:36:00 Onconfig parameter CONSOLE modified
from /infxprog/urs/console.log to
> /infxprog/usr/console.log.
> 14:36:00 Dataskip is now OFF for all dbspaces
> 14:36:00 Quiescent Mode
> 14:36:01 Checkpoint Completed: duration was 0 seconds.
> 14:36:03 (12) connection rejected - no calls allowed for sqlexec
> 14:36:03 listener-thread: err = -27002: oserr = 0: errstr = : Noconnections
> are allowed in Dynamic Server quiescent mode.
>
> 14:36:10 On-Line Mode
> 14:41:00 Checkpoint Completed: duration was 0 seconds.
> 14:46:00 Checkpoint Completed: duration was 0 seconds.
> 14:51:01 Checkpoint Completed: duration was 0 seconds.>
>
Sent via Deja.com http://www.deja.com/
Before you buy.
Good observation......
David Williams wrote:
> In article <8fa1ot$56q$1@nnrp1.deja.com>, irwin_goldstein@my-deja.com
> writes
> >I wouldn't swear to it, but I think the allocation of additional shared
> >memory segments by the engine is playing a role here. Try upping your
>
> AIX, big hint...
>
> AIX only allows 10 shared memory segments per process.
>
> DO onstat -g seg and count how many segments you have...
>
> >initial shared memory segment size (SHMVIRTSIZE) so Informix will
> >almost NEVER have to add more segments. Do this by adding up all of
> >the additional segments allocated from engine startup through a typical
> >cycle that includes some of your largest queries and/or busiest times
> >in terms of user sessions. It looks like you currently are adding new
> >segments 8MB at a time, so multiply the number of segments added by 8MB
> >and add this (converted to KB) to your current SHMVIRTSIZE. Consider
> >also increasing SHMADD so when the engine does need more memory, it
> >grabs bigger chunks (reducing the total number of additional segments
> >needed).
> >
>
> True, increase SHMVIRTSIZE.
>
> >HTH,
> >
> >Irwin Goldstein
> >Objective Software Systems, Inc.
> >http://www.objectsoft.com
> >
> >In article <20000509155914.22714.00004372@ng-cg1.aol.com>,
> > risudi@aol.com (Risudi) wrote:
> >> System HAS RDS(4gl) running with it.
> >>
> >> System appears to crash periodically (see logs at the bottom of this
> >message)
> >> According to the investigating tech the following symptoms are
> >occurring.
> >> The main log message is:
> >> 14:35:08 The Master Daemon Died
> >> 14:35:08 PANIC: Attempting to bring system down> >>
>
> I don't see why oninit does not give a stack trace and an af file
> when it panics. Seems very strange. Also oninit is not the name of
> the process which core dumps.
>
> ...
> >> Detail Data
> >> SIGNAL NUMBER
> >> 11
> >> USER'S PROCESS ID:
> >> 40770
> >> FILE SYSTEM SERIAL NUMBER
> >> 16
> >> INODE NUMBER
> >> 90112
> >> PROGRAM NAME
> >> fglgo
>
> Seems to be the fglgo which is core dumping. Sure you are not running
> out of swap space?
>
> --
> David Williams
--
Madison Pruet
===========================================
Enterprise Replication Product Developement
Dallas, Texas
Informix Software
===========================================
In article <8fa1ot$56q$1@nnrp1.deja.com>, irwin_goldstein@my-deja.com
writes
>I wouldn't swear to it, but I think the allocation of additional shared
>memory segments by the engine is playing a role here. Try upping your
AIX, big hint...
AIX only allows 10 shared memory segments per process.
DO onstat -g seg and count how many segments you have...
>initial shared memory segment size (SHMVIRTSIZE) so Informix will
>almost NEVER have to add more segments. Do this by adding up all of
>the additional segments allocated from engine startup through a typical
>cycle that includes some of your largest queries and/or busiest times
>in terms of user sessions. It looks like you currently are adding new
>segments 8MB at a time, so multiply the number of segments added by 8MB
>and add this (converted to KB) to your current SHMVIRTSIZE. Consider
>also increasing SHMADD so when the engine does need more memory, it
>grabs bigger chunks (reducing the total number of additional segments
>needed).
>
True, increase SHMVIRTSIZE.
>HTH,
>
>Irwin Goldstein
>Objective Software Systems, Inc.
>http://www.objectsoft.com
>
>In article <20000509155914.22714.00004372@ng-cg1.aol.com>,
> risudi@aol.com (Risudi) wrote:
>> System HAS RDS(4gl) running with it.
>>
>> System appears to crash periodically (see logs at the bottom of this
>message)
>> According to the investigating tech the following symptoms are
>occurring.
>> The main log message is:
>> 14:35:08 The Master Daemon Died
>> 14:35:08 PANIC: Attempting to bring system down>>
I don't see why oninit does not give a stack trace and an af file
when it panics. Seems very strange. Also oninit is not the name of
the process which core dumps.
...
>> Detail Data
>> SIGNAL NUMBER
>> 11
>> USER'S PROCESS ID:
>> 40770
>> FILE SYSTEM SERIAL NUMBER
>> 16
>> INODE NUMBER
>> 90112
>> PROGRAM NAME
>> fglgo
Seems to be the fglgo which is core dumping. Sure you are not running
out of swap space?
--
David Williams