Re: Informix halts for no obvious reason ?!
Posted in 1997
In article <5uk5r0$9s8@cssun.mathcs.emory.edu>, Milen Metodiev
<dbadmin@svilosa.bis.bg> writes
>
>Hi, folks.
>
>This is the umpteenth time this happens, so I wonder
>if anyone can help. Even a few hints will be more than
>welcome.
>
>When the system is busy enough ( 20-30 users of Informix )
>and people start issuing "heavy" queries, Informix falls into
>a strange deadlock. It looks like one process is in a critical
>section, and won't leave it. Informix waits for this process
>to get out of its critical section, so that it can complete
>the pending checkpoint. In the end, all users end up waiting
>for this checkpoint, which never occurs.
>
>It happens to random users, using random programs, i.e. there
>is no any consistency about this behaviour.
>
>Here is some information about our version of Informix, etc.
>
>======================================================
>Database: Informix OnLine 5.0.2.UC6
>OS : SCO OpenServer Enterprise System Release 3.0
>E-mail : dbadmin@svilosa.bis.bg
>======================================================
>
>The following is a listing of a file containing a snapshot of
>the system at the time of the failure:
>
>
>RSAM Version 5.01.UD2 -- On-Line (CKPT REQ) -- Up 01:14:09 -- 20288 Kbytes
>
>Message Log File: /usr/runtime/online.log
>ted
>08:10:22 Checkpoint Completed
>08:10:27 Checkpoint Completed
>08:14:11 Checkpoint Completed
>08:14:13 Checkpoint Completed
>08:14:16 Checkpoint Completed
>08:14:22 Checkpoint Completed
>08:14:26 Checkpoint Completed
>08:14:30 Checkpoint Completed
>08:14:34 Checkpoint Completed
>08:15:43 Checkpoint Completed
>08:17:30 Checkpoint Completed
>08:18:18 Checkpoint Completed
>08:18:35 Checkpoint Completed
>08:19:53 Checkpoint Completed
>08:25:04 Checkpoint Completed
>08:28:10 Checkpoint Completed>
>Configuration File: /usr/runtime/etc/tbconfig
>#**************************************************************************
>#
># INFORMIX SOFTWARE, INC.
>#
># Title: tbconfig.std
># Sccsid: @(#)tbconfig.std 9.1.1.1 1/11/92 19:15:16
># Description: INFORMIX-OnLine Configuration Parameters
>#
>#**************************************************************************
>
># Root Dbspace Configuration
>
>ROOTNAME rootdbs # Root dbspace name>ROOTPATH /usr/runtime/chunks/Chunk1
> # Path for device containing root dbspace
>ROOTOFFSET 50 # Offset of root dbspace into device (Kbytes)
>ROOTSIZE 683000 # Size of root dbspace (Kbytes)>
># Disk Mirroring Configuration
>
>MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
>MIRRORPATH # Path for device containing root dbspace>mirror
>MIRROROFFSET 0 # Offset into mirror device (Kbytes)
># Physical Log Configuration
>
>PHYSDBS rootdbs # Name of dbspace that contains physical log
>PHYSFILE 1000 # Physical log file size (Kbytes)>
Increase as others have suggested...
># Logical Log Configuration
>
>LOGFILES 50 # Number of logical log files
>LOGSIZE 1000 # Size of each logical log file (Kbytes)>
># Message Files
>
>MSGPATH /usr/runtime/online.log # OnLine message log pathname
>CONSOLE /dev/console # System console message pathname>
># Archive Tape Device
>
>TAPEDEV /dev/urStp1 # Archive tape device pathname
>TAPEBLK 16 # Archive tape block size (Kbytes)
>TAPESIZE 2000000 # Max amount of data to put on tape (Kbytes)>
># Logical Log Backup Tape Device
>
>LTAPEDEV /dev/rStp1 # Logical log tape device pathname
>LTAPEBLK 16 # Logical log tape block size (Kbytes)
>LTAPESIZE 2000000 # Max amount of data to put on log tape
>(Kbytes)>
># Identification Parameters
>
>SERVERNUM 1 # Unique id associated with this OnLine>instance
>DBSERVERNAME runtime # Unique name of this OnLine instance>
># Shared Memory Parameters
>
>RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)>USERS 40 # Maximum number of concurrent user processes
>TRANSACTIONS 150 # Maximum number of concurrent transactions
>LOCKS 50000 # Maximum number of locks
>BUFFERS 8000 # Maximum number of shared memory buffers Have you considered increasing BUFFERS? This seems low for an active
OLTP system. You will need to look at the OS level to see how much
pjysical memory is free but try increasing this to 16000, I normally
run 20000 but then when asked I plan for 40Mb of pjysical memory
allocated to BUFFERS.
>TBLSPACES 400 # Maximum number of active tblspaces
>CHUNKS 8 # Maximum number of chunks
>DBSPACES 8 # Maximum number of dbspaces and blobspaces
>PHYSBUFF 32 # Size of physical log buffers (Kbytes)
>LOGBUFF 32 # Size of logical log buffers (Kbytes)>LOGSMAX 50 # Maximum number of logical log files
>CLEANERS 1 # Number of page-cleaner processes Increase as Art S. Kagel suggested.
>SHMBASE 0x0 # Shared memory base address
>CKPTINTVL 300 # Checkpoint interval (in seconds)
>LRUS 8 # Number of LRU queues Again Increase as Art S. Kagel suggested.
>LRU_MAX_DIRTY 60 # LRU modified begin-cleaning limit (percent)
>LRU_MIN_DIRTY 50 # LRU modified end-cleaning limit (percent) Again Decrease as Art S. Kagel suggested.
>LTXHWM 65 # Long TX high-water mark (percent)
>LTXEHWM 75 # Long TX exclusive high-water mark (percent)>
># Machine- and Product-Specific Parameters
>
>DYNSHMSZ 0 # Dynamic shared memory size (Kbytes)
>GTRID_CMP_SZ 32 # Number of bytes to use in GTRID comparision
>DEADLOCK_TIMEOUT 60 # Max time to wait for lock in distributed
>env.
>TXTIMEOUT 300 # Transaction timeout for I-STAR (in seconds)>SPINCNT 0 # No. of times process tries for latch
>STAGEBLOB # Reserved for INFORMIX-OnLine/Optical
>
># System Page Size
>
>BUFFSIZE 2048 # Page size (do not change!)
>
>
>Users
>address flags pid user tty wait tout locks nreads nwrites
>804019f0 C-----D 385 root console 0 0 0 364 22
>80401a60 ------D 0 root console 0 0 0 0 0
>80401ad0 ------F 389 root 0 0 0 0 0
>80401b40 S------ 725 bos ttyp1 8040154c 0 0 172 1
>80401bb0 --A---M 730 informix tty05 0 0 0 0 0
>80401c20 S------ 1439 nadja ttyp0 8040154c 0 2 64 5
>80401c90 S------ 1350 s