Buffer Partitioning
Posted in 1999
A 7.31.UC4 shop on Solaris found 'onstat -P' showing 98% of its 120,000 buffers held by btree pages, with almost no data pages cached, high buffer waits and very slow batch jobs. Art Kagel attributed it to bug 115327, where index pages get mis-flagged as both leaf and node, so all index pages get MED-HIGH buffer priority and starve out data pages; the cure is to rebuild all indexes created under 7.2x or touched by deletes, on a patched version. Erik van Veen noted a related new defect (PTS 121515) about wide indexes filling the pool with branch nodes that the 115327 fix doesn't address, and Doug Agnew mentioned the NOLRUPRIO environment variable (from 7.30.UC7XB) to disable buffer priorities. Art also argued for ONCONFIG knobs to cap high-priority buffer percentages. No confirmation from the original poster is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Error Codes & Troubleshooting, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues
Hi, all
Our production OLTP server is 7.31.UC4 on Solaris 2.6 (Sun E-4500, 6 cpus, 4
gig memory). Our batch processing went at a snail's pace all weekend, and it
occurred to me only this morning to check the output of "onstat -P". Here is
what I found:
Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 3 days
20:34:16 -- 1311296 Kbytes
partnum total btree data other resident dirty
.
.
.
Totals: 120000 118099 1376 525 0 1836
Percentages:
Data 1.15
Btree 98.42
Other 0.44
PATROL is monitoring, among other things, our Kagel Bufwait Percentage, and
as you might expect, it has been going nuts, with the KBP well over 20%.
One of our main reasons for upgrading from 7.30.UC8 was poor performance,
and an Informix engineer told me (after the fact) that it was this very
condition that 7.31 was supposed to fix.
We did have the following Assert Failure Saturday morning:
01:01:08 Assert Failed: Page Check Error in phposition:isposition:bad page
01:01:08 Informix Dynamic Server Version 7.31.UC4
01:01:08 Who: Session(24627, fgrunner@great, 696, 651419288)
Thread(82079, srvinfx, 2b2cc9d8, 6)
File: rsdebug.c Line: 1036
01:01:08 Results: Possible inconsistencies in
'gr_prod:"fgrunner".stilocar'
01:01:08 Action: Run 'oncheck -cDI gr_prod:"fgrunner".stilocar'
01:01:28 See Also: /logs/af.44875e33
Naturally, we are checking this table, but could this have anything to do
with the high percentage of buffers being occupied by index pages? Would
running update statistics have anything to do with buffer partitioning
between Data and Index?
I have opened a case with Informix, but all I have gotten so far are the
usual articles about tuning. During our performance problems with 7.30, I
sent our $ONCONFIG file to Informix (along with a few onstats) for review,
but the only answer I got was "nothing totally out of whack". I therefore
submit to this august body the same data:
#**************************************************************************
#
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig.gr_prod
# Description: INFORMIX-OnLine Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace name
ROOTPATH /idev/rootdbs # Path for device containing root dbspace
ROOTOFFSET 64 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 500000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS phylog # Location (dbspace) of physical log
PHYSFILE 60000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 200 # Number of logical log files
LOGSIZE 9984 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /export/home/informix/prod.log # System message log file
path
CONSOLE /dev/console # System console message path
ALARMPROGRAM /export/home/informix/etc/log_full.sh # Alarm program path
# System Archive Tape Device
TAPEDEV /dev/rmt/3hn
TAPEBLK 2048 # Tape block size (Kbytes)
TAPESIZE 27262976 # Max amount of data to put on log tape
(Kbytes)
## Log Archive Tape Device
LTAPEDEV /dev/rmt/4hn
LTAPEBLK 2048 # Log tape block size (Kbytes)
LTAPESIZE 27262976 # Max amount of data to put on log tape
(Kbytes)
# Optical
STAGEBLOB # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 5 # Unique id corresponding to a OnLineinstance
DBSERVERNAME shm_prod # Name of default database server
DBSERVERALIASES tli_prod # List of alternate dbservernames
NETTYPE tlitcp,4,150,NET # Configure poll thread(s) for nettype
NETTYPE ipcshm,1,150,CPU # Configure poll thread(s) for nettype
DEADLOCK_TIMEOUT 300 # Max time to wait of lock in distributed
env.
RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 5 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps toone
NOAGE 1 # Process aging#AFF_SPROC 1 # Affinity start processor
#AFF_NPROCS 3 # Affinity number of processors
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 900000 # Maximum number of locks
BUFFERS 120000 # Maximum number of shared buffers
NUMAIOVPS 6 # Number of IO vps
PHYSBUFF 448 # Physical log buffer size (Kbytes)
LOGBUFF 6 # Logical log buffer size (Kbytes)LOGSMAX 300 # Maximum number of logical log files
CLEANERS 30 # Number of buffer cleaner processes
SHMBASE 0xa000000 # Shared memory base address
SHMVIRTSIZE 1000000 # initial virtual shared memory segment size
SHMADD 8192 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 1800000 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 240 # Check point interval (in sec)
LRUS 15 # Number of LRU queues
LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water markpercentage
LTXEHWM 60 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline workerthreads
ON_RECVRY_THREADS 1 # Default number of online worker threads
# Data Replication Variables
# DRAUTO: 0 manual, 1 retain type, 2 reverse type
DRAUTO 0 # DR automatic switchover
DRINTERVAL 30
uunet wrote:
>
> Hi, all
>
> Our production OLTP server is 7.31.UC4 on Solaris 2.6 (Sun E-4500, 6 cpus, 4
> gig memory). Our batch processing went at a snail's pace all weekend, and it
> occurred to me only this morning to check the output of "onstat -P". Here is
> what I found:
>
> Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 3 days
> 20:34:16 -- 1311296 Kbytes
> partnum total btree data other resident dirty
> .
> .
> .
> Totals: 120000 118099 1376 525 0 1836
>
> Percentages:
> Data 1.15
> Btree 98.42 <<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<
> Other 0.44
Stop right there. You arebeing eaten by bug 115327, though I thought that
7.31UC4 has the fix for that one. Hmmmm, have you rebuilt your indexes
since upgrading? Since the bug causes the index pages to get flagged as
BOTH leaf and node it confuses the buffer priority code and all index
pages get classed MED-HIGH and quickly starve out the data pages from the
buffer cache as data pages are only MEDIUM (except for RESIDENT tables).
The fix, besides getting a version which will continue the problem, is
to rebuild ALL indexes from which any rows may have been deleted and any
indexes created by 7.2x (since 7.2x did not maintain the flags properly
at all). You should not have to rebuild indexes built by 7.3x on tables
which are not deleted from since it is the BTREE cleaner threads which
mess up the flags, in the buggy versions, upon compressing sparse index
pages emtied by deletes.
Art S. Kagel
In article <38306E62.E77DC469@bloomberg.net>, Art S. Kagel
<kagel@bloomberg.net> writes
>
>
>uunet wrote:
>>
>> Hi, all
>>
>> Our production OLTP server is 7.31.UC4 on Solaris 2.6 (Sun E-4500, 6 cpus, 4
>> gig memory). Our batch processing went at a snail's pace all weekend, and it
>> occurred to me only this morning to check the output of "onstat -P". Here is
>> what I found:
>>
IDS 7.31.UC4-1 Workgroup Ed definitely has the patch, I have it!!!
I have installed it on three machine in the last week!!
Surely it is not XXX-1 ??
>> Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 3 days
>> 20:34:16 -- 1311296 Kbytes
>> partnum total btree data other resident dirty
>> .
>> .
>> .
>> Totals: 120000 118099 1376 525 0 1836
>>
>> Percentages:
>> Data 1.15
>> Btree 98.42 <<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<
>> Other 0.44
>
>Stop right there. You arebeing eaten by bug 115327, though I thought that
>7.31UC4 has the fix for that one. Hmmmm, have you rebuilt your indexes
>since upgrading? Since the bug causes the index pages to get flagged as
>BOTH leaf and node it confuses the buffer priority code and all index
>pages get classed MED-HIGH and quickly starve out the data pages from the
>buffer cache as data pages are only MEDIUM (except for RESIDENT tables).
>The fix, besides getting a version which will continue the problem, is
>to rebuild ALL indexes from which any rows may have been deleted and any
>indexes created by 7.2x (since 7.2x did not maintain the flags properly
>at all). You should not have to rebuild indexes built by 7.3x on tables
>which are not deleted from since it is the BTREE cleaner threads which
>mess up the flags, in the buggy versions, upon compressing sparse index
>pages emtied by deletes.
>
>Art S. Kagel
--
David Williams
Hi,
you may also want to consider PTS 121515 "CUSTOMERS WITH LOTS OF WIDE INDEXES
CAN SEE SEVERE PERFORMANCE PROBLEMS BECAUSE OVER 90% OF THE BUFFER POOL IS
FILLED WITH INDEX BRANCH NODES" - a variation on PTS 115327 mentioned by Art.
It was entered only a few days ago. The fix to PTS 115327 does not resolve this.
Cheers, Erik
uunet wrote:
> Hi, all
>
> Our production OLTP server is 7.31.UC4 on Solaris 2.6 (Sun E-4500, 6 cpus, 4
> gig memory). Our batch processing went at a snail's pace all weekend, and it
> occurred to me only this morning to check the output of "onstat -P". Here is
> what I found:
>
> Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 3 days
> 20:34:16 -- 1311296 Kbytes
> partnum total btree data other resident dirty
> .
> .
> .
> Totals: 120000 118099 1376 525 0 1836
>
> Percentages:
> Data 1.15
> Btree 98.42
> Other 0.44
>
> PATROL is monitoring, among other things, our Kagel Bufwait Percentage, and
> as you might expect, it has been going nuts, with the KBP well over 20%.
>
> One of our main reasons for upgrading from 7.30.UC8 was poor performance,
> and an Informix engineer told me (after the fact) that it was this very
> condition that 7.31 was supposed to fix.
>
> We did have the following Assert Failure Saturday morning:
>
> 01:01:08 Assert Failed: Page Check Error in phposition:isposition:bad page
> 01:01:08 Informix Dynamic Server Version 7.31.UC4
> 01:01:08 Who: Session(24627, fgrunner@great, 696, 651419288)
> Thread(82079, srvinfx, 2b2cc9d8, 6)
> File: rsdebug.c Line: 1036
> 01:01:08 Results: Possible inconsistencies in
> 'gr_prod:"fgrunner".stilocar'
> 01:01:08 Action: Run 'oncheck -cDI gr_prod:"fgrunner".stilocar'
> 01:01:28 See Also: /logs/af.44875e33>
> Naturally, we are checking this table, but could this have anything to do
> with the high percentage of buffers being occupied by index pages? Would
> running update statistics have anything to do with buffer partitioning
> between Data and Index?
>
> I have opened a case with Informix, but all I have gotten so far are the
> usual articles about tuning. During our performance problems with 7.30, I
> sent our $ONCONFIG file to Informix (along with a few onstats) for review,
> but the only answer I got was "nothing totally out of whack". I therefore
> submit to this august body the same data:
>
> #**************************************************************************
> #
> # INFORMIX SOFTWARE, INC.
> #
> # Title: onconfig.gr_prod
> # Description: INFORMIX-OnLine Configuration Parameters
> #
> #**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name
> ROOTPATH /idev/rootdbs # Path for device containing root dbspace
> ROOTOFFSET 64 # Offset of root dbspace into device
> (Kbytes)
> ROOTSIZE 500000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH # Path for device containing mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS phylog # Location (dbspace) of physical log
> PHYSFILE 60000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 200 # Number of logical log files
> LOGSIZE 9984 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /export/home/informix/prod.log # System message log file
> path
> CONSOLE /dev/console # System console message path
> ALARMPROGRAM /export/home/informix/etc/log_full.sh # Alarm program path>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt/3hn
> TAPEBLK 2048 # Tape block size (Kbytes)
> TAPESIZE 27262976 # Max amount of data to put on log tape
> (Kbytes)>
> ## Log Archive Tape Device
> LTAPEDEV /dev/rmt/4hn
> LTAPEBLK 2048 # Log tape block size (Kbytes)
> LTAPESIZE 27262976 # Max amount of data to put on log tape
> (Kbytes)>
> # Optical
>
> STAGEBLOB # INFORMIX-OnLine/Optical staging area
>
> # System Configuration
>
> SERVERNUM 5 # Unique id corresponding to a OnLine> instance
> DBSERVERNAME shm_prod # Name of default database server
> DBSERVERALIASES tli_prod # List of alternate dbservernames
> NETTYPE tlitcp,4,150,NET # Configure poll thread(s) for nettype
> NETTYPE ipcshm,1,150,CPU # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 300 # Max time to wait of lock in distributed
> env.
> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 5 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to> one
>
> NOAGE 1 # Process aging> #AFF_SPROC 1 # Affinity start processor
> #AFF_NPROCS 3 # Affinity number of processors
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 900000 # Maximum number of locks
> BUFFERS 120000 # Maximum number of shared buffers
> NUMAIOVPS 6 # Number of IO vps
> PHYSBUFF 448 # Physical log buffer size (Kbytes)
> LOGBUFF 6 # Logical log buffer size (Kbytes)> LOGSMAX 300 # Maximum number of logical log files
> CLEANERS 30 # Number of buffer cleaner processes
> SHMBASE 0xa000000 # Shared memory base address
> SHMVIRTSIZE 1000000 # initial virtual shared memory segment size
> SHMADD 8192 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 1800000 # Total shared memory (Kbytes). 0=>unlimited
> CKPTINTVL 240 # Check point interval (in sec)
> LRUS 15 # Number of LRU queues
> LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
> LTXHWM 50 # Long transaction high water mark> percentage
> LTXEHWM 60 # Long transaction high water mark
> (exclusive)
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 32 # Stack size (Kbytes)>
> # System Page Size
> # BUFFSIZE - OnLine no longer supports this configuration parameter.
> # To determine the page size used by OnLine on yo
With 7.30.UC7XB, Informix provided an environment variable (NOLRUPRIO) that,
when set, disabled buffer priorities -- everything (except resident tables)
went LOW.
Erik van Veen wrote in message <38325DAE.46797EDD@informix.com>...
>Hi,
>
>you may also want to consider PTS 121515 "CUSTOMERS WITH LOTS OF WIDE
INDEXES
>CAN SEE SEVERE PERFORMANCE PROBLEMS BECAUSE OVER 90% OF THE BUFFER POOL IS
>FILLED WITH INDEX BRANCH NODES" - a variation on PTS 115327 mentioned by
Art.
>It was entered only a few days ago. The fix to PTS 115327 does not resolve
this.
>
>Cheers, Erik
>
>uunet wrote:
>
>> Hi, all
>>
>> Our production OLTP server is 7.31.UC4 on Solaris 2.6 (Sun E-4500, 6
cpus, 4
>> gig memory). Our batch processing went at a snail's pace all weekend, and
it
>> occurred to me only this morning to check the output of "onstat -P". Here
is
>> what I found:
>>
>> Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 3 days
>> 20:34:16 -- 1311296 Kbytes
>> partnum total btree data other resident dirty
>> .
>> .
>> .
>> Totals: 120000 118099 1376 525 0 1836
>>
>> Percentages:
>> Data 1.15
>> Btree 98.42
>> Other 0.44
>>
>> PATROL is monitoring, among other things, our Kagel Bufwait Percentage,
and
>> as you might expect, it has been going nuts, with the KBP well over 20%.
>>
>> One of our main reasons for upgrading from 7.30.UC8 was poor performance,
>> and an Informix engineer told me (after the fact) that it was this very
>> condition that 7.31 was supposed to fix.
>>
>> We did have the following Assert Failure Saturday morning:
>>
>> 01:01:08 Assert Failed: Page Check Error in phposition:isposition:badpage
>> 01:01:08 Informix Dynamic Server Version 7.31.UC4
>> 01:01:08 Who: Session(24627, fgrunner@great, 696, 651419288)
>> Thread(82079, srvinfx, 2b2cc9d8, 6)
>> File: rsdebug.c Line: 1036
>> 01:01:08 Results: Possible inconsistencies in
>> 'gr_prod:"fgrunner".stilocar'
>> 01:01:08 Action: Run 'oncheck -cDI gr_prod:"fgrunner".stilocar'
>> 01:01:28 See Also: /logs/af.44875e33>>
>> Naturally, we are checking this table, but could this have anything to do
>> with the high percentage of buffers being occupied by index pages? Would
>> running update statistics have anything to do with buffer partitioning
>> between Data and Index?
>>
>> I have opened a case with Informix, but all I have gotten so far are the
>> usual articles about tuning. During our performance problems with 7.30, I
>> sent our $ONCONFIG file to Informix (along with a few onstats) for
review,
>> but the only answer I got was "nothing totally out of whack". I therefore
>> submit to this august body the same data:
>>
>>
#**************************************************************************
>> #
>> # INFORMIX SOFTWARE, INC.
>> #
>> # Title: onconfig.gr_prod
>> # Description: INFORMIX-OnLine Configuration Parameters
>> #
>>
#**************************************************************************
>>
>> # Root Dbspace Configuration
>>
>> ROOTNAME rootdbs # Root dbspace name
>> ROOTPATH /idev/rootdbs # Path for device containing root dbspace
>> ROOTOFFSET 64 # Offset of root dbspace into device
>> (Kbytes)
>> ROOTSIZE 500000 # Size of root dbspace (Kbytes)>>
>> # Disk Mirroring Configuration Parameters
>>
>> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
>> MIRRORPATH # Path for device containing mirroredroot
>> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>>
>> # Physical Log Configuration
>>
>> PHYSDBS phylog # Location (dbspace) of physical log
>> PHYSFILE 60000 # Physical log file size (Kbytes)>>
>> # Logical Log Configuration
>>
>> LOGFILES 200 # Number of logical log files
>> LOGSIZE 9984 # Logical log size (Kbytes)>>
>> # Diagnostics
>>
>> MSGPATH /export/home/informix/prod.log # System message log file
>> path
>> CONSOLE /dev/console # System console message path
>> ALARMPROGRAM /export/home/informix/etc/log_full.sh # Alarm programpath
>>
>> # System Archive Tape Device
>>
>> TAPEDEV /dev/rmt/3hn
>> TAPEBLK 2048 # Tape block size (Kbytes)
>> TAPESIZE 27262976 # Max amount of data to put on log tape
>> (Kbytes)>>
>> ## Log Archive Tape Device
>> LTAPEDEV /dev/rmt/4hn
>> LTAPEBLK 2048 # Log tape block size (Kbytes)
>> LTAPESIZE 27262976 # Max amount of data to put on log tape
>> (Kbytes)>>
>> # Optical
>>
>> STAGEBLOB # INFORMIX-OnLine/Optical staging area
>>
>> # System Configuration
>>
>> SERVERNUM 5 # Unique id corresponding to a OnLine>> instance
>> DBSERVERNAME shm_prod # Name of default database server
>> DBSERVERALIASES tli_prod # List of alternate dbservernames
>> NETTYPE tlitcp,4,150,NET # Configure poll thread(s) for nettype
>> NETTYPE ipcshm,1,150,CPU # Configure poll thread(s) for nettype
>> DEADLOCK_TIMEOUT 300 # Max time to wait of lock indistributed
>> env.
>> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>>
>> MULTIPROCESSOR 1 # 0 for single-processor, 1 for>> multi-processor
>> NUMCPUVPS 5 # Number of user (cpu) vps
>> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to>> one
>>
>> NOAGE 1 # Process aging>> #AFF_SPROC 1 # Affinity start processor
>> #AFF_NPROCS 3 # Affinity number of processors
>> AFF_SPROC 0 # Affinity start processor
>> AFF_NPROCS 0 # Affinity number of processors>>
>> # Shared Memory Parameters
>>
>> LOCKS 900000 # Maximum number of locks
>> BUFFERS 120000 # Maximum number of shared buffers
>> NUMAIOVPS 6 # Number of IO vps
>> PHYSBUFF 448 # Physical log buffer size (Kbytes)
>> LOGBUFF 6 # Logical log buffer size (Kbytes)>> LOGSMAX 300 # Maximum number of logical log files
>> CLEANERS 30 # Number of buffer cleaner processes
>> SHMBASE 0xa000000 # Shared memory base address
>> SHMVIRTSIZE 1000000 # initial virtual shared memory segmentsize
>> SHMADD 8192 # Size of new shared memory segments
>> (Kbytes)
>> SHMTOTAL 1800000 # Total shared memory (Kbytes).
0=>unlimited
>> CKPTINTVL 240 # Check point interval (in sec)
>> LRUS 15 # Number of LRU queues
>> LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
>> LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limi
Hi Erik,
You know what would solve this 'wide index' problem and ALL of the
similar problems which are going to crop up in the near future? My
original suggestion, to you in fact, that there be ONCONFIG parameters to
allow us DBAs to limit the % of buffers that can be gobbled by the
MED-HIGH and HIGH priority queue classes. This would put the control over
how much of the buffer queue can be snarfed up by resident table and index
node pages into the hands of the DBAs where it belongs and would keep you
guys in Menlo from having to develop more and more sophisticated
algorithms to control and age the buffer pages. If you remember I told
you when 7.31 was just a twinkle in your eye that buffer priority aging
was not going to solve this problem I foresaw which had not even happened
yet. Give us:
MAX_HIGH_PRIORITY 15
MAX_MEDHIGH_PRIORITY 20
and never visit this problem again!
Now if I could also get:
MAX_LRU_DIRTY_COUNT 10000
MIN_LRU_DIRTY_COUNT 1000
to control larger buffer pools with better granularity, I'd be a happy DBA.
Art S. Kagel
Erik van Veen wrote:
>
> Hi,
>
> you may also want to consider PTS 121515 "CUSTOMERS WITH LOTS OF WIDE INDEXES
> CAN SEE SEVERE PERFORMANCE PROBLEMS BECAUSE OVER 90% OF THE BUFFER POOL IS
> FILLED WITH INDEX BRANCH NODES" - a variation on PTS 115327 mentioned by Art.
> It was entered only a few days ago. The fix to PTS 115327 does not resolve this.
>
> Cheers, Erik
>
> uunet wrote:
>
> > Hi, all
> >
> > Our production OLTP server is 7.31.UC4 on Solaris 2.6 (Sun E-4500, 6 cpus, 4
> > gig memory). Our batch processing went at a snail's pace all weekend, and it
> > occurred to me only this morning to check the output of "onstat -P". Here is
> > what I found:
> >
> > Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 3 days
> > 20:34:16 -- 1311296 Kbytes
> > partnum total btree data other resident dirty
> > .
> > .
> > .
> > Totals: 120000 118099 1376 525 0 1836
> >
> > Percentages:
> > Data 1.15
> > Btree 98.42
> > Other 0.44
> >
> > PATROL is monitoring, among other things, our Kagel Bufwait Percentage, and
> > as you might expect, it has been going nuts, with the KBP well over 20%.
> >
> > One of our main reasons for upgrading from 7.30.UC8 was poor performance,
> > and an Informix engineer told me (after the fact) that it was this very
> > condition that 7.31 was supposed to fix.
> >
> > We did have the following Assert Failure Saturday morning:
> >
> > 01:01:08 Assert Failed: Page Check Error in phposition:isposition:bad page
> > 01:01:08 Informix Dynamic Server Version 7.31.UC4
> > 01:01:08 Who: Session(24627, fgrunner@great, 696, 651419288)
> > Thread(82079, srvinfx, 2b2cc9d8, 6)
> > File: rsdebug.c Line: 1036
> > 01:01:08 Results: Possible inconsistencies in
> > 'gr_prod:"fgrunner".stilocar'
> > 01:01:08 Action: Run 'oncheck -cDI gr_prod:"fgrunner".stilocar'
> > 01:01:28 See Also: /logs/af.44875e33> >
> > Naturally, we are checking this table, but could this have anything to do
> > with the high percentage of buffers being occupied by index pages? Would
> > running update statistics have anything to do with buffer partitioning
> > between Data and Index?
> >
> > I have opened a case with Informix, but all I have gotten so far are the
> > usual articles about tuning. During our performance problems with 7.30, I
> > sent our $ONCONFIG file to Informix (along with a few onstats) for review,
> > but the only answer I got was "nothing totally out of whack". I therefore
> > submit to this august body the same data:
> >
> > #**************************************************************************
> > #
> > # INFORMIX SOFTWARE, INC.
> > #
> > # Title: onconfig.gr_prod
> > # Description: INFORMIX-OnLine Configuration Parameters
> > #
> > #**************************************************************************
> >
> > # Root Dbspace Configuration
> >
> > ROOTNAME rootdbs # Root dbspace name
> > ROOTPATH /idev/rootdbs # Path for device containing root dbspace
> > ROOTOFFSET 64 # Offset of root dbspace into device
> > (Kbytes)
> > ROOTSIZE 500000 # Size of root dbspace (Kbytes)> >
> > # Disk Mirroring Configuration Parameters
> >
> > MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> > MIRRORPATH # Path for device containing mirrored root
> > MIRROROFFSET 0 # Offset into mirrored device (Kbytes)> >
> > # Physical Log Configuration
> >
> > PHYSDBS phylog # Location (dbspace) of physical log
> > PHYSFILE 60000 # Physical log file size (Kbytes)> >
> > # Logical Log Configuration
> >
> > LOGFILES 200 # Number of logical log files
> > LOGSIZE 9984 # Logical log size (Kbytes)> >
> > # Diagnostics
> >
> > MSGPATH /export/home/informix/prod.log # System message log file
> > path
> > CONSOLE /dev/console # System console message path
> > ALARMPROGRAM /export/home/informix/etc/log_full.sh # Alarm program path> >
> > # System Archive Tape Device
> >
> > TAPEDEV /dev/rmt/3hn
> > TAPEBLK 2048 # Tape block size (Kbytes)
> > TAPESIZE 27262976 # Max amount of data to put on log tape
> > (Kbytes)> >
> > ## Log Archive Tape Device
> > LTAPEDEV /dev/rmt/4hn
> > LTAPEBLK 2048 # Log tape block size (Kbytes)
> > LTAPESIZE 27262976 # Max amount of data to put on log tape
> > (Kbytes)> >
> > # Optical
> >
> > STAGEBLOB # INFORMIX-OnLine/Optical staging area
> >
> > # System Configuration
> >
> > SERVERNUM 5 # Unique id corresponding to a OnLine> > instance
> > DBSERVERNAME shm_prod # Name of default database server
> > DBSERVERALIASES tli_prod # List of alternate dbservernames
> > NETTYPE tlitcp,4,150,NET # Configure poll thread(s) for nettype
> > NETTYPE ipcshm,1,150,CPU # Configure poll thread(s) for nettype
> > DEADLOCK_TIMEOUT 300 # Max time to wait of lock in distributed
> > env.
> > RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
> >
> > MULTIPROCESSOR 1 # 0 for single-processor, 1 for> > multi-processor
> > NUMCPUVPS 5 # Number of user (cpu) vps
> > SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to> > one
> >
> > NOAGE 1 # Process aging> > #AFF_SPROC 1 # Affinity start processor
> > #AFF_NPROCS 3 # Affinity number of processors
> > AFF_SPROC 0 # Affinity start processor
> > AFF_NPROCS 0 # Affinity number of processors> >
> > # Shared Memory Parameters
> >
> > LOCKS 900000 # Maximum number of locks
> > BUFFERS 1200
In article <3834552B.2BB0A08F@bloomberg.net>, Art S. Kagel
<kagel@bloomberg.net> writes
>Hi Erik,
> You know what would solve this 'wide index' problem and ALL of the
>similar problems which are going to crop up in the near future? My
>original suggestion, to you in fact, that there be ONCONFIG parameters to
>allow us DBAs to limit the % of buffers that can be gobbled by the
>MED-HIGH and HIGH priority queue classes. This would put the control over
>how much of the buffer queue can be snarfed up by resident table and index
>node pages into the hands of the DBAs where it belongs and would keep you
>guys in Menlo from having to develop more and more sophisticated
>algorithms to control and age the buffer pages. If you remember I told
True. Also true that improving the intelligence will help. Surely
index nodes pages higher up in the tree should be higher priority. Each
'top' node in the index should have highest priority since it is always
needed to access the other nodes. Surely being able to keep the top node
in each index used frequently + the 1st level of nodes,probably <100Kb
per index (??) would help. i.e. lock these down and make them resident.
Being able to avoid all the processing used to age them would help..
Any thoughts?? New onconfig parameter lock index heads? Perhaps
even have them out of the buffer cache + in a separate pool loaded
when online boots + expanded/shrunk when indexes are dropped/created.
These means they are guaranteed to be in memory??
C code currently is
if need top node in index
lock + examine LRU queues (complex code)
rather than
if need top node in index
if onconfig parameter set
find top node in hashed table of indexes
follow pointer index_head_node->.. for data!!
else
complex LRU code
Similar for 1st level in index which are a few pages that are
typically access a lot. In this day of 100's of Mb's of memory I like
the idea of locking down 20Nb or so to speed things up! After all
index lookups are down a lot on OLTP systems.
>you when 7.31 was just a twinkle in your eye that buffer priority aging
>was not going to solve this problem I foresaw which had not even happened
>yet. Give us:
>
>MAX_HIGH_PRIORITY 15
>MAX_MEDHIGH_PRIORITY 20
>
>and never visit this problem again!
>
>Now if I could also get:
>
>MAX_LRU_DIRTY_COUNT 10000
>MIN_LRU_DIRTY_COUNT 1000
>
>to control larger buffer pools with better granularity, I'd be a happy DBA.
>
>Art S. Kagel
>
>Erik van Veen wrote:
>>
>> Hi,
>>
>> you may also want to consider PTS 121515 "CUSTOMERS WITH LOTS OF WIDE INDEXES
>> CAN SEE SEVERE PERFORMANCE PROBLEMS BECAUSE OVER 90% OF THE BUFFER POOL IS
>> FILLED WITH INDEX BRANCH NODES" - a variation on PTS 115327 mentioned by Art.
>> It was entered only a few days ago. The fix to PTS 115327 does not resolve
>this.
>>
>> Cheers, Erik
>>
>> uunet wrote:
>>
>> > Hi, all
>> >
>> > Our production OLTP server is 7.31.UC4 on Solaris 2.6 (Sun E-4500, 6 cpus, 4
>> > gig memory). Our batch processing went at a snail's pace all weekend, and it
>> > occurred to me only this morning to check the output of "onstat -P". Here is
>> > what I found:
>> >
>> > Informix Dynamic Server Version 7.31.UC4 -- On-Line -- Up 3 days
>> > 20:34:16 -- 1311296 Kbytes
>> > partnum total btree data other resident dirty
>> > .
>> > .
>> > .
>> > Totals: 120000 118099 1376 525 0 1836
>> >
>> > Percentages:
>> > Data 1.15
>> > Btree 98.42
>> > Other 0.44
>> >
>> > PATROL is monitoring, among other things, our Kagel Bufwait Percentage, and
>> > as you might expect, it has been going nuts, with the KBP well over 20%.
>> >
>> > One of our main reasons for upgrading from 7.30.UC8 was poor performance,
>> > and an Informix engineer told me (after the fact) that it was this very
>> > condition that 7.31 was supposed to fix.
>> >
>> > We did have the following Assert Failure Saturday morning:
>> >
>> > 01:01:08 Assert Failed: Page Check Error in phposition:isposition:bad page
>> > 01:01:08 Informix Dynamic Server Version 7.31.UC4
>> > 01:01:08 Who: Session(24627, fgrunner@great, 696, 651419288)
>> > Thread(82079, srvinfx, 2b2cc9d8, 6)
>> > File: rsdebug.c Line: 1036
>> > 01:01:08 Results: Possible inconsistencies in
>> > 'gr_prod:"fgrunner".stilocar'
>> > 01:01:08 Action: Run 'oncheck -cDI gr_prod:"fgrunner".stilocar'
>> > 01:01:28 See Also: /logs/af.44875e33>> >
>> > Naturally, we are checking this table, but could this have anything to do
>> > with the high percentage of buffers being occupied by index pages? Would
>> > running update statistics have anything to do with buffer partitioning
>> > between Data and Index?
>> >
>> > I have opened a case with Informix, but all I have gotten so far are the
>> > usual articles about tuning. During our performance problems with 7.30, I
>> > sent our $ONCONFIG file to Informix (along with a few onstats) for review,
>> > but the only answer I got was "nothing totally out of whack". I therefore
>> > submit to this august body the same data:
>> >
>> > #**************************************************************************
>> > #
>> > # INFORMIX SOFTWARE, INC.
>> > #
>> > # Title: onconfig.gr_prod
>> > # Description: INFORMIX-OnLine Configuration Parameters
>> > #
>> > #**************************************************************************
>> >
>> > # Root Dbspace Configuration
>> >
>> > ROOTNAME rootdbs # Root dbspace name
>> > ROOTPATH /idev/rootdbs # Path for device containing root dbspace
>> > ROOTOFFSET 64 # Offset of root dbspace into device
>> > (Kbytes)
>> > ROOTSIZE 500000 # Size of root dbspace (Kbytes)>> >
>> > # Disk Mirroring Configuration Parameters
>> >
>> > MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
>> > MIRRORPATH # Path for device containing mirrored root
>> > MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>> >
>> > # Physical Log Configuration
>> >
>> > PHYSDBS phylog # Location (dbspace) of physical log
>> > PHYSFILE 60000 # Physical log file size (Kbytes)>> >
>> > # Logical Log Configuration
>> >
>> > LOGFILES 200 # Number of logical log files
>> > LOGSIZE 9984 # Logical log size (Kbytes)>> >
>> > # Diagnostics
>> >
>> > MSGPATH /export/home/informix/prod.log # System message log file
>> > path
>> > CONSOLE /dev/console # System console message path
>> > ALARMPROGRAM /export/home/informix/etc/log_full.sh # Alarm program path>> >
>> > # System Archive Tape Device
>> >
>> > TAPEDEV /dev/rmt/3hn
>> > TAPEBLK 2048 # Tape block size (Kbytes)
>> > TAPESIZE 27262976 # Max amount of data to put on log tape
>> > (Kbytes)>> >
>> > ## Log Archive Tape D