Write cache hits slipping !
Posted in 1999
Topics: Installation, Setup & Upgrades, Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues
Our environment is a new HP 9000 (R390) server running HP-UX 11.0. We are
on Dynamic Server, Workgroups Edition, version 7.30.UC9. Our legacy
application is running on it - it is written in Informix 4GL, and we're on
version 7.20 of the 4GL/RDS.
We connect to shared memory for our production instance. We have three
physical disks, mirrored by hardware. All IDS chunks are raw disk logical
volumes. The first disk has the rootdbs on it (it includes the physical
log). The second disk is set aside for another application and database.
The third disk has a number of dbspaces on it. There are 3 three databases
in the production instance, and there is a dbspace for each database. Usage
peaks at about 80 users, and generally usage is between 6:30 am and 10pm,
Monday to Friday.
We have ported the system from Standard Engine on an old server, and we went
live on Monday,
Our problem (and not a huge one).
On the whole, our new Dynamic Server system has been running great since it
went live Monday the 18th October. There's been no downtime since we've
been live, and no one has been denied access. The onstat commands I have
run appear to show we have lots and lots of all of the major resources.
Our read cache hits have settled in at just under 99%, which is fantastic.
However, our write cache hits started at just over 57% when we started up
IDS on the 18th. They climbed to a peak of just under 92% on Wednesday the
20th, but have slipped away. They're sitting just over 84% now and still
falling.
The level of slippage would indicate that it's probably not just the
application at issue here.
I enclose our current onconfig file and the output of onstat -p and
onstat -l.
However, my questions are:
What resources should I be looking at first.
Which onstat commands should I be running and what should I be looking out
for.
Thanks in advance.
*** HERE COMES THE ONCONFIG FILE ***
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig.clr
# Description: Informix Dynamic Server Configuration Parameters
#
# Alnis Bajars 22/9/99. Modify template for Colorcorp installation.
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace name
ROOTPATH /dev/chunk_01 # Path for device containing root dbspace
ROOTOFFSET 0 # Offset of root dbspace into device
(Kbytes)
ROOTSIZE 990000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS rootdbs # Location (dbspace) of physical log
PHYSFILE 20000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 10 # Number of logical log files
LOGSIZE 5000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /opt/informix720/online.log # System message log file path
CONSOLE /dev/console # System console message path
ALARMPROGRAM /opt/informix720/etc/log_full.sh # Alarm program pathSYSALARMPROGRAM /opt/informix720/etc/evidence.sh # System Alarm program path
TBLSPACE_STATS 1
# System Archive Tape Device
TAPEDEV /dev/rmt/0m # Tape device path
TAPEBLK 32 # Tape block size (Kbytes)
TAPESIZE 24000000 # Maximum amount of data to put on tape
(Kbytes)
# Log Archive Tape Device
LTAPEDEV /var/dblogs/saturn.log/log_prod # Log tape device path
LTAPEBLK 32 # Log tape block size (Kbytes)
LTAPESIZE 500000 # Max amount of data to put on log tape
(Kbytes)
# Optical
STAGEBLOB # Informix Dynamic Server/Optical staging
area
# System Configuration
SERVERNUM 1 # Unique id corresponding to a DynamicServer instance
DBSERVERNAME saturn # Name of default database server
DBSERVERALIASES # List of alternate dbservernames
NETTYPE ipcshm,240,120,CPU # Configure poll thread(s) for nettype
DEADLOCK_TIMEOUT 120 # Max time to wait of lock in distributed
env.
RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 0 # 0 for single-processor, 1 formulti-processor
NUMCPUVPS 3 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps toone
NOAGE 1 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 100000 # Maximum number of locks
BUFFERS 100000 # Maximum number of shared buffers
NUMAIOVPS 30 # Number of IO vps
PHYSBUFF 64 # Physical log buffer size (Kbytes)
LOGBUFF 64 # Logical log buffer size (Kbytes)LOGSMAX 12 # Maximum number of logical log files
CLEANERS 2 # Number of buffer cleaner processes
SHMBASE 0x0 # Shared memory base address
SHMVIRTSIZE 32768 # Initial virtual shared memory segment size
SHMADD 8092 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 600 # Check point interval (in sec)
LRUS 32 # Number of LRU queues
LRU_MAX_DIRTY 60 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 40 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water markpercentage
LTXEHWM 60 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 64 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - Dynamic Server no longer supports this configuration parameter.
# To determine the page size used by Dynamic Server on your
platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline
restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 20 # Default number of offline workerthreads
ON_RECVRY_THREADS 10 # Default number of online worker threads
# Data Replication Variables
# DRAUTO: 0 manual, 1 retain type, 2 reverse type
DRAUTO 0 # DR automatic switchover
DRINTERVAL 30 # DR max time between DR buffer flushes (in
sec)
DRTIMEOUT 30 # DR network timeout (in sec)
DRLOSTFOUND /opt/informi
Ok, here goes my first EVER attempt at advising on someone elses onconfig,
I'd go with any advice Art gives you ;o))
See below...
--
---------------------------------------
Tony Flaherty aef@mfs.misys.co.uk
Analyst Programmer
Misys Financial Systems
All statements and opinions are my own,
Misys don't pay me enough to have opinions
on their behalf
.
Alnis Bajars wrote in message ...
>Our environment is a new HP 9000 (R390) server running HP-UX 11.0. We are
>on Dynamic Server, Workgroups Edition, version 7.30.UC9. Our legacy
>application is running on it - it is written in Informix 4GL, and we're on
>version 7.20 of the 4GL/RDS.
>
>We connect to shared memory for our production instance. We have three
>physical disks, mirrored by hardware. All IDS chunks are raw disk logical
>volumes. The first disk has the rootdbs on it (it includes the physical
>log). The second disk is set aside for another application and database.
>The third disk has a number of dbspaces on it. There are 3 three databases
>in the production instance, and there is a dbspace for each database.
Usage
>peaks at about 80 users, and generally usage is between 6:30 am and 10pm,
>Monday to Friday.
>
>We have ported the system from Standard Engine on an old server, and we
went
>live on Monday,
>
>Our problem (and not a huge one).
>
>On the whole, our new Dynamic Server system has been running great since it
>went live Monday the 18th October. There's been no downtime since we've
>been live, and no one has been denied access. The onstat commands I have
>run appear to show we have lots and lots of all of the major resources.
>
>Our read cache hits have settled in at just under 99%, which is fantastic.
>
>However, our write cache hits started at just over 57% when we started up
>IDS on the 18th. They climbed to a peak of just under 92% on Wednesday the
>20th, but have slipped away. They're sitting just over 84% now and still
>falling.
>
>The level of slippage would indicate that it's probably not just the
>application at issue here.
>
>I enclose our current onconfig file and the output of onstat -p and
>onstat -l.>
>However, my questions are:
>What resources should I be looking at first.
>Which onstat commands should I be running and what should I be looking out
>for.
>
>Thanks in advance.
>
>*** HERE COMES THE ONCONFIG FILE ***
>
># INFORMIX SOFTWARE, INC.
>#
># Title: onconfig.clr
># Description: Informix Dynamic Server Configuration Parameters
>#
># Alnis Bajars 22/9/99. Modify template for Colorcorp installation.
>#
>#**************************************************************************
>
># Root Dbspace Configuration
>
>ROOTNAME rootdbs # Root dbspace name
>ROOTPATH /dev/chunk_01 # Path for device containing root dbspace
>ROOTOFFSET 0 # Offset of root dbspace into device
>(Kbytes)
>ROOTSIZE 990000 # Size of root dbspace (Kbytes)>
># Disk Mirroring Configuration Parameters
>
>MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
>MIRRORPATH # Path for device containing mirrored root
>MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
># Physical Log Configuration
>
>PHYSDBS rootdbs # Location (dbspace) of physical log
>PHYSFILE 20000 # Physical log file size (Kbytes)>
># Logical Log Configuration
>
>LOGFILES 10 # Number of logical log files
>LOGSIZE 5000 # Logical log size (Kbytes)>
># Diagnostics
>
>MSGPATH /opt/informix720/online.log # System message log file path
>CONSOLE /dev/console # System console message path
>ALARMPROGRAM /opt/informix720/etc/log_full.sh # Alarm program path>SYSALARMPROGRAM /opt/informix720/etc/evidence.sh # System Alarm program
path
>TBLSPACE_STATS 1>
># System Archive Tape Device
>
>TAPEDEV /dev/rmt/0m # Tape device path
>TAPEBLK 32 # Tape block size (Kbytes)
>TAPESIZE 24000000 # Maximum amount of data to put on tape
>(Kbytes)>
># Log Archive Tape Device
>
>LTAPEDEV /var/dblogs/saturn.log/log_prod # Log tape device path
>LTAPEBLK 32 # Log tape block size (Kbytes)
>LTAPESIZE 500000 # Max amount of data to put on log tape
>(Kbytes)>
># Optical
>
>STAGEBLOB # Informix Dynamic Server/Optical staging
>area
>
># System Configuration
>
>SERVERNUM 1 # Unique id corresponding to a Dynamic>Server instance
>DBSERVERNAME saturn # Name of default database server
>DBSERVERALIASES # List of alternate dbservernames
>NETTYPE ipcshm,240,120,CPU # Configure poll thread(s) for nettype
>DEADLOCK_TIMEOUT 120 # Max time to wait of lock in distributed
>env.
>RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
>MULTIPROCESSOR 0 # 0 for single-processor, 1 for>multi-processor
>NUMCPUVPS 3 # Number of user (cpu) vps
>SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to>one
>
>NOAGE 1 # Process aging
>AFF_SPROC 0 # Affinity start processor
>AFF_NPROCS 0 # Affinity number of processors>
># Shared Memory Parameters
>
>LOCKS 100000 # Maximum number of locks
>BUFFERS 100000 # Maximum number of shared buffers
>NUMAIOVPS 30 # Number of IO vps
>PHYSBUFF 64 # Physical log buffer size (Kbytes)
>LOGBUFF 64 # Logical log buffer size (Kbytes)>LOGSMAX 12 # Maximum number of logical log files
>CLEANERS 2 # Number of buffer cleaner processes
try setting this to the same as the number of LRU's
>SHMBASE 0x0 # Shared memory base address
>SHMVIRTSIZE 32768 # Initial virtual shared memory segmentsize
>SHMADD 8092 # Size of new shared memory segments
>(Kbytes)
>SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
>CKPTINTVL 600 # Check point interval (in sec)
>LRUS 32 # Number of LRU queues
>LRU_MAX_DIRTY 60 # LRU percent dirty begin cleaning limit
>LRU_MIN_DIRTY 40 # LRU percent dirty end cleaning limit
I think general advice is to run with these at much lower values I'm running
at 10,5
>LTXHWM 50 # Long transaction high water mark>percentage
>LTXEHWM 60 # Long transaction high water mark
>(exclusive)
>TXTIMEOUT 0x12c # Transaction timeout (in sec)
>STACKSIZE 64 # Stack size (Kbytes)>
># System Page Size
># BUFFSIZE - Dynamic Server no longer supports this configuration
parameter.
># To determine the page size used by Dynamic Server on yo
Alnis Bajars <alnisb@colorcorp.com.au> wrote in message
news:zc6R3.343$8F2.4196@nsw.nnrp.telstra.net...
Hi Alnis ! See below ... I'll pay your attention only to the params, which
can potentially increase your write caching.
> Our environment is a new HP 9000 (R390) server running HP-UX 11.0. We are
> on Dynamic Server, Workgroups Edition, version 7.30.UC9. Our legacy
> application is running on it - it is written in Informix 4GL, and we're on
> version 7.20 of the 4GL/RDS.
>
> We connect to shared memory for our production instance. We have three
> physical disks, mirrored by hardware. All IDS chunks are raw disk logical
> volumes. The first disk has the rootdbs on it (it includes the physical
> log). The second disk is set aside for another application and database.
> The third disk has a number of dbspaces on it. There are 3 three
databases
> in the production instance, and there is a dbspace for each database.
Usage
> peaks at about 80 users, and generally usage is between 6:30 am and 10pm,
> Monday to Friday.
>
> We have ported the system from Standard Engine on an old server, and we
went
> live on Monday,
>
> Our problem (and not a huge one).
>
> On the whole, our new Dynamic Server system has been running great since
it
> went live Monday the 18th October. There's been no downtime since we've
> been live, and no one has been denied access. The onstat commands I have
> run appear to show we have lots and lots of all of the major resources.
>
> Our read cache hits have settled in at just under 99%, which is fantastic.
>
> However, our write cache hits started at just over 57% when we started up
> IDS on the 18th. They climbed to a peak of just under 92% on Wednesday
the
> 20th, but have slipped away. They're sitting just over 84% now and still
> falling.
>
> The level of slippage would indicate that it's probably not just the
> application at issue here.
>
> I enclose our current onconfig file and the output of onstat -p and
> onstat -l.>
> However, my questions are:
> What resources should I be looking at first.
> Which onstat commands should I be running and what should I be looking out
> for.
>
> Thanks in advance.
>
> *** HERE COMES THE ONCONFIG FILE ***
>
> # INFORMIX SOFTWARE, INC.
> #
> # Title: onconfig.clr
> # Description: Informix Dynamic Server Configuration Parameters
> #
> # Alnis Bajars 22/9/99. Modify template for Colorcorp installation.
> #
>
#**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name
> ROOTPATH /dev/chunk_01 # Path for device containing root dbspace
> ROOTOFFSET 0 # Offset of root dbspace into device
> (Kbytes)
> ROOTSIZE 990000 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH # Path for device containing mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS rootdbs # Location (dbspace) of physical log
> PHYSFILE 20000 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 10 # Number of logical log files
> LOGSIZE 5000 # Logical log size (Kbytes)>
> # Diagnostics
>
> MSGPATH /opt/informix720/online.log # System message log file path
> CONSOLE /dev/console # System console message path
> ALARMPROGRAM /opt/informix720/etc/log_full.sh # Alarm program path> SYSALARMPROGRAM /opt/informix720/etc/evidence.sh # System Alarm program
path
> TBLSPACE_STATS 1>
> # System Archive Tape Device
>
> TAPEDEV /dev/rmt/0m # Tape device path
> TAPEBLK 32 # Tape block size (Kbytes)
> TAPESIZE 24000000 # Maximum amount of data to put on tape
> (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV /var/dblogs/saturn.log/log_prod # Log tape device path
> LTAPEBLK 32 # Log tape block size (Kbytes)
> LTAPESIZE 500000 # Max amount of data to put on log tape
> (Kbytes)>
> # Optical
>
> STAGEBLOB # Informix Dynamic Server/Optical staging
> area
>
> # System Configuration
>
> SERVERNUM 1 # Unique id corresponding to a Dynamic> Server instance
> DBSERVERNAME saturn # Name of default database server
> DBSERVERALIASES # List of alternate dbservernames
> NETTYPE ipcshm,240,120,CPU # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 120 # Max time to wait of lock in distributed
> env.
> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 0 # 0 for single-processor, 1 for> multi-processor
> NUMCPUVPS 3 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to> one
>
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 100000 # Maximum number of locks
> BUFFERS 100000 # Maximum number of shared buffers
> NUMAIOVPS 30 # Number of IO vps
> PHYSBUFF 64 # Physical log buffer size (Kbytes)
> LOGBUFF 64 # Logical log buffer size (Kbytes)> LOGSMAX 12 # Maximum number of logical log files
> CLEANERS 2 # Number of buffer cleaner processes
> SHMBASE 0x0 # Shared memory base address
> SHMVIRTSIZE 32768 # Initial virtual shared memory segmentsize
> SHMADD 8092 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
> CKPTINTVL 600 # Check point interval (in sec)
Increase this to at minimum 1800. Respectively, increase size of your
physical log. Your checkpoints must be triggered only by this param. This
also has other side (long recovery).
> LRUS 32 # Number of LRU queues
> LRU_MAX_DIRTY 60 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 40 # LRU percent dirty end cleaning limit
For example, this two. If you have many LRU writes (onstat -F), increase
this params. Try to avoid LRU writes at all if you need to increase your
write caching. This may have impact on server processing (long checkpoints).
> *** HERE COMES THE OUTPUT OF onstat -p ***
>
>
> Informix Dynamic Server Version 7.30.UC9 -- On-Line -- Up 10 days
> 01:13:23 -- 316312 Kbytes
>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 1967417 1688298 181506257 98.92 362295 692137
Alnis Bajars wrote:
>
> Our environment is a new HP 9000 (R390) server running HP-UX 11.0. We are
> on Dynamic Server, Workgroups Edition, version 7.30.UC9. Our legacy
> application is running on it - it is written in Informix 4GL, and we're on
> version 7.20 of the 4GL/RDS.
[SNIP]
> Our read cache hits have settled in at just under 99%, which is fantastic.
>
> However, our write cache hits started at just over 57% when we started up
Not to worry that was probably due to data loading in to multiple tables
such results in that situation is normal. Did you zero the stats after
the loading to get a clear production picture?
> IDS on the 18th. They climbed to a peak of just under 92% on Wednesday the
> 20th, but have slipped away. They're sitting just over 84% now and still
> falling.
Informix says anything over 75% is good and I used to like 85% myself but
truth is in a well tuned OLTP system with enough buffers you should be
seeing around 90% write cache in 7.xx. IDS's buffer management is just
that good.
> The level of slippage would indicate that it's probably not just the
> application at issue here.
Except for bulk loads application design can rarely affect write cache
percentages, now read cache is a different story. Poor application and
database design can trash read cache ratios.
> I enclose our current onconfig file and the output of onstat -p and
> onstat -l.
An onstat -d and onstat -g iof (or onstat -D) would also help to see if
heavy activity on a single chunk is trashing the ratio by flushing pages
from other chunks.
> However, my questions are:
> What resources should I be looking at first.
BUFFERS is the first and most important resource affecting the cache
ratios.
> Which onstat commands should I be running and what should I be looking out
> for.
> Thanks in advance.
>
> *** HERE COMES THE ONCONFIG FILE ***
>
> # INFORMIX SOFTWARE, INC.
> #
> # Title: onconfig.clr
> # Description: Informix Dynamic Server Configuration Parameters
> #
> # Alnis Bajars 22/9/99. Modify template for Colorcorp installation.
> #
> #**************************************************************************
>
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name
> ROOTPATH /dev/chunk_01 # Path for device containing root dbspace
I hope that /dev/chunk_01 is a link and not the device file's actual name.
Not using links for chunks is a can of worms that will cause you to lose
sleep eventually, and not from worry either. Also I suggest you place the
links somewhere OTHER than /dev so an overeager sysadmin does not remove
them (experience speaking).
> TBLSPACE_STATS 1
I know that this is on by default for 7.30, but, collecting tablespace
stats is VERY expensive. You should turn it off if you are not currently
monitoring the results.
[SNIP]
> SERVERNUM 1 # Unique id corresponding to a Dynamic> Server instance
> DBSERVERNAME saturn # Name of default database server
> DBSERVERALIASES # List of alternate dbservernames
> NETTYPE ipcshm,240,120,CPU # Configure poll thread(s) for nettype
You really should have a network connection DBSERVERALIAS, even if all of
your users are local using the shared memory connection. There are
certain circumstances that will come up where you will thank the foresight
that added the option. If nothing else if you have a program with memory
problems it cannot crash the engine on a network connection, it can over
shared memory, so you can shift it to that connection on the fly while you
recode the app to fix the problem!
> DEADLOCK_TIMEOUT 120 # Max time to wait of lock in distributed
> env.
> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 0 # 0 for single-processor, 1 for> multi-processor
Do you REALLY have only one physical CPU? If not you MUST turn
MULTIPROCESSOR on to enable MSP protection code to avoid CPU VP deadlocks.
> NUMCPUVPS 3 # Number of user (cpu) vps
If you have only one CPU then an HP can usefully handle two CPU VPs per
CPU. I think that the third one is probably wasted.
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to> one
>
> NOAGE 1 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 100000 # Maximum number of locks
> BUFFERS 100000 # Maximum number of shared buffers
Look at the onstat -p stats and analyse. Your app runs 14 hours a day
over 10 days it read in 1.7MM pages to buffers which turns over all of the
100,000 buffers in 8 hours. Discount this with the buffers written which
turns over the buffers in 6 hours. Looks like you are turning the entire
buffer cache 4-5 times during the day. Increase the BUFFERS to at least
200,000 (maybe as much as 400,000 but go slow start with 200,000).
> NUMAIOVPS 30 # Number of IO vps
Are you using COOKED chunks? Is KAIO enabled? KAIO on HP is hard to get
working but it does work well. If you are using KAIO (see onstat -g iov
for the presence of kio threads and see how they are being used versus the
aio threads) and RAW chunks then you only need a few AIO VPS say 2-6. If
you are using COOKED chunks you only need 1.5 times the number of COOKED
chunks so you can support 20 chunks efficiently with 30 AIO VPS.
> PHYSBUFF 64 # Physical log buffer size (Kbytes)
> LOGBUFF 64 # Logical log buffer size (Kbytes)> LOGSMAX 12 # Maximum number of logical log files
> CLEANERS 2 # Number of buffer cleaner processes
CLEANER >= LRUS! Make this at least 32 for efficient cleaning.
> SHMBASE 0x0 # Shared memory base address
> SHMVIRTSIZE 32768 # Initial virtual shared memory segment size
Watch onstat -g seg for additional Virtual segments. HP can only support
ONE virtual segment (there are only 4 special purpose registers for shared
memory handles and the Resident segment uses one and the shared memory
connection uses another). Fold the size of any additional, dynamically
allocated, virtual segments into SHMVIRTSIZE for the next restart.
> SHMADD 8092 # Size of new shared memory segments
> (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
> CKPTINTVL 600 # Check point interval (in sec)
> LRUS 32 # Number of LRU queues
> LRU_MAX_DIRTY 60 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 40 # LRU percent dirty end cleaning limit
These values are fine for a DSS system but for OLTP, unless there is a lot
of I/O by other applications on the controllers, you want to have most of
the I/Os be LRU Writes (at least 2/3s) NOT Chunk Writes as suggested in
the admin guide. This spreads the
In article <3815DB5C.E8471068@bloomberg.net>, Art S. Kagel
<kagel@bloomberg.net> writes
>Alnis Bajars wrote:
>>
>> Our environment is a new HP 9000 (R390) server running HP-UX 11.0. We are
>> on Dynamic Server, Workgroups Edition, version 7.30.UC9. Our legacy
>> application is running on it - it is written in Informix 4GL, and we're on
>> version 7.20 of the 4GL/RDS.
>[SNIP]
>
>> Our read cache hits have settled in at just under 99%, which is fantastic.
>>
>> However, our write cache hits started at just over 57% when we started up
>Not to worry that was probably due to data loading in to multiple tables
>such results in that situation is normal. Did you zero the stats after
>the loading to get a clear production picture?
>
>> IDS on the 18th. They climbed to a peak of just under 92% on Wednesday the
>> 20th, but have slipped away. They're sitting just over 84% now and still
>> falling.
>
>Informix says anything over 75% is good and I used to like 85% myself but
>truth is in a well tuned OLTP system with enough buffers you should be
>seeing around 90% write cache in 7.xx. IDS's buffer management is just
>that good.
>
>> The level of slippage would indicate that it's probably not just the
>> application at issue here.
>
>Except for bulk loads application design can rarely affect write cache
>percentages, now read cache is a different story. Poor application and
>database design can trash read cache ratios.
>
>> I enclose our current onconfig file and the output of onstat -p and
>> onstat -l.>
>An onstat -d and onstat -g iof (or onstat -D) would also help to see if
>heavy activity on a single chunk is trashing the ratio by flushing pages
>from other chunks.
>
>> However, my questions are:
>> What resources should I be looking at first.
>
>BUFFERS is the first and most important resource affecting the cache
>ratios.
>
>> Which onstat commands should I be running and what should I be looking out
>> for.
>
>> Thanks in advance.
>>
>> *** HERE COMES THE ONCONFIG FILE ***
>>
>> # INFORMIX SOFTWARE, INC.
>> #
>> # Title: onconfig.clr
>> # Description: Informix Dynamic Server Configuration Parameters
>> #
>> # Alnis Bajars 22/9/99. Modify template for Colorcorp installation.
>> #
>> #**************************************************************************
>>
>> # Root Dbspace Configuration
>>
>> ROOTNAME rootdbs # Root dbspace name
>> ROOTPATH /dev/chunk_01 # Path for device containing root dbspace>
>I hope that /dev/chunk_01 is a link and not the device file's actual name.
>Not using links for chunks is a can of worms that will cause you to lose
>sleep eventually, and not from worry either. Also I suggest you place the
>links somewhere OTHER than /dev so an overeager sysadmin does not remove
>them (experience speaking).
>
>> TBLSPACE_STATS 1>
>I know that this is on by default for 7.30, but, collecting tablespace
>stats is VERY expensive. You should turn it off if you are not currently
>monitoring the results.
>
>[SNIP]
>> SERVERNUM 1 # Unique id corresponding to a Dynamic>> Server instance
>> DBSERVERNAME saturn # Name of default database server
>> DBSERVERALIASES # List of alternate dbservernames
>> NETTYPE ipcshm,240,120,CPU # Configure poll thread(s) for nettype>
>You really should have a network connection DBSERVERALIAS, even if all of
>your users are local using the shared memory connection. There are
>certain circumstances that will come up where you will thank the foresight
>that added the option. If nothing else if you have a program with memory
>roblems it cannot crash the engine on a network connection, it can over
Sometimes even with a core dump and shared memory dump Informix cannot
tell what caused it and advise moving to network connections.
>shared memory, so you can shift it to that connection on the fly while you
>recode the app to fix the problem!
>
>> DEADLOCK_TIMEOUT 120 # Max time to wait of lock in distributed
>> env.
>> RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>>
>> MULTIPROCESSOR 0 # 0 for single-processor, 1 for>> multi-processor
>
>Do you REALLY have only one physical CPU? If not you MUST turn
>MULTIPROCESSOR on to enable MSP protection code to avoid CPU VP deadlocks.
>
>> NUMCPUVPS 3 # Number of user (cpu) vps>
>If you have only one CPU then an HP can usefully handle two CPU VPs per
>CPU. I think that the third one is probably wasted.
>
>> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to>> one
>>
>> NOAGE 1 # Process aging
>> AFF_SPROC 0 # Affinity start processor
>> AFF_NPROCS 0 # Affinity number of processors
AFF_SPROC=0 AFF_NPROCS=3
>>
>> # Shared Memory Parameters
>>
>> LOCKS 100000 # Maximum number of locks
>> BUFFERS 100000 # Maximum number of shared buffers>
>Look at the onstat -p stats and analyse. Your app runs 14 hours a day
>over 10 days it read in 1.7MM pages to buffers which turns over all of the
>100,000 buffers in 8 hours. Discount this with the buffers written which
>turns over the buffers in 6 hours. Looks like you are turning the entire
>buffer cache 4-5 times during the day. Increase the BUFFERS to at least
>200,000 (maybe as much as 400,000 but go slow start with 200,000).
>
Agreed, increaswe buffers
>> NUMAIOVPS 30 # Number of IO vps>
>Are you using COOKED chunks? Is KAIO enabled? KAIO on HP is hard to get
>working but it does work well. If you are using KAIO (see onstat -g iov
Use 7.31 with has it fixed.
>for the presence of kio threads and see how they are being used versus the
>aio threads) and RAW chunks then you only need a few AIO VPS say 2-6. If
>you are using COOKED chunks you only need 1.5 times the number of COOKED
>chunks so you can support 20 chunks efficiently with 30 AIO VPS.
>
>> PHYSBUFF 64 # Physical log buffer size (Kbytes)
>> LOGBUFF 64 # Logical log buffer size (Kbytes)>> LOGSMAX 12 # Maximum number of logical log files
>> CLEANERS 2 # Number of buffer cleaner processes>
>CLEANER >= LRUS! Make this at least 32 for efficient cleaning.
>
>> SHMBASE 0x0 # Shared memory base address
>> SHMVIRTSIZE 32768 # Initial virtual shared memory segment size>
>Watch onstat -g seg for additional Virtual segments. HP can only support
>ONE virtual segment (there are only 4 special purpose registers for shared
>memory handles and the Resident segment uses one and the shared memory
>connection uses another). Fold the size of any additional, dynamically
>allocated, virtual segments into SHMVIRTSIZE for the next restart.
>
>> SHMADD 8092 # Size of new shared memory segments
>> (Kbytes)
>> SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
>> CKPTINTVL 600 # Ch