Periodic extremely slow ontape archives
Posted in 2000
Topics: Backup & Restore, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues, Versions, Editions & End-of-Life
Greetings. We have been experiencing intermittently long archives
using ontape for several months. These long runs have increased
in duration and frequency during this time. The box doesn't look
pregnant! Most nights we finish in under 3 hours, sometimes over 21 hours.
We average around 2 long archives per week. Our production batch is
scheduled to run after this, so a long run adversely effects us the next day.
There is no obvious correlation between system activity during the archive
and it's duration. By day we do mostly OLTP and run batch at night. We
schedule the archive before the nightly batch and after most daily
OLTP work.
Here is an interesting example of what
has happened the last 2 days.
Monday's archive started at 9 PM
Time to complete = 21 hours 35 minutes
Average processor time / 15 minute increment = 2 - 3 seconds
Tuesday's archive started just 2 1/2 hrs after the long archive:
Time to complete = 2 hours 23 minutes
Average processor time / 15 minute increment = 23 - 24 seconds
** Note: The processor times were obtained from a unix 'ps' command
taken at 15 minute intervals.
CPU Utilization was < 50% both days
Memory Utilization < 70% both days
No apparent difference in jobs running on the two days
The backup is done via an Informix Utility and uses compression on a
DLT 7000 tape drive.
Disk IO never spiked on Monday, but it spiked for the entire 2.5 hours on
Tuesday (as it should have).
It looks like we had plenty of available resources during this time. ontape
was just t a a a k i ng i t ' s t i i i i m e.
I'd really appreciate some help with this one as we are getting hit pretty hard
and i'm stumped for now. Thanks!
Here's our system specs:
Hardware: Sun Enterprise 6000, 14 cpus @ 248 MHz, 7GB memory
OS: Solaris 2.6
Database: IDS 7.30.UC6
Our database size is around 80 GB. of RAID 5. I really do want to change, Art.
This Sun host supports 2 Informix instances. One is a magnetic cache for
blobs and is archived to /dev/null.
Other subsystems: Plexus Floware
The Floware system allocates shared memory and has background
daemons. It uses an Informix database that it manages. API calls to this
system are made by client applications utilizing workflow functionality.
Here's our onconfig:
Informix Dynamic Server Version 7.30.UC6 -- On-Line -- Up 11 days 22:49:41 --
1180224 Kbytes
Configuration File: /usr/informix/etc/onconfig.onl
#**************************************************************************
#
# INFORMIX SOFTWARE, INC.
#
# Title: onconfig for wil_online
# Description: INFORMIX-OnLine Configuration Parameters
#
#**************************************************************************
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace nameROOTPATH /dev/vx/rdsk/rawdg/d10001 # Path for root dbspace device
ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 500000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path device containing mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS phys # Location (dbspace) of physical log
PHYSFILE 30000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 50 # Number of logical log files
LOGSIZE 10000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /usr/informix/onl_online.log # System message log file path
CONSOLE /dev/console # System console message path
ALARMPROGRAM /usr/informix/etc/no_log.sh # Alarm program path# SYSALARMPROGRAM /usr/informix/etc/evidence.sh # System Alarm program path
# TBLSPACE_STATS 1
# System Archive Tape Device
TAPEDEV /dev/rmt/1c # Tape device path
TAPEBLK 64 # Tape block size (Kbytes)
TAPESIZE 60000000 # Maximum amount of data to put on tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV /dev/rmt/0c # Log tape device path
LTAPEBLK 64 # Log tape block size (Kbytes)
LTAPESIZE 35000000 # Max amount of data to put on log tape (Kbytes)
# Optical
STAGEBLOB # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 30 # Unique id corresponding to a OnLine instanceDBSERVERNAME wil_online_shm # Name of default database server
DBSERVERALIASES wil_online,wil_online_2,wil_online_3
# List of alternate dbservernames
NETTYPE ipcshm,3,200,CPU
NETTYPE tlitcp,10,250,NET
#changed 12/16/1998 tpp
#NETTYPE tlitcp,8,250,NET
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-processor
NUMCPUVPS 8 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
NOAGE 1 # Process aging
AFF_SPROC 1 # Affinity start processor
AFF_NPROCS 8 # Affinity number of processors
# Shared Memory Parameters
LOCKS 100000 # Maximum number of locks
BUFFERS 80000 # Maximum number of shared buffers
NUMAIOVPS 2 # Number of IO vps
PHYSBUFF 128 # Physical log buffer size (Kbytes)
LOGBUFF 34 # Logical log buffer size (Kbytes)LOGSMAX 500 # Maximum number of logical log files
CLEANERS 24 # Number of buffer cleaner processes
SHMBASE 0xa000000 # Shared memory base address
SHMVIRTSIZE 1000000 # initial virtual shared memory segment size
SHMADD 65536 # Size of new shared memory segments (Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 32 # Number of LRU queues
LRU_MAX_DIRTY 3 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water mark percentage
LTXEHWM 60 # Long transaction high water mark (exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker th
In article <850e1p$gu7$1@news.xmission.com>, Dick.Brieck@chase.com
writes
>
>
>
>Greetings. We have been experiencing intermittently long archives
>using ontape for several months. These long runs have increased
>in duration and frequency during this time. The box doesn't look
>pregnant! Most nights we finish in under 3 hours, sometimes over 21 hours.
>We average around 2 long archives per week. Our production batch is
>scheduled to run after this, so a long run adversely effects us the next day.
>There is no obvious correlation between system activity during the archive
>and it's duration. By day we do mostly OLTP and run batch at night. We
>schedule the archive before the nightly batch and after most daily
>OLTP work.
>
What does
sar -u 1 1
sar -d 1 1
several times and
onstat -u
give?
--
David Williams
Have you had the tapes checked and the tape drive cleaned?
Dick.Brieck@chase.com wrote in message <850e1p$gu7$1@news.xmission.com>...
>
>
>
>Greetings. We have been experiencing intermittently long archives
>using ontape for several months. These long runs have increased
>in duration and frequency during this time. The box doesn't look
>pregnant! Most nights we finish in under 3 hours, sometimes over 21 hours.
>We average around 2 long archives per week. Our production batch is
>scheduled to run after this, so a long run adversely effects us the next
day.
>There is no obvious correlation between system activity during the archive
>and it's duration. By day we do mostly OLTP and run batch at night. We
>schedule the archive before the nightly batch and after most daily
>OLTP work.
>
>Here is an interesting example of what
>has happened the last 2 days.
>
>Monday's archive started at 9 PM
> Time to complete = 21 hours 35 minutes
> Average processor time / 15 minute increment = 2 - 3 seconds
>
>Tuesday's archive started just 2 1/2 hrs after the long archive:
> Time to complete = 2 hours 23 minutes
> Average processor time / 15 minute increment = 23 - 24 seconds
>
>** Note: The processor times were obtained from a unix 'ps' command
>taken at 15 minute intervals.
>
>CPU Utilization was < 50% both days
>Memory Utilization < 70% both days
>
>No apparent difference in jobs running on the two days
>
>The backup is done via an Informix Utility and uses compression on a
> DLT 7000 tape drive.
>
>Disk IO never spiked on Monday, but it spiked for the entire 2.5 hours on
>Tuesday (as it should have).
>
>It looks like we had plenty of available resources during this time.
ontape>was just t a a a k i ng i t ' s t i i i i m e.
>
>I'd really appreciate some help with this one as we are getting hit pretty
hard
>and i'm stumped for now. Thanks!
>
>Here's our system specs:
>
>Hardware: Sun Enterprise 6000, 14 cpus @ 248 MHz, 7GB memory
>OS: Solaris 2.6
>Database: IDS 7.30.UC6
>Our database size is around 80 GB. of RAID 5. I really do want to change,
Art.
>
>This Sun host supports 2 Informix instances. One is a magnetic cache for
>blobs and is archived to /dev/null.
>
>Other subsystems: Plexus Floware
>
>The Floware system allocates shared memory and has background
>daemons. It uses an Informix database that it manages. API calls to this
>system are made by client applications utilizing workflow functionality.
>
>Here's our onconfig:
>
>Informix Dynamic Server Version 7.30.UC6 -- On-Line -- Up 11 days
22:49:41 --
>1180224 Kbytes
>
>Configuration File: /usr/informix/etc/onconfig.onl
>#**************************************************************************
>#
># INFORMIX SOFTWARE, INC.
>#
># Title: onconfig for wil_online
># Description: INFORMIX-OnLine Configuration Parameters
>#
>#**************************************************************************
>
># Root Dbspace Configuration
>
>ROOTNAME rootdbs # Root dbspace name>ROOTPATH /dev/vx/rdsk/rawdg/d10001 # Path for root dbspace device
>ROOTOFFSET 0 # Offset of root dbspace into device
(Kbytes)
>ROOTSIZE 500000 # Size of root dbspace (Kbytes)>
># Disk Mirroring Configuration Parameters
>
>MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
>MIRRORPATH # Path device containing mirrored root
>MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
># Physical Log Configuration
>
>PHYSDBS phys # Location (dbspace) of physical log
>PHYSFILE 30000 # Physical log file size (Kbytes)>
># Logical Log Configuration
>
>LOGFILES 50 # Number of logical log files
>LOGSIZE 10000 # Logical log size (Kbytes)>
># Diagnostics
>
>MSGPATH /usr/informix/onl_online.log # System message log file path
>CONSOLE /dev/console # System console message path
>ALARMPROGRAM /usr/informix/etc/no_log.sh # Alarm program path># SYSALARMPROGRAM /usr/informix/etc/evidence.sh # System Alarm program
path
># TBLSPACE_STATS 1
>
>
># System Archive Tape Device
>
>TAPEDEV /dev/rmt/1c # Tape device path
>TAPEBLK 64 # Tape block size (Kbytes)
>TAPESIZE 60000000 # Maximum amount of data to put on tape
(Kbytes)>
># Log Archive Tape Device
>
>LTAPEDEV /dev/rmt/0c # Log tape device path
>LTAPEBLK 64 # Log tape block size (Kbytes)
>LTAPESIZE 35000000 # Max amount of data to put on log tape
(Kbytes)>
># Optical
>
>STAGEBLOB # INFORMIX-OnLine/Optical staging area
>
># System Configuration
>
>SERVERNUM 30 # Unique id corresponding to a OnLineinstance
>DBSERVERNAME wil_online_shm # Name of default database server
>DBSERVERALIASES wil_online,wil_online_2,wil_online_3
> # List of alternate dbservernames
>NETTYPE ipcshm,3,200,CPU
>NETTYPE tlitcp,10,250,NET
>#changed 12/16/1998 tpp
>#NETTYPE tlitcp,8,250,NET
>DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed
env.
>RESIDENT 1 # Forced residency flag (Yes = 1, No = 0)
>
>MULTIPROCESSOR 1 # 0 for single-processor, 1 formulti-processor
>NUMCPUVPS 8 # Number of user (cpu) vps
>SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps toone
>
>NOAGE 1 # Process aging
>AFF_SPROC 1 # Affinity start processor
>AFF_NPROCS 8 # Affinity number of processors>
># Shared Memory Parameters
>
>LOCKS 100000 # Maximum number of locks
>BUFFERS 80000 # Maximum number of shared buffers
>NUMAIOVPS 2 # Number of IO vps
>PHYSBUFF 128 # Physical log buffer size (Kbytes)
>LOGBUFF 34 # Logical log buffer size (Kbytes)>LOGSMAX 500 # Maximum number of logical log files
>CLEANERS 24 # Number of buffer cleaner processes
>SHMBASE 0xa000000 # Shared memory base address
>SHMVIRTSIZE 1000000 # initial virtual shared memory segmentsize
>SHMADD 65536 # Size of new shared memory segments
(Kbytes)
>SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
>CKPTINTVL 300 # Check point interval (in sec)
>LRUS 32 # Number of LRU queues
>LRU_MAX_DIRTY 3 # LRU percent dirty begin cleaning limit
>LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
>LTXHWM 50 # Long transaction high water markpercentage
>LTXEHWM 60 # Long transaction high water mark
(exclusive)
>TXTIMEOUT 0x12c # Transaction timeou
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix