Slowest restore Ever via onbar need help
Posted in 2014
Topics: High Availability & Replication, Backup & Restore, Performance & Tuning, Storage & Space Management, SQL Development & Query Writing, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues, Versions, Editions & End-of-Life
Hi Gurus,
I need your help in my onbar restore. I am doing redirected restore from
machine 1 to machine2(DEV). This is the first time i encountered so so so so
slow restoration. It's working but a 10 GB restore will took me 10 hours..
Here's the onconfig file.
I hope someone can analyze and share their knowledge. IDS 9.4UC6 on Solaris 5.x
=======================================
ROOTNAME rootdbs0000 # Root dbspace nameROOTPATH /ids_05/links/ids0127_device0000 #Path for device containing root
dbspace
ROOTOFFSET 4 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 500000 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS physdbs0000 # Location (dbspace) of physical log
PHYSFILE 500000 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 44 # Number of logical log files
LOGSIZE 10000 # Logical log size (Kbytes)
# Diagnostics
MSGPATH /appl/informix/log/online.log # System message log file path
CONSOLE /appl/informix/log/online.log # System console message path
SYSALARMPROGRAM /appl/informix/idsxsw02/etc/evidence.sh # System Alarm program
path
TBLSPACE_STATS 1
# System Archive Tape Device
#TAPEDEV /dev/null
TAPEDEV /appl/informix/tape
TAPEBLK 1024 # Tape block size (Kbytes)
TAPESIZE 20000000 # Maximum amount of data to put on tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV /appl/informix/log/ltapedev # Log tape device path
#LTAPEDEV /dev/null # Log tape device path
LTAPEBLK 16 # Log tape block size (Kbytes)
LTAPESIZE 20000000 # Max amount of data to put on log tape (Kbytes)
# Optical
STAGEBLOB # INFORMIX-OnLine/Optical staging area
# System Configuration
SERVERNUM 127 # Unique id corresponding to a OnLine instance
DBSERVERNAME ids0127 # Name of default database server
DBSERVERALIASES # List of alternate dbservernames
NETTYPE ipcshm,1,200,CPU # Configure poll thread(s) for nettype
NETTYPE soctcp,2,200,NET # Configure poll thread(s) for nettype
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-processor
NUMCPUVPS 1 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
NOAGE 0 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 100000 # Maximum number of locks
BUFFERS 10000 # Maximum number of shared buffers
NUMAIOVPS 1 # Number of IO vps
PHYSBUFF 64 # Physical log buffer size (Kbytes)
LOGBUFF 40 # Logical log buffer size (Kbytes)LOGSMAX 64 # Maximum number of logical log files
CLEANERS 2 # Number of buffer cleaner processes
SHMBASE 0x30000000 # Shared memory base address
SHMVIRTSIZE 80000 # initial virtual shared memory segment size
SHMADD 8000 # Size of new shared memory segments (Kbytes)
SHMTOTAL 300000 # Total shared memory (Kbytes). 0=>unlimited
#SHMTOTAL 200000 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 900 # Check point interval (in sec)
LRUS 2 # Number of LRU queues
LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water mark percentage
LTXEHWM 60 # Long transaction high water mark (exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
DD_HASHSIZE 511 # Length of Data Dictionary Hash
DD_HASHMAX 20 # Width of Data Dictionary Hash# System Page Size
# BUFFSIZE - OnLine no longer supports this configuration parameter.
# To determine the page size used by OnLine on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline worker threads
ON_RECVRY_THREADS 1 # Default number of online worker threads
# Data Replication Variables
# DRAUTO: 0 manual, 1 retain type, 2 reverse type
DRAUTO 0 # DR automatic switchover
DRINTERVAL 30 # DR max time between DR buffer flushes (in sec)
DRTIMEOUT 30 # DR network timeout (in sec)DRLOSTFOUND /informix/etc/dr.lostfound # DR lost+found file path
# CDR Variables
CDR_LOGBUFFERS 2048 # size of log reading buffer pool (Kbytes)
CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-vp,additional)
CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR queue (Kbytes)
# Backup/Restore variables
BAR_BSALIB_PATH /usr/lib/ibsad001.so
### BAR_BSALIB_PATH /usr/lib/ibsad001.a
BAR_ACT_LOG /appl/informix/log/bar_act_onlineipc_qua.log
BAR_MAX_BACKUP 5
BAR_RETRY 1
BAR_NB_XPORT_COUNT 10
BAR_XFER_BUF_SIZE 15
# Read Ahead Variables
RA_PAGES 32 # Number of pages to attempt to read ahead
RA_THRESHOLD 16 # Number of pages left before next group
# DBSPACETEMP:
# OnLine equivalent of DBTEMP for SE. This is the list of dbspaces
# that the OnLine SQL Engine will use to create temp tables etc.
# If specified it must be a colon separated list of dbspaces that exist
# when the OnLine system is brought online. If not specified, or if
# all dbspaces specified are invalid, various ad hoc queries will create
# temporary files in /tmp instead.
DBSPACETEMP tempdb1,tempdb2 # Default temp dbspaces
# DUMP*:
# The following parameters control the type of diagnostics information which
# is preserved when an unanticipated error condition (assertion failure) occurs
# during OnLine operations.
# For DUMPSHMEM, DUMPGCORE and DUMPCORE 1 means Yes, 0 means No.
DUMPDIR /tmp # Preserve diagnostics in this directory
DUMPSHMEM 0 # Dump a copy of shared memory
DUMPGCORE 0 # Dump a core image using 'gcore'
DUMPCORE 0 # Dump a core image (Warning:this aborts OnLine)
DUMPCNT 1 # Number of shared memory or gcore dumps forStandard input
FILLFACTOR 90 # Fill factor for building indexes
# method for OnLine to use when determining current time
USEOSTIME 1 # 0: use internal time(fast), 1: get time from OS(slow)
# Parallel Database Queries (pdq)
MAX_PDQPRIORITY 50 # Maximum allowed pdqpriority
DS_MAX_QUERIES # Maximum number of decision support queries
DS_TOTAL_MEMORY # Decision support memory (Kbytes)
DS_MAX_SCANS 1048576 # Maximum number of decision support scans
DATASKIP off# OPTCOMPIND
# 0 => Nested loop joins will be preferred (where
# possible) over sortmerge joins and hash joins.
# 1 => If the transaction isolation mode is not
# "repeatable read", optimizer behaves as in (2)
# below. Otherwise it behaves as in (0) above.
# 2 => Use costs regardless of the transaction isolation
# mode. Nested loop joins are not necessarily
# preferred. Optimizer bases its decision purely
# on costs.
OPTCO
I'm also not sure if the "checkpoint req" message is causing the
blocking...can i execute 'onmode -c' to force checkpoint?
Jack,
which kind of execution time did you have formerly for this operation? Or is
this the first time you do it between both servers?
From what I can see in you onconfig file, it has extremely few buffers
(although onbar writes directly to disk), and only 1 CPU VP which will prevent
you from taking benefit of the parallel restore of onbar. This may be the
cause of the slowness.
Did you run iostat on the target server, is the server in wait-io state ?
Check onstat -pr, DskWrits and Pagwritscolumn and see how many pages are
written every 5 seconds. This should be more than a few thousands writes per
seconds for Pagwrits.
and also check for any complaints in the BAR_ACT_LOG file
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix