How to avoid blocking checkpoints?
Posted in 2017
A user on IDS 7.31.TD6 under Windows 2008 R2 reported checkpoints blocking for 4-6 minutes, and posted his ONCONFIG. Respondents noted too few page cleaners (CLEANERS 7 against LRUS 128 and a 129-chunk dbspace) and too-lax LRU settings, leaving many dirty buffers at checkpoint time; advice was to raise CLEANERS to 128, lower LRU_MAX_DIRTY to about 2, add an AIO VP, watch onstat -R/-F/-D/ioq, check for bloated indexes, and move off RAID5 to RAID10 (plus suggestions to upgrade to 12.10, which has non-blocking checkpoints). After applying the changes the poster saw shorter but still long (~4 minute) checkpoints during a dbimport, so no full resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Backup & Restore, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Versions, Editions & End-of-Life
Dear:
I have a problem with the checkpoints of an IDS 7.31 TD6 engine running on a
Windows 2008 Server R2. The server has an array of 4 disks in RAID 5 for the
data, and 3 arrays more than 2 disks in RAID 1 for the dbspaces temorial, for
the Logdbs, and for the Physical Log, respectively.
When they occur (the checkpoints), sometimes they take a long time (between 4
and 6 minutes).
I have tried changing the size of the physical log, as they say in some IBM
documents (taking it to 110%) of the total size of Logical Logs, but the
problem persists.
I transcribe my onconfig to see if any one notice any parameters that have to
be modified .:
###########################################################################
# Root Dbspace Configuration
ROOTNAME rootdbs # Root dbspace name
ROOTPATH C:\\\\IFMXDATA\\\\ol_abril\\\\rootdbs_dat.000
# Path for device containing root dbspace
ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 122880 # Size of root dbspace (Kbytes)
# Disk Mirroring Configuration Parameters
MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
MIRRORPATH # Path for device containing mirrored root
MIRROROFFSET 0 # Offset into mirrored device (Kbytes)
# Physical Log Configuration
PHYSDBS physdbs # Location (dbspace) of physical log
PHYSFILE 281600 # Physical log file size (Kbytes)
# Logical Log Configuration
LOGFILES 20 # Number of logical log files
LOGSIZE 51200 # Logical log size (Kbytes)LOG_BACKUP_MODE CONT # Logical log backup mode (MANUAL, CONT)
# Diagnostics
MSGPATH C:\\\\Informix\\\\ol_abril.log # System message log file path
CONSOLE C:\\\\Informix\\\\conol_abril.log # System console message path
ALARMPROGRAM C:\\\\Informix\\\\etc\\
o_log.bat # Alarm program path
# System Diagnostic Script.
# SYSALARMPROGRAM - Full path of the system diagnostic script (e.g.
# c:\\\\informix\\\\etc\\\\evidence.bat.) Set this parameter
# if you want a different Diagnostic Script than
# {INFORMIXDIR}\\\\etc\\\\evidence.bat, which is default.
# System Archive Tape Device
TAPEDEV NUL # Tape device path
#TAPEDEV \\\\\\\\.\\\\TAPE0 # Tape device path
TAPEBLK 16 # Tape block size (Kbytes)
TAPESIZE 10240 # Maximum amount of data to put on tape (Kbytes)
# Log Archive Tape Device
LTAPEDEV NUL # Log tape device path
#LTAPEDEV \\\\\\\\.\\\\TAPE1 # Log tape device path
LTAPEBLK 16 # Log tape block size (Kbytes)
LTAPESIZE 10240 # Max amount of data to put on log tape (Kbytes)
# Optical
STAGEBLOB # Informix Dynamic Server/Optical staging area
OPTICAL_LIB_PATH # Location of Optical Subsystem driver DLL
# System Configuration
SERVERNUM 0 # Unique id corresponding to a server instance
DBSERVERNAME ol_abril # Name of default Dynamic Server
DBSERVERALIASES # List of alternate dbservernames
NETTYPE soctcp,1,,NET # Override sqlhosts nettype parameters
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-processor
NUMCPUVPS 7 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
NOAGE 0 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
# Shared Memory Parameters
LOCKS 200000 # Maximum number of locks
BUFFERS 300000 # Maximum number of shared buffers
NUMAIOVPS 1 # Number of IO vps
PHYSBUFF 128 # Physical log buffer size (Kbytes)
LOGBUFF 128 # Logical log buffer size (Kbytes)LOGSMAX 20 # Maximum number of logical log files
CLEANERS 7 # Number of buffer cleaner processes
SHMBASE 0xc000000 # Shared memory base address
SHMVIRTSIZE 131072 # initial virtual shared memory segment size
SHMADD 16384 # Size of new shared memory segments (Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 128 # Number of LRU queues
LRU_MAX_DIRTY 5 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water mark percentage
LTXEHWM 60 # Long transaction high water mark (exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
# System Page Size
# BUFFSIZE - Dynamic Server no longer supports this configuration parameter.
# To determine the page size used by Dynamic Server on your platform
# see the last line of output from the command, 'onstat -b'.
# Recovery Variables
# OFF_RECVRY_THREADS:
# Number of parallel worker threads during fast recovery or an offline restore.
# ON_RECVRY_THREADS:
# Number of parallel worker threads during an online restore.
OFF_RECVRY_THREADS 10 # Default number of offline worker threads
ON_RECVRY_THREADS 1 # Default number of online worker threads
# Data Replication Variables
# DRAUTO: 0 manual, 1 retain type, 2 reverse type
DRAUTO 0 # DR automatic switchover
DRINTERVAL 30 # DR max time between DR buffer flushes (in sec)
DRTIMEOUT 30 # DR network timeout (in sec)
DRLOSTFOUND \\\\tmp # DR lost+found file path
# CDR Variables
CDR_LOGBUFFERS 2048 # size of log reading buffer pool (Kbytes)
CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-vp,additional)
CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR queue (Kbytes)CDR_LOGDELTA 30 # % of log space allowed in queue memory
CDR_NUMCONNECT 16 # Expected connections per server
CDR_NIFRETRY 300 # Connection retry (seconds)
CDR_NIFCOMPRESS 0 # Link level compression (-1 never, 0 none, 9 max)
# Backup/Restore variables
BAR_ACT_LOG C:\\\\Informix\\\\bar_ol_abril.log #Path of log file for onbar.exe
BAR_MAX_BACKUP 0
BAR_RETRY 1
BAR_NB_XPORT_COUNT 10
BAR_XFER_BUF_SIZE 15
BAR_BSALIB_PATH C:\\\\ISM\\\\2.20\\\\bin\\\\libbsa.dll # Location of ISM XBSA DLL
RESTARTABLE_RESTORE on #To support restartable restore..values on/off
# Informix Storage Manager variables
ISM_DATA_POOL ISMData
ISM_LOG_POOL ISMLogs
# Read Ahead Variables
RA_PAGES 60 # Number of pages to attempt to read ahead
RA_THRESHOLD 64 # Number of pages left before next group
# DBSPACETEMP:
# Dynamic Server equivalent of DBTEMP for SE. This is the list of dbspaces
# that the Dynamic Server SQL Engine will use to create temp tables etc.
# If specified it must be a colon separated list of dbspaces that exist
# when the Dynamic Server system is brought online. If not specified, or if
# all dbspaces specified are invalid, various ad hoc queries will create
# temporary files in /tmp instead.
DBSPACETEMP
tmp0,tmp1,tmp2,tmp3,tmp4,tmp5,tmp6,tmp7,tmp8,tmp9,tmp10,tmp11,tmp12,tmp13,tmp14,
tmp15,tmp16,tmp17,tmp18,tmp19,tmp20,tmp21,tmp22,tmp23,tmp24,tmp25,tmp26,tmp27,tm
p28,tmp29
# Default temp dbspaces
# DUMP*:
# The following parameters control the type of diagnostics information which
# is preserved when an unanticipated error condition (assertion failure) occurs
# during Dynamic Server operations.
# For DUMPSHMEM, DUMPGCORE and DUMPCORE 1 means Yes, 0 means No.
DUMPDIR \\\\tmp # Preserve diagnostics in this directory
DUMPSHMEM 1 # Dump a copy of shared memory@@NL
Original post:
<stuff removed>
CLEANERS 7 # Number of buffer cleaner processes
SHMBASE 0xc000000 # Shared memory base address
SHMVIRTSIZE 131072 # initial virtual shared memory segment size
SHMADD 16384 # Size of new shared memory segments (Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 128 # Number of LRU queues
LRU_MAX_DIRTY 5 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
<stuff removed>
Response:
I removed most the $ONCONFIG file and left in the main factors in checkpoint
flush times in 7.31. It's hard to say without seeing at least onstat
information from the instance while some of the longer checkpoints were
happening...but I'll still throw out two things to consider.
First, in that old a version, to reduce checkpoint duration, you'll want to
have as few dirty buffers as possible to flush come your checkpoint interval.
So you want to have your system doing mostly LRU writes and not checkpoint
writes. You could lower your LRU_MAX_DIRTY from 5 down to 2 (your version is
so old I'm not sure if it allows non-integer numbers for the LRU settings). I
don't recall if really aggressive LRU settings for this version was 2 and 1
(for max/min) or 1 and 0, but for starters I'd maybe try 2 and 1 for the
max/min lru settings.
Next, you only have 7 page cleaners configured (CLEANERS in the $ONCONFIG) but
you have 128 LRUS configured. So once you have more then 7 LRU's that need to
get cleaned, the other 121 LRUS are going to possibly go above LRU_MAX_DIRTY.
So there is a reasonable chance that come checkpoint interval, many of your
LRU's could have more then 5% dirty buffers on them. You should probably
increase the value or CLEANERS...possibly going all the way up to the number
of LRU's you have so that each LRU could have a cleaner working on it when it
exceeds LRU_MAX_DIRTY.
If you'd like some verification of this before changing, I would suggest
monitoring onstat -R output on your system and get some record of it right
before checkpoints kick off. That way you can see how many dirty buffers the
LRU's are carrying before the checkpoint starts. I would expect that on longer
checkpoints you would have more LRU's that would be >= LRU_MAX_DIRTY% buffers
then checkpoints you have that perform more quickly.
Gustavo:
Are the chunks RAW or COOKED? If cooked then you should have at least 1.5
AIO VPs per active (non-temp) chunk and you only have one. You have seven
CLEANER threads, now many non-temp chunks do you have? You should have at
least one CLEANER per actively modified chunk (those with lots of writes in
onstat -D output). Post your onstat -g iov, onstat -F and onstat -R output.Maybe we'll see something.
Art
Art S. Kagel, President and Principal Consultant
ASK Database Management
www.askdbmgt.com
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on the IIUG, nor any other organization with which I am
associated either explicitly, implicitly, or by inference. Neither do
those opinions reflect those of other individuals affiliated with any
entity with which I am affiliated nor those of the entities themselves.
On Mon, May 15, 2017 at 12:01 PM, GUSTAVO ECHENIQUE <
gustavo.echenique@cemdo.com.ar> wrote:
> Dear:
>
> I have a problem with the checkpoints of an IDS 7.31 TD6 engine running on
> a
> Windows 2008 Server R2. The server has an array of 4 disks in RAID 5 for
> the
> data, and 3 arrays more than 2 disks in RAID 1 for the dbspaces temorial,
> for
> the Logdbs, and for the Physical Log, respectively.
>
> When they occur (the checkpoints), sometimes they take a long time
> (between 4
> and 6 minutes).
>
> I have tried changing the size of the physical log, as they say in some IBM
> documents (taking it to 110%) of the total size of Logical Logs, but the
> problem persists.
>
> I transcribe my onconfig to see if any one notice any parameters that have
> to
> be modified .:
>
> ############################################################
> ###############
> # Root Dbspace Configuration
>
> ROOTNAME rootdbs # Root dbspace name
> ROOTPATH C:\\\\IFMXDATA\\\\ol_abril\\\\rootdbs_dat.000>
> # Path for device containing root dbspace
> ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
> ROOTSIZE 122880 # Size of root dbspace (Kbytes)>
> # Disk Mirroring Configuration Parameters
>
> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> MIRRORPATH # Path for device containing mirrored root
> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> # Physical Log Configuration
>
> PHYSDBS physdbs # Location (dbspace) of physical log
> PHYSFILE 281600 # Physical log file size (Kbytes)>
> # Logical Log Configuration
>
> LOGFILES 20 # Number of logical log files
> LOGSIZE 51200 # Logical log size (Kbytes)> LOG_BACKUP_MODE CONT # Logical log backup mode (MANUAL, CONT)
>
> # Diagnostics
>
> MSGPATH C:\\\\Informix\\\\ol_abril.log # System message log file path
> CONSOLE C:\\\\Informix\\\\conol_abril.log # System console message path
> ALARMPROGRAM C:\\\\Informix\\\\etc\\
o_log.bat # Alarm program path>
> # System Diagnostic Script.
> # SYSALARMPROGRAM - Full path of the system diagnostic script (e.g.
> # c:\\\\informix\\\\etc\\\\evidence.bat.) Set this parameter
> # if you want a different Diagnostic Script than
> # {INFORMIXDIR}\\\\etc\\\\evidence.bat, which is default.
>
> # System Archive Tape Device
>
> TAPEDEV NUL # Tape device path
> #TAPEDEV \\\\\\\\.\\\\TAPE0 # Tape device path
> TAPEBLK 16 # Tape block size (Kbytes)
> TAPESIZE 10240 # Maximum amount of data to put on tape (Kbytes)>
> # Log Archive Tape Device
>
> LTAPEDEV NUL # Log tape device path
> #LTAPEDEV \\\\\\\\.\\\\TAPE1 # Log tape device path
> LTAPEBLK 16 # Log tape block size (Kbytes)
> LTAPESIZE 10240 # Max amount of data to put on log tape (Kbytes)>
> # Optical
>
> STAGEBLOB # Informix Dynamic Server/Optical staging area
> OPTICAL_LIB_PATH # Location of Optical Subsystem driver DLL
>
> # System Configuration
>
> SERVERNUM 0 # Unique id corresponding to a server instance
> DBSERVERNAME ol_abril # Name of default Dynamic Server
> DBSERVERALIASES # List of alternate dbservernames
> NETTYPE soctcp,1,,NET # Override sqlhosts nettype parameters
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
> RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-processor
> NUMCPUVPS 7 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
>
> NOAGE 0 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors>
> # Shared Memory Parameters
>
> LOCKS 200000 # Maximum number of locks
> BUFFERS 300000 # Maximum number of shared buffers
> NUMAIOVPS 1 # Number of IO vps
> PHYSBUFF 128 # Physical log buffer size (Kbytes)
> LOGBUFF 128 # Logical log buffer size (Kbytes)> LOGSMAX 20 # Maximum number of logical log files
> CLEANERS 7 # Number of buffer cleaner processes
> SHMBASE 0xc000000 # Shared memory base address
> SHMVIRTSIZE 131072 # initial virtual shared memory segment size
> SHMADD 16384 # Size of new shared memory segments (Kbytes)
> SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
> CKPTINTVL 300 # Check point interval (in sec)
> LRUS 128 # Number of LRU queues
> LRU_MAX_DIRTY 5 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
> LTXHWM 50 # Long transaction high water mark percentage
> LTXEHWM 60 # Long transaction high water mark (exclusive)
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 32 # Stack size (Kbytes)>
> # System Page Size
> # BUFFSIZE - Dynamic Server no longer supports this configuration
> parameter.
> # To determine the page size used by Dynamic Server on your platform
> # see the last line of output from the command, 'onstat -b'.
>
> # Recovery Variables
> # OFF_RECVRY_THREADS:
> # Number of parallel worker threads during fast recovery or an offline
> restore.
> # ON_RECVRY_THREADS:
> # Number of parallel worker threads during an online restore.
>
> OFF_RECVRY_THREADS 10 # Default number of offline worker threads
> ON_RECVRY_THREADS 1 # Default number of online worker threads>
> # Data Replication Variables
> # DRAUTO: 0 manual, 1 retain type, 2 reverse type
> DRAUTO 0 # DR automatic switchover
> DRINTERVAL 30 # DR max time between DR buffer flushes (in sec)
> DRTIMEOUT 30 # DR network timeout (in sec)
> DRLOSTFOUND \\\\tmp # DR lost+found file path>
> # CDR Variables
> CDR_LOGBUFFERS 2048 # size of log reading buffer pool (Kbytes)
> CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-vp,additional)
> CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
> CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR queue (Kbytes)> CDR_LOGDELTA 30 # % of log space allowed in queue memory
> CDR_NUMCONNECT 16 # Expected connections per server
> CDR_NIFRETRY 300 # Connection retry (seconds)
> CDR_NIFCOMPRESS 0 # Link level compression (-1 never, 0 none, 9 max)>
> # Backup/Restore variables
> BAR_ACT_LOG C:\\\\Informix\\\\bar_ol_abril.log #Path of log file for onbar.exe
> BAR_MAX_BACKUP 0
> BAR_RETRY 1
> BAR_NB_XPORT_COUNT 10
> BAR_XFER_BUF_SIZE 15
> BAR_BSALIB_PATH C:\\\\ISM\\\\2.20\\\\bin\\\\libbsa.dll # Location of ISM XBS
You also might want to watch out for large, uncleaned indices (with=20
possibly large contiguous areas of deleted keys).
If you have large tables undergoing large deletes (or updates modifying=20
indexed columns) you might want to use 'oncheck -pT db:tabname' to detect =
overly large indices and recreate them.
From: "JACQUES RENAUT" <jrenaut@us.ibm.com>
To: ids@iiug.org
Date: 15.05.2017 19:47
Subject: Re: How to avoid blocking checkpoints? [39182]
Sent by: ids-bounces@iiug.org
Original post:=20
<stuff removed>=20
CLEANERS 7 # Number of buffer cleaner processes=20
SHMBASE 0xc000000 # Shared memory base address=20
SHMVIRTSIZE 131072 # initial virtual shared memory segment size=20
SHMADD 16384 # Size of new shared memory segments (Kbytes)=20
SHMTOTAL 0 # Total shared memory (Kbytes). 0=3D>unlimited=20
CKPTINTVL 300 # Check point interval (in sec)=20
LRUS 128 # Number of LRU queues=20LRU=5FMAX=5FDIRTY 5 # LRU percent dirty begin cleaning limit=20
LRU=5FMIN=5FDIRTY 1 # LRU percent dirty end cleaning limit=20
<stuff removed>=20
Response:=20
I removed most the $ONCONFIG file and left in the main factors in=20
checkpoint=20
flush times in 7.31. It's hard to say without seeing at least onstat=20
information from the instance while some of the longer checkpoints were=20
happening...but I'll still throw out two things to consider.=20
First, in that old a version, to reduce checkpoint duration, you'll want=20
to=20
have as few dirty buffers as possible to flush come your checkpoint=20
interval.=20
So you want to have your system doing mostly LRU writes and not checkpoint =
writes. You could lower your LRU=5FMAX=5FDIRTY from 5 down to 2 (your versi=
on=20
is=20
so old I'm not sure if it allows non-integer numbers for the LRU=20
settings). I=20
don't recall if really aggressive LRU settings for this version was 2 and=20
1=20
(for max/min) or 1 and 0, but for starters I'd maybe try 2 and 1 for the=20
max/min lru settings.=20
Next, you only have 7 page cleaners configured (CLEANERS in the $ONCONFIG) =
but=20
you have 128 LRUS configured. So once you have more then 7 LRU's that need =
to=20
get cleaned, the other 121 LRUS are going to possibly go above=20
LRU=5FMAX=5FDIRTY.=20
So there is a reasonable chance that come checkpoint interval, many of=20
your=20
LRU's could have more then 5% dirty buffers on them. You should probably=20
increase the value or CLEANERS...possibly going all the way up to the=20
number=20
of LRU's you have so that each LRU could have a cleaner working on it when =
it=20
exceeds LRU=5FMAX=5FDIRTY.=20
If you'd like some verification of this before changing, I would suggest=20
monitoring onstat -R output on your system and get some record of it right =
before checkpoints kick off. That way you can see how many dirty buffers=20
the=20
LRU's are carrying before the checkpoint starts. I would expect that on=20
longer=20
checkpoints you would have more LRU's that would be >=3D LRU=5FMAX=5FDIRTY%=
=20
buffers=20
then checkpoints you have that perform more quickly.=20
***************************************************************************=
****=20
Forum Note: Use "Reply" to post a response in the discussion forum.=20
Hi Art!
I send you the copy of the commands that you requested.
C:\\\\Informix>onstat -g iov
Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:20:41 -- 1379136
Kbytes
AIO I/O vps:
class/vp s io/s totalops dskread dskwrite dskcopy wakeups io/wup errors
kio 0 s 54.2 1041585 125284 916301 0 205791447 0.0 0
kio 1 i 61.3 1177393 239878 937515 0 3581357382 -0.0 0
kio 2 i 67.3 1294468 329623 964845 0 468431746 0.0 0
kio 3 i 69.8 1340824 373310 967514 0 736248903 0.0 0
kio 4 i 331.2 6366742 5332708 1034034 0 3516245753 -0.0 0
kio 5 i 78.6 1510065 357187 1152878 0 595532308 0.0 0
kio 6 i 61.5 1181848 263363 918485 0 459739069 0.0 0
msc 0 i 0.0 27 0 0 0 28 1.0 0
aio 0 i 0.0 332 64 0 0 333 1.0 0
pio 0 i 0.0 0 0 0 0 1 0.0 0
lio 0 i 0.0 0 0 0 0 1 0.0 0
C:\\\\Informix>onstat -F
Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:37:59 -- 1379136
Kbytes
Fg Writes LRU Writes Chunk Writes
82305 6231910 2447839
address flusher state data
58320510 0 L 179 = 0Xb3
58320a08 1 L 243 = 0Xf3
58320f00 2 L 71 = 0X47
583213f8 3 L 45 = 0X2d
583218f0 4 L 173 = 0Xad
58321de8 5 L 121 = 0X79
583222e0 6 L 247 = 0Xf7
states: Exit Idle Chunk Lru
C:\\\\Informix>onstat -D
Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:39:09 -- 1379136
Kbytes
Dbspaces
address number flags fchunk nchunks flags owner name
5831e150 1 1 1 1 N informix rootdbs
586e7568 2 2001 2 1 N T informix tmp0
586e7628 3 2001 3 1 N T informix tmp1
586e76e8 4 2001 4 1 N T informix tmp2
586e77a8 5 2001 5 1 N T informix tmp3
586e7868 6 2001 6 1 N T informix tmp4
586e7928 7 2001 7 1 N T informix tmp5
586e79e8 8 2001 8 1 N T informix tmp6
586e7aa8 9 2001 9 1 N T informix tmp7
586e7b68 10 2001 10 1 N T informix tmp8
586e7c28 11 2001 11 1 N T informix tmp9
586e7ce8 12 2001 12 1 N T informix tmp10
586e7da8 13 2001 13 1 N T informix tmp11
586e7e68 14 2001 14 1 N T informix tmp12
586e7f28 15 2001 15 1 N T informix tmp13
586ea018 16 2001 16 1 N T informix tmp14
586ea0d8 17 2001 17 1 N T informix tmp15
586ea198 18 2001 18 1 N T informix tmp16
586ea258 19 2001 19 1 N T informix tmp17
586ea318 20 2001 20 1 N T informix tmp18
586ea3d8 21 2001 21 1 N T informix tmp19
586ea498 22 2001 22 1 N T informix tmp20
586ea558 23 2001 23 1 N T informix tmp21
586ea618 24 2001 24 1 N T informix tmp22
586ea6d8 25 2001 25 1 N T informix tmp23
586ea798 26 2001 26 1 N T informix tmp24
586ea858 27 2001 27 1 N T informix tmp25
586ea918 28 2001 28 1 N T informix tmp26
586ea9d8 29 2001 29 1 N T informix tmp27
586eaa98 30 2001 30 1 N T informix tmp28
586eab58 31 2001 31 1 N T informix tmp29
586eac18 32 1 32 1 N informix logdbs
586eacd8 33 1 33 1 N informix physdbs
586ead98 34 1 34 129 N informix cemdo
34 active, 2047 maximum
Hi,
Where is the rest of onstat -D? We need to see writes per chunk.
Also can you send onstat -g ioq?
I would increase CLEANERS to 32 and check again.
This looks that there is only 1 dbspace for data/index/blobs is that correct?
Could it be that all dirty pages at the time the checkpoint starts are all in
the same chunk?
That would explain the large checkpoint times.
Also go into Windows Performance Monitor and enable the following counters
Physical Disk -> Avg. Disk sec/Read
Physical Disk -> Avg. Disk sec/Transfer
Physical Disk -> Avg. Disk sec/Write
Physical Disk -> Avg. Disk Write Queue Length
and check what you get during checkpoints.
Regards,
David.
> On 15 May 2017 at 22:30 GUSTAVO ECHENIQUE <gustavo.echenique@cemdo.com.ar>
wrote:
>
>
> Hi Art!
>
> I send you the copy of the commands that you requested.
>
> C:\\\\Informix>onstat -g iov
>
> Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:20:41 -- 1379136
> Kbytes
>
> AIO I/O vps:
> class/vp s io/s totalops dskread dskwrite dskcopy wakeups io/wup errors
> kio 0 s 54.2 1041585 125284 916301 0 205791447 0.0 0
> kio 1 i 61.3 1177393 239878 937515 0 3581357382 -0.0 0
> kio 2 i 67.3 1294468 329623 964845 0 468431746 0.0 0
> kio 3 i 69.8 1340824 373310 967514 0 736248903 0.0 0
> kio 4 i 331.2 6366742 5332708 1034034 0 3516245753 -0.0 0
> kio 5 i 78.6 1510065 357187 1152878 0 595532308 0.0 0
> kio 6 i 61.5 1181848 263363 918485 0 459739069 0.0 0
> msc 0 i 0.0 27 0 0 0 28 1.0 0
> aio 0 i 0.0 332 64 0 0 333 1.0 0
> pio 0 i 0.0 0 0 0 0 1 0.0 0
> lio 0 i 0.0 0 0 0 0 1 0.0 0
>
> C:\\\\Informix>onstat -F
>
> Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:37:59 -- 1379136
> Kbytes
>
> Fg Writes LRU Writes Chunk Writes
> 82305 6231910 2447839
>
> address flusher state data
> 58320510 0 L 179 = 0Xb3
> 58320a08 1 L 243 = 0Xf3
> 58320f00 2 L 71 = 0X47
> 583213f8 3 L 45 = 0X2d
> 583218f0 4 L 173 = 0Xad
> 58321de8 5 L 121 = 0X79
> 583222e0 6 L 247 = 0Xf7
>
> states: Exit Idle Chunk Lru
>
> C:\\\\Informix>onstat -D
>
> Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:39:09 -- 1379136
> Kbytes
>
> Dbspaces
> address number flags fchunk nchunks flags owner name
> 5831e150 1 1 1 1 N informix rootdbs
> 586e7568 2 2001 2 1 N T informix tmp0
> 586e7628 3 2001 3 1 N T informix tmp1
> 586e76e8 4 2001 4 1 N T informix tmp2
> 586e77a8 5 2001 5 1 N T informix tmp3
> 586e7868 6 2001 6 1 N T informix tmp4
> 586e7928 7 2001 7 1 N T informix tmp5
> 586e79e8 8 2001 8 1 N T informix tmp6
> 586e7aa8 9 2001 9 1 N T informix tmp7
> 586e7b68 10 2001 10 1 N T informix tmp8
> 586e7c28 11 2001 11 1 N T informix tmp9
> 586e7ce8 12 2001 12 1 N T informix tmp10
> 586e7da8 13 2001 13 1 N T informix tmp11
> 586e7e68 14 2001 14 1 N T informix tmp12
> 586e7f28 15 2001 15 1 N T informix tmp13
> 586ea018 16 2001 16 1 N T informix tmp14
> 586ea0d8 17 2001 17 1 N T informix tmp15
> 586ea198 18 2001 18 1 N T informix tmp16
> 586ea258 19 2001 19 1 N T informix tmp17
> 586ea318 20 2001 20 1 N T informix tmp18
> 586ea3d8 21 2001 21 1 N T informix tmp19
> 586ea498 22 2001 22 1 N T informix tmp20
> 586ea558 23 2001 23 1 N T informix tmp21
> 586ea618 24 2001 24 1 N T informix tmp22
> 586ea6d8 25 2001 25 1 N T informix tmp23
> 586ea798 26 2001 26 1 N T informix tmp24
> 586ea858 27 2001 27 1 N T informix tmp25
> 586ea918 28 2001 28 1 N T informix tmp26
> 586ea9d8 29 2001 29 1 N T informix tmp27
> 586eaa98 30 2001 30 1 N T informix tmp28
> 586eab58 31 2001 31 1 N T informix tmp29
> 586eac18 32 1 32 1 N informix logdbs
> 586eacd8 33 1 33 1 N informix physdbs
> 586ead98 34 1 34 129 N informix cemdo
> 34 active, 2047 maximum
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
OK, I see that you have KAIO threads running so that and the fact the the
io/wkup for the single aio vp is 1.0 would indicate that your chunks are
all RAW. That's good. You could take advantage of a second aio vp, but that
has nothing to do with checkpoint performance.
Looking at the onstat -F output I see that about 71% or your writes are LRU
writes and 27% are chunk writes at checkpoint time. Normally that's a good
balance. But I also see that about 1% of your writes are foreground writes
and those are slowing your server down. The cause is not enough CLEANERS.
You have 128 LRU queues that are flushing almost constantly and only 7
CLEANER threads to process the IO requests. This is indirectly affecting
checkpoint processing times be leaving more dirty pages at checkpoint time
than should be there. Increase CLEANERS to the maximum value of 128.
Another thing that is slowing you down in that your data dbspace has 129
chunks. If more than 7 of those are actively being written to there are not
enough CLEANER threads to efficiently flush those chunks at checkpoint time
and THAT IS directly affecting checkpoint duration. The increase in
CLEANERS to 128 should alleviate this problem somewhat. However, the data
disk is RAID5 which beside being inherently unsafe (see my paper on RAID5
on my web site at:
http://www.askdbmgt.com/why-raid5-should-be-avoided-at-all-costs.html ) is
also more than 50% slower than RAID10 to write to! Investing in more and
faster drives configured for RAID10 would help tremendously.
Finally, as suggested by others, reduce the LRU_MAX_DIRTY from 5 to 2. I
would not set LRU_MIN_DIRTY to zero. That has cause problems on some sites.
Finally, I would make two strong recommendations:
1. Move your database server to Linux which will outperform Windows
every time, especially for databases.
2. Move your data to write optimized flash drives rather than spindles.
They are 3 or more times faster and speed is what you need.
3. Upgrade to the latest Informix release (12.10.FC8W2). Based on the
size of your current system, you could upgrade to version12 Workgroup
edition. With a sockets license which is only about $18,000 per socket for
up to two sockets, it is affordable for most organizations. The current
versions are also faster than v7.31 was.
Art
Art S. Kagel, President and Principal Consultant
ASK Database Management
www.askdbmgt.com
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on the IIUG, nor any other organization with which I am
associated either explicitly, implicitly, or by inference. Neither do
those opinions reflect those of other individuals affiliated with any
entity with which I am affiliated nor those of the entities themselves.
On Mon, May 15, 2017 at 4:30 PM, GUSTAVO ECHENIQUE <
gustavo.echenique@cemdo.com.ar> wrote:
> Hi Art!
>
> I send you the copy of the commands that you requested.
>
> C:\\\\Informix>onstat -g iov
>
> Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:20:41 --
> 1379136
> Kbytes
>
> AIO I/O vps:
> class/vp s io/s totalops dskread dskwrite dskcopy wakeups io/wup errors
> kio 0 s 54.2 1041585 125284 916301 0 205791447 0.0 0
> kio 1 i 61.3 1177393 239878 937515 0 3581357382 -0.0 0
> kio 2 i 67.3 1294468 329623 964845 0 468431746 0.0 0
> kio 3 i 69.8 1340824 373310 967514 0 736248903 0.0 0
> kio 4 i 331.2 6366742 5332708 1034034 0 3516245753 -0.0 0
> kio 5 i 78.6 1510065 357187 1152878 0 595532308 0.0 0
> kio 6 i 61.5 1181848 263363 918485 0 459739069 0.0 0
> msc 0 i 0.0 27 0 0 0 28 1.0 0
> aio 0 i 0.0 332 64 0 0 333 1.0 0
> pio 0 i 0.0 0 0 0 0 1 0.0 0
> lio 0 i 0.0 0 0 0 0 1 0.0 0
>
> C:\\\\Informix>onstat -F
>
> Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:37:59 --
> 1379136
> Kbytes
>
> Fg Writes LRU Writes Chunk Writes
> 82305 6231910 2447839
>
> address flusher state data
> 58320510 0 L 179 = 0Xb3
> 58320a08 1 L 243 = 0Xf3
> 58320f00 2 L 71 = 0X47
> 583213f8 3 L 45 = 0X2d
> 583218f0 4 L 173 = 0Xad
> 58321de8 5 L 121 = 0X79
> 583222e0 6 L 247 = 0Xf7
>
> states: Exit Idle Chunk Lru
>
> C:\\\\Informix>onstat -D
>
> Informix Dynamic Server Version 7.31.TD6 -- On-Line -- Up 05:39:09 --
> 1379136
> Kbytes
>
> Dbspaces
> address number flags fchunk nchunks flags owner name
> 5831e150 1 1 1 1 N informix rootdbs
> 586e7568 2 2001 2 1 N T informix tmp0
> 586e7628 3 2001 3 1 N T informix tmp1
> 586e76e8 4 2001 4 1 N T informix tmp2
> 586e77a8 5 2001 5 1 N T informix tmp3
> 586e7868 6 2001 6 1 N T informix tmp4
> 586e7928 7 2001 7 1 N T informix tmp5
> 586e79e8 8 2001 8 1 N T informix tmp6
> 586e7aa8 9 2001 9 1 N T informix tmp7
> 586e7b68 10 2001 10 1 N T informix tmp8
> 586e7c28 11 2001 11 1 N T informix tmp9
> 586e7ce8 12 2001 12 1 N T informix tmp10
> 586e7da8 13 2001 13 1 N T informix tmp11
> 586e7e68 14 2001 14 1 N T informix tmp12
> 586e7f28 15 2001 15 1 N T informix tmp13
> 586ea018 16 2001 16 1 N T informix tmp14
> 586ea0d8 17 2001 17 1 N T informix tmp15
> 586ea198 18 2001 18 1 N T informix tmp16
> 586ea258 19 2001 19 1 N T informix tmp17
> 586ea318 20 2001 20 1 N T informix tmp18
> 586ea3d8 21 2001 21 1 N T informix tmp19
> 586ea498 22 2001 22 1 N T informix tmp20
> 586ea558 23 2001 23 1 N T informix tmp21
> 586ea618 24 2001 24 1 N T informix tmp22
> 586ea6d8 25 2001 25 1 N T informix tmp23
> 586ea798 26 2001 26 1 N T informix tmp24
> 586ea858 27 2001 27 1 N T informix tmp25
> 586ea918 28 2001 28 1 N T informix tmp26
> 586ea9d8 29 2001 29 1 N T informix tmp27
> 586eaa98 30 2001 30 1 N T informix tmp28
> 586eab58 31 2001 31 1 N T informix tmp29
> 586eac18 32 1 32 1 N informix logdbs
> 586eacd8 33 1 33 1 N informix physdbs
> 586ead98 34 1 34 129 N informix cemdo
> 34 active, 2047 maximum
>
>
> ************************************************************
> *******************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
Dear Friends:
Many thanks to all who helped me, such as Art Kagel, Jacques Renaut, Andreas
Legner, etc.
My configuration was as follows:
#############################################################
# Shared Memory Parameters
LOCKS 200000 # Maximum number of locks
BUFFERS 300000 # Maximum number of shared buffers
NUMAIOVPS 2 # Number of IO vps
PHYSBUFF 128 # Physical log buffer size (Kbytes)
LOGBUFF 128 # Logical log buffer size (Kbytes)LOGSMAX 20 # Maximum number of logical log files
CLEANERS 128 # Number of buffer cleaner processes
SHMBASE 0xc000000 # Shared memory base address
SHMVIRTSIZE 131072 # initial virtual shared memory segment size
SHMADD 16384 # Size of new shared memory segments (Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 128 # Number of LRU queues
LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 1 # LRU percent dirty end cleaning limit
LTXHWM 50 # Long transaction high water mark percentage
LTXEHWM 60 # Long transaction high water mark (exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
##################################################################
I added an additional AIO, as Art and Jacques told me, I also changed the
amount of cleaners to 128 as Art suggested, and changed the configuration of
the RAID 5 array to RAID 10.
Blocking checkpoints continue to appear, less durable than before, but the
maximum was almost 4 minutes.
I clarify that I tested this configuration with a dbimport of a base. Could
this have been the problem?
Anyway, I can not understand that the IDS does not "know" that it has 15K rpm,
and 6Gbps disks, however old the version of it.
I say this because on a server that has 10K rpm disks, and only two RAID 5
disk arrays (one for the operating system and the other for the database,
temporary files, logical logs, physical log), it takes practically the same
thing To perform the dbimport.
But that's what least worries me. What really worries me is that certain
processes take exactly the same thing as on a Windows 2000 server, with 2Gbps
disks.
I am disconcerted.
Gustavo, how about considering going version 12.10 ? You are losing a bunch of new features that definately are very interesting. You missed 9.30, 9.40 10.00 11.10, 11.50, 11.70 and 12.10 which is at fixpack # 9 In all this, non-blocking checkpoints have been introduced long time ago. But now you have great new features like deep embedding of MongoDB flavor NoSql, query sharding, a great and very complete replication model and so many more. Pay a visit to https://www.ibm.com/support/knowledgecenter/en/SSGU8G_12.1.0/com.ibm.po.doc/new_ features_ce.htm No Informix is not dead at all, you will see Rgds Eric
Related threads
- onbar -c -F in Windows Informix instance
- Anyone... SQLCODE=-668, ISAM error=-1
- Not using the 100% logical log page size alloacted to informix