Re: Long checkpoints
Posted in 2008
Topics: High Availability & Replication, Backup & Restore, Installation, Setup & Upgrades, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues, Versions, Editions & End-of-Life
On Sep 10, 4:47 pm, Obnoxio The Clown <obno...@serendipita.com> wrote:
> askel wrote:
> > Hello there,
>
> > I've been reading about long checkpoints issue for a while but haven't
> > found solution for our case.
>
> > We're using IDS 9.40.FC6 on Solaris 9 for OLTP. It is Sun Fire V1280
> > with 8 UltraSPARC-III+ 1.2 GHz and 16 Gb of RAM on board. The database
> > is not large (compressed level 0 backup is about 6 Gb). Some time ago
> > checkpoint duration went from 1-3 sec up to 30-70 sec. It is happening
> > during high load on database. During night and slow time checkpoints
> > still stay at low 2-3 sec. Anything that overlaps with checkpoint is
> > slow as hell. Update of few records selected by primary key takes few
> > seconds. More complicated queries takes much longer during checkpoint
> > time.
>
> > Could anybody help me with solving that. We tried to renew our support
> > contract but we were asked to upgrade Informix to at least IDS 10
> > which is not possible at the moment for several reasons.
>
> > I'd appreciate any help.
>
> > Here is onstat -p output (last onstat -z was executed about a day
> > ago):
>
> > IBM Informix Dynamic Server Version 9.40.FC6 -- On-Line -- Up 3
> > days 12:13:19 -- 2750464 Kbytes>
> > Profile
> > dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> > 43922482 68439349 11047096683 99.60 2249362 6528629 774619170
> > 99.71
>
> > isamtot open start read write rewrite delete
> > commit rollbk
> > 9137221980 176193695 605900359 6361264370 412085885 2959249 8019697
> > 518554 512
>
> > gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> > 2 0 0 0 0 0 0
>
> > ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> > 0 0 0 167642.70 25625.88 295 590
>
> > bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress
> > seqscans
> > 3463504 580 2288152479 0 0 3236 3094299
> > 6164708
>
> > ixda-RA idx-RA da-RA RA-pgsused lchwaits
> > 3923244 2397751 5165461 11407339 1532876
>
> > Here is onconfig content:
>
> > ROOTNAME rootdbs # Root dbspace name> > ROOTPATH /opt/informix/links/rootdbs1
> > # Path for device containing root
> > dbspace
> > ROOTOFFSET 0 # Offset of root dbspace into device
> > (Kbytes)
> > ROOTSIZE 2048000 # Size of root dbspace (Kbytes)>
> > # Disk Mirroring Configuration Parameters
>
> > MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
> > MIRRORPATH # Path for device containing mirrored> > root
> > MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>
> > # Physical Log Configuration
>
> > PHYSDBS rootdbs # Location (dbspace) of physical log
> > PHYSFILE 204800 # Physical log file size (Kbytes)>
> > # Logical Log Configuration
>
> > LOGFILES 76 # Number of logical log files
> > LOGSIZE 25000 # Logical log size (Kbytes)>
> > # Diagnostics
>
> > MSGPATH /var/log/online.log # System message log file path
> > CONSOLE /dev/console # System console message path>
> > # To automatically backup logical logs, edit alarmprogram.sh and set
> > # BACKUPLOGS=Y
> > ALARMPROGRAM /opt/informix/etc/alarmprogram.sh # Alarm program path
> > TBLSPACE_STATS 0 # Maintain tblspace statistics>
> > # System Archive Tape Device
>
> > LTAPEDEV /opt/informix/links/ltapedev # Log tape device path
> > LTAPEBLK 16 # Log tape block size (Kbytes)
> > LTAPESIZE 102400000 # Max amount of data to put on log
> > tape (Kbytes)>
> > # Optical
>
> > STAGEBLOB # Informix Dynamic Server staging
> > area
>
> > # System Configuration
>
> > SERVERNUM 0 # Unique id corresponding to a OnLine> > instance
> > DBSERVERNAME ca2 # Name of default database server
> > DBSERVERALIASES ca2_tcp,scdb2 # List of alternate dbservernames
> > NETTYPE ipcshm,1,50,CPU # Configure poll thread(s) for nettype
> > NETTYPE tlitcp,4,100,NET # Configure poll thread(s) for> > nettype
> > DEADLOCK_TIMEOUT 60 # Max time to wait of lock in> > distributed env.
> > RESIDENT 0 # Forced residency flag (Yes = 1, No =
> > 0)
>
> > MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-> > processor
> > NUMCPUVPS 7 # Number of user (cpu) vps
> > SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> > to one
>
> > NOAGE 1 # Process aging
> > AFF_SPROC 1 # Affinity start processor
> > AFF_NPROCS 7 # Affinity number of processors>
> > # Shared Memory Parameters
>
> > LOCKS 64000 # Maximum number of locks
> > BUFFERS 1024000 # Maximum number of shared buffers
> > NUMAIOVPS 11 # Number of IO vps
> > PHYSBUFF 2048 # Physical log buffer size (Kbytes)
> > LOGBUFF 2048 # Logical log buffer size (Kbytes)
> > CLEANERS 128 # Number of buffer cleaner processes
> > SHMBASE 0xa000000 # Shared memory base address
> > SHMVIRTSIZE 512000 # initial virtual shared memory> > segment size
> > SHMADD 8192 # Size of new shared memory segments
> > (Kbytes)
> > SHMTOTAL 0 # Total shared memory (Kbytes).
> > 0=>unlimited
> > CKPTINTVL 600 # Check point interval (in sec)
> > LRUS 128 # Number of LRU queues
> > LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
> > LRU_MIN_DIRTY 0 # LRU percent dirty end cleaning limit
> > TXTIMEOUT 0x12c # Transaction timeout (in sec)
> > STACKSIZE 64 # Stack size (Kbytes)
>
> > DYNAMIC_LOGS 2
> > LTXHWM 70
> > LTXEHWM 80
> > OFF_RECVRY_THREADS 10 # Default number of offline worker> > threads
> > ON_RECVRY_THREADS 1 # Default number of online worker> > threads
>
> > # Data Replication Variables
> > DRINTERVAL 30 # DR max time between DR buffer
> > flushes (in sec)
> > DRTIMEOUT 30 # DR network timeout (in sec)
> > DRLOSTFOUND /opt/informix/etc/dr.lostfound # DR lost+found file> > path
>
> > # CDR Variables
> > CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-
> > vp,additional)
> > CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
> > CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
> > queue (Kbytes)
> > CDR_NIFCOMPRESS 0 # Link level compression (-1 never, 0
> > none, 9 max)
>
> > CDR_SERIAL 0,0 # Serial Column Sequence
> > CDR_DBSPACE # dbspace for syscdr d
askel wrote:
> On Sep 10, 4:47 pm, Obnoxio The Clown <obno...@serendipita.com> wrote:
>
>> askel wrote:
>>
>>> Hello there,
>>>
>>> I've been reading about long checkpoints issue for a while but haven't
>>> found solution for our case.
>>>
>>> We're using IDS 9.40.FC6 on Solaris 9 for OLTP. It is Sun Fire V1280
>>> with 8 UltraSPARC-III+ 1.2 GHz and 16 Gb of RAM on board. The database
>>> is not large (compressed level 0 backup is about 6 Gb). Some time ago
>>> checkpoint duration went from 1-3 sec up to 30-70 sec. It is happening
>>> during high load on database. During night and slow time checkpoints
>>> still stay at low 2-3 sec. Anything that overlaps with checkpoint is
>>> slow as hell. Update of few records selected by primary key takes few
>>> seconds. More complicated queries takes much longer during checkpoint
>>> time.
>>>
>>> Could anybody help me with solving that. We tried to renew our support
>>> contract but we were asked to upgrade Informix to at least IDS 10
>>> which is not possible at the moment for several reasons.
>>>
>>> I'd appreciate any help.
>>>
>>> Here is onstat -p output (last onstat -z was executed about a day
>>> ago):
>>>
>>> IBM Informix Dynamic Server Version 9.40.FC6 -- On-Line -- Up 3
>>> days 12:13:19 -- 2750464 Kbytes>>>
>>> Profile
>>> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
>>> 43922482 68439349 11047096683 99.60 2249362 6528629 774619170
>>> 99.71
>>>
>>> isamtot open start read write rewrite delete
>>> commit rollbk
>>> 9137221980 176193695 605900359 6361264370 412085885 2959249 8019697
>>> 518554 512
>>>
>>> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
>>> 2 0 0 0 0 0 0
>>>
>>> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
>>> 0 0 0 167642.70 25625.88 295 590
>>>
>>> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress
>>> seqscans
>>> 3463504 580 2288152479 0 0 3236 3094299
>>> 6164708
>>>
>>> ixda-RA idx-RA da-RA RA-pgsused lchwaits
>>> 3923244 2397751 5165461 11407339 1532876
>>>
>>> Here is onconfig content:
>>>
>>> ROOTNAME rootdbs # Root dbspace name>>> ROOTPATH /opt/informix/links/rootdbs1
>>> # Path for device containing root
>>> dbspace
>>> ROOTOFFSET 0 # Offset of root dbspace into device
>>> (Kbytes)
>>> ROOTSIZE 2048000 # Size of root dbspace (Kbytes)>>>
>>> # Disk Mirroring Configuration Parameters
>>>
>>> MIRROR 0 # Mirroring flag (Yes = 1, No = 0)
>>> MIRRORPATH # Path for device containing mirrored>>> root
>>> MIRROROFFSET 0 # Offset into mirrored device (Kbytes)>>>
>>> # Physical Log Configuration
>>>
>>> PHYSDBS rootdbs # Location (dbspace) of physical log
>>> PHYSFILE 204800 # Physical log file size (Kbytes)>>>
>>> # Logical Log Configuration
>>>
>>> LOGFILES 76 # Number of logical log files
>>> LOGSIZE 25000 # Logical log size (Kbytes)>>>
>>> # Diagnostics
>>>
>>> MSGPATH /var/log/online.log # System message log file path
>>> CONSOLE /dev/console # System console message path>>>
>>> # To automatically backup logical logs, edit alarmprogram.sh and set
>>> # BACKUPLOGS=Y
>>> ALARMPROGRAM /opt/informix/etc/alarmprogram.sh # Alarm program path
>>> TBLSPACE_STATS 0 # Maintain tblspace statistics>>>
>>> # System Archive Tape Device
>>>
>>> LTAPEDEV /opt/informix/links/ltapedev # Log tape device path
>>> LTAPEBLK 16 # Log tape block size (Kbytes)
>>> LTAPESIZE 102400000 # Max amount of data to put on log
>>> tape (Kbytes)>>>
>>> # Optical
>>>
>>> STAGEBLOB # Informix Dynamic Server staging
>>> area
>>>
>>> # System Configuration
>>>
>>> SERVERNUM 0 # Unique id corresponding to a OnLine>>> instance
>>> DBSERVERNAME ca2 # Name of default database server
>>> DBSERVERALIASES ca2_tcp,scdb2 # List of alternate dbservernames
>>> NETTYPE ipcshm,1,50,CPU # Configure poll thread(s) for nettype
>>> NETTYPE tlitcp,4,100,NET # Configure poll thread(s) for>>> nettype
>>> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in>>> distributed env.
>>> RESIDENT 0 # Forced residency flag (Yes = 1, No =
>>> 0)
>>>
>>> MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi->>> processor
>>> NUMCPUVPS 7 # Number of user (cpu) vps
>>> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps>>> to one
>>>
>>> NOAGE 1 # Process aging
>>> AFF_SPROC 1 # Affinity start processor
>>> AFF_NPROCS 7 # Affinity number of processors>>>
>>> # Shared Memory Parameters
>>>
>>> LOCKS 64000 # Maximum number of locks
>>> BUFFERS 1024000 # Maximum number of shared buffers
>>> NUMAIOVPS 11 # Number of IO vps
>>> PHYSBUFF 2048 # Physical log buffer size (Kbytes)
>>> LOGBUFF 2048 # Logical log buffer size (Kbytes)
>>> CLEANERS 128 # Number of buffer cleaner processes
>>> SHMBASE 0xa000000 # Shared memory base address
>>> SHMVIRTSIZE 512000 # initial virtual shared memory>>> segment size
>>> SHMADD 8192 # Size of new shared memory segments
>>> (Kbytes)
>>> SHMTOTAL 0 # Total shared memory (Kbytes).
>>> 0=>unlimited
>>> CKPTINTVL 600 # Check point interval (in sec)
>>> LRUS 128 # Number of LRU queues
>>> LRU_MAX_DIRTY 2 # LRU percent dirty begin cleaning limit
>>> LRU_MIN_DIRTY 0 # LRU percent dirty end cleaning limit
>>> TXTIMEOUT 0x12c # Transaction timeout (in sec)
>>> STACKSIZE 64 # Stack size (Kbytes)
>>>
>>> DYNAMIC_LOGS 2
>>> LTXHWM 70
>>> LTXEHWM 80
>>> OFF_RECVRY_THREADS 10 # Default number of offline worker>>> threads
>>> ON_RECVRY_THREADS 1 # Default number of online worker>>> threads
>>>
>>> # Data Replication Variables
>>> DRINTERVAL 30 # DR max time between DR buffer
>>> flushes (in sec)
>>> DRTIMEOUT 30 # DR network timeout (in sec)
>>> DRLOSTFOUND /opt/informix/etc/dr.lostfound # DR lost+found file>>> path
>>>
>>> # CDR Variables
>>> CDR_EVALTHREADS 1,2 # evaluator threads (per-cpu-
>>> vp,additional)
>>> CDR_DSLOCKWAIT 5 # DS lockwait timeout (seconds)
>>> CDR_QUEUEMEM 4096 # Maximum amount of memory for any CDR
>>> queue (Kbytes)
>>> CDR_NIFCOMPRESS 0 # Link level compression (-1 never, 0
>>> none, 9 max)
>>>
>>> CDR_SERIAL
Obnoxio, Thank you for the advice. It doesn't seem to be dangerous but could you tell me how do I reverse the effect of executing that command or how do I check what is the currently used value? We do have backups but restoring the database is last thing I'd want to do about live database/system. Cheers, Alexander
askel wrote:
> Obnoxio,
>
> Thank you for the advice. It doesn't seem to be dangerous but could
> you tell me how do I reverse the effect of executing that command or
> how do I check what is the currently used value? We do have backups
> but restoring the database is last thing I'd want to do about live
> database/system.
>
onstat -C tells you that the value was 500.
onmode -C threshold 500 will re-set the value.
--
Cheers,
Obnoxio the Clown
http://obotheclown.blogspot.com
askel wrote:
> Obnoxio,
>
> Thank you for the advice. It doesn't seem to be dangerous but could
> you tell me how do I reverse the effect of executing that command or
> how do I check what is the currently used value? We do have backups
> but restoring the database is last thing I'd want to do about live
> database/system.
>
> Cheers,
> Alexander
You have several interesting things in your configuration like:
LRU_MIN_DIRTY 07 CPUVPs and 11 AIOVPs for a compressed 6GB database where you have about 2GB
in buffer cache, 2MB for phys and logical log buffer...
I believe that what OTC is saying it that only the Btree scanner could mess up
that instance to the point of getting checkpoints up to 30-70...
Id like to advise you to check your I/O, your system logs, how many dirty
buffers you have at checkpoint start time etc.
Also, are you using cooked files? Are your AIOVPs really working, or just the
first ones (onstat -g iov)?
Also... You don't have "DEFAULT_ATTACH" variable set up do you?
Regards, and tell us if OTC suggestion helped...
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...