AIX 5.3 IDS 11.50 FC5 - Performance Advisory
Posted in 2010
A DBA on AIX 5.3 / IDS 11.50.FC5 saw a "Performance Advisory: Logical log file size might be too small for a checkpoint to complete" in online.log, but only during the first Auto Update Statistics refresh following the weekly system reboot (logs were filling within seconds during that burst of activity). Posters discussed CKPTINTVL and checkpoint behaviour; an IBM specialist noted that from 11.10 onward checkpointing was reworked (no more fuzzy checkpoints, now blocking/non-blocking), and non-blocking checkpoints need more log space, so larger logical logs avoid the warning — see chapter 15 of the Administrator's Guide. No confirmation that the poster applied a fix is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Storage & Space Management, Logging & Checkpoints
Have any of you experienced this? In reviewing the online.log, I discovered that AFTER a system reboot, and ONLY AFTER a system reboot, the first time the "Auto Update Statistics Refresh" begins, the following performance advisory is recorded. The Backup occurs nightly, the AUS Evaluation & Refresh occurs nightly and the reboot occurs weekly. Is this something I should be concerned about? 22:31:40 Level 0 Archive started on rootdbs, ordbs, oridxdbs, clerkdbs, ccisdbs, ccisidxdbs, criminalidxdbs, casemgtdbs, casemgtidxdbs, trafficdbs, jurydbs, imagesdbs, criminaldbs, reportsdbs, crimdbs, crimidxdbs, physlogdbs, logdbs01, logdbs02, logdbs03, clerkidxdbs, imagesidxdbs, juryidxdbs, sbdefaultdbs, reportsbs 22:47:30 Archive on rootdbs, ordbs, oridxdbs, clerkdbs, ccisdbs, ccisidxdbs, criminalidxdbs, casemgtdbs, casemgtidxdbs, trafficdbs, jurydbs, imagesdbs, criminaldbs, reportsdbs, crimdbs, crimidxdbs, physlogdbs, logdbs01, logdbs02, logdbs03, clerkidxdbs, imagesidxdbs, juryidxdbs, sbdefaultdbs, reportsbs Completed. Wed Mar 31 01:00:57 2010 01:00:57 Logical Log 76980 Complete, timestamp: 0x937cfb82. 01:00:58 Logical Log 76980 - Backup Started 01:00:58 Logical Log 76980 - Backup Completed 01:16:03 Logical Log 76981 Complete, timestamp: 0x937e4c5b. 01:16:03 Logical Log 76981 - Backup Started 01:16:04 Logical Log 76981 - Backup Completed 01:16:10 Performance Advisory: Logical log file size might be too small for a checkpoint to complete. 01:16:10 Results: The size of individual logical log files is too small for the current workload, resulting in each log file filling very quickly. If log files fill in less than 30 seconds, the checkpoint might remain blocked because the last log file fills during the time needed to perform the checkpoint. 01:16:10 Action: Increase the size of the individual logical log files so that it takes at least 30 seconds to fill each one. Look at the online log to determine how quickly the log files are filling, and then increase the size of the files proportionately. 01:16:10 Logical Log 76982 Complete, timestamp: 0x937f1b7e. 01:16:11 Logical Log 76982 - Backup Started 01:16:11 Logical Log 76982 - Backup Completed 01:16:34 Logical Log 76983 Complete, timestamp: 0x938038a2. 01:16:34 Logical Log 76983 - Backup Started 01:16:35 Logical Log 76983 - Backup Completed 01:46:43 Logical Log 76984 Complete, timestamp: 0x938193b3. 01:46:43 Logical Log 76984 - Backup Started 01:46:43 Logical Log 76984 - Backup Completed 02:31:49 SCHAPI: Warning!!! dbWorker thread is cleaning up temp tables for task 21.
What?.. Wow!.. I've heard of log files growing too big to where they threaten available space, but not guaranteeing a checkpoint becuase they're too small, needing a bigger size (>30 seconds)? Hmm!.. I'm wondering if IDS should provide some kind of dynamic log file management. CJIS is a mission-critical app which needs full transaction integrity!
The funny thing is the archive and the AUS Refresh execute every night. The only time this message occurs is at the time the AUS Refresh executes after a system reboot. I'm trying to understand why the system reboot affects the logs in this way. :/
Could it be that the database server marks the log as Deleted (D) and drops it when you take a level-0 backup of all the dbspaces?
CKPTINTVL specifies the frequency, expressed in seconds, at which the database server checks to determine whether a checkpoint is needed. When a full checkpoint occurs, all pages in the shared-memory buffer pool are written to disk. When a fuzzy checkpoint occurs, nonfuzzy pages are written to disk, and the page numbers of fuzzy pages are recorded in the logical log. If you set CKPTINTVL to an interval that is too short, the system spends too much time performing checkpoints, and the performance of other work suffers. If you set CKPTINTVL to an interval that is too long, fast recovery might take too long. In practice, 30 seconds is the smallest interval that the database server checks. If you specify a checkpoint interval of 0, the database server does not check if the checkpoint interval has elapsed. However, the database server still performs checkpoints. Other conditions, such as the physical log becoming 75 percent full, also cause the database server to perform checkpoints.
Hi, Your description is correct for versions before 11.10. At that point the checkpointing algorithm was completely reworked. There are no more "fuzzy" pages or checkpoints. There are non-blocking and blocking checkpoints, depending on a number of factors. One very important factor is that non-blocking checkpoints typically require more log space than ordinary checkpoints had required in earlier releases. If that space is available, then checkpoints are usually not a noticeable interruption, even if they take several seconds. The new algorithms are described in chapter 15 of the IDS Administrator's Guide. Cheers, Dick Snoke Executive IT Specialist IBM Software Group - ChannelWorks Tel: (404) 487-1595 Email: dsnoke@us.ibm.com From: "FRANK COMPUTER" <frank@frankcomputer.com> To: ids@iiug.org Date: 03/31/10 04:46 PM Subject: Re: Performance Advisory - CKPTINTVL ? [19474] Sent by: ids-bounces@iiug.org CKPTINTVL specifies the frequency, expressed in seconds, at which the database server checks to determine whether a checkpoint is needed. When a full checkpoint occurs, all pages in the shared-memory buffer pool are written to disk. When a fuzzy checkpoint occurs, nonfuzzy pages are written to disk, and the page numbers of fuzzy pages are recorded in the logical log. If you set CKPTINTVL to an interval that is too short, the system spends too much time performing checkpoints, and the performance of other work suffers. If you set CKPTINTVL to an interval that is too long, fast recovery might take too long. In practice, 30 seconds is the smallest interval that the database server checks. If you specify a checkpoint interval of 0, the database server does not check if the checkpoint interval has elapsed. However, the database server still performs checkpoints. Other conditions, such as the physical log becoming 75 percent full, also cause the database server to perform checkpoints. ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.