RE: Checkpoint Question - need feedback
Posted in 2006
Topics: High Availability & Replication, SQL Development & Query Writing, Server Administration, Logging & Checkpoints
Hi, everybody, Sorry for joining somewhat late to this discussion... As for the physical log size, I don't see any problem at all. In a real environment, physical log usually resides on a dedicated hard drive. It's not so easy to buy a drive less then 70GB now days... And most of it is wasted anyway. What really bothers me is an amount of atomic I/O operation. New logging algorithm should not increase the I/O rate (operation/sec), while it can increase the overall write speed (bytes/sec). With the new implementation, it should be possible to choose between old non-fuzzy checkpoint mechanism and a new one. Also, along with fuzzy checkpoints and fractional lru min/max, checkpoint duration can be improved through the use of a write-back cache in teh disk array... The real checkpoint problem is a checkpoint in HDR environment. With HDR, checkpoints are always synchronous and never fuzzy. With HDR, it is very easy to run into a long checkpoint by simply running some read-only queries on Secondary, that consume a lot of I/O bandwitht, that can otherwise be used for data synchronization between checkpoints. In my opinion, it is very important to address the HDR checkpont issue in the new non-blocking checkpoint algorithm. For me, in non-HDR environment checkpoint duration is not a problem at all: one should just properly configure caching in the disk array (now many people are still using bare SCSI drives nowdays?) and properly protect array's write-back cache. -Alexey -----Original Message----- From: informix-list-bounces@iiug.org [mailto:informix-list-bounces@iiug.org] On Behalf Of Madison Pruet Sent: Wednesday, July 05, 2006 6:22 PM To: informix-list@iiug.org Subject: Checkpoint Question - need feedback I've got a question for the user community. Checkpoints are a real pain because there is a period of time in which user threads are going to be blocked. I've been playing around with an idea in which I think I could do non-blocking checkpoints. The cost of doing this, however, would be that there would probably need to be a significant increase in the size of the physical log file. I know that implementing the fractional LRU min/max has helped reduce the impact of the checkpoint, but the fractional LRU min/max is not a guarantee - since it is possible that the LRU page writers won't be able to keep up with the current activity. Also - this idea would eliminate fuzzy checkpoints, which is probably a good thing since fuzzy checkpoints do impact the recovery time of the server. So - basic question --- Is the cost of the increased physical log file (maybe 3-4 times larger in some cases) totally outweigh the benefit of non-blocking checkpoints? _______________________________________________ Informix-list mailing list Informix-list@iiug.org http://www.iiug.org/mailman/listinfo/informix-list
Alexey Sonkin wrote:
> Hi, everybody,
>
> Sorry for joining somewhat late to this discussion...
>
>
> So - basic question --- Is the cost of the increased physical log file
> (maybe 3-4 times larger in some cases) totally outweigh the benefit of
> non-blocking checkpoints?
>
Yes.
And for us now it would be nice if the checkpoint code.
a) tried to flush all dirty pages before starting the checkpoint.
(onmode -B anyone?).
b) checked if the number of dirty buffers was > a configurable value
if so flush again - value set to 0 to disable this functionality
c) then start the normal checkpoint i.e. block all threads etc.
If I have 100,000 dirty buffers at least try to flush them and maybe
try again
in case you got unlucky then do a checkpoint.
That way hopefully users with 5-15 minute checkpoint intervals on
OLTP systems might avoid getting 10 second checkpoints whilst user
screens are frozen!
Step a) would flush most stuff and step b) catches an unlucky situation
where users do a lot of updates in one go. Hopefully users cannot type
so
fast that c) hangs around for too long!