RE: CLEANERS
Posted in 2000
A DBA found that raising CLEANERS to 128 made checkpoints take about a minute every six hours. Andrew Hamm argued the cause is disk I/O contention: total throughput per disk rises only up to roughly 5-7 concurrent I/O threads and then collapses, so too many page cleaners hammering chunks hurts checkpoint time. He posted Informix's 'pfread' C benchmark plus a ksh wrapper to measure per-device throughput at varying thread counts, and advice on balancing chunks/extents across disks. The original poster reported that cutting CLEANERS from 128 to 80 eliminated the long checkpoints; his follow-up question about running pfread safely on a live production box went unanswered.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Logging & Checkpoints
I am not aware of pwrite can you point me in the right direction Cheers Barry > -----Original Message----- > From: Andrew Hamm [SMTP:ahamm@sanderson.net.au] > Sent: 07 December 2000 00:22 > To: informix-list@iiug.org > Subject: Re: CLEANERS > > hanna_shaw@my-deja.com wrote in message <90m20c$m3k$1@nnrp1.deja.com>... > > > >But I found that checkpoint duration is increased after CLEANERS > >increased to 128. It looks like that checkpoint take 1 minute every 6 > >hours. I never had the problem before. It's really scary for > >performance is the most important thing in the server.(The reponse time > >should be 2-3 seconds.) > > You'll probably find that the outrageous checkpoint time is due to massive > contention for disk activity. Are you aware of the little Informix utility > pwrite that is handed out in tuning courses? Pay attention to the > consequences of the measures from that... It's immutable. > >
Barry Tomlinson wrote in message <90nrfq$ng2$1@news.xmission.com>... > >I am not aware of pwrite >can you point me in the right direction Strictly speaking, pwrite is offered as pread, but if you are brave you change a read() into a write() and you can also test write performance. Here's the poo: pread is a little C program which deliberately hits a disk division with as much read activity as is humanely possible (should that be androidly possible?). [I hope you can all see the danger of changing the process into pwrite] This simple process allows you to measure the I/O performance of the disk. Now the really clever part goes like this: <SNIP out of an internal document> A (perhaps) strange but true fact about disks is that total throughput on a disk actually increases as you increase the number of I/O threads upto a certain limit. This does not mean that it improves the throughput of an individual thread, in fact the performance of each thread drops as you would expect - we are all familiar with the concept of reducing the amount of seek time by physically sequential writes. However, typical measurements I've seen show that between 5 and 7 threads are the optimum number, and that's for all machines I've tested so far. That includes HP machinery as well as Alphas, IBM and even lowly SCO boxes. Why? Read on: Typical relative numbers for throughput are: 1 200 # 1 thread, thruput: 200 .. 5 750 # 5 threads, ... 6 800 7 800 8 700 9 500 If you divide the total thoughput up for individual threads, you will indeed see that at 6 or 7, each thread is getting a bit over 100, which is clearly worse than the 200 available to a single thread. However, that result of 800 TOTAL throughput is what we are can exploit. The technique is, divide the disk into 6 or 7 divisions. I prefer to choose on the low side (ie 5), because there is often also a UNIX partition on any single disk, and also one day in the future you may need to add another division in the leftover space, and you will note that the performance falls of quite seriously once you go over the peak, so it's definitely a good idea to plan for that contingency. Now that we have measured each of our disks and have made decisions about how many divisions, the next task is to keep each of those divisions as busy as each other, across all the disks. This is a challenging operation, but thankfully the Online 7 engines provide a set of statistics which record the actual disk activity. Over about 18 months I have been working on a script which uses this actual information to supply a set of table distributions. The difficult part of the process was the compromise between table activity vs. table size. Although you might imagine that a big table would get more activity, in practice (with the statistics I have collected) it's not completely true although there is some correlation. Another tricky aspect of the task is trying to predict future growth, so that you may allocate appropriate table extents that allow for a few years activity. One big performance killer with Informix is to have tables which are broken up into many small regions on disk, so it's very important to predict good extent sizes. I've attempted to integrate the existing statistics about extent scattering into the algorithm, but it turns out that it's a fairly unreliable factor so recently I've reduced its weighting in the algorithm. Further improvements can come from considering "lazy" tables - what I mean by this is merely tables with minimal throughput. Thoughtful fragmentation of large tables is also exploitable. If you have fragmented some tables so that the "busy area" is small and the non-busy area is large (a classical example: a working table with live documents that quickly become historical and yet hang around forever due to legal or reporting requirements) then you could consider those fragments to be "lazy" too. If you can identify your lazy tables or fragments, you could probably add a few lazy chunks to take the really sleepy tables or table fragments. This is an important point considering that disks these days are well and truly exceeding the 2Gb limits currently applied to informix chunks. Another way of exploiting fragmented tables is to consider large tables that could be off-loaded to PDQ activity, so that they won't interfere too much with the ordinary buffers. I suppose a bunch of PDQ threads will kick in every so often with adverse affects on the total number of threads on a disk or 3, but you can judge the potentials of your own system. [Which leads on to the assertions I made about splitting off DSS type activity into warehouse machinery, thereby allowing more purging of the OLTP and better general OLTP performance. This is my exciting new task over the next several months... If you are on a principally DSS system..... sorry! you can't split off any further. But the general thread balancing will still apply] </SNIP> So that's the poo from Informix and the origins of the pwrite (sorry, pread) process. The immutable results of this test is, if you get too many busy threads burning away on a disk, your total throughput will fall through the floor again. The bit I like about checkpoint writes is that the number of threads is highly predictable. LRU writes (and FG writes also) are far less predictable, since the number of active ones depends on the number of LRU queues, and I don't think the cleaners doing LRU writes are polite enough to concentrate on one chunk only. Don't they do work against all chucks represented on the LRU queue? Somebody please contradict me before I am forced to grab that bloody manual again. In the worst case, one disk may be doing most of the work. When you reorder, you may get your other 3, 4, 5, 10 (whatever) disks working almost equally. Given that you can improve the throughput on a disk by a factor of approx 5, and can improve total throughput by the factor of how many more disks actually get busy, it's fairly easy to predict the improvement in checkpoint speeds. Lets guess you get 4 times the number of busy disks, each with an improvement of 3. So we are talking about a 12-fold improvement in checkpoints, and general read activity too. And no doubt LRU writes if you can carefully manage the number of threads. If you already have all your disks busy, but there happen to be too many busy chunks, then you may well have fallen over the "hump" in the performance curve. Reducing the busy chunks should improve performance. If you already have these factors thru dilligence and hard work, congratulations and would you like a job in sunny Australia? Once this layout is sorted, other tuning exercises can be confidently applied knowing that the I/O system is not letting you down. So, I imagine there would be good interest in getting a hold of pread. Does this mailing list support attachments? Or shall I dump the C file and associated scripts directly into another message?
Andrew Hamm wrote in message <3a3020cf$1@news.iprimus.com.au>... >Barry Tomlinson wrote in message <90nrfq$ng2$1@news.xmission.com>... Hmmm - everyone has gone quiet. Is it because the weekend? For a change, Australia is ahead of the world. I'm hoping for some contradictions or challenging questions. Or maybe you folks are eagerly awaiting a copy of pfread. I'll embed a copy and details of usage (that's the slow part that's holding me back) tomorrow. I'd be most interested to get people's opinions of the subject.
Andrew Hamm wrote in message <3a34743e$1@news.iprimus.com.au>... Here's the C code. It's a single file (narrow ledge?) Get him compiled first. Probably trivial exercise for most of you. There is an associated shell script wrapper, which I admit I've been messing with. I'll make sure it's not damaged and send it on with some usage notes. I've been messing with the script so that it .... ummm, .... does a few more things, and if I'm happy with that I'll pass it on too. Stay tooned... <FILE pfread.c> /* * Program to do disk reads for performance measurements * Usage: pfread [-b bufsize] [-n nloops] [-m min] [-M max] * [-v] [-s seed|=seq|file] file * The three read modes are: * -s seed --do random reads, seed the generator using seed * -s =seq --do sequential reads * -s file --read file to get buffer numbers to read * The default is random reads with seed of 0. */ #include <stdio.h> #include <time.h> int bufsize = 2048; /* read buffer size */ long nloops = 1000; /* number of reads */ long min = 0; /* minimum buffer offset to read */ long max = 51200; /* maximum buffer offset to read */ char *file; /* file to read from */ int verbose; /* if set, print buffer numbers being read */ int sequential; /* if set, do sequential reads */ FILE *sfile; /* if set, file sfile to get buffer numbers */ int seed; /* otherwise use seed to do random reads */ int wflag; extern long atol(); void perror_exit(s) char *s; { perror(s); exit(2); } void usage() { fprintf(stderr, "usage: pfread [-b bufsize] [-n nloops] [-m min] [-M max] [-v] [-s seed|=seq|file] file\\n"); exit(1); } long scaled_atol(s) char *s; { int sl; long n; n = atol(s); sl = strlen(s); if (sl) { switch (s[sl-1]) { case 'b': n *= 512; break; case 'k': n *= 1024; break; } } return n; } void getargs(argc, argv) char **argv; { extern char *optarg; extern int optind; int i; while ((i = getopt(argc, argv, "b:n:m:M:vs:")) != -1) { switch (i) { case 'h': case '?': usage(); break; case 'b': bufsize = scaled_atol(optarg); break; case 'n': nloops = atol(optarg); break; case 'm': min = atol(optarg); break; case 'M': max = atol(optarg); break; case 'v': verbose = 1; break; case 'w': wflag = 1; break; case 's': if (sequential || sfile) usage(); if (strcmp(optarg, "=seq") == 0) sequential = 1; else if (sfile = fopen(optarg, "r")) ; else seed = atoi(optarg); break; } } if (optind != (argc-1)) usage(); file = argv[optind]; } main(argc, argv) char **argv; { long i, n; int fd; char *buf, *valloc(); time_t start_time, end_time, elapsed_time; getargs(argc, argv); if ((fd = open(file, 0)) < 0) perror_exit("open"); if ((buf = valloc(bufsize)) == 0) perror_exit("valloc"); if (!sequential && !sfile) srand(seed); n = min; start_time = time(0); for (i = 0; i < nloops; i++) { int nread; /* * Do a read */ if (verbose) printf("%ld\\n", n); if (lseek(fd, (long) n * bufsize, 0) < 0) perror_exit("lseek"); nread = read(fd, buf, bufsize); if (nread < 0) perror_exit("read"); if (nread != bufsize) perror_exit("short read"); /* * Determine next buffer to read */ if (sequential) { n = (n+1) % max; if (n == 0) n = min; } else if (sfile) { if (fscanf(sfile, "%d\\n", &n) != 1) break; } else { n = rand() % max + min; } } end_time = time(0); elapsed_time = end_time - start_time; printf("%s\\t%d\\t%d\\t%d seconds\\t%d KB/sec\\n", argv[1], bufsize, nloops, elapsed_time, (int)((double) (nloops * ( bufsize / 1024 )) / (double) elapsed_time) ); return 0; }
Andrew Hamm wrote in message <3a34743e$1@news.iprimus.com.au>... Here's the shell scipt. It seems clean again. Note that the compiled C program should be called pfread, not a.out. Usage is quite simple for this one: pick the name of your raw device - eg /dev/ronline1a (of course, you can use a symlink name directly if you are smart enough to be using symlinks) and pick the number of threads you want to throw at it - say, 1 to 10. Make sure the disk is totally quiet and the machine generally quiet also. Run this: pfread.sh N /dev/roneline1a for N = 1 to 10 and see what numbers pops out. Notice the dramatic fall-off at the end when you get too many threads running. The pfread program itself takes various options, such as sequential write directives etc, which the shell script does not take, but of course you can patch them into the shell script. One of the improvements I was making to the script was to make it accept those arguments on the command line. After running this process for a while, you'll start to wish the process automatically applied FOR I = 1 to N. That was another improvement. Finally, I figured, OK, so we can test a single disk, but how can we test the effects of I/O controller channels as well? Obviously the script should take more than one device name as well. Improvement number 3. Finally, testing of write performance would be very useful, especially with RAID systems which skew the read vs. write speeds etc. I won't show people how to modify for write tests - after all, it's really trivial, and if you accidentally apply a write test to an important piece of disk, you are in deep poo. Caveat Emptor. At this time, you have all the tools you need for some interesting tests, but stay tooned for a production quality modified script - one I'm not ashamed to show the world. Once you know how much pressure a disk can take, you can make considered decisions about the number of chunks. Basically, one chunk is flushed per page cleaner during checkpoints, and ideally you want at least enough cleaners to flush the right number of chunks simultaneously, and you would also want to make the chunks need around about the same amount of flushing - ie balance the write loads to the chunks. LRU writes have different rules for flushing, and I think it would be more random how many threads get busy when LRU flushes kick in. Any information about predicting LRU write activity would be greatly appreciated. Also, it would be really interesting to see results from people testing raid systems, especially the dreaded RAID 5. <FILE pfread.sh>
Andrew Hamm wrote in message <3a34743e$1@news.iprimus.com.au>... Andrew Hamm wrote in message <3a34743e$1@news.iprimus.com.au>... Here's the shell scipt. It seems clean again. Note that the compiled C program should be called pfread, not a.out. Usage is quite simple for this one: pick the name of your raw device - eg /dev/ronline1a (of course, you can use a symlink name directly if you are smart enough to be using symlinks) and pick the number of threads you want to throw at it - say, 1 to 10. Make sure the disk is totally quiet and the machine generally quiet also. Run this: pfread.sh N /dev/roneline1a for N = 1 to 10 and see what numbers pops out. Notice the dramatic fall-off at the end when you get too many threads running. The pfread program itself takes various options, such as sequential write directives etc, which the shell script does not take, but of course you can patch them into the shell script. One of the improvements I was making to the script was to make it accept those arguments on the command line. After running this process for a while, you'll start to wish the process automatically applied FOR I = 1 to N. That was another improvement. Finally, I figured, OK, so we can test a single disk, but how can we test the effects of I/O controller channels as well? Obviously the script should take more than one device name as well. Improvement number 3. Finally, testing of write performance would be very useful, especially with RAID systems which skew the read vs. write speeds etc. I won't show people how to modify for write tests - after all, it's really trivial, and if you accidentally apply a write test to an important piece of disk, you are in deep poo. Caveat Emptor. At this time, you have all the tools you need for some interesting tests, but stay tooned for a production quality modified script - one I'm not ashamed to show the world. Once you know how much pressure a disk can take, you can make considered decisions about the number of chunks. Basically, one chunk is flushed per page cleaner during checkpoints, and ideally you want at least enough cleaners to flush the right number of chunks simultaneously, and you would also want to make the chunks need around about the same amount of flushing - ie balance the write loads to the chunks. LRU writes have different rules for flushing, and I think it would be more random how many threads get busy when LRU flushes kick in. Any information about predicting LRU write activity would be greatly appreciated. Also, it would be really interesting to see results from people testing raid systems, especially the dreaded RAID 5. <FILE pfread.sh> #!/bin/ksh # pfread.sh # Simple shell script to produce an average total KB/sec read # performance for multiple concurrent pfread processes # # USAGE: # pfread.sh n device function usage { echo " usage: $0 n device where: n = number of concurrent processes device = device to measure, ie /dev/rdsk/c0t4d0s1 " exit 1 } # usage # Verify that the correct number of arguments were provided if [[ $# -ne 2 ]]; then usage fi # Verify that the devices is readable by this process if [[ ! -r $2 ]]; then echo "Cannot read device: $2" exit -1 fi # initialize a counter variable let c=0 # create a temp file to hold pfread output outfile=/tmp/$$ # execute the specified number of concurrent pfread processes in the background while [[ $c < $1 ]]; do pfread $options $2 >> $outfile & let c=$c+1 done # on Sun OS, a wait with no arguement waits for all background processes # This behavior may differ on AT&T and other platforms wait let i=0 let j=0 ## get a sum of the KB/sec read for each thread for k in `awk '{print $6}' $outfile`; do let i=$i+$k done # Since threads are executing concurrently and we have already divided the # total KB read/thread by the duration of execution for a KB/sec/thread number # here we are only interested in the KB/sec sum. These is no need to divide # by the number of threads. # echo "$2:\\t$1 concurrent read threads\\t$i KB/sec."
I set 'CLEANERS' in our box from 128 to 80 this weekend and didn't get long checkpoint time again. It seems that overweight CLEANERS affects checkpoint time indeed. Andrew, will 'pfread' add much workload to disks? Is it usually used for setting up a production box at the beginning? I'm thinking that maybe it's not suitable for a running production box. In article <3a356eb7@news.iprimus.com.au>, "Andrew Hamm" <ahamm@sanderson.net.au> wrote: > Andrew Hamm wrote in message <3a34743e$1@news.iprimus.com.au>... > > Here's the shell scipt. It seems clean again. > > Note that the compiled C program should be called pfread, not a.out. > > Usage is quite simple for this one: pick the name of your raw device - eg > /dev/ronline1a (of course, you can use a symlink name directly if you are > smart enough to be using symlinks) and pick the number of threads you want > to throw at it - say, 1 to 10. Make sure the disk is totally quiet and the > machine generally quiet also. Run this: > > pfread.sh N /dev/roneline1a > > for N = 1 to 10 and see what numbers pops out. Notice the dramatic fall-off > at the end when you get too many threads running. The pfread program itself > takes various options, such as sequential write directives etc, which the > shell script does not take, but of course you can patch them into the shell > script. One of the improvements I was making to the script was to make it > accept those arguments on the command line. > > After running this process for a while, you'll start to wish the process > automatically applied FOR I = 1 to N. That was another improvement. > > Finally, I figured, OK, so we can test a single disk, but how can we test > the effects of I/O controller channels as well? Obviously the script should > take more than one device name as well. Improvement number 3. > > Finally, testing of write performance would be very useful, especially with > RAID systems which skew the read vs. write speeds etc. I won't show people > how to modify for write tests - after all, it's really trivial, and if you > accidentally apply a write test to an important piece of disk, you are in > deep poo. Caveat Emptor. > > At this time, you have all the tools you need for some interesting tests, > but stay tooned for a production quality modified script - one I'm not > ashamed to show the world. > > Once you know how much pressure a disk can take, you can make considered > decisions about the number of chunks. Basically, one chunk is flushed per > page cleaner during checkpoints, and ideally you want at least enough > cleaners to flush the right number of chunks simultaneously, and you would > also want to make the chunks need around about the same amount of flushing - > ie balance the write loads to the chunks. > > LRU writes have different rules for flushing, and I think it would be more > random how many threads get busy when LRU flushes kick in. Any information > about predicting LRU write activity would be greatly appreciated. > > Also, it would be really interesting to see results from people testing raid > systems, especially the dreaded RAID 5. > > <FILE pfread.sh> > > Sent via Deja.com http://www.deja.com/