Re: How to tell where to tune
Posted in 1993
Andrew Burt writes: |> I've been spending considerable time with tbstat and sar recently, trying |> to pin down exactly where our bottleneck is. I'd like to eek out every drop |> of performance I can get (since a particular application is going to need to |> run, at current estimates, for six months; if I can cut that down... well, |> I don't think I need to say more). |> |> Platform: Online 5.0, Data General Aviion 5225 (2 processor), 192Mb ram, |> 20Gb raid array split 10Gb as raw partitions for online, 10Gb unix fs (not |> used by online and quiescent during the runs). [stuff deleted] |> |> My goal was to keep the rootdbs in the middle of the disks since eventually |> the data will use all the space and I realize having things like the system |> tables in the middle is optimal; Well, some system catalog data is locally cached, and the rest is likely to be sitting in shared memory buffers if it is accessed frequently, so system catalogs should not be your primary concern. Your LOGS should be the big concern, at least as far as physical i/o goes. If you are not logging your database, your physical log is still used, so give priority to the physical log first, logical logs second, in terms of disk location (better yet, keep the two types of logs on different disks). |> The SPINCNT param is set to 750 per suggestion in release notes. The |> nature of the program is to read all the records sequentially in one table |> and insert them into proper places in a bunch of others (this is a data |> conversion effort, flat file -> RDBMS).) |> |> The problem I'm having is in finding the bottleneck. Sar -d (device |> requests to the disks) shows minimal disk activity, almost none; sar -c |> (system calls, esp. read/write) shows some activity, but nowhere near |> the peaks I've seen (avgs about 60 reads/sec, about 120k/sec, almost no |> writes). tbstat -p shows read caching about 95% and write caching about |> 88% [This leads me to believe the buffers are doing their job fine.] Cacheing tells only a partial story. FOrget write cacheing, since all data buffer writes are asynchronous to the servers anyway (page cleaners take care of that). Of that 95% read cache, the 5% can be significant. How many dskreads are you seeing? If you can decrease that, you will buy performance (i/o is teh most costly operation, and reads are synchronous to the server). Try increasing your buffers. |> sar -u shows cpu usage (summed for the two cpus) at about 95% idle. |> |> To see if they had any effect, I've tried: |> Upped #buffers (to 20,000 = 40Mb worth) [this changed the mix of |> writes to 100% chunk writes (tbstat -F) up from 50%; but I |> didn't see any major performance change in the long run] |> More/less frequent checkpoints (1/5 min) [actually it did this by itself |> since the 32k physical log would fill up every 90 seconds; |> but there was no performance change when it was doing 5 min |> checkpoints] I suspect this is a large part of your problem. Checkpoints are VERY costly, since they block all transaction activity. Your goal should be to make them as infrequent as is palatible considering recovery issues. 100% chunk writes is death to a system! 1. Tune CLEANERS == LRUS (you may want to up LRUS from the default 8, too) 2. Increase your physical log size as much as you can. Increase CKPTINTVL (maybe even max it out). Let the plog size determine checkpoint frequency. 3. Tune LRU_MAX_DIRTY down to 20 or less. For very high throughput systems, I've gone down as low as 5. Let idle/LRU writes keep that buffer pool clean, so checkpoints are very quick when they happen. 4. Tune LRU_MIN_DIRTY 5-10 less than LRU_MAX_DIRTY. You don't want to over-work the cleaners. Once the system gets humming, they will be busy enough keeping the queues down to the max dirty value. What all the above winds up doing is spreading out your writes more evenly, in terms of when it gets done. By doing 100% chunk writes, all your writes occur at checkpoint time, which means loooong checkpoints during which no transactions are occurring. The above gives you transactions and writes occuring together (more or less), with minimal "no transactions" time. |> I don't think the SPINCNT param is the problem, as tbstat -p "lokwaits" is 0, |> and, indeed, this is the only informix process running. None of the other |> sar or tbstat info look odd (e.g., no swapping, lock contention, etc.) If you have 95% idle time, you can go WAY up with that SPINCNT value. Don't look at lokwaits, though, since they have nothing to do with SPINCNT; it relates to lchwaits. The general rule you can follow is, if there is cpu to burn, up that SPINCNT value (2000+ is not unusual). |> My main question is... where's the bottleneck? CPU time shows 95% idle. |> Disks show minimal activity. It's like it's just going to sleep for a |> few microseconds here and there every few microseconds... (But I assume |> context switches are charged to user or system time and aren't the culprit.) |> |> Methinks one or the other should be peaking... |> |> Any thoughts? Things to look at? To try? See above. Dave Disclaimer: These opinions are not those of Informix Software, Inc. ************************************************************************** "I look back with some satisfaction on what an idiot I was when I was 25, but when I do that, I'm assuming I'm no longer an idiot." - Andy Rooney