How to tell where to tune
Posted in 1993
I've been spending considerable time with tbstat and sar recently, trying to pin down exactly where our bottleneck is. I'd like to eek out every drop of performance I can get (since a particular application is going to need to run, at current estimates, for six months; if I can cut that down... well, I don't think I need to say more). Platform: Online 5.0, Data General Aviion 5225 (2 processor), 192Mb ram, 20Gb raid array split 10Gb as raw partitions for online, 10Gb unix fs (not used by online and quiescent during the runs). (FYI, the raid array is set up as 4 groups of 5 drives, so Unix thinks it has 4 disks of about 5Gb each; these are partitioned into three 1.5ishGb partitions each since Informix can't handle a chunk larger than 2Gb (boo, hiss). The informix chunks are in three dbspaces of two chunks each, like so: Raid unit, Partition Informix dbspace ("phys disk/partition") 0, 1 1 (rootdbs) 0, 2 1 1, 1 2 (data2) 1, 2 2 0, 0 3 (temporarily unused) 1, 0 3 My goal was to keep the rootdbs in the middle of the disks since eventually the data will use all the space and I realize having things like the system tables in the middle is optimal; I did this by making the first third of each disks in a temporarily unused chunk; assume no I/O to that; when the 2nd & 3rd chunks are full I'll add in the 1st. The raid units are each on their own dedicated scsi-2 bus (seperate cards even). The main unix disks are on yet another bus and have no non-informix activity during this run. The SPINCNT param is set to 750 per suggestion in release notes. The nature of the program is to read all the records sequentially in one table and insert them into proper places in a bunch of others (this is a data conversion effort, flat file -> RDBMS).) The problem I'm having is in finding the bottleneck. Sar -d (device requests to the disks) shows minimal disk activity, almost none; sar -c (system calls, esp. read/write) shows some activity, but nowhere near the peaks I've seen (avgs about 60 reads/sec, about 120k/sec, almost no writes). tbstat -p shows read caching about 95% and write caching about 88% [This leads me to believe the buffers are doing their job fine.] sar -u shows cpu usage (summed for the two cpus) at about 95% idle. To see if they had any effect, I've tried: Upped #buffers (to 20,000 = 40Mb worth) [this changed the mix of writes to 100% chunk writes (tbstat -F) up from 50%; but I didn't see any major performance change in the long run] More/less frequent checkpoints (1/5 min) [actually it did this by itself since the 32k physical log would fill up every 90 seconds; but there was no performance change when it was doing 5 min checkpoints] I don't think the SPINCNT param is the problem, as tbstat -p "lokwaits" is 0, and, indeed, this is the only informix process running. None of the other sar or tbstat info look odd (e.g., no swapping, lock contention, etc.) My main question is... where's the bottleneck? CPU time shows 95% idle. Disks show minimal activity. It's like it's just going to sleep for a few microseconds here and there every few microseconds... (But I assume context switches are charged to user or system time and aren't the culprit.) Methinks one or the other should be peaking... Any thoughts? Things to look at? To try? -- Andrew Burt aburt@du.edu "But if he was dying he wouldn't bother to carve "Aaaaargh", he'd just say it."