Better speed w/large C-ISAM files
Posted in 1991
Path: emory!swrinde!cs.utexas.edu!uunet!decwrl!sgi!cdp!scott From: scott@igc.org (Scott Weikart) Newsgroups: comp.databases.informix Message-ID: <1561600004@igc.org> Date: 29 Nov 91 22:48:00 GMT Sender: Notesfile to Usenet Gateway <notes@igc.org> Nf-ID: #N:cdp:1561600004:004:2955 Nf-From: cdp.UUCP!scott Nov 29 14:48:00 1991 > I was wondering if there has been any effort to parrallelize/distribute > c-isam. I am having problems with the size of my C-ISAM files (several 100 > Megs) and I was thinking that such a big file could be handled by several > machines (cpu/disk). First, I'll talk about the disk speed aspect. If you're running under SunOS 4.* (or is it 4.1.*), you effectively have a disk cache that's as large as main memory. So, the more DRAM you add, the more disk cache you'll have. You can put 128MB in a SparcStation 2, and much more in some of the SparcServer boxes. I think System V Release 4 is the same way, and I know there are '486 motherboards that will hold 128MB. So, if you don't need all of your C-ISAM files at all times (or if you buy a SunOS/SVR4 box that will hold hundreds of megabytes), then it's possible that your C-ISAM data will always be in DRAM. Remember, DRAM these days costs less than $50/MB, so this is a pretty cheap way to get amazing improvements in "disk" throughput. Another thought is to try to figure out if you can get ISQL-SE to work with a larger C-ISAM node size. bcheck will change the C-ISAM node size, but a techie at Informix told me that this is only for porting between dissimilar architectures, that ISQL-SE in general expects only one node size. Anyway, if you could increase the node size, you'd need fewer kernel calls and disk accesses when walking the indexes. Another idea is to try to reduce the seek time on the disks that hold the C-ISAM files. One simple method is to buy a faster disk (these days you can get 10-12ms average seek time and 5ms average rotational latency). A more subtle method is to buy a *bigger* disk, and put the C-ISAM files in as small a filing system (partition, actually) as possible. By making the partition a very small percentage of the total disk size, you can get actual seek times that are 1/2 or 1/3 of the "average" seek time. To make it easier to make small disk partitions, you can split the C-ISAM files across multiple disks, by putting absolute paths into the systables table. Or, even simpler, use disk striping to spread the C-ISAM partition across multiple disks, to keep the partition size per disk small. If you're in a multi-user rather than a batch environment, multiple disks will also helped if your controller supports overlapped seek. As for the CPU overhead, the obvious thing is to buy a faster CPU. If you have the fastest CPU already for your current architecture, you could buy a multi-processor. Even if there is no Informix product that uses fine-grained multi-processing, you could effectively use a small number of processors. One processor for each Informix back-end process, as much as one processor for each Informix user to run the front-end applications, and processor(s) for all the other Unix jobs that could be running (cron, I/O if the multi-processing is symmetric, program development, networking/email, etc). -scott