Re: Performance in chunk vs. chunks?
Posted in 1992
>From: uunet!neptune.dsc.com!jbh (J. B. Hendrix) >Message-Id: <jbh.705287859@neptune> >Subject: Performance in chunk vs. chunks? >Date: 8 May 92 01:17:39 GMT >X-Informix-List-Id: <news.1216> > >My boss from the old school is concerned that a heavily used >DB spread across multiple chunks on separate partitions in the >same disk will suffer performance degradation as compared to >a single chunk for the whole DB (obviously a large chunk, about >150 Meg). > >My cohorts and I have no information to confirm or debunk his >suspicions as this is a new product for us. Has anyone experienced >said degradation (or perhaps enhanced) performance from spreading >a DB across multiple chunks in separate partitions? Any experiences >or even opinions are very welcome. > >Thanks in advance, Jason B. Hendrix >jbh@dsc.com This is basically a question of physics and logic, rather than being specifically about OnLine. Nevertheless, the results are moderately important, and your boss certainly isn't completely wrong, so here goes. Magnetic disk normally have a single moving head. This head moves in and out over various different tracks, all of which are circles (not spirals as in a gramophone record, if you can remember what they were:-). In a book, you can now draw nice circular pictures; in e-mail, that's a trifle difficult, and also unnecessary, since a diagram like the one below will do: 0 63 v v +----------------------------------------------------------------+ | | +----------------------------------------------------------------+ This represents a section through half a disk, with track zero nearer the centre of the disk than track 63. Now, when a disk is partitioned, certain tracks are allocated to different file systems, and some of them might be allocated to an Online database. If there is a single disk on the machine, the split-up might look something like this: 0 63 v v +----------------------------------------------------------------+ A |swap|root |tmp |user |dbase | +----------------------------------------------------------------+ Here we have a single database space. We could allocate a couple of chunks to the database by modifying the layout: 0 63 v v +----------------------------------------------------------------+ B |swap|root |tmp |dbase1 |user |dbase2 | +----------------------------------------------------------------+ Now, when the disk has to read some information from the disk, it has to mechanically move the heads to the correct track, and then wait for the right portion of the track to arrive under the head -- this is the seek time and the spin or read latency periods. Seek times are widely quoted, and very important. A fast disk has about 10 millisecond seek time from inside end to outside end; a slowish disk has more than 20 milliseconds, a slow disk more than 40 milliseconds. Suppose that OnLine is working in configuration A. To access the different parts of of the database, it only has to move the head back and forth over the space labelled dbase -- not very far. Contrast this with configuration B where if it needs to access something in chunk dbase1, then something in dbase2, it has to move the heads much further. Thus, in principle, a single big chunk will be more efficient. However, notice that there are a root, swap and tmp partition, not to mention the user partition, on the same disk. There will always be other activity on the machine requiring access to files on these parts of the disk, so regardless of which configuration is used, the disk heads are going to be far busier than just the activity caused by accessing the database. Because of this, it is very difficult to be categorical about which layout will be better. Note that OnLine will minimise the number of physical reads by keeping as much of the database in shared memory as you have allowed it to, which will reduce the amount of disk traffic -- that is its raison d'etre. When you have multiple disk drives, the situation may become clearer; you can allocate one drive to the database, and one to the operating system. Or several drives to the database, and several to the O/S. Bear in mind that if you are mirroring to a single disk, you will automatically cause the disk heads to move between the two separate areas where the data has to be written to. And you may improve performance in a multi-disk mirrored system if the primary chunk is on one disk on one controller and the mirror chunk on a second disk on a second controller. This shows you the sort of considerations that are required to see what the performance of a disk system muight be. Predicting the actual performance is not easy -- nay, downright difficult -- and is best handled by empirical measurement anyway. I hope this is some help to you, Yours, Jonathan Leffler (johnl@obelix.informix.com) #include <disclaimer.h>