Re: Largest chunk size < largest extent size?!?
Posted in 1993
Let me preface this followup with the following: I cannot speak for Informix as a whole, only my personal perspective; what I have to say does NOT equate to company policy. I recently had some discussions with folks on the issue of RAID technology, and the conclusion I must draw is that we really have not taken the unique aspects of such devices into account with our current servers. This translates into some inconvenience for our users when working with RAID devices. I do not know for sure if we have done anything in our 6.0 product to change this. Andrew Burt writes: |> 1) Is this pure horse-puckey or what? What *is* the size of |> the max chunk (in #pages)? The max chunk size is 2Gb for 2K pages, 4Gb for 4K pages. This is due to the number of bits we use to track chunk size. That the allowable extent size is larger than this is unfortunate; you can never actually get an extent size larger than the above numbers. There is clearly a need for a documentation change. |> 2) If it's 16M pages, not 1M, then how can I convince tbmonitor |> to let me have a rootdbs with 4M pages? You cannot. 2Gb/4Gb is as big as you can go with a single chunk. To make your rootdbs bigger, you must add additional chunks after initialization. |> 3) If it's really 1M pages, why?!?!? Shouldn't it be 16M as |> the docs imply? Are we being misled here? The number of bytes available for tracking extent size is larger than that for chunk size. The chunk size forms the practical limitation. You are not being misled so much as misinformed. |> 4) How does informix allocate space among multiple partitions? |> That is, I could split the 4.6Gb into two 2Gb (or three 1.5s) |> but if online is going to try to spread the work evenly among |> them, that would *decrease* performance, since it would make the |> heads travel quite a bit more since they're on the same physical |> units. I'm afraid this issue is a little too complex to discuss here (i.e. I don't have the time to type it all in). There is the potential for increased head movement. |> 5) How would *you* set up a 4.6Gb disk to work with online? (Bear |> in mind, it will be mostly full of data very soon, so any scheme |> that relies on it not using it all right now, planning for growth, |> etc., is pointless). Also bear in mind the raid array can't |> be partitioned efficiently to use less than the 5 drives, so |> saying "use one drive for rootdbs and..." won't cut it. You will, out of necessity, have to partition the drive into 2Gb partitions. When calculating seeks, OnLine will take the chunk offset specified and add to it the offset into the chunk for the page it is after. In other words, it will try to get to the location with a single lseek(). Unfortunately, the offset arg to lseek() is a signed long, specifying bytes, so the max single seek is 2^31. By creating partitions > 2^31 bytes and then breaking it up into max sized chunks, you are in for trouble. When a seek location is formed for the second chunk, it will overflow the long and cause lseek() to return an error. This means that a chunk in a partition of > 2^31 bytes beyond the first 2^31 bytes will be "unreachable". Thus the need to keep your partitions down to 2Gb max. |> 6) We actually have two stripes of these (9Gb total, but with an |> estimated load of about 7Gb usage), so my plan (unless I get |> better advice from netlanders) will be to |> (a) assume there's no way to get a larger chunk; |> (b) partition the 2 4.6Gb's into smaller units (say, 1.5Gb); |> (c) fill them one-chunk-from-raid1, one-chunk-from-raid2, |> raid1, raid2, etc.; This, based on my limited knowledge of RAIDs, should work ok, provided that by "partition" you don't mean "break up into chunks". |> (d) Hope informix doesn't try to be too clever trying to get |> chunks 1,3,5 or 2,4,6 going in parallel. :-( The only time we would do anything close to parallel i/o (at least in pre-6.0) would be during checkpoints, when page cleaners will be assigned to chunks and work at the same time. If thrashing becomes a big performance problem, you could decrease your page cleaners. Also bear in mind that chunks are flushed, at checkpoint time, in the order in which they were created. By alternating the raids as you mention above, you can encourage a minimum of interference by keeping the page cleaners down to 2. You may find that even with the thrashing, performance would be better with > 2. The best suggestion I can give is to use your LRU tuneables to minimize chunk writes altogether (maximizing LRU and Idle writes). |> (e) Hope I can fit enough chunks I suggest keeping your device names as small as possible, since one limitation on the number of chunks is the number of chunk description structures that can fit on one page. Since the device name is the one variable length element of that structure, it becomes the deciding factor. If necessary, use links to create short device names (e.g. /dev/a, /dev/b, etc.). The other limitation is the number of open files your system can support, which is typically tuneable. On startup, OnLine (sqlturbo) will have 6 open files plus one for each chunk defined (mirror chunks count separately, of course). Dave