Random musings on fragmentation vs. interleaving
Posted in 2008
To counter the effects of boredom I decided it's time to (locally) resolve the business about the term 'fragmentation' as applied to Informix database tables. Informix use the term 'fragmentation' to describe the deliberate placement of data in certain areas of disk so as to help the query optimiser ("optimizer" if you're a merkin). But when a database table grows and Informix needs to allocate it some more disk, this can lead to 'bad' fragmentation, which Informix refers to as 'interleaving' (i.e. if a table takes up 100k of space and the engine needs to allocate another 100k but it can't make it so that both the 100k extents are in a sequential bit of disk it 'interleaves' the second 100k extent anywhere it can sandwich it in. Sometimes it does this by breaking it up into smaller bits, which is even worse!). The data is thus no longer contiguous and inhabits bits of disk in an uncontrolled and therefore inefficient manner. So there's 'good' and 'bad' fragmentation and I propose [(c)2008 me] that from now on we call tables that exhibit 'bad' fragmentation 'High Entropy' tables; a table which exists in only one extent can therefore be termed an 'entropy 0' table. As I said, I was bored (!)