MORE RAM - BIG CONTROLLER CACHE / DB ON RAW - DB ON FS
Posted in 1999
Topics: Performance & Tuning, Server Administration, Logging & Checkpoints, Platform-Specific Issues, Jobs, Consulting & Announcements
MOPSOMER@raychem.com wrote: > > I recently had a discussion with a consultant. The following is a > replay of this discussion. I just wanted to see your opinions on it. snip.... By and large he/she is right. Staged / asynchronous writes are about all controller memory is good for (and it's *really* great for that...), assuming you have enough cache on the database for reads. Admittedly, Solaris 2.5/6 w/ Informix is getting a bit tight for big systems, but you are not far out from 64 bit OS/DBMS and maybe even SAP, then life should be pretty good for a long time. Hopefully, you won't have to do anything too drastic in the interim. greg
I recently had a discussion with a consultant. The following is a replay of this discussion. I just wanted to see your opinions on it. I : Why do you only sell storage arrays with no cache or a small one ? X : Disk subsystems with GBs of NVRAM come from the mainframe world, where memory prices are high and memory capacities quite low. In the open systems world, RAM is not expensive, so it is better to spend your money on extra GBs of RAM as this is much faster and the OS, DB and application can use that RAM more intelligent than that the disk subsystem controller can do with its big cache. I : Okay, but our BUFFERS are already set at the maximum (768.000) and shared memory in Solaris 2.5.1 is limited to 3.75 GB anyway. X : The biggest benefit from controller caches comes from enabling fast-writes and colescing writes. The effect of the controller cache on reads is very limited, as your database is over 300 GB. For the fast-writes, you do not need a cache which is several GB. I : Okay, but what about the RAM that I can't use for the BUFFERS ? In Informix 7.30, the maximum number of BUFFERS is still the same. And even if I could set BUFFERS much higher, my checkpoints would probably become much too long for an OLTP environment. I already have checkpoints which take between 10 and 15 seconds nowadays. X : You could also put your db in filesystems instead of using raw. I : This is new to me. I thought filesystems were much slower ! X : The difference is not that big anymore. With the direct I/O improvements in Solaris 2.6 and the usage of Veritas VxFS instead of UFS, you will get performance comparable to raw disk, but the investment in additional RAM can be used as filesystem cache, so you can indirectly increase your Informix BUFFERS, avoiding a lot of expensive disk I/O or I/Os satisfied by the controller cache. I/Os satisfied by RAM are much faster than the controller cache. I : [End of discussion; surprised by the statements on raw versus fs] Any ideas or comments on this ? Did somebody already convert their Informix databases from raw disk to UFS or VxFS (for performance reasons) ? Is it indeed a better idea to invest in RAM than going for a big controller cache (4-16 GB) ? Where does that 768,000 limit for the ONCONFIG parameter BUFFERS come from ? Would it indeed make no sense (no performance improvements) to have 4 GB of BUFFERS instead of the 1.5 GB which is the maximum in Informix ? Mario Opsomer
Hi,
how about a test-program ? Sure, I will write a test program
where your FS will always be faster than a raw device:
Create two dbspaces, "dbs1" on a raw device and "dbs2" on a FS.
create database funny in EITHER_dbs1_OR_dbs2;
create table t1 ( f1 int, f2 char(1300) );
insert into t1 select colno, colname from syscolumns;-- repeat the last insert until you have appr. 5,000 rows
create index idx01 on t1 ( f1 );
update statistics on t1(f1);
------------ now a small ESQL/C program ------------------
#include <stdio.h>
main()
{
$int i;
$char name_vc[2200];
EXEC SQL WHENEVER ERROR STOP;
EXEC SQL DATABASE funny;
for ( i = 1; i < 5000; i++ )
{
EXEC SQL SELECT name into :name_vc FROM t1 where nr = :i;
if ( ( i % 10 ) == 0 )
{
printf("\\r%4d rows retrieved", i );
fflush( stdout );
}
}
printf("\\n");
}
-----------------------------------------------------------
Compile and run the ESQL/C program. Check it once with
dbs1 and once with dbs2. I'm sure, the file system will
always be faster than the raw device. I tried it on
a SCO box. The difference was 5.6sec for the FS, 6.9sec for
the raw device.
But, this test was designed to make the filesystem access
faster. If I would use a cursor instead of the static SELECT,
the raw device would be faster than the filesystem.
I never found a person who could explain, how a filesystem
can be faster than a raw device. It's not neccessarily
much slower than a raw device, but it cannot be faster.
The funny discussion about the filesystem cache is not
interesting, I think. Informix uses the OS call "open()"
to open a chunk. The second step is the "stat()" call where
they look for the "device type" of the chunk. If the chunk
is a block device or an ordinary file in the filesystem,
they reopen the file with the flag "O_SYNC". Read your
UNIX manual and learn the meaning of this flag. In other
words, Informix does not use the filesystem cache for
write operations.
I suggest, run a few represantive tests.
Finally, let's have a look at your current disk speed.
You have about 3GB of BUFFERS, that is, if you set your
LRU_MAX_DIRTY value to 1 your system has to write up to 30MB
during a checkpoint. (onstat -R -r ) This takes appr. 10sec
-> your write performance is about 3MB/sec. If you would use
a write cache of 64MB, the checkpoint duration must
reduce to <= 1sec. You can monitor the disk speed by the
"sar -d 5 1000" command. Look at the "blks/s" ( Blocks
per seconds ) column. If the disk speed is fast enough,
your checkpoint duration has another reason, i.e. thread
synchronization, a running archive process a.s.o.
Otherwise the disk controllers' write cache is not enabled.
Any comments ?
Best regards,
Stefan Weideneder
PS: How about more disks or simply several machines ?
why don't you just give me a few GB of your ram, I'm sure it will make you feel better, and it will certainly make me feel better. Oh, go on, you know it makes sense. Languishing in a pityfull 256MB of ram Paul Blamire
In article <36A46A1E.2A04D4D3@weideneder.de>, Stefan Weideneder
<stefan@weideneder.de> writes
>Hi,
>
>how about a test-program ? Sure, I will write a test program
>where your FS will always be faster than a raw device:
>
>
>Create two dbspaces, "dbs1" on a raw device and "dbs2" on a FS.
>
>create database funny in EITHER_dbs1_OR_dbs2;>
>create table t1 ( f1 int, f2 char(1300) );
>insert into t1 select colno, colname from syscolumns;>-- repeat the last insert until you have appr. 5,000 rows
>create index idx01 on t1 ( f1 );
>update statistics on t1(f1);>
>------------ now a small ESQL/C program ------------------
>#include <stdio.h>
>main()
>{
> $int i;
> $char name_vc[2200];
> EXEC SQL WHENEVER ERROR STOP;
> EXEC SQL DATABASE funny;
>
> for ( i = 1; i < 5000; i++ )
> {
> EXEC SQL SELECT name into :name_vc FROM t1 where nr = :i;
> if ( ( i % 10 ) == 0 )
> {
> printf("\\r%4d rows retrieved", i );
> fflush( stdout );
> }
> }
> printf("\\n");
>}
>-----------------------------------------------------------
>Compile and run the ESQL/C program. Check it once with
>dbs1 and once with dbs2. I'm sure, the file system will
>always be faster than the raw device. I tried it on
>a SCO box. The difference was 5.6sec for the FS, 6.9sec for
>the raw device.
>
>But, this test was designed to make the filesystem access
>faster. If I would use a cursor instead of the static SELECT,
>the raw device would be faster than the filesystem.
>
>I never found a person who could explain, how a filesystem
>can be faster than a raw device. It's not neccessarily
Read ahead? Try increasing the online read ahead parameters..
>much slower than a raw device, but it cannot be faster.
>The funny discussion about the filesystem cache is not
>interesting, I think. Informix uses the OS call "open()"
>to open a chunk. The second step is the "stat()" call where
>they look for the "device type" of the chunk. If the chunk
>is a block device or an ordinary file in the filesystem,
>they reopen the file with the flag "O_SYNC". Read your
>UNIX manual and learn the meaning of this flag. In other
>words, Informix does not use the filesystem cache for
>write operations.
>
>I suggest, run a few represantive tests.
>
>Finally, let's have a look at your current disk speed.
>You have about 3GB of BUFFERS, that is, if you set your
>LRU_MAX_DIRTY value to 1 your system has to write up to 30MB
>during a checkpoint. (onstat -R -r ) This takes appr. 10sec
>-> your write performance is about 3MB/sec. If you would use
Surely you should have CLEANERS = min (number of disks,10)
and have one chunk per disks, hence you write to several
disks in parallel? Does 3Mb/sec seem low in this case?
How many disks does this system have if it needs 3Gb of
BUFFERS?
>a write cache of 64MB, the checkpoint duration must
>reduce to <= 1sec. You can monitor the disk speed by the
>"sar -d 5 1000" command. Look at the "blks/s" ( Blocks
>per seconds ) column. If the disk speed is fast enough,
>your checkpoint duration has another reason, i.e. thread
>synchronization, a running archive process a.s.o.
>Otherwise the disk controllers' write cache is not enabled.
>
>Any comments ?
>
>Best regards,
>
>Stefan Weideneder
>
>PS: How about more disks or simply several machines ?
--
David Williams