IDS 10 on RHEL4 Linux: I/O benchmark of RAW DEVICES vs BLOCK DEVICES vs COOKED FILES
Posted in 2006
Topics: Performance & Tuning, Installation, Setup & Upgrades, Storage & Space Management, Server Administration, Migration, Import/Export & Data Conversion, Platform-Specific Issues
Hello there,
this is a long message (sorry in advance) which contains some hopefully
interesting information. There are also some questions at the end.
I have this nice new Red Hat Enterprise Linux 4 machine, it's an HP DL 385
with 8 internal 10K RPM SAS disks, 2 dual-core AMD Opteron CPUs and 4 GB RAM.
Machine features 256 Mb of cache on the disks controller. Here it is:
http://tinyurl.com/p6k6y
I have installed Informix Dynamic Server 10.00.FC4 (64 bit).
I bumped into this IBM article: http://tinyurl.com/gfovb, where it says that
Informix IDS 10 on RHEL4 Linux is able to access block devices in unbuffered
fashion -like it does for the raw devices- using KAIO (Kernel Asynchronous I/O).
The server is brand new, and I can play with it a bit before it goes into
production, so I decided to do some performance testing of RAW vs BLOCK
vs COOKED, to see which one was better from the I/O performance point of view.
Some definitions, first:
* RAW DEVICE or CHARACTER DEVICE = devices through which data is transmitted
one character at a time, using unbuffered input and output routines; each
character is read from, or written to, the device immediately; such devices
are disk partitions without a file system on them, and are therefore managed
from outside the operating system using direct access and bypassing the OS
layer and its cache
* BLOCK DEVICE = devices through which data is transmitted in the form of
blocks, the most significant difference between block and character devices
is that block devices use buffered input and output routines; normally a block
device is a mounted partition with a filesystem
Here is how to distinguish, listing with 'ls -l', if a device is a file or a
raw/character device or a block device:
1) FILE: -rw-rw-rw- (prefix is "-")
2) RAW/CHAR: crw-rw-rw- (prefix is "c")
3) BLOCK: brw-rw-rw- (prefix is "b")
Given in all cases the same server and the same identical set of physical disks,
here are the tests I have performed:
1. time for dbimport of a certain database (around 500 Mb big once dbexport-ed
on a filesystem)
2. execution of a raw benchmark program that creates 1 table with 1 row and
forces 1000 INSERT/UPDATE/DELETE operations in a prepared and non-prepared
fashion
Here are the results of the two tests:
1. Time to do dbimport of the same database:
- RAW: 14 min 45 sec
- BLOCK: 14min 54 sec
- COOKED: 12 min 20 sec
Raw devices and block devices are in line. Cooked file gives around 20%
better performance. Dropping the DB and repeating dbimport when using cooked
files entails even better results: there must be some filesystem caching
being used, and that is not available when the dbspace is configured using
raw or block devices. For the record, the filesystem type used is EXT3,
default for RHEL4 Linux.
2. Average figures (upon three runs) for 1,000 INSERT/UPDATE/DELETE operations:
RAW BLOCK COOKED
- average INSERT transactions per second 9403 9413 8866
- average UPDATE transactions per second 258 256 254
- average DELETE transactions per second 1371 1442 1342
- average PREPARED INSERT trans per sec 17995 19262 13443
- average PREPARED UPDATE trans per sec 304 309 317
- average PREPARED DELETE trans per sec 1326 1239 1316
Apart INSERT operations, which seems to be 20-30% slower, all results of
COOKED files are in line with those of RAW and BLOCK devices.
I always knew that COOKED files should have been "much" slower than RAW devices.
Latest Informix Admin Guide, http://tinyurl.com/m2of2, does even state this
clearly (page 241 and following ones): <<The database server can use regular
operating-system files or raw disk devices to store data. On UNIX, you should
use raw disk devices to store data whenever performance is important. [...]
When dbspaces reside on raw disk devices (also called character-special
devices), the database server uses unbuffered disk access. Performance is much
better when you use raw disk devices than cooked files because the database
server has direct I/O access to the devices. A raw disk directly transfers data
between the database server memory and disk without also copying data.>>
So, here are my questions:
- The machine that have been used features 256Mb of cache on the disks
controller: how much do you think this influenced the results, especially in
giving advantage to the filesystem-based (COOKED) dbspace compared to the
low-level access (RAW / BLOCK) to the devices?
- Given the same combination of OS and Informix version, are similar results
to be always and reasonably expected, independently from the underlying
hardware (with/without disk cache etc.)?
- Test #1 was clearly won by COOKED files configuration, but that test is not
representative of a typical OLTP kind of I/O: having an OLTP application that
needs to run on the machine, would you trust what Informix says in its
Admin Guide (see above) and results of #2 or go for COOKED files instead?
- In which proven cases the sentence from Informix documentation "On UNIX, you
should use raw disk devices to store data whenever performance is important"
is to be taken in consideration? OLTP-style usage or what?
- Any obvious reason why in test #2 only INSERT (not UPDATE nor DELETE)
operations were significantly slower when using COOKED files based dbspaces?
Hope my findings might be of some interest for someone.
Waiting for your feedback on my questions,
Ciao,
Rupan3rd (from Italy)
Rupan3rd wrote:
> - Test #1 was clearly won by COOKED files configuration, but that
> test is not representative of a typical OLTP kind of I/O: having an
> OLTP application that needs to run on the machine, would you trust
> what Informix says in its Admin Guide (see above) and results of #2
> or go for COOKED files instead?
Test 1 involves loading from the filesystem which will use OS caching. I
may be on shaky ground here but it may be that writing to a cooked file
that also uses the same OS cache is faster than a raw device in this
instance. However, as you say, unless you do a lot of dbimports this
won't matter to you very much.
Test 2 works on a single table with one row initially. This is a small
amount of data and all the operations could easily take place inside
Informix's buffers with only occasional LRU writes or maybe a checkpoint
making any difference. 1000 operations may not be enough to cause a
checkpoint. You would need to tune your server to have a very small
bufferpool to force all the changes to be written to disc more quickly
to see any significant differences between filesystems in this test. For
this reason I would be interested to see your ONCONFIG and your cached
read and write rate from "onstat -p" after each test. I would also
bounce the server before every test (if you didn't do this before) to
clear out the bufferpool and reset the statistics ("onstat -z").
Ben.