Re: Raw vs. Cooked files
Posted in 2005
Topics: Performance & Tuning, Installation, Setup & Upgrades, Storage & Space Management, Platform-Specific Issues, Versions, Editions & End-of-Life
Hi,
I'd say the correct answer to the question is:
"It depends."
--------------------------------------------------------------------
The implementation of file system I/O buffering by the
OS can be very efficient on Linux. But for that the OS
needs sufficient memory for the buffering. And this is
what makes the "It depends".
Therefore there are 2 extremes:
a) almost all physical memory is used by IDS in the
form of shared memory (allocated SHM segments
listed by "onstat -g seg"). A good indication for this
scenario is a machine that uses little swap space,
because it already has swapped out some minor
"applications". The system is at the (sensible) limit
of its memory usage.
In this scenario, raw devices and the utilization of
KAIO (available with IDS 10.00.UC1 on Linux) will
in general perform better than cooked files.
b) there's a lot of unused physical memory left over,
even though IDS is configured optimally (e.g.
BUFFERS is big enough, etc.). A recent example
of such a scenario was 32-bit IDS installed on a
64-bit Linux, where IDS can use only ~4 GB of the
available 24 GB of physical memory. In such a
scenario, the OS can allocate a lot of memory for
file system I/O buffering and this will make file
system I/O (i.e. for chunks on cooked files) very
efficient. In this scenario it will normally perform
better than raw devices and KAIO.
This is the general experience on Linux, where
cooked file chunks are on a non-journalling file
system. Not all journalling file systems are un-suitable
for IDS chunks, but "reiserfs" certainly is un-suitable,
so make sure you stay clear of that. E.g. use the
"ext2" file system instead.
Unix flavours other than Linux may exhibit completely
different behaviour regarding raw devices vs. cooked
files. This is mainly because they use different
implementations of file systems and their I/O buffering.
These things cannot be compared across different
operating systems.
--------------------------------------------------------------------
I think I've written an e-mail to this effect before to this
list, but I'm not sure.
So whoever feels responsible for some certain FAQ
pages, please feel free to cut&paste the above text
between the "---------" separators into the FAQ (with or
without citing my name ...). But make sure that the text
is complete !
Regards,
Martin
--
Martin Fuerderer
IBM Informix Development Munich, Germany
Information Management
owner-informix-list@iiug.org wrote on 18.05.2005 00:40:08:
> Rajesh Kapur wrote:
> > Hello,
>
> Chec out last weeks postings here where this subject was beaten to
death. I
> say YES! Others say NO! Bottom line? Test it yourself. The article
you
> quote in the FAQ tells you how I tested it.
>
> Art S. Kagel
>
> > I am installing IDS 10 on RHEL3 on an intel x86 server. I need to make
the
> > choice between raw devices and cooked files.
> > I found the following information in the FAQ's. Is this information
still
> > accurate? Do cooked files continue to be significantly slower than the
raw
> > files? I have heard that UNIX file systems have improved over the
years and
> > the performance divide between raw and cooked files has narrowed
> > considerabley. Is it true?
> >
> > Thanks!
> > - Rajesh
> >
--------------------------------------------------------------------------------------
> >
> > 6.38 Is raw disk faster than cooked files?
> > On 22nd Jun 1998 kagel@bloomberg.net (Art S. Kagel) wrote:-
> >
> > ....................the safety issue of cooked files is no longer a
problem.
> > The big problem with cooked files still is performance. All writes and
reads
> > to/from cooked files MUST go through the UNIX buffer cache. This means
an
> > additional copy from the server's output buffer to the UNIX buffer
page then
> > a synchronous write to disk. This is opposed to a write to a way file
where
> > the server's output buffer is written directly to disk without the
> > intervening copy. This is just faster. Anyone who has written anything
that
> > can test this can attest to the difference in speed. Here are my test
> > results:
> >
> >
> > FileType Sync? Times (real/user/system) 2run avg
> > --------------- ----- -----------------------------------------
> > Filesystem file N 14.40/3.70/2.52
> > Y 15.02/3.61/2.63
> > Cooked disk N 12.81/3.74/2.24
> > Y 13.42/3.84/2.43
> > Raw disk N 9.32/3.67/1.52
> > Y 9.40/3.66/1.44
> >
> > From this you can clearly see the cost of Cooked files and of synced
cooked
> > files. The tests were done with a version of my ul.ec utility modified
to
> > optionally open the output file O_SYNC. Cooked disk partition is
almost 50%
> > slower than raw disk partition and cooked filesystem files are almost
60%
> > slower. The penalty for O_SYNC is an additional 5% for cooked files
and
> > negligible for RAW files (as expected). The test file was 2.85MB
written
> > using 4K I/O pages (the default fopen application buffer size) which
should
> > simulate Informix performance. The Cooked and Raw disk partition tests
were
> > conducted to a singleton 9GB Fast Wide SCSI II drive using the raw and
> > cooked device files corresponding to the same drive.
> >
> >
sending to informix-list
Martin Fuerderer wrote:
> Hi,
I emailed a simple response yesterday, but the link seems to be lost.
Here's another more detailed one:
It does not matter HOW much memory the OS has for a buffer cache. There are
still several reasons why RAW will ALWAYS be faster than COOKED devices or
OS filesystem files:
1) If the OS cache is too small bulk server IO like checkpoints and peak
load LRU flushing will thrash the cache killing OS and server performance.
- Thanks for pointing that out Jake.
2) All IO to a COOKED device or FS file MUST first be copied from the IDS
buffers to the OS buffers before the actual IO operation. This MUST be
extra overhead not needed on a RAW device - by definition!
3) If you are writing to an FS file chunk instead of a COOKED device you
have the additional overhead of the filesystem and directory and inode
maintenance. While so called 'fast filesystems' reduce these particular
overhead items they can do nothing to eliminate it or to eliminate the
effects of #2!
4) The cache buys you nothing for writing because IDS correctly opens all
COOKED chunks in O_SYNC mode which means that the buffer is immediately
flushed to disk as soon as it is copied from the IDS buffer and IDS still
has to wait for this write to complete as it does for RAW. However, for RAW
IDS does not have to wait for the second buffer copy.
5) On the read side, it's counterintuitive, but, there is little gained from
the OS cache. Why? Reasons:
- Everything that IDS needs is already in its own cache and also in the
SAN cache and also on the disk drives' caches! You're just adding one more
redundant cache!
- Each block read has to be copied first to the OS buffer cache before
it can be handed to IDS if it's not already in OS cache. Raw doesn't have
this overhead.
- Because IDS's cache management is RDBMS behavior specific, chances
are if a page is in the OS cache, then it's already in the IDS cache and
will never be read from OS cache again. Is this redundant to the first
point? Maybe but so is the OS cache ;-) [OK cheap shot. Back to serious
points...]
6) Filesystem file chunks will be fragmented on disk and not contiguous like
COOKED device chunks and RAW chunks. Some filesystems are better than
others at reducing file fragmentation, but none are perfect.
You state that on Linux with lots of cache you find AIO to FS files will
outperform KAIO on RAW device. Prove it Martin. I posted my test results,
post yours (and don't forget to post O_SYNC #'s if you are using a direct to
disk test as I did rather than a DB timing -- admittedly difficult to time
accurately which is why I use a direct-to-disk test). My test was performed
using my ul.ec utility to read a large IDS table writting to disk files, raw
and cooked devices using the identical query with and without the -U flag to
enable/disable O_SYNC IO (with -U is how IDS does it, but I was curious what
the real cost of that O_SYNC setting was). A pre-test run of ul.ec with the
same query primed the cache of whatever records would be likely to remain
after a completed run. The source for ul.ec is in my utils2_ak package in
the IIUG Repository, so you can repeat that test or devise your own.
Art S. Kagel
> I'd say the correct answer to the question is:
> "It depends."
>
> --------------------------------------------------------------------
> The implementation of file system I/O buffering by the
> OS can be very efficient on Linux. But for that the OS
> needs sufficient memory for the buffering. And this is
> what makes the "It depends".
>
> Therefore there are 2 extremes:
>
> a) almost all physical memory is used by IDS in the
> form of shared memory (allocated SHM segments
> listed by "onstat -g seg"). A good indication for this
> scenario is a machine that uses little swap space,
> because it already has swapped out some minor
> "applications". The system is at the (sensible) limit
> of its memory usage.
>
> In this scenario, raw devices and the utilization of
> KAIO (available with IDS 10.00.UC1 on Linux) will
> in general perform better than cooked files.
>
> b) there's a lot of unused physical memory left over,
> even though IDS is configured optimally (e.g.
> BUFFERS is big enough, etc.). A recent example
> of such a scenario was 32-bit IDS installed on a
> 64-bit Linux, where IDS can use only ~4 GB of the
> available 24 GB of physical memory. In such a
> scenario, the OS can allocate a lot of memory for
> file system I/O buffering and this will make file
> system I/O (i.e. for chunks on cooked files) very
> efficient. In this scenario it will normally perform
> better than raw devices and KAIO.
>
> This is the general experience on Linux, where
> cooked file chunks are on a non-journalling file
> system. Not all journalling file systems are un-suitable
> for IDS chunks, but "reiserfs" certainly is un-suitable,
> so make sure you stay clear of that. E.g. use the
> "ext2" file system instead.
>
> Unix flavours other than Linux may exhibit completely
> different behaviour regarding raw devices vs. cooked
> files. This is mainly because they use different
> implementations of file systems and their I/O buffering.
> These things cannot be compared across different
> operating systems.
> --------------------------------------------------------------------
>
> I think I've written an e-mail to this effect before to this
> list, but I'm not sure.
>
> So whoever feels responsible for some certain FAQ
> pages, please feel free to cut&paste the above text
> between the "---------" separators into the FAQ (with or
> without citing my name ...). But make sure that the text
> is complete !
>
> Regards,
> Martin
<Previous SNIPPED>
Art S. Kagel schrieb: > Martin Fuerderer wrote: > >> Hi, > > > I emailed a simple response yesterday, but the link seems to be lost. > Here's another more detailed one: > > It does not matter HOW much memory the OS has for a buffer cache. There > are still several reasons why RAW will ALWAYS be faster than COOKED > devices or OS filesystem files: > > 1) If the OS cache is too small bulk server IO like checkpoints and peak > load LRU flushing will thrash the cache killing OS and server > performance. - Thanks for pointing that out Jake. > > 2) All IO to a COOKED device or FS file MUST first be copied from the > IDS buffers to the OS buffers before the actual IO operation. This MUST > be extra overhead not needed on a RAW device - by definition! > > 3) If you are writing to an FS file chunk instead of a COOKED device you > have the additional overhead of the filesystem and directory and inode > maintenance. While so called 'fast filesystems' reduce these particular > overhead items they can do nothing to eliminate it or to eliminate the > effects of #2! > > 4) The cache buys you nothing for writing because IDS correctly opens > all COOKED chunks in O_SYNC mode which means that the buffer is > immediately flushed to disk as soon as it is copied from the IDS buffer > and IDS still has to wait for this write to complete as it does for > RAW. However, for RAW IDS does not have to wait for the second buffer > copy. > > 5) On the read side, it's counterintuitive, but, there is little gained > from the OS cache. Why? Reasons: > - Everything that IDS needs is already in its own cache and also in > the SAN cache and also on the disk drives' caches! You're just adding > one more redundant cache! > - Each block read has to be copied first to the OS buffer cache > before it can be handed to IDS if it's not already in OS cache. Raw > doesn't have this overhead. > - Because IDS's cache management is RDBMS behavior specific, chances > are if a page is in the OS cache, then it's already in the IDS cache and > will never be read from OS cache again. Is this redundant to the first > point? Maybe but so is the OS cache ;-) [OK cheap shot. Back to > serious points...] > > 6) Filesystem file chunks will be fragmented on disk and not contiguous > like COOKED device chunks and RAW chunks. Some filesystems are better > than others at reducing file fragmentation, but none are perfect. > > You state that on Linux with lots of cache you find AIO to FS files will > outperform KAIO on RAW device. Prove it Martin. I posted my test > results, post yours (and don't forget to post O_SYNC #'s if you are > using a direct to disk test as I did rather than a DB timing -- > admittedly difficult to time accurately which is why I use a > direct-to-disk test). My test was performed using my ul.ec utility to > read a large IDS table writting to disk files, raw and cooked devices > using the identical query with and without the -U flag to enable/disable > O_SYNC IO (with -U is how IDS does it, but I was curious what the real > cost of that O_SYNC setting was). A pre-test run of ul.ec with the same > query primed the cache of whatever records would be likely to remain > after a completed run. The source for ul.ec is in my utils2_ak package > in the IIUG Repository, so you can repeat that test or devise your own. Art, this is all corresponding to my findings. Still, when using a 32bit IFX engine in a 64bit OS it can pay off - if you have tons of memory and if one cannot upgrade to a 64bit IFX version - to put tables with high READ access rate onto cooked files. Only so much that the OS is able to cache them in the memory not seen by IFX because of the 32bit addrspace limitations. Dic_k [ ... snipped ... (to hide the TOFU style posting) ... ] -- Richard Kofler SOLID STATE EDV Dienstleistungen GmbH Vienna/Austria/Europe
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g