Re: raw vs. cooked files under linux
Posted in 2005
A follow-up debate (no single question being fixed) on whether Informix chunks under Linux should use raw devices or cooked files. Art Kagel argues strongly for raw: his benchmarks show raw devices roughly 15-25% faster than cooked devices and cooked devices 10-15% faster than filesystem files (figures reproduced from the FAQ), says setup is trivial (symlink to the device plus correct permissions), and that this holds even on SANs, alongside his usual RAID10-over-RAID5 advice. Data Goob counters that raw adds admin risk for inexperienced/one-person shops and that modern SAN/cooked performance is adequate, later conceding partially. Others note I/O gains don't translate one-for-one into response time. No firm resolution beyond the shared advice to benchmark in your own environment.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Backup & Restore, Storage & Space Management, Stored Procedures & SPL, Server Administration, Platform-Specific Issues
Data Goob wrote:
> Art S. Kagel wrote:
>
>> Heinz Weitkamp wrote:
>>
>>> Hi all,
>>
>>
>>
>> Go RAW all the way. Contrary to the Goob's understanding my own
>> testing shows RAW devices to be 15-25% faster than COOKED devices
>> which are in turn 10-15% faster than filesystem files, even on the
>> best filesystem implementations that adds up to a 25% improvement.
>> Well worth any minor inconveniences in my book.
>
>
> OK so you get some speed increases, but also the additional management
SOME speed increase? You call 25-40% improvement SOME increase? Ask your
users if they'd like their apps to run 25-40% faster and see if they think
it's insignificant!
> of linking raw files. The original post indicated he was a newbie.
> Raw files are not necessarily difficult, but they ARE optional and not
> something I would entrust to a newbie. It's a great learning experience,
> but who's paying for it??
This is NOT rocket science. The RAW device files are created at the time
the disk partitions are created. All any DBA has to do is create a symbolic
link to the RAW device (and he'd want to do that anyway so he can
move/rename the actual file/device even if it's COOKED or a FS file) and ask
the SA's (assuming he's not doubling as SA in which case the newbie moniker
is probably misplaced) to set the permissions on the device and make sure
they stay set. POOF! RAW chunks! And as John points out there is FAR more
risk that some well intentioned SA will delete those HUGE chunk files.
> But another point to make is that in many SAN environments, raw files are
> a complete waste of time, with the data already being spread out over many
> drives, and, with speeds typically in the 10-15K rpm range, or even faster,
> you get speeds today on cooked files that are faster than raw was even a
> couple of years ago.
You are completely misinformed. We got SANs, hundreds of SANs, the very
best SANS in existence. We also have dozens of IDS server machines. All
using RAW disk because they still test out to be MUCH faster than COOKED
devices or COOKED filesystem files, even on SANs.
It sounds like you think I'm proposing using singleton disk drives when I
say RAW. NOOO. As always I promote using RAID10 everywhere for anything
that's not absolutely trivial. Not only that, we tend to plaid those RAID10
arrays into massive logical devices covering many dozens of mirrored pairs
across multiple SANs. Either way you have to carve up those massive logical
devices into smaller logical devices (still spread over many many spindles).
For servers earlier than 9.40 I have the SAs carve out 2GB chunks, for
9.40+ they are carving several different sizes so that the DBAs have the
flexibility to add a 2, 4, 8, or 16GB chunk where needed. This is easy stuff.
Yes, the intricacies of EMC's SANs and VMs are a bit esoteric, and I leave
that for the SAs, but you cannot say, "Ahh but what if the DBA is the SA?",
because if he is, then he should know his jobs, both of them, and that part
becomes routine very quickly.
You assert that cooked files are as fast today as RAW on SANs. Have YOU
tested it? Are you comparing RAW to COOKED with O_SYNC? Or are you
depending on the vendor stats? Not me, not with my reputation, user
satisfaction, and my job at stake.
> Another point, if **YOU** are the one managing the system, do you
> want to create extra work for yourself, and extra risk if you don't
> have to or don't have the experience to work with them? In a
> lot of shops where ONE DBA is the rule, dear god, make sure you know
> what to do with raw systems before taking on the additional management.
> ( Your tips below illustrate the point )
There's NOTHING to do with RAW devices that's different or harder than using
COOKED devices and you save having to create filesystems if you are
comparing to COOKED FS files instead of devices. As to my 'tips' those are
not tips. That's how one MUST perform an external archive regardless of
whether the chunks are RAW device, COOKED device, or FS files. That's the
same anyway. It's your suggestion to archive externally rather than use the
simple ontape or to integrate the Informix archive with the other system
archives using onbar that adds more work for the DBA, not any of mine. Did
you READ what I wrote?
> Do you know what to do to restore? This will typically happen right
> as you were about to go on your vacation, so be prepared.
Most likely as I sitting down to dinner at a REALLY nice place. But, of
course I know how to restore, so do the system admins, the operators, and
all of the other DBAs so I'd better not get a call for something as trivial
and routine as a restore.
> Lastly, many environments change so rapidly that implementing raw
> disks can actually impede repurposing or modifying architectures
> at the pace necessary for the environment/business. To be sure you
> can argue all my points to the opposing view, but the bottom line is to
> do what is right for your environment and the best overall risk for you
> and your career. I would argue RAW is great if you can, but I'd argue
> more that it isn't worth the risk in most circumstances, even if you
> can.
You're kidding right? I discovered a month or so ago that some dense junior
SA had set up a bunch of new IDS server machines with JBOD instead of a
RAID10 array (Boy was I pissed) for servers that have been in production for
almost a year (a YEAR with no mirrors - AAAAHHH). I discovered it because I
was seeing slow response times for writes on IDS servers that seemed to be
tuned fine. Do you know what it took to get all that data (many GBs) moved
onto RAID10 arrays? Do you know how much pain it caused? None! It took
the SAs about 6 hours to copy the JBOD disk contents to a new RAID10 array
while the server was offline, and repoint the chunk links to the new RAW
devices and we were back online. Could it have been done with no downtime?
Sure, well with only 5 minutes downtime anyway, just add the new chunks as
IDS level mirrors, let them catch up, shutdown, swap the links between
primary and mirror chunks, restart, and drop the mirrors. Trivial really,
and I suggested doing just that but we had the downtime available, so why
not let the SAs do all the work.
> If you want to assume the risks, go for it, making doggone sure you
> do the backups, and perform an actual restore test to your satisfaction,
> something very few DBAs are willing to do. You will be quite surprised
If you're not doing the periodic restore tests and running periodic archive
verifies using archecker then you deserve to be fired for incompetence.
What does that have to do with using RAW devices? Or with how using RAW
devices affects how one performs the archives in the first place? And BTW,
while I was running the overall DBA group here, for almost six years, our
restore tests never failed and neither did any live restore. Since OL5.08
Informix archives have been utterly reliable. I trust ontape archives with
my company's data, with its business, with my own job.
> what your options are at restore-time, make sure you communicate them
> to management ahead of a disaster--you **will** be surprised what you can
> and can
On Thu, 12 May 2005 14:33:52 -0400, "Art S. Kagel" <kagel@bloomberg.net> wrote: >Data Goob wrote: >> Art S. Kagel wrote: >> > >> what your options are at restore-time, make sure you communicate them >> to management ahead of a disaster--you **will** be surprised what you can >> and cannot do if you never do a restore test--reading documentation and >> never restoring at least once is begging the boss to fire you for >> incompetence. Of course with the almost zero number of Informix people >> available to replace you, your job is probably pretty secure even if you >> screw it up. :-) > >Almost zero? Will you be in Denver? Come see how many Informix DBAs attend >my own sessions or Jonathan's or Lester's. There are a lot more of us out >there then you imagine. > Well said, Art. As for the job security comments . . . just remember that there are still a goodnumber of DBAs who will not be able to go to Denver. > >I see you ignored these last two paragraphs where I explicitely stated that >you cannot do what you propose this newbie do. I'm sorry if I'm flaming >you. I only flame about once a year, but you're spreading false information >to a newbie and asserting it's the truth. I can't let it go unchallenged. > >Art S. Kagel
"Art S. Kagel" <kagel@bloomberg.net> wrote in message news:4283A190.9060802@bloomberg.net... > SOME speed increase? You call 25-40% improvement SOME increase? Ask your > users if they'd like their apps to run 25-40% faster and see if they think > it's insignificant! Of course, i/o rates are just one of many factors contributing to overall user response time, and it is unlikely in the extreme that x% improvement in i/o times will translate to x% improvement in response times.
Art,
Your points are well made, and I'm willing to admit I could be
misinformed, but quite honestly you are working in an exceptional
environment compared to a lot of environments out there. You
make the assertion that everyone should simply get it when it
comes to disk drives, and quite honestly very few of us can keep
up with this. You also make the assumption that it is easy when
you should see what I have to work with when using sys admins--you
even allude to the jr DBA who got it wrong. Imagine that this is
95% of the environments out there. And as far as misinforming
people bang on EMC, they even preach their RAID5 is 'better',
so we even have to deal with that as well as a lot of other
vendor misinformation.
Also imagine how Informix is less than 3% of the market
so those of you out there that are quite comfortable with raw
files and Informix are in a very small minority. It just ain't
like Informix-on-raw is everywhere, it isn't! I used to preach
raw files everywhere but became weary dealing with people who
don't know the difference, or aren't knowledgable, can't do it,
etc., want to just get the system up, want to use MSSQL, etc.
Put your notes in the FAQ, I'm willing to refer people to the
FAQ not only for Informix but any other DB out there and your
time spent on this is even more valuable to the community at
large. I already refer to your RAID5 comments, you can bet
that I'll add your notes on raw-vs-cooked to my knowledge base.
Not all of us get to work in the same kind of environment you do,
or with Informix for that matter. You just have no idea what
some of the environments are like out there, just unbelievable.
Art S. Kagel wrote:
> Data Goob wrote:
>
>> Art S. Kagel wrote:
>>
>>> Heinz Weitkamp wrote:
>>>
>>>> Hi all,
>>>
>>>
>>>
>>>
>>> Go RAW all the way. Contrary to the Goob's understanding my own
>>> testing shows RAW devices to be 15-25% faster than COOKED devices
>>> which are in turn 10-15% faster than filesystem files, even on the
>>> best filesystem implementations that adds up to a 25% improvement.
>>> Well worth any minor inconveniences in my book.
>>
>>
>>
>> OK so you get some speed increases, but also the additional management
>
>
> SOME speed increase? You call 25-40% improvement SOME increase? Ask
> your users if they'd like their apps to run 25-40% faster and see if
> they think it's insignificant!
>
>> of linking raw files. The original post indicated he was a newbie.
>> Raw files are not necessarily difficult, but they ARE optional and not
>> something I would entrust to a newbie. It's a great learning experience,
>> but who's paying for it??
>
>
> This is NOT rocket science. The RAW device files are created at the
> time the disk partitions are created. All any DBA has to do is create a
> symbolic link to the RAW device (and he'd want to do that anyway so he
> can move/rename the actual file/device even if it's COOKED or a FS file)
> and ask the SA's (assuming he's not doubling as SA in which case the
> newbie moniker is probably misplaced) to set the permissions on the
> device and make sure they stay set. POOF! RAW chunks! And as John
> points out there is FAR more risk that some well intentioned SA will
> delete those HUGE chunk files.
>
>> But another point to make is that in many SAN environments, raw files are
>> a complete waste of time, with the data already being spread out over
>> many
>> drives, and, with speeds typically in the 10-15K rpm range, or even
>> faster,
>> you get speeds today on cooked files that are faster than raw was even a
>> couple of years ago.
>
>
> You are completely misinformed. We got SANs, hundreds of SANs, the very
> best SANS in existence. We also have dozens of IDS server machines.
> All using RAW disk because they still test out to be MUCH faster than
> COOKED devices or COOKED filesystem files, even on SANs.
>
> It sounds like you think I'm proposing using singleton disk drives when
> I say RAW. NOOO. As always I promote using RAID10 everywhere for
> anything that's not absolutely trivial. Not only that, we tend to plaid
> those RAID10 arrays into massive logical devices covering many dozens of
> mirrored pairs across multiple SANs. Either way you have to carve up
> those massive logical devices into smaller logical devices (still spread
> over many many spindles). For servers earlier than 9.40 I have the SAs
> carve out 2GB chunks, for 9.40+ they are carving several different sizes
> so that the DBAs have the flexibility to add a 2, 4, 8, or 16GB chunk
> where needed. This is easy stuff.
>
> Yes, the intricacies of EMC's SANs and VMs are a bit esoteric, and I
> leave that for the SAs, but you cannot say, "Ahh but what if the DBA is
> the SA?", because if he is, then he should know his jobs, both of them,
> and that part becomes routine very quickly.
>
> You assert that cooked files are as fast today as RAW on SANs. Have YOU
> tested it? Are you comparing RAW to COOKED with O_SYNC? Or are you
> depending on the vendor stats? Not me, not with my reputation, user
> satisfaction, and my job at stake.
>
>> Another point, if **YOU** are the one managing the system, do you
>> want to create extra work for yourself, and extra risk if you don't
>> have to or don't have the experience to work with them? In a
>> lot of shops where ONE DBA is the rule, dear god, make sure you know
>> what to do with raw systems before taking on the additional management.
>> ( Your tips below illustrate the point )
>
>
> There's NOTHING to do with RAW devices that's different or harder than
> using COOKED devices and you save having to create filesystems if you
> are comparing to COOKED FS files instead of devices. As to my 'tips'
> those are not tips. That's how one MUST perform an external archive
> regardless of whether the chunks are RAW device, COOKED device, or FS
> files. That's the same anyway. It's your suggestion to archive
> externally rather than use the simple ontape or to integrate the
> Informix archive with the other system archives using onbar that adds
> more work for the DBA, not any of mine. Did you READ what I wrote?
>
>> Do you know what to do to restore? This will typically happen right
>> as you were about to go on your vacation, so be prepared.
>
>
> Most likely as I sitting down to dinner at a REALLY nice place. But, of
> course I know how to restore, so do the system admins, the operators,
> and all of the other DBAs so I'd better not get a call for something as
> trivial and routine as a restore.
>
>> Lastly, many environments change so rapidly that implementing raw
>> disks can actually impede repurposing or modifying architectures
>> at the pace necessary for the environment/business. To be sure you
>> can argue all my points to the opposing view, but the bottom line is to
>> do what is right for your environment and the best overall risk for you
>> and your career. I would argue RAW is great if you can, but I'd argue
>> more that it isn't worth the risk in most circumstances, even if you
>> can.
>
>
> You're kidding right? I discovered a month or so ago that some dense
> junior SA had set u
John Carlson wrote: > On Thu, 12 May 2005 14:33:52 -0400, "Art S. Kagel" > <kagel@bloomberg.net> wrote: > > >>Data Goob wrote: >> >>>Art S. Kagel wrote: <SNIP>> > As for the job security comments . . . just remember that there are > still a goodnumber of DBAs who will not be able to go to Denver. Truth. This is the first time Bloomberg is allowing me to attend, and I'm here twelve years. And your main point is a good one - that as many as attend the conference there are probably a few times that number who can not, will not, or simply do not attend. Heck, I'm the only one attending from my company, none of the Informix or DB2 DBAs will be joining me. As to being able to replace us ... when I asked to get back to coding most of the time Bloomberg was able to find, train, and hire nine DBAs and a manager to replace me. That's more a statement of how easy it was to find qualified people [and to how dumb I was for working so hard and so quietly that no one realized that I could have used eight more people on my team] than it is a statement as to whether it really takes nine people to replace me. Art S. Kagel
Data Goob wrote: > Art, > > Your points are well made, and I'm willing to admit I could be > misinformed, but quite honestly you are working in an exceptional > environment compared to a lot of environments out there. You > make the assertion that everyone should simply get it when it > comes to disk drives, and quite honestly very few of us can keep > up with this. You also make the assumption that it is easy when > you should see what I have to work with when using sys admins--you > even allude to the jr DBA who got it wrong. Imagine that this is > 95% of the environments out there. And as far as misinforming > people bang on EMC, they even preach their RAID5 is 'better', > so we even have to deal with that as well as a lot of other > vendor misinformation. There actually is one implementation of RAID5 that stands up to testing as far superior to others and that is Apple Computer's XServe RAID. These units have two XOR engines and come remarkably close in performance to 0+1 (within a few percentage points) and also have the advantage of having minimal loss of performance when running in degraded mode. So far this year most of the storage I have worked with has been Apple's and it is a poorly kept secret that Oracle Corp. is now using them in their data center in Redwood Shores. -- Daniel A. Morgan University of Washington damorgan@x.washington.edu (replace 'x' with 'u' to respond)
Data Goob wrote: > Art, > > Your points are well made, and I'm willing to admit I could be > misinformed, but quite honestly you are working in an exceptional > environment compared to a lot of environments out there. You No argument. I am blessed. > make the assertion that everyone should simply get it when it > comes to disk drives, and quite honestly very few of us can keep Actually, while I STRONGLY recommend RAW disk and RAID10, I do not expect anyone to 'get it'. Indeed, every single time I post on either the subject of RAW-vs-COOKED or RAID5-vs-RAID10 I ALWAYS recommend MOST strongly that everyone test it for themselves, because, as you correctly point out, every environment is different. Even if a particular reader's RAID5 or COOKED FILE implementation is no better than the ones I've tested, their RAID10 or RAW implementation could be FAR WORSE. Nope, don't expect anyone to get anything except my own credo: Test it for yourself! Don't believe what's touted by the vendors, the Data Goob ;-), nor by me. > up with this. You also make the assumption that it is easy when > you should see what I have to work with when using sys admins--you It IS easy. But as I am fond of saying (until everyone around me is staring at the ceiling and shuffling their feet) "The problem with making things foolproof is that fools are so devious!". The reason that the newbie SA (not DBA) didn't allocate a RAID10 array for the new machines is because noone told him he should and the senior SAs didn't realize that they hadn't clued him in. > even allude to the jr DBA who got it wrong. Imagine that this is > 95% of the environments out there. And as far as misinforming > people bang on EMC, they even preach their RAID5 is 'better', > so we even have to deal with that as well as a lot of other > vendor misinformation. Ya, I know. But their 'better' RAID5 came from the DG Clariion purchase and it was those arrays that we were comparing RAID5 to RAID10 on originally! > Also imagine how Informix is less than 3% of the market > so those of you out there that are quite comfortable with raw > files and Informix are in a very small minority. It just ain't > like Informix-on-raw is everywhere, it isn't! I used to preach > raw files everywhere but became weary dealing with people who > don't know the difference, or aren't knowledgable, can't do it, > etc., want to just get the system up, want to use MSSQL, etc. Right, Oracle for instance is less affected > Put your notes in the FAQ, I'm willing to refer people to the > FAQ not only for Informix but any other DB out there and your > time spent on this is even more valuable to the community at > large. I already refer to your RAID5 comments, you can bet > that I'll add your notes on raw-vs-cooked to my knowledge base. > Not all of us get to work in the same kind of environment you do, > or with Informix for that matter. You just have no idea what > some of the environments are like out there, just unbelievable. > <SNIP> This is in the FAQ. Here's an extract: ...the safety issue of cooked files is no longer a problem. The big problem with cooked files still is performance. All writes and reads to/from cooked files MUST go through the UNIX buffer cache. This means an additional copy from the server's output buffer to the UNIX buffer page then a synchronous write to disk. This is opposed to a write to a way file where the server's output buffer is written directly to disk without the intervening copy. This is just faster. Anyone who has written anything that can test this can attest to the difference in speed. Here are my test results: FileType Sync? Times (real/user/system) 2run avg --------------- ----- ----------------------------------------- Filesystem file N 14.40/3.70/2.52 Y 15.02/3.61/2.63 Cooked disk N 12.81/3.74/2.24 Y 13.42/3.84/2.43 Raw disk N 9.32/3.67/1.52 Y 9.40/3.66/1.44 From this you can clearly see the cost of Cooked files and of synced cooked files. The tests were done with a version of my ul.ec utility modified to optionally open the output file O_SYNC. Cooked disk partition is almost 50% slower than raw disk partition and cooked filesystem files are almost 60% slower. The penalty for O_SYNC is an additional 5% for cooked files and negligible for RAW files (as expected). The test file was 2.85MB written using 4K I/O pages (the default fopen application buffer size) which should simulate Informix performance. The Cooked and Raw disk partition tests were conducted to a singleton 9GB Fast Wide SCSI II drive using the raw and cooked device files corresponding to the same drive.
Neil Truby wrote: > "Art S. Kagel" <kagel@bloomberg.net> wrote in message > news:4283A190.9060802@bloomberg.net... > > >>SOME speed increase? You call 25-40% improvement SOME increase? Ask your >>users if they'd like their apps to run 25-40% faster and see if they think >>it's insignificant! > > > Of course, i/o rates are just one of many factors contributing to overall > user response time, and it is unlikely in the extreme that x% improvement in > i/o times will translate to x% improvement in response times. > > Granted Neil. Perhaps my point was too broad to penetrate to the heart of what I wanted to say - so no point at all ;-( I did not mean to imply that a 25% increase in IO speed would translate into a 25% increase in application throughput or response times. My point was just that - just as our users would recognize a 25-40% improvement as something significant, recognizable, and valuable - so should we not devalue the impact that a 25-40% increase would have on overall server performance. Art S. Kagel
DA Morgan wrote: > Data Goob wrote: > >> Art, >> >> Your points are well made, and I'm willing to admit I could be >> misinformed, but quite honestly you are working in an exceptional >> environment compared to a lot of environments out there. You >> make the assertion that everyone should simply get it when it >> comes to disk drives, and quite honestly very few of us can keep >> up with this. You also make the assumption that it is easy when >> you should see what I have to work with when using sys admins--you >> even allude to the jr DBA who got it wrong. Imagine that this is >> 95% of the environments out there. And as far as misinforming >> people bang on EMC, they even preach their RAID5 is 'better', >> so we even have to deal with that as well as a lot of other >> vendor misinformation. > > > There actually is one implementation of RAID5 that stands up to testing > as far superior to others and that is Apple Computer's XServe RAID. > These units have two XOR engines and come remarkably close in > performance to 0+1 (within a few percentage points) and also have the > advantage of having minimal loss of performance when running in degraded > mode. So far this year most of the storage I have worked with has been > Apple's and it is a poorly kept secret that Oracle Corp. is now using > them in their data center in Redwood Shores. Interesting. Got to keep that in mind. But my big question, which you are unlikely to have an answer to, is did Apple solve the partial media failure problem in RAID5? Art S. Kagel