RAW devices - a thing of the past ?
Posted in 1999
Topics: Performance & Tuning, Installation, Setup & Upgrades, Platform-Specific Issues
Hi All I've been mulling this one over for a little while now, and I'd thought I'd get your thoughts on it. The discussion centers around using raw devices on Solaris 2.6 and upwards. We all know and acknowlegde that using raw devices gives us the best possible thruput for db i/o, because operations to raw devices do not have to go thru the o/s's buffer - the reads/writes are *direct*, not indirect as with cooked files. Bypassing the o/s buffercache is even more important when related to the way that Solaris handles virtual memory. What ends up happending, if you use cooked files, is that the data for those files ends up *dominating* the o/s buffercache - why, because there is usually a lot of it (the data), and it's accessed very frequently. Consequently free vm page availability suffers, and the whole vm management is subverted, at the expense of all other present and future processes. So, assume (for whatever reason - we know that they exists and are valid), you just want to stick with cooked files. What can you do to avoid this problem of your db data repeatedly and constantly filling up the buffercache ? Until recently I was only aware of one option - and that is to employ the new Priority Paging algorithm (available of a patch for Sol2.6, and comes standard with Sol7+) which patches the vm manager is Solaris to "be smarter" about what it keeps in the cache and what it flushes out - depending on whether it's "data" or executable code. The aim is that, when it comes time to free some pages, the page stealing deamon will free up data pages in preference to executable pages, on the basis that we're always going to have more data in cache than executable, but it's probably most important to cache as much executable as possible. Now, as I see it, there are a couple is issues with this approach. Firstly it can only be a band-aid measure - ie, it will relieve the symptom but not the cause. The page scanner is still going to have to work just as hard in examining the pages and types of pages it may or not not free. Secondly it relies crucially on the permission bits on the files that get cached. If your data files have an x bit set, then there going to be treated as high prioroity executables. So where does this all leave us ? Well, at this point, along comes a little gem by the name of "forcedirectio". I was told about this by a guy (forget his name, but whoever you are - Kudos ! :) on another unix-based newsgroup, in reply to a guy who was having trouble comprehending how his SAP installation was able to chew up GB after GB of memory, no matter how much more he added. The answer was he was using cooked files, and they were just going straight into, and remaining in, the cache - hurrah for Solaris vm management - it's doing it's job and making use of our expensive memory, but in this case it's doing it's job a little too well. I followed this up and the guy tells me about the apparently little known mount option forcedirectio, which makes all i/o to that f/s direct - ie, *bypassing the buffercache*... man mount_ufs(1M): ... noforcedirectio | forcedirectio If forcedirectio is specified and supported by the file system, then for the duration of the mount forced direct I/O will be used. If the filesystem is mounted using forcedirectio, then data is transferred directly between user address space and the disk. If the filesystem is mounted using nofor- cedirectio, then data is buffered in kernel address space when data is transferred between user address space and the disk. forceddirectio is a performance option that bene- fits only from large sequencial data transfers. The default behavior is noforcedirectio. ... This option can be used to remount a live f/s on the fly. The best way to actually see how memory is being allocated on a very general level, IMHO, is with the prtmem tool in the RMCmem package. Using this package to sample data over serveral days, I was able to get a picture of the effects on using and not using forcedirectio, and the results were just as expected: before - buffercache leaps up as soon as people start using the db, and stays there until some time after they stop using it; after - buffercache is used and then freed when other data processing processes (eg a perl script chewing thru many megs of data) fire up, and then drops back when other processes need some room. Basically, normal memory management in the presence of many gigs of db data on cooked files. The bottom line: memory is available when and how it is needed, and, you can get a clear and accuratre picture of your machine's memory usage - when, where, by what, how much, for how long, etc, instead of just seeing free memory flatlining as soon as the db gets u! ! sed. CONCLUSION ---------- Raw devices provide the greatest i/o thruput becuase they bypass the buffercache. Raw devices are not as easy to maintain for a dynamic system than cooked files. Cooked files are slow because they go theu the buffercahce. Priority paging can help this, but only to a degree. Mounting a f/s holding cooked db files with forcedirectio gives you most of the benefits of using raw space, without the drawbacks. Any comments/discussion most welcome. Thanks for your time in reading this ! :) Andrew Reardon UNIX/Informix Administrator Aircraft Systems, Boeing Aust. Ltd. Ph: +61 7 3306 3346 Mob: +61 0419 745 831
Hi, I'm an somewhat experienced administrator of SUN Solaris and Sybase. I think that the problem you mention, is unique to each system and its usage. In general I agree on what you say. but there is more to it, as I see it. On Sybase for instance, the direct IO from RAW devices is faster, only because the Sybase dataserver is also caching data itself. My assumption is, that the speed of any OS's buffered IO is probably mostly the same as for a dataserver of the same technical quality. Developers aim at the best IO and normally acheive it for each individual system. Typically, I think, tuning can do some also. For a dataserver doing caching for itself, buffers IO is not the best - double buffering. For that reason you can tell the dataserver not to cache specific segments of its dataareas. In the end, it must be up to the system administrator to run tests to check if the buffering within the dataserver or within the OS is faster. In general, I would leave it up to the dataserver to cache data, since the dataserver writers most certainly knows best, what should be cached to get the best throughput. An example: For most types of large databases, it is more important to cache indexdata than real data. This is clear from the point of view, if you know where to find data, it is faster to find it. And now we are having more types of data. Data and indexes. The OS'es are unable to separate these in different caches, but the better dataservers normally are! Regards Henrik.
Henrik K. Larsen wrote in message ... >Hi, > >I'm an somewhat experienced administrator of SUN Solaris and Sybase... AND you also manage to fit being a striker for Glasgow Celtic into your spare time?
I would want to benchmark this feature to see how close to RAW performance the filesystem can get. Keep in mind that you still have to get to each page of the cooked file through the filesystem driver, on top of and VM and raw IO drivers, and via the inode block list. Also there is no guarantee, in a filesytem, that consecutive file blocks will represent contiguous disk space. This is more important than it would seem because of Informix's readahead and bigwrites features as well as buffer cleaning algorithms, especially during checkpoints. These optimizations all assume that consecutive pages within a single chunk will be physically contiguous. I suspect that you are correct that performance is far better than a 'normal' filesystem but you will have to benchmark it using Informix. IBM, HP, & EMC all insist that their filesystems are so sophisticated that RAW device use is no longer neccessary but benchmarks have not born this out to date. Art S. Kagel andrew@andrew.bal.bna.boeing.com wrote: > > Hi All > > I've been mulling this one over for a little while now, and I'd thought I'd get your thoughts on it. > > The discussion centers around using raw devices on Solaris 2.6 and upwards. > > We all know and acknowlegde that using raw devices gives us the best possible thruput for db i/o, because operations to raw devices do not have to go thru the o/s's buffer - the reads/writes are *direct*, not indirect as with cooked files. > > Bypassing the o/s buffercache is even more important when related to the way that Solaris handles virtual memory. What ends up happending, if you use cooked files, is that the data for those files ends up *dominating* the o/s buffercache - why, because there is usually a lot of it (the data), and it's accessed very frequently. Consequently free vm page availability suffers, and the whole vm management is subverted, at the expense of all other present and future processes. > > So, assume (for whatever reason - we know that they exists and are valid), you just want to stick with cooked files. What can you do to avoid this problem of your db data repeatedly and constantly filling up the buffercache ? Until recently I was only aware of one option - and that is to employ the new Priority Paging algorithm (available of a patch for Sol2.6, and comes standard with Sol7+) which patches the vm manager is Solaris to "be smarter" about what it keeps in the cache and what it flushes out - depending on whether it's "data" or executable code. The aim is that, when it comes time to free some pages, the page stealing deamon will free up data pages in preference to executable pages, on the basis that we're always going to have more data in cache than executable, but it's probably most important to cache as much executable as possible. > > Now, as I see it, there are a couple is issues with this approach. Firstly it can only be a band-aid measure - ie, it will relieve the symptom but not the cause. The page scanner is still going to have to work just as hard in examining the pages and types of pages it may or not not free. Secondly it relies crucially on the permission bits on the files that get cached. If your data files have an x bit set, then there going to be treated as high prioroity executables. > > So where does this all leave us ? Well, at this point, along comes a little gem by the name of "forcedirectio". I was told about this by a guy (forget his name, but whoever you are - Kudos ! :) on another unix-based newsgroup, in reply to a guy who was having trouble comprehending how his SAP installation was able to chew up GB after GB of memory, no matter how much more he added. The answer was he was using cooked files, and they were just going straight into, and remaining in, the cache - hurrah for Solaris vm management - it's doing it's job and making use of our expensive memory, but in this case it's doing it's job a little too well. > > I followed this up and the guy tells me about the apparently little known mount option forcedirectio, which makes all i/o to that f/s direct - ie, *bypassing the buffercache*... man mount_ufs(1M): > > ... > noforcedirectio | forcedirectio > If forcedirectio is specified and > supported by the file system, then > for the duration of the mount > forced direct I/O will be used. If > the filesystem is mounted using > forcedirectio, then data is > transferred directly between user > address space and the disk. If the > filesystem is mounted using nofor- > cedirectio, then data is buffered > in kernel address space when data > is transferred between user address > space and the disk. forceddirectio > is a performance option that bene- > fits only from large sequencial > data transfers. The default > behavior is noforcedirectio. > ... > > This option can be used to remount a live f/s on the fly. > > The best way to actually see how memory is being allocated on a very general level, IMHO, is with the prtmem tool in the RMCmem package. Using this package to sample data over serveral days, I was able to get a picture of the effects on using and not using forcedirectio, and the results were just as expected: before - buffercache leaps up as soon as people start using the db, and stays there until some time after they stop using it; after - buffercache is used and then freed when other data processing processes (eg a perl script chewing thru many megs of data) fire up, and then drops back when other processes need some room. Basically, normal memory management in the presence of many gigs of db data on cooked files. The bottom line: memory is available when and how it is needed, and, you can get a clear and accuratre picture of your machine's memory usage - when, where, by what, how much, for how long, etc, instead of just seeing free memory flatlining as soon as the db gets u! > ! > sed. > > CONCLUSION > ---------- > > Raw devices provide the greatest i/o thruput becuase they bypass the buffercache. > Raw devices are not as easy to maintain for a dynamic system than cooked files. > Cooked files are slow because they go theu the buffercahce. > Priority paging can help this, but only to a degree. > > Mounting a f/s holding cooked db files with forcedirectio gives you most of the > benefits of using raw space, without the drawbacks. > > Any comments/discussion most welcome. > > Thanks for your time in reading this ! :) > > Andrew Reardon > UNIX/Informix Administrator > Aircraft Systems, Boeing Aust. Ltd. > Ph: +61 7 3306 3346 Mob: +61 0419 745 831
In article <7o3lrs$gov$1@taliesin.netcom.net.uk>, "Neil Truby" <ntruby@netcomuk.co.uk> wrote: > > Henrik K. Larsen wrote in message ... > >Hi, > > > >I'm an somewhat experienced administrator of SUN Solaris and Sybase... > > AND you also manage to fit being a striker for Glasgow Celtic into your > spare time? > > Way to confuse our American friends, Neil, viz. soccer/football ? Tick VG Glyn Balmer (UK) -- If it always works, why don't parachutists pull the emergency 'chute first? Sent via Deja.com http://www.deja.com/ Share what you know. Learn what you don't.
Glynnie wrote in message <7o9c9m$9p8$1@nnrp1.deja.com>... >> AND you also manage to fit being a striker for Glasgow Celtic into >your >> spare time? >> >> >Way to confuse our American friends, Neil, viz. soccer/football ? >Tick VG Obviously if I'd just put "Celtic" everyone would have thought I meant basketball or something!