Re: Can Disk mirroring/db snapshots replace the usage of taking database backups?
Posted in 2008
Topics: Storage & Space Management, Server Administration, Logging & Checkpoints
Neil Truby wrote:
> "Fernando Nunes" <spam@onlinedomus.net> wrote in message
> news:ftjj02$apt$1@aioe.org...
>
>> If a fast recovery needs human intervention (and nothing wrong happened to
>> the disks) than you're facing a bug... This should be true to any
>> relational database system with logging (or redo? logs).
>>
>> Database must know exactly what was written and what wasn't... Things like
>> disk controllers cache can confuse this..., but in any case is a bug...
>
> I agree.
> So, it begs the question: "What is onmode -c block for?" (in the context of
> "external" backups).
>
>
Any time the engine is started, we go through
1) physical recovery
2) roll forward of the contents of the logical log files
3) roll back of any uncommitted transactions after step 2
Physical recovery gets us to the point of a physically consistent disk
image as of the last successful checkpoint. This is performed by
restoring the pages in the physical log to the chunks. The page is
written to the physical log only if the page existed as of the last
checkpoint and only if it has not been written already to the physical
log file.
The key to being able to recover the instance is to have a consistent
starting point. With IDS that consistent starting point is the last
checkpoint. With an online backup, we initiate a 'backup checkpoint' as
the backup is started and then as the page would be written to the
physical log file, we also check to see if the page existed prior start
of the backup checkpoint. If it did then we will write the prior image
of the page to the backup. In this way as part of the restore, we are
able to logically extend the functionality of the physical log file to
include these 'prior image pages' and thus are able to restore the
instance as of the backup checkpoint.
Onmode -c block prevents IOs (writes) from being performed. That makes
it possible to perform an external backup and to get a consistent image
of the disks as of that point in time. This is similar to what would
happen with a system crash because with a system crash, we would get a
physically consistent image as of the last successful checkpoint from
the physical log contents and then use the logical log to roll forward
to the most recent contents to the logical log. Since we have no way to
cause the external backup to accept the 'prior image pages' like we do
with an online backup, our recourse is to block writes while the
external backup is underway.
If writes to the chunks were not blocked, then while the external copy
was being made, then there would not be a consistent image of the disks.
The first part of the copy could belong to one checkpoint and the last
part to another.
"Madison Pruet" <mpruet1@verizon.net> wrote in message news:47FE20AE.2050404@verizon.net... > If writes to the chunks were not blocked, then while the external copy was > being made, then there would not be a consistent image of the disks. The > first part of the copy could belong to one checkpoint and the last part to > another. Sorry to labour this, but I just don't get it! In a disk snapshot using consistency groups - that is, all the pages in all the LUNs in the consistency group are snapped to exactly the same point in time - fast recovery should work just fine. Shouldn't it? If the asertion is that Fast Recovery cannot be relied upon from an exact, time-self-consistent copy of the live chunks, isn't that the same as saying that it can't be relied upon from the primary chunks themselves?!
Neil Truby wrote: > "Madison Pruet" <mpruet1@verizon.net> wrote in message > news:47FE20AE.2050404@verizon.net... > >> If writes to the chunks were not blocked, then while the external copy was >> being made, then there would not be a consistent image of the disks. The >> first part of the copy could belong to one checkpoint and the last part to >> another. > > Sorry to labour this, but I just don't get it! > In a disk snapshot using consistency groups - that is, all the pages in all > the LUNs in the consistency group are snapped to exactly the same point in > time - fast recovery should work just fine. Shouldn't it? If we only wanted to support external backup in a disk snapshot using consistency groups and the disk snapshot software could support multiple disk subsystems, then that would probably work. However, we want to be able to support external backups in other environments as well. > > If the asertion is that Fast Recovery cannot be relied upon from an exact, > time-self-consistent copy of the live chunks, isn't that the same as saying > that it can't be relied upon from the primary chunks themselves?! That is not a logical deduction from what I posted. I described how when performing recovery that we obtained a point of consistency by first applying all of the pages from the physical log - and how that got us to the consistent point in time as of the last checkpoint. I also described how we logically extended the physical log during the online archive by saving the before images of any page which existed prior to the archive checkpoint, but which was updated during the backup. > >
On 10 Apr, 17:00, "Neil Truby" <neil.tr...@ardenta.com> wrote: > "Madison Pruet" <mpru...@verizon.net> wrote in message > > news:47FE20AE.2050404@verizon.net... > > > If writes to the chunks were not blocked, then while the external copy was > > being made, then there would not be a consistent image of the disks. The > > first part of the copy could belong to one checkpoint and the last part to > > another. > > Sorry to labour this, but I just don't get it! > In a disk snapshot using consistency groups - that is, all the pages in all > the LUNs in the consistency group are snapped to exactly the same point in > time - fast recovery should work just fine. Shouldn't it? > > If the asertion is that Fast Recovery cannot be relied upon from an exact, > time-self-consistent copy of the live chunks, isn't that the same as saying > that it can't be relied upon from the primary chunks themselves?! If the order of writes to the chunks does not change i.e. writes to the log happen before other writes and no corruption ever get written to disk then yes disk snapshots should be ok. If someone adds a device to the database and not to the consistency group then this all fails...and guess who gets the blame..yes the DBA and Informix. Or if the host has to be rebuilt and the consistency group is not recreated .... You create dependency between the DBA and the UNIX SA/Storage Team....if this dependency is forgotten then it will not be noticed until a restore from a disk snapshot fails and we all know how often those are done in general. You may find you need to disable some level of caching in the storage to make this work correctly as well. I remember with EMC Symmetrix and SRDF there was a flag to disable write re-ordering. What did the SA do? Not set the flag as..performance tanks with it set. The UNIX SA said to me "well if the Oracle database breaks tough we live with it and I did not tell anyone until now!!"...and that 'hidden' issue was in production as well....