Duration of RSS init (ontape -p) vs. ontape -r?
Posted in 2017
User experienced 20x slower restore time with 'ontape -r' (80 min) versus 'ontape -p' (4 min) for a 34GB database on identically configured Solaris/Informix systems. Root cause was asymmetrical mirrored SSDs created with mismatched cylinder/block geometry after replacing drives with larger ones. After detaching newer, larger slices to eliminate asymmetric mirroring, restore time dropped to under 4 minutes, resolving the issue.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Backup & Restore, Storage & Space Management, Server Administration
Informix 12.10
Solaris 10 1/13
Oracle T4-1, with SSDs.
I experience the following, which seems strange to me:
We use RSS between two identically configured T4-1 boxes. If we take an ontape
archive from the primary, and use 'ontape -p' to restore it on the RSS (after
which we run 'onmode -d RSS', everything works fine. The archive from ontape
is about 34GB, and it took about 4 minutes for the 'ontape -p' to run on the
RSS box.
We created a 2nd Informix instance on that same machine (the RSS), with
exactly the same chunks, and also restored that same ontape archive to the new
instance, using 'ontape -r'. This created a separate, independent version of
the same database, which we will use as a disposable, experimental version of
the "real database" that exists on the primary and RSS. Of course, the path
names to the chunks in the second instance could not be exactly the same as
the first instance, so we used the -rename option of 'ontape -r'.
It took about 80 minutes for the 'ontape -r' on the 2nd instance to run-- as
opposed to the 3 minutes 45 seconds for the 'ontape -p' on the first instance.
Why is there such a difference? We tried it several times, and got the same
results.
The difference seems so huge (factor of about 20).
Thank you for any comments.
DG
P.S. Just for grins, we also redid the test with just the second Informix
instance live. (I.e., the first instance of Informix-- the RSS-- was offline).
Same result (the restore took 80 minutes).
blockquote, div.yahoo_quoted { margin-left: 0 !important; border-left:1px
#715FFA solid !important; padding-left:1ex !important; background-color:white
!important; } Probably because the chunks had to be created and expanded. I'm
guessing that this is a brand new instance and using cooked file chunks.
Sent from Yahoo Mail for iPad
On Monday, January 30, 2017, 5:31 PM, DAVID GROVE <david.grove@alaska.gov>
wrote:
Informix 12.10
Solaris 10 1/13
Oracle T4-1, with SSDs.
I experience the following, which seems strange to me:
We use RSS between two identically configured T4-1 boxes. If we take an ontape
archive from the primary, and use 'ontape -p' to restore it on the RSS (after
which we run 'onmode -d RSS', everything works fine. The archive from ontape
is about 34GB, and it took about 4 minutes for the 'ontape -p' to run on the
RSS box.
We created a 2nd Informix instance on that same machine (the RSS), with
exactly the same chunks, and also restored that same ontape archive to the new
instance, using 'ontape -r'. This created a separate, independent version of
the same database, which we will use as a disposable, experimental version of
the "real database" that exists on the primary and RSS. Of course, the path
names to the chunks in the second instance could not be exactly the same as
the first instance, so we used the -rename option of 'ontape -r'.
It took about 80 minutes for the 'ontape -r' on the 2nd instance to run-- as
opposed to the 3 minutes 45 seconds for the 'ontape -p' on the first instance.
Why is there such a difference? We tried it several times, and got the same
results.
The difference seems so huge (factor of about 20).
Thank you for any comments.
DG
P.S. Just for grins, we also redid the test with just the second Informix
instance live. (I.e., the first instance of Informix-- the RSS-- was offline).
Same result (the restore took 80 minutes).
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Original post:
Informix 12.10
Solaris 10 1/13
Oracle T4-1, with SSDs.
I experience the following, which seems strange to me:
We use RSS between two identically configured T4-1 boxes. If we take an ontape
archive from the primary, and use 'ontape -p' to restore it on the RSS (after
which we run 'onmode -d RSS', everything works fine. The archive from ontape
is about 34GB, and it took about 4 minutes for the 'ontape -p' to run on the
RSS box.
We created a 2nd Informix instance on that same machine (the RSS), with
exactly the same chunks, and also restored that same ontape archive to the new
instance, using 'ontape -r'. This created a separate, independent version of
the same database, which we will use as a disposable, experimental version of
the "real database" that exists on the primary and RSS. Of course, the path
names to the chunks in the second instance could not be exactly the same as
the first instance, so we used the -rename option of 'ontape -r'.
It took about 80 minutes for the 'ontape -r' on the 2nd instance to run-- as
opposed to the 3 minutes 45 seconds for the 'ontape -p' on the first instance.
Why is there such a difference? We tried it several times, and got the same
results.
The difference seems so huge (factor of about 20).
Thank you for any comments.
DG
P.S. Just for grins, we also redid the test with just the second Informix
instance live. (I.e., the first instance of Informix-- the RSS-- was offline).
Same result (the restore took 80 minutes).
Response:
My guess would be the following. When you do ontape -p, we just restore the
data, and the restore doesn't get to the part where it clears the physical and
logical logs. When you do an ontape -r, it does clear the physical and logical
logs. Easiest way to see if that clearing the logs is the difference look for
the following message in your MSGPATH file (you won't see this message
directly after an ontape -p finishes, but you should see it as part of the
messages before ontape -r finishes...gets back to the unix prompt):
08:35:30 Clearing the physical and logical logs has started
08:35:30 Cleared 107 MB of the physical and logical logs in 0 seconds
Jacques Renaut
IBM Informix
You also need to test the IO performance on the two secondary machines. It
could be that one machine is not performing IO nearly as fast as the other.
Madison Pruet
Retired and Loving it
On Monday, January 30, 2017 8:17 PM, Madison Pruet <madison_pruet@yahoo.com>
wrote:
#yiv8669123438 blockquote, #yiv8669123438 div.yiv8669123438yahoo_quoted
{margin-left:0 !important;border-left:1px #715FFA solid
!important;padding-left:1ex !important;background-color:white;}Probably
because the chunks had to be created and expanded. I'm guessing that this is a
brand new instance and using cooked file chunks.
Sent from Yahoo Mail for iPad
On Monday, January 30, 2017, 5:31 PM, DAVID GROVE <david.grove@alaska.gov>
wrote:
Informix 12.10
Solaris 10 1/13
Oracle T4-1, with SSDs.
I experience the following, which seems strange to me:
We use RSS between two identically configured T4-1 boxes. If we take an ontape
archive from the primary, and use 'ontape -p' to restore it on the RSS (after
which we run 'onmode -d RSS', everything works fine. The archive from ontape
is about 34GB, and it took about 4 minutes for the 'ontape -p' to run on the
RSS box.
We created a 2nd Informix instance on that same machine (the RSS), with
exactly the same chunks, and also restored that same ontape archive to the new
instance, using 'ontape -r'. This created a separate, independent version of
the same database, which we will use as a disposable, experimental version of
the "real database" that exists on the primary and RSS. Of course, the path
names to the chunks in the second instance could not be exactly the same as
the first instance, so we used the -rename option of 'ontape -r'.
It took about 80 minutes for the 'ontape -r' on the 2nd instance to run-- as
opposed to the 3 minutes 45 seconds for the 'ontape -p' on the first instance.
Why is there such a difference? We tried it several times, and got the same
results.
The difference seems so huge (factor of about 20).
Thank you for any comments.
DG
P.S. Just for grins, we also redid the test with just the second Informix
instance live. (I.e., the first instance of Informix-- the RSS-- was offline).
Same result (the restore took 80 minutes).
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Thank you, Madison.
Madison nailed it.
I had recently replaced one side of some of the mirrors on this box. But,
because the economics of SSDs had changed so much from the time of original
purchase, I replaced the drives with larger drives. My thinking was that I
would create slices that would be compatible with the existing drives (for
mirroring purposes) and also gain some extra space (for other purposes). Then,
later on, i could replace the other side of the mirror, also with larger
drives, and, once again, have identical drives-- or repeat the process with
larger drives, thus continuing this "leapfroging" process with the disk space
game. I want to keep one side of the mirrors significantly newer than the
other side because they are SSDs, and I want to make sure all my mirrored SSDs
have at least one side that is less than two years old.
Well, the cylinder/block geometry (of course) is not an exact match, and I
made the smallest slices on the new drives that were greater than or equal to
the slices on the existing SSDs, and made mirrors of them.
The asymmetrical mirrored drives work fine, but, there is a write performance
to pay for the mismatch.
After Madison called it to my attention, I detached the newer, larger slices
from these asymmetrical mirrors (made "1-way" mirrors) and retried the 'ontape
-r' on the second Informix instance (that had been using the asymmetric
mirrors).
Voila! Problem solved. The restore ran in less than 4 minutes.
Thank you, Madison and Jacques, for your helpful responses.
Regards,
DG
P.S. I might add that I created the asymmetric mirrors only after consulting the documentation, which affirmed SVM's ability to accommodate the same. If any performance degradation were mentioned, I missed it.