I/O errors during backup
Posted in 2009
Topics: Backup & Restore, Installation, Setup & Upgrades, Storage & Space Management, Logging & Checkpoints, Migration, Import/Export & Data Conversion
Hi Guys,
on an old installation I inherited, I have the following issue :
Informix Dynamic Server 7.31 UC6
Mips Platform running Reliant Unix.
2 I/O errors (with a history) appear during the backup with ontape.
But backup seems to complete anyway.
116942 20:02:26 Level 0 Archive started on rootdbs, patient
116943 20:05:03 I/O read chunk 1, pagenum 14047, pagecnt 1 --> errno = 5
116944 20:05:03 I/O read chunk 1, pagenum 14047, pagecnt 1 --> errno = 5
116945 20:07:57 Checkpoint Completed: duration was 10 seconds.
116946 20:09:27 Logical Log 26683 Complete.
116947 20:09:28 Logical Log 26683 - Backup Started
116948 20:09:38 Logical Log 26683 - Backup Completed
116964 21:09:42 Logical Log 26684 Complete.
116965 21:09:44 Logical Log 26684 - Backup Started
116966 21:09:53 Logical Log 26684 - Backup Completed
116975 21:50:52 Archive on rootdbs, patient Completed.
First question, is this backup reliable ?
(onstat -g arc indicates 20u02 which is the startime of the backup)
The history of the I/O errors :
Complete database (only 2 dbspaces) is on slice 0 of a LUN (RAID5)
Some time ago the OS reported read errors on D200 (part of RAID5),
but no actions were taken, application ran fine.
A couple of days ago, the disk D500 (part of same RAID5) completely failed,
raid was in degraded mode.
After a while informix was blocked, last logical log in use, logical logs
could not be backed up (read error).
Got informix running again by changing LTAPEDEV to /dev/null (lost some
logical logs), removed the logical log causing the problem, and took a level0
backup.
OK informix running again.
Then raid-intervention replace D500, got 3 read errors but raid was rebuild to
optimal state.
During 2 days no error-messages.
Yesterday evening D200 (with the original read-error) was also replaced and
raid was rebuild to optimal state.
But backup afterwards generates error-messages again.
To exclude the OS and HW, we will try a dd-command to read the entire slice
If the problem is informix-related, are there any suggestions, other than
unloading and re-initialising the database ?
Thx for any reactions,
Jacques Lapeire
I know that you inherited this, but I can't miss the opportunity to
reiterate the the community as a whole:
RAID5 DOES NOT WORK!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! Your data
is not safe if it is stored using RAID5 technology, not anybodies!
NO RAID5!!! NO RAID5!!! NO RAID5!!! NO RAID5!!! NO RAID5!!! NO
RAID5!!! NO RAID5!!! NO RAID5!!!
This is EXACTLY the scenario I, and others, have been warning about for over
10 years!
As to your question about the integrity of the archive, use archecker to
validate the archive. You should do the follow in any case:
1. Export the entire database (all databases) to flat files using
dbexport as a last ditch safety net.
2. If the archive is good, create new chunks FOR ALL CHUNKS (not just the
ones that have already failed!) on clean good NEW non-RAID5 drives, link the
new disk space to the original chunk path links (I do hope they used links
for the chunk paths and not the actual device paths) and restore the
archive.
3. If the archive is not good, then create new clean non-RAID5 chunks FOR
ALL CHUNKS (not just the ones that have already failed!, reinitialize the
engine in a new ROOT chunk, add the other dbspaces and chunks that you will
need and restore the data using dbimport.
Just a suggestion, but since IDS v 7.31 is going out of support in 2009 your
employer should strongly consider upgrading to IDS 11.50 (which is over 40%
faster than 7.31 anyway and sports many new features) and moving the server
to newer, faster, cheaper hardware WITHOUT RAID5! BTW, I don't care how
"optimal" your vendor thinks that the RAID5 setup is now, is sucks! You are
getting 40% of the write performance of a comparable RAID10 array and your
data is still not safe!
Art
On Fri, Jan 16, 2009 at 8:54 AM, JACQUES LAPEIRE <
jacques.lapeire@fujitsu-siemens.com> wrote:
> Hi Guys,
>
> on an old installation I inherited, I have the following issue :
>
> Informix Dynamic Server 7.31 UC6
> Mips Platform running Reliant Unix.
>
> 2 I/O errors (with a history) appear during the backup with ontape.
> But backup seems to complete anyway.
>
> 116942 20:02:26 Level 0 Archive started on rootdbs, patient
> 116943 20:05:03 I/O read chunk 1, pagenum 14047, pagecnt 1 --> errno = 5
> 116944 20:05:03 I/O read chunk 1, pagenum 14047, pagecnt 1 --> errno = 5
> 116945 20:07:57 Checkpoint Completed: duration was 10 seconds.
> 116946 20:09:27 Logical Log 26683 Complete.
> 116947 20:09:28 Logical Log 26683 - Backup Started
> 116948 20:09:38 Logical Log 26683 - Backup Completed
> 116964 21:09:42 Logical Log 26684 Complete.
> 116965 21:09:44 Logical Log 26684 - Backup Started
> 116966 21:09:53 Logical Log 26684 - Backup Completed
> 116975 21:50:52 Archive on rootdbs, patient Completed.
>
> First question, is this backup reliable ?
> (onstat -g arc indicates 20u02 which is the startime of the backup)
>
> The history of the I/O errors :
>
> Complete database (only 2 dbspaces) is on slice 0 of a LUN (RAID5)
> Some time ago the OS reported read errors on D200 (part of RAID5),
> but no actions were taken, application ran fine.
> A couple of days ago, the disk D500 (part of same RAID5) completely failed,
> raid was in degraded mode.
> After a while informix was blocked, last logical log in use, logical logs
> could not be backed up (read error).
> Got informix running again by changing LTAPEDEV to /dev/null (lost some
> logical logs), removed the logical log causing the problem, and took a
> level0
> backup.
> OK informix running again.
> Then raid-intervention replace D500, got 3 read errors but raid was rebuild
> to
> optimal state.
> During 2 days no error-messages.
> Yesterday evening D200 (with the original read-error) was also replaced and
> raid was rebuild to optimal state.
>
> But backup afterwards generates error-messages again.
>
> To exclude the OS and HW, we will try a dd-command to read the entire slice
>
> If the problem is informix-related, are there any suggestions, other than
> unloading and re-initialising the database ?
>
> Thx for any reactions,
>
> Jacques Lapeire
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g