chunk offline
Posted in 2009
A chunk holding part of an index went offline (marked PD) during disk work; the DBA used onspaces -s to bring it back online and the data looked queryable. Art Kagel explained onspaces -s itself doesn't corrupt anything, but IDS performs no consistency check when re-enabling an unmirrored chunk, so any DML done while the chunk was down leaves missing/orphaned keys — hence run oncheck -cI (capital I) or rebuild the index before releasing the system. That's what happened: oncheck -cI later reported "Bad key information in TBLspace description" and rebuilding the index fixed it. Advice: try onspaces -s only if the disk is sound, otherwise replace the chunk and restore/roll forward logs; also replace the drive, since SCSI errors suggest failing media.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Storage & Space Management, Server Administration, Logging & Checkpoints, Clustering, Grid & MACH11
Hi All,
We had a situation were a non-critical chunk went offline.
This was a cluster setup and during some testing/disk alignment by
Unix, this chunk went offline. It was marked PD.
This chunk contains only a portion of one Index.( Checked before using
oncheck –pe ).
Used onspaces –s to bring the chunk online. This operation was
successful & checkpoint was done.
Using oncheck –pe now able to see the index fragment in this chunk.
Question : Is it possible in any way that data in chunk will corrupt
due to this operation onspaces –s? ( because now able to see the index
fragment using oncheck –pe , and able to query table). Unfortunately
couldn’t complete the oncheck –ci because of time constraints.
What are the consequence of doing onspaces –s ?
Informix version : 9.21 ( aware that this version is no more supported
by IBM)
OS version : 11.0
Thanks & regards,
Jose.
The onspaces -s won't corrupt anything, however, if the underlying table was
updated/inserted/deleted to/from while the chunk was offline the keys in
that index fragment may be missing keys or have orphaned keys belonging to
deleted or modified rows. You should run the oncheck -cI (capital 'I' to
check key values and structure not lower 'i'which only checks index
structure) as soon as possible or just rebuild the index.
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
See you at the 2010 IIUG Informix Conference
April 25-28, 2010
Overland Park (Kansas City), KS
www.iiug.org/conf
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
On Wed, Dec 16, 2009 at 1:07 AM, Jo <josephska@gmail.com> wrote:
> Hi All,
>
> We had a situation were a non-critical chunk went offline.
> This was a cluster setup and during some testing/disk alignment by
> Unix, this chunk went offline. It was marked PD.
> This chunk contains only a portion of one Index.( Checked before using
> oncheck –pe ).
> Used onspaces –s to bring the chunk online. This operation was
> successful & checkpoint was done.
> Using oncheck –pe now able to see the index fragment in this chunk.
>
> Question : Is it possible in any way that data in chunk will corrupt
> due to this operation onspaces –s? ( because now able to see the index
> fragment using oncheck –pe , and able to query table). Unfortunately
> couldn’t complete the oncheck –ci because of time constraints.
>
> What are the consequence of doing onspaces –s ?
>
> Informix version : 9.21 ( aware that this version is no more supported
> by IBM)
> OS version : 11.0
>
> Thanks & regards,
> Jose.
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
Thanks for the reply.
Continuing the issue:
Later after some basic check like query table, oncheck -cc, -cr & UAT
by application team the system was released. (As told earlier
unfortunately couldn’t complete the oncheck –cI because of time
constraints.)
After some time few SCSI errors were reported on UNIX side related to
this disk. And Informix message log also reported some errors.
Application side transactions were failing. From Informix message log,
it was pointing to insert on one table, (that had index portion on the
chunk that went down & later online).
Now checked oncheck – cI & it gave error “ERROR: Bad key
information in TBLspace description.”
So ultimately had to do rebuild the index. And this solved the
issues.
Question: What would be the best approach for this kind of recovery?
Doing onspaces –s as initial step for recovery, the correct step?
As I understand using onspaces command to recover a down chunk should
be tried first before any other recovery methods are attempted.
Thank you.
Regards,
Jose
It depends on the situation. If you know that there is no damage to the
disk/partition on which the chunk resides, then trying an onspace -s to set
the chunk live again is a good first try. The fallback if this is not
working out is to replace the chunk and restore from an archive rolling
forward the logical logs.
I am concerned, however, because you say that the system was reporting SCSI
errors. This can mean the the chunk is losing integrity. If the drive is
older, it may be that the drive media are beginning to fail and the drive
has run out of replacement sectors and is now returning reconstructed data
using the CRC codes. If this is what's going on, and even if there are
still some replacement sectors left, it may be time to replace the drive
before the data is no longer recoverable.
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
See you at the 2010 IIUG Informix Conference
April 25-28, 2010
Overland Park (Kansas City), KS
www.iiug.org/conf
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
On Wed, Dec 16, 2009 at 9:47 AM, Jo <josephska@gmail.com> wrote:
> Thanks for the reply.
>
> Continuing the issue:
> Later after some basic check like query table, oncheck -cc, -cr & UAT
> by application team the system was released. (As told earlier
> unfortunately couldn’t complete the oncheck –cI because of time
> constraints.)
>
> After some time few SCSI errors were reported on UNIX side related to
> this disk. And Informix message log also reported some errors.
> Application side transactions were failing. From Informix message log,
> it was pointing to insert on one table, (that had index portion on the
> chunk that went down & later online).
>
> Now checked oncheck – cI & it gave error “ERROR: Bad key
> information in TBLspace description.”
>
> So ultimately had to do rebuild the index. And this solved the
> issues.
>
> Question: What would be the best approach for this kind of recovery?
> Doing onspaces –s as initial step for recovery, the correct step?
>
> As I understand using onspaces command to recover a down chunk should
> be tried first before any other recovery methods are attempted.
>
> Thank you.
>
> Regards,
> Jose
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
Hi Art,
Thanks for the reply.
Suppose underlying table was updated/inserted/deleted to/from while
the chunk was offline, then after bringing the chunk online (using
onspaces –s) shouldn’t Informix report some sort of error? (suppose if
this chunk was not consistent with rest of the chunks)
Didn’t find any error message either in the prompt or Informix logs.
Thanks & regards,
Jose.
Forgot to add: SCSI errors are not seen any more since few days. -Jose.
The errors you would see are exactly what you did see, invalid key in the
index(es) located on that chunk. IDS does not make any consistency check
when it marks a chunk back online after it has been marked down. That is
why on earlier releases you could not mark a down chunk back online unless
the chunk had a mirror chunk that was not offline so that the chunk could be
refreshed from the mirror when it was enabled.
Art
Art S. Kagel
Oninit (www.oninit.com)
IIUG Board of Directors (art@iiug.org)
See you at the 2010 IIUG Informix Conference
April 25-28, 2010
Overland Park (Kansas City), KS
www.iiug.org/conf
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Oninit, the IIUG, nor any other organization
with which I am associated either explicitly or implicitly. Neither do
those opinions reflect those of other individuals affiliated with any entity
with which I am affiliated nor those of the entities themselves.
On Wed, Dec 16, 2009 at 11:19 AM, Jo <josephska@gmail.com> wrote:
> Hi Art,
>
> Thanks for the reply.
> Suppose underlying table was updated/inserted/deleted to/from while
> the chunk was offline, then after bringing the chunk online (using
> onspaces –s) shouldn’t Informix report some sort of error? (suppose if
> this chunk was not consistent with rest of the chunks)
>
> Didn’t find any error message either in the prompt or Informix logs.
>
>
> Thanks & regards,
> Jose.
>
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
Doesn't mean anything. I'd still worry and want that drive replaced as soon as possible. Art Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) See you at the 2010 IIUG Informix Conference April 25-28, 2010 Overland Park (Kansas City), KS www.iiug.org/conf Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, Dec 16, 2009 at 11:34 AM, Jo <josephska@gmail.com> wrote: > Forgot to add: SCSI errors are not seen any more since few days. > > -Jose. > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list >
Hello Jose,
the fact that you could bring online an unmirrored chunk is a bug.
and yes you will have corruption.
i would probably gofor a restore... hmmmm first on a test box though.
(
dono if you have the logs from the point where it went down... you
could do a warm restore of
only the affected dbspace....)
you also could try and repair... i guess i would drop and recreate the
index.
hmm i would drop the index, drop the dbspace containing the bad chunk
(unload all first)
and then recreate.
Superboer.
On 16 dec, 07:07, Jo <joseph...@gmail.com> wrote:
> Hi All,
>
> We had a situation were a non-critical chunk went offline.
> This was a cluster setup and during some testing/disk alignment by
> Unix, this chunk went offline. It was marked PD.
> This chunk contains only a portion of one Index.( Checked before using
> oncheck –pe ).
> Used onspaces –s to bring the chunk online. This operation was
> successful & checkpoint was done.
> Using oncheck –pe now able to see the index fragment in this chunk.
>
> Question : Is it possible in any way that data in chunk will corrupt
> due to this operation onspaces –s? ( because now able to see the index
> fragment using oncheck –pe , and able to query table). Unfortunately
> couldn’t complete the oncheck –ci because of time constraints.
>
> What are the consequence of doing onspaces –s ?
>
> Informix version : 9.21 ( aware that this version is no more supported
> by IBM)
> OS version : 11.0
>
> Thanks & regards,
> Jose.
On 16 Dec, 16:19, Jo <joseph...@gmail.com> wrote:
> Hi Art,
>
> Thanks for the reply.
> Suppose underlying table was updated/inserted/deleted to/from while
> the chunk was offline, then after bringing the chunk online (using
> onspaces –s) shouldn’t Informix report some sort of error? (suppose if
> this chunk was not consistent with rest of the chunks)
>
> Didn’t find any error message either in the prompt or Informix logs.
>
> Thanks & regards,
> Jose.
Informix will not report an error unless some of the pages it has to
read to bring the chunk online are corrupt.
It does not need to read all data/index pages to bring the chunk
online,probably only reserved pages,chunk free list and maybe some
partition pages.
It is your responsibility to consistency check all pages in the chunk
after you bring it online and before allowing user access to the
chunk.