ontape or IDS problem
Posted in 2007
Topics: Backup & Restore, Installation, Setup & Upgrades, Storage & Space Management, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
Dear all,
I need you opinion and help on a problem that nobody (even IBM support) can
identify it source yet. The problem is with IDS 9.40.FC8 (on HP-UX 11) and
ontape. The ontape cause the engine to hang/freeze at about over 90%
complete. The instance is consisting of 2 databases and they are 90%
blobdbs. The level-0 archive starts well but when it reaches over 90%
complete (could be anywhere between the 90 and 100 % mark) after about 12-13
hours it cause the whole Instance to hang for no apparent reasons. The
logical logs are not full; we rolled out any corruption as per IBM support.
This problem is going for 5 weeks now, and no valid backup since it last
run.
The change since the problem start happening:
- Upgrade from FC4 to FC8 and that to fix a bug in FC4 when dropping a chunk
that is not the last or at the end of the dbs.
- Adding over 3 TB of chunks to the blobdbs.
- Significant growth of data which is about 20GB/week (it can be more if it
is not for this problem).
When the problem start happening it used to happen only if archive-0 starts
on a weekday, So thinking maybe because of user load and the insertion job
that is keep adding new blobs (Sunday backup complete successfully). For the
last 3 weeks all backup are failing. We even tried level-1 or 2. It does the
same it run for 4-5 hours and when its 95-99% full the whole instance hang.
So far the PMR is escalated to engineering and the next step it trying onbar
and set IDS to debug mode to trace it while it is doing the ontape. IBM
support is now suspecting that this can be a new bug.
Have anybody seen a problem like this before and can provide and comments or
answers?
Thanks
Basim
Chafik, Basim wrote:
> Dear all,
>
> I need you opinion and help on a problem that nobody (even IBM support) can
> identify it source yet. The problem is with IDS 9.40.FC8 (on HP-UX 11) and
> ontape. The ontape cause the engine to hang/freeze at about over 90%
> complete. The instance is consisting of 2 databases and they are 90%
> blobdbs. The level-0 archive starts well but when it reaches over 90%
> complete (could be anywhere between the 90 and 100 % mark) after about 12-13
> hours it cause the whole Instance to hang for no apparent reasons. The
> logical logs are not full; we rolled out any corruption as per IBM support.
> This problem is going for 5 weeks now, and no valid backup since it last
> run.
>
> The change since the problem start happening:
> - Upgrade from FC4 to FC8 and that to fix a bug in FC4 when dropping a chunk
> that is not the last or at the end of the dbs.
> - Adding over 3 TB of chunks to the blobdbs.
> - Significant growth of data which is about 20GB/week (it can be more if it
> is not for this problem).
>
> When the problem start happening it used to happen only if archive-0 starts
> on a weekday, So thinking maybe because of user load and the insertion job
> that is keep adding new blobs (Sunday backup complete successfully). For the
> last 3 weeks all backup are failing. We even tried level-1 or 2. It does the
> same it run for 4-5 hours and when its 95-99% full the whole instance hang.
>
> So far the PMR is escalated to engineering and the next step it trying onbar
> and set IDS to debug mode to trace it while it is doing the ontape. IBM
> support is now suspecting that this can be a new bug.
>
> Have anybody seen a problem like this before and can provide and comments or
> answers?
>
> Thanks
> Basim
Do you have enough temporary space to do the backup? How full are your
temporary dbspaces when it blocks?
Are the ontape sessions consuming memory?
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
> - Adding over 3 TB of chunks to the blobdbs.
potentially it does something really inefficient trying to locate
blobs in empty blobspace chunks.
how many of those 3TB chunks are empty??
Suppose you have 20 empty chunks, i would drop in that case 18 chunks
and look if your problem
is the same or less.
How big are they???
maybe an onstat -D ( hopefully nothing else runs then ) may at least
show how much i/o is done on what chunk.
If onstat does not work, have a look with sar or iostat.
Superboer.
On 11 Jul, 14:07, "Chafik, Basim" <Basim.Cha...@BancTec.ca> wrote:
> Dear all,
>
> I need you opinion and help on a problem that nobody (even IBM support) can
> identify it source yet. The problem is with IDS 9.40.FC8 (on HP-UX 11) and
> ontape. The ontape cause the engine to hang/freeze at about over 90%
> complete. The instance is consisting of 2 databases and they are 90%
> blobdbs. The level-0 archive starts well but when it reaches over 90%
> complete (could be anywhere between the 90 and 100 % mark) after about 12-13
> hours it cause the whole Instance to hang for no apparent reasons. The
> logical logs are not full; we rolled out any corruption as per IBM support.
> This problem is going for 5 weeks now, and no valid backup since it last
> run.
>
> The change since the problem start happening:
> - Upgrade from FC4 to FC8 and that to fix a bug in FC4 when dropping a chunk
> that is not the last or at the end of the dbs.
> - Adding over 3 TB of chunks to the blobdbs.
> - Significant growth of data which is about 20GB/week (it can be more if it
> is not for this problem).
>
> When the problem start happening it used to happen only if archive-0 starts
> on a weekday, So thinking maybe because of user load and the insertion job
> that is keep adding new blobs (Sunday backup complete successfully). For the
> last 3 weeks all backup are failing. We even tried level-1 or 2. It does the
> same it run for 4-5 hours and when its 95-99% full the whole instance hang.
>
> So far the PMR is escalated to engineering and the next step it trying onbar
> and set IDS to debug mode to trace it while it is doing the ontape. IBM
> support is now suspecting that this can be a new bug.
>
> Have anybody seen a problem like this before and can provide and comments or
> answers?
>
> Thanks
> Basim
Any errors in the online.log or OS error log?
What does onstat -g ath give when it hangs?
What about onstat -g stk <tid> for the ontape related threads when it
hangs?
David.
Related threads
- IDS 10 table-level restore
- Informix Development Webinar December 11, 2007
- ontape -p/r with changed ROOTPATH
- Migrate from HP PA-RISC to HP ITANIUM by ontape