Informix 11.10FC2W5 hangs & crashes
Posted in 2008
A site running IDS 11.10FC2W5 on a Sun T5120 reported repeated hangs on blocked checkpoints plus assertion failures/crashes, starting after a workload increase; IBM support had them patch to W5 and enable btree scanner ALICE mode. Suggestions from the list: check onstat -R and onstat -u for threads in critical sections, watch for huge/fragmented indexes causing long checkpoints, consider DataBlade issues, and one site said W3 cured their hangs. Mark Jamison explained that a single BTSCANNER thread using leaf scanning can fall behind and hold resources needed for checkpoints, so more scanner threads plus ALICE should address the hangs — but not the crashes, which deserve a separate PMR. No confirmed fix is recorded; the poster said IBM was preparing a fix to try.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Installation, Setup & Upgrades, Storage & Space Management, Server Administration, Logging & Checkpoints
We have about 6 or 7 MAJOR failures of Infmormix 11.10FC2W5 in the past
few days. I have a pmr going and a tech support person working on it for
several days now.
I am hoping someone else might have seen this problem and fixed it them
selves. I am all Googled out.
We have had Assertions and HANGS on Blocked Check Points. Sometimes we
can stay up for only 3 hours other times we stay up for a whole day.
This morning alone we Hang at 6am and then again at 9. Now we are quite
till 17:30. We have restricted usage to Medical Emergencies Only to
prevent any unnecessary usage. We move some stuff over to M/S Sql
(something some people around here looked for excuses to do).
We have been running 11.10FC2 since Aug 16th no problems. Then all of a
sudden we have an assertion at 3 oclock last Thursday monring. Since
then we have had about 6/7 Informix Failures.
IBM Support had us upgrade to W5 and since we did that we have had 4
more problems. IBM has added and onmode -C alice 6 for some reason I do
not understand. We can not continue like this.
ALL of these problems start off with about an hour of the informix log
saying Txns blocked:1 and checkpoint durations greater than 0. We used
to have zero all the time for these. During the hour before lock up we
have slow periods of up to 5 minutes then all of a sudden things free up
and we go back to normal, on and off hangs. Then after about an hour of
this the system freezes on blocked check point.
Physical log is at 1% used and there are 7 logical logs and only one is
in use. The others are backed up.
I have tons of onstat -a's. Has anyone had some experience with blocked
check points that might be able to shed some light on this.
Why would informix go belly up all of a sudden like this?
Running on a Sun Sparc T5120 32gb memory, 64 virtual processors (cores
or what ever they call them). Ton's of disk space now. We upgraded Aug
16th and things have been great, then crash!
Thanks for any suggestions.
We may need to hire special help!
,,,Tim
Tim Ertl wrote:
> We have about 6 or 7 MAJOR failures of Infmormix 11.10FC2W5 in the past
> few days. I have a pmr going and a tech support person working on it for
> several days now.
>
> I am hoping someone else might have seen this problem and fixed it them
> selves. I am all Googled out.
>
> We have had Assertions and HANGS on Blocked Check Points. Sometimes we
> can stay up for only 3 hours other times we stay up for a whole day.
> This morning alone we Hang at 6am and then again at 9. Now we are quite
> till 17:30. We have restricted usage to Medical Emergencies Only to
> prevent any unnecessary usage. We move some stuff over to M/S Sql
> (something some people around here looked for excuses to do).
>
> We have been running 11.10FC2 since Aug 16th no problems. Then all of a
> sudden we have an assertion at 3 oclock last Thursday monring. Since
> then we have had about 6/7 Informix Failures.
> IBM Support had us upgrade to W5 and since we did that we have had 4
> more problems. IBM has added and onmode -C alice 6 for some reason I do
> not understand. We can not continue like this.
>
> ALL of these problems start off with about an hour of the informix log
> saying Txns blocked:1 and checkpoint durations greater than 0. We used
> to have zero all the time for these. During the hour before lock up we
> have slow periods of up to 5 minutes then all of a sudden things free up
> and we go back to normal, on and off hangs. Then after about an hour of
> this the system freezes on blocked check point.
>
> Physical log is at 1% used and there are 7 logical logs and only one is
> in use. The others are backed up.
>
> I have tons of onstat -a's. Has anyone had some experience with blocked
> check points that might be able to shed some light on this.
>
> Why would informix go belly up all of a sudden like this?
>
> Running on a Sun Sparc T5120 32gb memory, 64 virtual processors (cores
> or what ever they call them). Ton's of disk space now. We upgraded Aug
> 16th and things have been great, then crash!
>
> Thanks for any suggestions.
> We may need to hire special help!
>
What has changed?
--
Cheers,
Obnoxio the Clown
http://obotheclown.blogspot.com
Obnoxio the Clown,
Wow the best question in the world is the one nobody can answer. The
customer (me) always say, Nothing in the past 6 weeks? But even I have to
ask what changed. We have had an increase in database activity and one of
the tables had a new field added. Three programs had some business logic
changed but even after backing off all changes prior to the first Assertion
we still had 5 more problems. So it looks like the increase in loading I
guess.
I am told by IBM that the fix is to add a Btree scanner called alice. I read
a little about it and it looks more like a performance tweek. I can not
understand how that can crash twice and hang 4 times a whole server. All 6
times on a Blocked Checkpoint. I am one of the "admin free" kinda sites. We
set things up and they run for 8 years. Then I upset the cart by installing
11.10FC2 on a new Sun 6 weeks ago and now it goes bump, mostly after the
onbar backups have finished (3 times it did that).
If I get a full nights sleep tonight I might be willing to accept Alice but
I am still skeptical. I wish I ask more about the one line fix to the
onconfig file. I may have to go to another user group meeting to get that
answer.
Tim Ertl
413-442-9000 x6211
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Obnoxio The Clown
Sent: Monday, September 29, 2008 7:07 PM
To: ids@iiug.org
Subject: Re: Informix 11.10FC2W5 hangs & crashes [13521]
Tim Ertl wrote:
> We have about 6 or 7 MAJOR failures of Infmormix 11.10FC2W5 in the past
> few days. I have a pmr going and a tech support person working on it for
> several days now.
>
> I am hoping someone else might have seen this problem and fixed it them
> selves. I am all Googled out.
>
> We have had Assertions and HANGS on Blocked Check Points. Sometimes we
> can stay up for only 3 hours other times we stay up for a whole day.
> This morning alone we Hang at 6am and then again at 9. Now we are quite
> till 17:30. We have restricted usage to Medical Emergencies Only to
> prevent any unnecessary usage. We move some stuff over to M/S Sql
> (something some people around here looked for excuses to do).
>
> We have been running 11.10FC2 since Aug 16th no problems. Then all of a
> sudden we have an assertion at 3 oclock last Thursday monring. Since
> then we have had about 6/7 Informix Failures.
> IBM Support had us upgrade to W5 and since we did that we have had 4
> more problems. IBM has added and onmode -C alice 6 for some reason I do
> not understand. We can not continue like this.
>
> ALL of these problems start off with about an hour of the informix log
> saying Txns blocked:1 and checkpoint durations greater than 0. We used
> to have zero all the time for these. During the hour before lock up we
> have slow periods of up to 5 minutes then all of a sudden things free up
> and we go back to normal, on and off hangs. Then after about an hour of
> this the system freezes on blocked check point.
>
> Physical log is at 1% used and there are 7 logical logs and only one is
> in use. The others are backed up.
>
> I have tons of onstat -a's. Has anyone had some experience with blocked
> check points that might be able to shed some light on this.
>
> Why would informix go belly up all of a sudden like this?
>
> Running on a Sun Sparc T5120 32gb memory, 64 virtual processors (cores
> or what ever they call them). Ton's of disk space now. We upgraded Aug
> 16th and things have been great, then crash!
>
> Thanks for any suggestions.
> We may need to hire special help!
>
What has changed?
--
Cheers,
Obnoxio the Clown
http://obotheclown.blogspot.com
****************************************************************************
***
Forum Note: Use "Reply" to post a response in the discussion forum.
Could you tell me your PMR number?
Tim,
Are you using any DataBlades ? Eg: Spatial, rtree ... ??
You may need to locate these into your new $INFORMIXDIR/extend directory
as they may not be included in DBServer Distro.
Do you have an af files or message log errors to post ??
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Tim Ertl
Sent: Tuesday, 30 September 2008 10:47 AM
To: ids@iiug.org
Subject: RE: Informix 11.10FC2W5 hangs & crashes [13522]
Obnoxio the Clown,
Wow the best question in the world is the one nobody can answer. The
customer (me) always say, Nothing in the past 6 weeks? But even I have
to
ask what changed. We have had an increase in database activity and one
of
the tables had a new field added. Three programs had some business logic
changed but even after backing off all changes prior to the first
Assertion
we still had 5 more problems. So it looks like the increase in loading I
guess.
I am told by IBM that the fix is to add a Btree scanner called alice. I
read
a little about it and it looks more like a performance tweek. I can not
understand how that can crash twice and hang 4 times a whole server. All
6
times on a Blocked Checkpoint. I am one of the "admin free" kinda sites.
We
set things up and they run for 8 years. Then I upset the cart by
installing
11.10FC2 on a new Sun 6 weeks ago and now it goes bump, mostly after the
onbar backups have finished (3 times it did that).
If I get a full nights sleep tonight I might be willing to accept Alice
but
I am still skeptical. I wish I ask more about the one line fix to the
onconfig file. I may have to go to another user group meeting to get
that
answer.
Tim Ertl
413-442-9000 x6211
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Obnoxio The Clown
Sent: Monday, September 29, 2008 7:07 PM
To: ids@iiug.org
Subject: Re: Informix 11.10FC2W5 hangs & crashes [13521]
Tim Ertl wrote:
> We have about 6 or 7 MAJOR failures of Infmormix 11.10FC2W5 in the
past
> few days. I have a pmr going and a tech support person working on it
for
> several days now.
>
> I am hoping someone else might have seen this problem and fixed it
them
> selves. I am all Googled out.
>
> We have had Assertions and HANGS on Blocked Check Points. Sometimes we
> can stay up for only 3 hours other times we stay up for a whole day.
> This morning alone we Hang at 6am and then again at 9. Now we are
quite
> till 17:30. We have restricted usage to Medical Emergencies Only to
> prevent any unnecessary usage. We move some stuff over to M/S Sql
> (something some people around here looked for excuses to do).
>
> We have been running 11.10FC2 since Aug 16th no problems. Then all of
a
> sudden we have an assertion at 3 oclock last Thursday monring. Since
> then we have had about 6/7 Informix Failures.
> IBM Support had us upgrade to W5 and since we did that we have had 4
> more problems. IBM has added and onmode -C alice 6 for some reason I
do
> not understand. We can not continue like this.
>
> ALL of these problems start off with about an hour of the informix log
> saying Txns blocked:1 and checkpoint durations greater than 0. We used
> to have zero all the time for these. During the hour before lock up we
> have slow periods of up to 5 minutes then all of a sudden things free
up
> and we go back to normal, on and off hangs. Then after about an hour
of
> this the system freezes on blocked check point.
>
> Physical log is at 1% used and there are 7 logical logs and only one
is
> in use. The others are backed up.
>
> I have tons of onstat -a's. Has anyone had some experience with
blocked
> check points that might be able to shed some light on this.
>
> Why would informix go belly up all of a sudden like this?
>
> Running on a Sun Sparc T5120 32gb memory, 64 virtual processors (cores
> or what ever they call them). Ton's of disk space now. We upgraded Aug
> 16th and things have been great, then crash!
>
> Thanks for any suggestions.
> We may need to hire special help!
>
What has changed?
--
Cheers,
Obnoxio the Clown
http://obotheclown.blogspot.com
************************************************************************
****
***
Forum Note: Use "Reply" to post a response in the discussion forum.
************************************************************************
*******
Forum Note: Use "Reply" to post a response in the discussion forum.
***************************************************************
This message is intended for the addressee named and may contain confidential
information. If you are not the intended recipient, please delete it and
notify the sender. Views expressed in this message are those of the individual
sender, and are not necessarily the views of the Department of Lands. This
email message has been swept by MIMEsweeper for the presence of computer
viruses.
***************************************************************
Please consider the environment before printing this email.
Hi,
first I have to say that I don't know IDS11 (only up to IDS10) and checkpoint
algorithms and monitoring has changed between these versions.
But I would have a look at two things
onstat -R |tailThat will show if there is IO writing dirty pages to disk.
More interesting should be
onstat -u | grep XThat will show threads in critical sections - maybe you can find some that are
responsible for the block.
If it is not an internal thread (btree cleaner , backup , ...) you can
further look at the corresponding session (column 3 of onstat -u output).
With IDS10 we have some issues with long checkpoints due to INSERTs on tables
with many changes (insert/delete) and indexes of huge size. We have to rebuild
these indices quite often to avoid excessive checkpoints.
Regards,
Andreas Kutsche
>
-------------------------------------------
SPAR Österreichische Warenhandels-AG
Hauptzentrale
A - 5015 Salzburg, Europastrasse 3
FN 34170 a
Tel: +43 662 4470 24223
Mobile: +43 664 6259575
E-Mail: Andreas.KUTSCHE@spar.at
Internet: http://www.spar.at
Wichtiger Hinweis: Der Inhalt dieser E-Mail kann vertrauliche und rechtlich
geschützte Informationen, insbesondere Betriebs- oder Geschäftsgeheimnisse,
enthalten, zu deren Geheimhaltung der Empfänger verpflichtet ist. Die
Informationen in dieser E-Mail sind ausschließlich für den Adressaten
bestimmt. Sollten Sie die E-Mail irrtümlich erhalten haben so ersuchen wir
Sie, die Nachricht von Ihrem System zu löschen und sich mit uns in Verbindung
zu setzen.
Über das Internet versandte E-Mails können leicht manipuliert oder unter
fremdem Namen erstellt werden. Daher schließen wir die rechtliche
Verbindlichkeit der in dieser Nachricht enthaltenen Informationen aus. Der
Inhalt der E-Mail ist nur rechtsverbindlich, wenn er von uns schriftlich
bestätigt und gezeichnet wird.
Sollte trotz der von uns verwendeten Virus-Schutzprogramme durch die Zusendung
von E-Mails ein Virus in Ihre Systeme gelangen, haften wir nicht für evtl.
hieraus entstehende Schäden.
Wir danken für Ihr Verständnis.
Important notice: The contents of this e-mail may contain confidential and
legally protected information that is in particular related to operational and
trade secrets, which the recipient is obliged to treat as confidential. The
information in this e-mail is made available exclusively for use by the
addressee. In the event that the e-mail may have been sent to you in error, we
would ask you to kindly delete this communication from your system and to
contact us.
E-mails sent via the Internet can be easily manipulated or sent out under
someone else's name. We therefore do not accept legal liability for the
information contained in this communication. The contents of the e-mail are
only legally binding if they have been confirmed and signed by us in writing.
If, in spite of our using Antivirus protection software, a virus may have
penetrated your system through the sending of this e-mail, we do not accept
liability for any damage that may possibly arise as a result of this.
We trust that you appreciate our position.
-------------------------------------------
-----Ursprüngliche Nachricht-----
> Von: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] Im Auftrag von Tim
> Ertl
> Gesendet: Dienstag, 30. September 2008 02:47
> An: ids@iiug.org
> Betreff: RE: Informix 11.10FC2W5 hangs & crashes [13522]
>
> Obnoxio the Clown,
> Wow the best question in the world is the one nobody can answer. The
> customer (me) always say, Nothing in the past 6 weeks? But even I have to
> ask what changed. We have had an increase in database activity and one of
> the tables had a new field added. Three programs had some business logic
> changed but even after backing off all changes prior to the first
> Assertion
> we still had 5 more problems. So it looks like the increase in loading I
> guess.
>
> I am told by IBM that the fix is to add a Btree scanner called alice. I
> read
> a little about it and it looks more like a performance tweek. I can not
> understand how that can crash twice and hang 4 times a whole server. All 6
> times on a Blocked Checkpoint. I am one of the "admin free" kinda sites.
> We
> set things up and they run for 8 years. Then I upset the cart by
> installing
> 11.10FC2 on a new Sun 6 weeks ago and now it goes bump, mostly after the
> onbar backups have finished (3 times it did that).
>
> If I get a full nights sleep tonight I might be willing to accept Alice
> but
> I am still skeptical. I wish I ask more about the one line fix to the
> onconfig file. I may have to go to another user group meeting to get that
> answer.
>
> Tim Ertl
> 413-442-9000 x6211
>
> -----Original Message-----
> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
> Obnoxio The Clown
> Sent: Monday, September 29, 2008 7:07 PM
> To: ids@iiug.org
> Subject: Re: Informix 11.10FC2W5 hangs & crashes [13521]
>
> Tim Ertl wrote:
> > We have about 6 or 7 MAJOR failures of Infmormix 11.10FC2W5 in the past
> > few days. I have a pmr going and a tech support person working on it for
> > several days now.
> >
> > I am hoping someone else might have seen this problem and fixed it them
> > selves. I am all Googled out.
> >
> > We have had Assertions and HANGS on Blocked Check Points. Sometimes we
> > can stay up for only 3 hours other times we stay up for a whole day.
> > This morning alone we Hang at 6am and then again at 9. Now we are quite
> > till 17:30. We have restricted usage to Medical Emergencies Only to
> > prevent any unnecessary usage. We move some stuff over to M/S Sql
> > (something some people around here looked for excuses to do).
> >
> > We have been running 11.10FC2 since Aug 16th no problems. Then all of a
> > sudden we have an assertion at 3 oclock last Thursday monring. Since
> > then we have had about 6/7 Informix Failures.
> > IBM Support had us upgrade to W5 and since we did that we have had 4
> > more problems. IBM has added and onmode -C alice 6 for some reason I do
> > not understand. We can not continue like this.
> >
> > ALL of these problems start off with about an hour of the informix log
> > saying Txns blocked:1 and checkpoint durations greater than 0. We used
> > to have zero all the time for these. During the hour before lock up we
> > have slow periods of up to 5 minutes then all of a sudden things free up
> > and we go back to normal, on and off hangs. Then after about an hour of
> > this the system freezes on blocked check point.
> >
> > Physical log is at 1% used and there are 7 logical logs and only one is
> > in use. The others are backed up.
> >
> > I have tons of onstat -a's. Has anyone had some experience with blocked
> > check points that might be able to shed some light on this.
> >
> > Why would informix go belly up all of a sudden like this?
> >
> > Running on a Sun Sparc T5120 32gb memory, 64 virtual processors (cores
> > or what ever they call them). Ton's of disk space now. We upgraded Aug
> > 16th and things have been great, then crash!
> >
> > Thanks for any suggestions.
> > We may need to hire special help!
> >
>
> What has cha
We had experienced hang issues when we upgraded our databases to IDS=0D=0A1=
1=2E10FC2, we had to apply the W3 patch (ids 11=2E10FC2W3)to resolve the=0D=
=0Aissue=2E No hangs after W3 patch was installed, also we have noticed=0D=
=0Abetter performance with 11=2E10fc2w3 as compared to 11=2E10fc2=2E=0D=0A =
=0D=0A=0D=0A-----Original Message-----=0D=0AFrom: ids-bounces@iiug=2Eorg [m=
ailto:ids-bounces@iiug=2Eorg] On Behalf Of=0D=0ATim Ertl=0D=0ASent: Monday,=
September 29, 2008 5:01 PM=0D=0ATo: ids@iiug=2Eorg=0D=0ASubject: Informix =
11=2E10FC2W5 hangs & crashes [13520]=0D=0A=0D=0AWe have about 6 or 7 MAJOR =
failures of Infmormix 11=2E10FC2W5 in the past =0D=0Afew days=2E I have a p=
mr going and a tech support person working on it for=0D=0A=0D=0Aseveral day=
s now=2E =0D=0A=0D=0AI am hoping someone else might have seen this problem =
and fixed it them =0D=0Aselves=2E I am all Googled out=2E =0D=0A=0D=0AWe ha=
ve had Assertions and HANGS on Blocked Check Points=2E Sometimes we =0D=0Ac=
an stay up for only 3 hours other times we stay up for a whole day=2E =0D=
=0AThis morning alone we Hang at 6am and then again at 9=2E Now we are quit=
e =0D=0Atill 17:30=2E We have restricted usage to Medical Emergencies Only =
to =0D=0Aprevent any unnecessary usage=2E We move some stuff over to M/S Sq=
l =0D=0A(something some people around here looked for excuses to do)=2E =0D=
=0A=0D=0AWe have been running 11=2E10FC2 since Aug 16th no problems=2E Then=
all of a =0D=0Asudden we have an assertion at 3 oclock last Thursday monri=
ng=2E Since =0D=0Athen we have had about 6/7 Informix Failures=2E =0D=0AIBM=
Support had us upgrade to W5 and since we did that we have had 4 =0D=0Amor=
e problems=2E IBM has added and onmode -C alice 6 for some reason I do =0D=
=0Anot understand=2E We can not continue like this=2E =0D=0A=0D=0AALL of th=
ese problems start off with about an hour of the informix log =0D=0Asaying =
Txns blocked:1 and checkpoint durations greater than 0=2E We used =0D=0Ato =
have zero all the time for these=2E During the hour before lock up we =0D=
=0Ahave slow periods of up to 5 minutes then all of a sudden things free up=
=0D=0A=0D=0Aand we go back to normal, on and off hangs=2E Then after about =
an hour of =0D=0Athis the system freezes on blocked check point=2E =0D=0A=
=0D=0APhysical log is at 1% used and there are 7 logical logs and only one =
is =0D=0Ain use=2E The others are backed up=2E =0D=0A=0D=0AI have tons of o=
nstat -a's=2E Has anyone had some experience with blocked =0D=0Acheck point=
s that might be able to shed some light on this=2E =0D=0A=0D=0AWhy would in=
formix go belly up all of a sudden like this? =0D=0A=0D=0ARunning on a Sun =
Sparc T5120 32gb memory, 64 virtual processors (cores =0D=0Aor what ever th=
ey call them)=2E Ton's of disk space now=2E We upgraded Aug =0D=0A16th and =
things have been great, then crash! =0D=0A=0D=0AThanks for any suggestions=
=2E =0D=0AWe may need to hire special help! =0D=0A=0D=0A,,,Tim =0D=0A=0D=0A=
=0D=0A*********************************************************************=
***=0D=0A******* =0D=0A Forum Note: Use "Reply" to post a response in the =
discussion forum=2E =0D=0A=0D=0AConsider our environment; please print this=
e-mail only if truly=0Anecessary=2E Thank you!
We needed a patch from W4 for onbar backup memory corruption and W5 was
available so we used W5.
I did get some wonderful help today and I think we have an IBM fix in the
works. As soon as we get it we will try it!
I want to thank everyone who offered their thoughts. I was hoping someone
had the same problem but it looks like not to many did, it looks like it is
a very current problem.
Tim Ertl
413-442-9000 x6211
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Savio
Pinto (s)
Sent: Tuesday, September 30, 2008 3:37 PM
To: ids@iiug.org
Subject: RE: Informix 11.10FC2W5 hangs & crashes [13536]
We had experienced hang issues when we upgraded our databases to IDS=0D=0A1=
1=2E10FC2, we had to apply the W3 patch (ids 11=2E10FC2W3)to resolve the=0D=
=0Aissue=2E No hangs after W3 patch was installed, also we have noticed=0D=
=0Abetter performance with 11=2E10fc2w3 as compared to 11=2E10fc2=2E=0D=0A =
=0D=0A=0D=0A-----Original Message-----=0D=0AFrom: ids-bounces@iiug=2Eorg [m=
ailto:ids-bounces@iiug=2Eorg] On Behalf Of=0D=0ATim Ertl=0D=0ASent: Monday,=
September 29, 2008 5:01 PM=0D=0ATo: ids@iiug=2Eorg=0D=0ASubject: Informix =
11=2E10FC2W5 hangs & crashes [13520]=0D=0A=0D=0AWe have about 6 or 7 MAJOR =
failures of Infmormix 11=2E10FC2W5 in the past =0D=0Afew days=2E I have a p=
mr going and a tech support person working on it for=0D=0A=0D=0Aseveral day=
s now=2E =0D=0A=0D=0AI am hoping someone else might have seen this problem =
and fixed it them =0D=0Aselves=2E I am all Googled out=2E =0D=0A=0D=0AWe ha=
ve had Assertions and HANGS on Blocked Check Points=2E Sometimes we =0D=0Ac=
an stay up for only 3 hours other times we stay up for a whole day=2E =0D=
=0AThis morning alone we Hang at 6am and then again at 9=2E Now we are quit=
e =0D=0Atill 17:30=2E We have restricted usage to Medical Emergencies Only =
to =0D=0Aprevent any unnecessary usage=2E We move some stuff over to M/S Sq=
l =0D=0A(something some people around here looked for excuses to do)=2E =0D=
=0A=0D=0AWe have been running 11=2E10FC2 since Aug 16th no problems=2E Then=
all of a =0D=0Asudden we have an assertion at 3 oclock last Thursday monri=
ng=2E Since =0D=0Athen we have had about 6/7 Informix Failures=2E =0D=0AIBM=
Support had us upgrade to W5 and since we did that we have had 4 =0D=0Amor=
e problems=2E IBM has added and onmode -C alice 6 for some reason I do =0D=
=0Anot understand=2E We can not continue like this=2E =0D=0A=0D=0AALL of th=
ese problems start off with about an hour of the informix log =0D=0Asaying =
Txns blocked:1 and checkpoint durations greater than 0=2E We used =0D=0Ato =
have zero all the time for these=2E During the hour before lock up we =0D=
=0Ahave slow periods of up to 5 minutes then all of a sudden things free up=
=0D=0A=0D=0Aand we go back to normal, on and off hangs=2E Then after about =
an hour of =0D=0Athis the system freezes on blocked check point=2E =0D=0A=
=0D=0APhysical log is at 1% used and there are 7 logical logs and only one =
is =0D=0Ain use=2E The others are backed up=2E =0D=0A=0D=0AI have tons of o=
nstat -a's=2E Has anyone had some experience with blocked =0D=0Acheck point=
s that might be able to shed some light on this=2E =0D=0A=0D=0AWhy would in=
formix go belly up all of a sudden like this? =0D=0A=0D=0ARunning on a Sun =
Sparc T5120 32gb memory, 64 virtual processors (cores =0D=0Aor what ever th=
ey call them)=2E Ton's of disk space now=2E We upgraded Aug =0D=0A16th and =
things have been great, then crash! =0D=0A=0D=0AThanks for any suggestions=
=2E =0D=0AWe may need to hire special help! =0D=0A=0D=0A,,,Tim =0D=0A=0D=0A=
=0D=0A*********************************************************************=
***=0D=0A******* =0D=0A Forum Note: Use "Reply" to post a response in the =
discussion forum=2E =0D=0A=0D=0AConsider our environment; please print this=
e-mail only if truly=0Anecessary=2E Thank you!
****************************************************************************
***
Forum Note: Use "Reply" to post a response in the discussion forum.
Whoops,
that should have said
"Sounds like you certainly should have 2 PMR's open, as the hangs
mostly likely don't cause the crashes, and the crashes don't cause the hangs.
Alice mode is certainly a performance tuning option, but performance tuning
can quite often resolve certain things like checkpoint hangs.
----- Original Message ----
From: Mark Jamison <majp51@yahoo.com>
To: ids@iiug.org
Sent: Tuesday, September 30, 2008 8:27:27 AM
Subject: Re: Informix 11.10FC2W5 hangs & crashes [13520]
Hi Tim,
The reason alice, and/or numbr of BTREE scanners have been added would only be
in response to a hang, not the crashes.
Sounds like you certainly should have 2 PMR's open, as the hangs mostly likely
cause the crashes, and the crashes don't cause the hangs.
Getting back to the hangs though. Are they continuing to occur after the
changes to BTSCANNER?
The reason I ask, is that until 11.50 the default for BTSCANNER is 1 thread
and using Leaf scanning. On a busy system, this combination could, although by
no means is guaranteed, to cause system slowness and blocked checkpoints. With
leaf scanning, it is very easy for a single BTSCANNER thread to get behind in
cleaning and never be able to catch up, the worst case scenario with a
BTSCANNER thread in that condition with leaf scanning doesn't allow one index
to finish cleaning during the time allotted for an active hot list, and
meanwhile the BTSCANNER thread is holding resources needed for a checkpoint to
complete. Increasing the number of BTSCANNER threads, would help mitigate this
problem, and enabling Alice mode would likely eliminate the hangs. The reason
why is that Alice is much faster than leaf scanning. By much faster, I mean
that on your average 1 MB Index, a leaf scan can do as many as 680 I/O's,
while Alice will do 4.
Don't suppose you have an onstat -C all from your last hang? For that matter,
do you happen to have an onstat -C all now?
Again BTSCANNER was only likely addressing the hangs.
----- Original Message ----
From: Tim Ertl <tim@lmrgroup.com>
To: ids@iiug.org
Sent: Monday, September 29, 2008 5:00:49 PM
Subject: Informix 11.10FC2W5 hangs & crashes [13520]
We have about 6 or 7 MAJOR failures of Infmormix 11.10FC2W5 in the past
few days. I have a pmr going and a tech support person working on it for
several days now.
I am hoping someone else might have seen this problem and fixed it them
selves. I am all Googled out.
We have had Assertions and HANGS on Blocked Check Points. Sometimes we
can stay up for only 3 hours other times we stay up for a whole day.
This morning alone we Hang at 6am and then again at 9. Now we are quite
till 17:30. We have restricted usage to Medical Emergencies Only to
prevent any unnecessary usage. We move some stuff over to M/S Sql
(something some people around here looked for excuses to do).
We have been running 11.10FC2 since Aug 16th no problems. Then all of a
sudden we have an assertion at 3 oclock last Thursday monring. Since
then we have had about 6/7 Informix Failures.
IBM Support had us upgrade to W5 and since we did that we have had 4
more problems. IBM has added and onmode -C alice 6 for some reason I do
not understand. We can not continue like this.
ALL of these problems start off with about an hour of the informix log
saying Txns blocked:1 and checkpoint durations greater than 0. We used
to have zero all the time for these. During the hour before lock up we
have slow periods of up to 5 minutes then all of a sudden things free up
and we go back to normal, on and off hangs. Then after about an hour of
this the system freezes on blocked check point.
Physical log is at 1% used and there are 7 logical logs and only one is
in use. The others are backed up.
I have tons of onstat -a's. Has anyone had some experience with blocked
check points that might be able to shed some light on this.
Why would informix go belly up all of a sudden like this?
Running on a Sun Sparc T5120 32gb memory, 64 virtual processors (cores
or what ever they call them). Ton's of disk space now. We upgraded Aug
16th and things have been great, then crash!
Thanks for any suggestions.
We may need to hire special help!
,,,Tim
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Hi Tim,
The reason alice, and/or numbr of BTREE scanners have been added would only be
in response to a hang, not the crashes.
Sounds like you certainly should have 2 PMR's open, as the hangs mostly likely
cause the crashes, and the crashes don't cause the hangs.
Getting back to the hangs though. Are they continuing to occur after the
changes to BTSCANNER?
The reason I ask, is that until 11.50 the default for BTSCANNER is 1 thread
and using Leaf scanning. On a busy system, this combination could, although by
no means is guaranteed, to cause system slowness and blocked checkpoints. With
leaf scanning, it is very easy for a single BTSCANNER thread to get behind in
cleaning and never be able to catch up, the worst case scenario with a
BTSCANNER thread in that condition with leaf scanning doesn't allow one index
to finish cleaning during the time allotted for an active hot list, and
meanwhile the BTSCANNER thread is holding resources needed for a checkpoint to
complete. Increasing the number of BTSCANNER threads, would help mitigate this
problem, and enabling Alice mode would likely eliminate the hangs. The reason
why is that Alice is much faster than leaf scanning. By much faster, I mean
that on your average 1 MB Index, a leaf scan can do as many as 680 I/O's,
while Alice will do 4.
Don't suppose you have an onstat -C all from your last hang? For that matter,
do you happen to have an onstat -C all now?
Again BTSCANNER was only likely addressing the hangs.
----- Original Message ----
From: Tim Ertl <tim@lmrgroup.com>
To: ids@iiug.org
Sent: Monday, September 29, 2008 5:00:49 PM
Subject: Informix 11.10FC2W5 hangs & crashes [13520]
We have about 6 or 7 MAJOR failures of Infmormix 11.10FC2W5 in the past
few days. I have a pmr going and a tech support person working on it for
several days now.
I am hoping someone else might have seen this problem and fixed it them
selves. I am all Googled out.
We have had Assertions and HANGS on Blocked Check Points. Sometimes we
can stay up for only 3 hours other times we stay up for a whole day.
This morning alone we Hang at 6am and then again at 9. Now we are quite
till 17:30. We have restricted usage to Medical Emergencies Only to
prevent any unnecessary usage. We move some stuff over to M/S Sql
(something some people around here looked for excuses to do).
We have been running 11.10FC2 since Aug 16th no problems. Then all of a
sudden we have an assertion at 3 oclock last Thursday monring. Since
then we have had about 6/7 Informix Failures.
IBM Support had us upgrade to W5 and since we did that we have had 4
more problems. IBM has added and onmode -C alice 6 for some reason I do
not understand. We can not continue like this.
ALL of these problems start off with about an hour of the informix log
saying Txns blocked:1 and checkpoint durations greater than 0. We used
to have zero all the time for these. During the hour before lock up we
have slow periods of up to 5 minutes then all of a sudden things free up
and we go back to normal, on and off hangs. Then after about an hour of
this the system freezes on blocked check point.
Physical log is at 1% used and there are 7 logical logs and only one is
in use. The others are backed up.
I have tons of onstat -a's. Has anyone had some experience with blocked
check points that might be able to shed some light on this.
Why would informix go belly up all of a sudden like this?
Running on a Sun Sparc T5120 32gb memory, 64 virtual processors (cores
or what ever they call them). Ton's of disk space now. We upgraded Aug
16th and things have been great, then crash!
Thanks for any suggestions.
We may need to hire special help!
,,,Tim
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Hi Mark, I did not know these problems were not related. We ran on the old
server (9.3) for 8 years+/- then we upgrade to 11.10 and we run fine for 6
weeks. Then we had a crash that was determined to be a memory corruption
(anything can happen problem) followed by a HANG a little later. My best
guess was that both were one problem. 20x20 hind sight we had two problems
at almost the same time. What are the odds in that? Especially with
Informix. We applied W5 and things went bizzerk! 1 more crash and 4 more
hangs in two days.
We are currently stable for the past 27 hours (+/-) but everything is
running 5 to 10 times slower. That is understandable for the change IBM
made. I think it must have been a work around rather than a fix. I dunno and
was not able to find out. IBM added a BTSCANNER num=2, threshold=200,
rangesize=500, alice=6. threshold 5,000 is the default and increasing 10x or
100x the manual says should improve performance. So when IBM made it 200
(yes 200) we felt the impact with a tremendous slow down.
Apparently W5 created a new problem when it fixed an old one. IBM has a fix
coming. They have identified my problem with some magic they have and a
patch will be out soon (I HOPE, our performance is bad right now).
I was looking for others with the same problem and I found one other person
but they had not resolved yet either.
Now I think we have a fix coming so I am happy.
Tim Ertl
413-442-9000 x6211
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Mark
Jamison
Sent: Tuesday, September 30, 2008 4:00 PM
To: ids@iiug.org
Subject: Re: Informix 11.10FC2W5 hangs & crashes [13539]
Hi Tim,
The reason alice, and/or numbr of BTREE scanners have been added would only
be
in response to a hang, not the crashes.
Sounds like you certainly should have 2 PMR's open, as the hangs mostly
likely
cause the crashes, and the crashes don't cause the hangs.
Getting back to the hangs though. Are they continuing to occur after the
changes to BTSCANNER?
The reason I ask, is that until 11.50 the default for BTSCANNER is 1 thread
and using Leaf scanning. On a busy system, this combination could, although
by
no means is guaranteed, to cause system slowness and blocked checkpoints.
With
leaf scanning, it is very easy for a single BTSCANNER thread to get behind
in
cleaning and never be able to catch up, the worst case scenario with a
BTSCANNER thread in that condition with leaf scanning doesn't allow one
index
to finish cleaning during the time allotted for an active hot list, and
meanwhile the BTSCANNER thread is holding resources needed for a checkpoint
to
complete. Increasing the number of BTSCANNER threads, would help mitigate
this
problem, and enabling Alice mode would likely eliminate the hangs. The
reason
why is that Alice is much faster than leaf scanning. By much faster, I mean
that on your average 1 MB Index, a leaf scan can do as many as 680 I/O's,
while Alice will do 4.
Don't suppose you have an onstat -C all from your last hang? For that
matter,
do you happen to have an onstat -C all now?
Again BTSCANNER was only likely addressing the hangs.
----- Original Message ----
From: Tim Ertl <tim@lmrgroup.com>
To: ids@iiug.org
Sent: Monday, September 29, 2008 5:00:49 PM
Subject: Informix 11.10FC2W5 hangs & crashes [13520]
We have about 6 or 7 MAJOR failures of Infmormix 11.10FC2W5 in the past
few days. I have a pmr going and a tech support person working on it for
several days now.
I am hoping someone else might have seen this problem and fixed it them
selves. I am all Googled out.
We have had Assertions and HANGS on Blocked Check Points. Sometimes we
can stay up for only 3 hours other times we stay up for a whole day.
This morning alone we Hang at 6am and then again at 9. Now we are quite
till 17:30. We have restricted usage to Medical Emergencies Only to
prevent any unnecessary usage. We move some stuff over to M/S Sql
(something some people around here looked for excuses to do).
We have been running 11.10FC2 since Aug 16th no problems. Then all of a
sudden we have an assertion at 3 oclock last Thursday monring. Since
then we have had about 6/7 Informix Failures.
IBM Support had us upgrade to W5 and since we did that we have had 4
more problems. IBM has added and onmode -C alice 6 for some reason I do
not understand. We can not continue like this.
ALL of these problems start off with about an hour of the informix log
saying Txns blocked:1 and checkpoint durations greater than 0. We used
to have zero all the time for these. During the hour before lock up we
have slow periods of up to 5 minutes then all of a sudden things free up
and we go back to normal, on and off hangs. Then after about an hour of
this the system freezes on blocked check point.
Physical log is at 1% used and there are 7 logical logs and only one is
in use. The others are backed up.
I have tons of onstat -a's. Has anyone had some experience with blocked
check points that might be able to shed some light on this.
Why would informix go belly up all of a sudden like this?
Running on a Sun Sparc T5120 32gb memory, 64 virtual processors (cores
or what ever they call them). Ton's of disk space now. We upgraded Aug
16th and things have been great, then crash!
Thanks for any suggestions.
We may need to hire special help!
,,,Tim
****************************************************************************
***
Forum Note: Use "Reply" to post a response in the discussion forum.
****************************************************************************
***
Forum Note: Use "Reply" to post a response in the discussion forum.
I have a question on dbexport. I accidentally deleted the .sql script that
goes with the dbexport. Is it possible to recreate it with dbschema or some
other tool or am I just SOL? I know that the dbschema utility will create a
script to recreate the database structure, but does the dbexport put extra
data specific information in the script for import?
Larry
Larry,
I am way behind on my email, so you may have already gotten a response on
this. I know that years ago, when we were moving data from version 5.x to 7.x
via dbexport/dbimport we tried to perform some edits in the .sql file. We
found that the .sql file created by dbexport contains information necessary
for the dbimport. dbimport is sensitive to the row count and the unload file
name that is actually contained as a comment in the .sql file. I assume that
is still the case. The file name would at least be a requirement, so dbschema
output will not do you any good.
However, if you can capture the schema with dbschema and you still have the
unload files, you could conceivably create the destination database and
tables, load the tables manually, then create the indexes, etc. You may have
to do some work to map the file names to the tables. I believe the file names
are the first few characters of the table name followed by the tabid of the
table. If a lot of your table names begin with the same characters, you'll
have to sort out which file maps to each table.
Our company has not allowed us to upgrade in quite some time, so my
information may be a bit stale. Hopefully this helps.
Rob Schmitz
Embarq Data Management
rob.b.schmitz@embarq.com
www.embarq.com
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of LARRY
SORENSEN
Sent: Thursday, October 02, 2008 11:35 AM
To: ids@iiug.org
Subject: RE: Informix 11.10FC2W5 hangs & crashes [13555]
I have a question on dbexport. I accidentally deleted the .sql script that
goes with the dbexport. Is it possible to recreate it with dbschema or some
other tool or am I just SOL? I know that the dbschema utility will create a
script to recreate the database structure, but does the dbexport put extra
data specific information in the script for import?
Larry
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.