Urgent help is needed...
Posted in 2007
Topics: Platform-Specific Issues
Hi Gurus,
We are having issues with the database (IDS: 7.31 UD8 on Solaris 8) where we
are seeing lots of waits on buffers "B" -- sometimes applications fail with
something like "cannot position within table" -- some processes that hold
exclusive locks do not release causing other ones to wait on and fail (time
out) -- I run oncheck -cc, this is what I'm getting
oncheck -cc customerValidating database customer
Validating systables for database customer
Unknown error message 0.
IBM says it is because something holds on to systables which I'm not convinced.
Can someone who has experience with this shreds some lights.
Thanks
Kern --
________________________________________________________________________________
____
Get easy, one-click access to your favorites.
Make Yahoo! your homepage.
http://www.yahoo.com/r/hs
Hi,
It is normally not a problem if you have bufwaits, but these should not
take long.
Is this a new problem, or has the database been stable in this setup for
a while ? Was any config param changed ?
Was the application changed ? If you have exclusive locks not being
released, you should check these first (onstat -k to find out the owner,
onstat -g sql to see what is the owner doing).Maybe you have a problem with wrong lock-levels ?
Pls send more details on your problem. Onstat -p output in the first
place, maybe you have remarkable output in online.log, what is your
current buffer size ?
A little more info would help ..
Marcus
-----Original Message-----
From: Kern Doe [mailto:kern_doe@yahoo.com]
Sent: Sunday, December 02, 2007 9:33 PM
To: ids@iiug.org
Subject: Urgent help is needed... [10569]
Hi Gurus,
We are having issues with the database (IDS: 7.31 UD8 on Solaris 8)
where we are seeing lots of waits on buffers "B" -- sometimes
applications fail with something like "cannot position within table" --
some processes that hold exclusive locks do not release causing other
ones to wait on and fail (time
out) -- I run oncheck -cc, this is what I'm getting
oncheck -cc customerValidating database customer
Validating systables for database customer Unknown error message 0.
IBM says it is because something holds on to systables which I'm not
convinced.
Can someone who has experience with this shreds some lights.
Thanks
Kern --
________________________________________________________________________
____________
Get easy, one-click access to your favorites.
Make Yahoo! your homepage.
http://www.yahoo.com/r/hs
************************************************************************
*******
Forum Note: Use "Reply" to post a response in the discussion forum.
Thanks for your repply Marcus.
Ok, I would think this is normal -- but I've seen output from oncheck -cc
beore, locks (or xlocks) will cause a different error (ISAM error: record is
locked.)
The db had been stable for a while, no change to config or params. The big
issues are xlocks, yes I do them -- they used to come and go, but now they stay
onstat -k | grep Xaea8504 0 5c676d08 b7adef0 HDR+IX 1e00050 0 0
b32e954 0 5c676d08 b572980 HDR+IX 1e00044 0 0
b33166c 0 5c676d08 b32e954 HDR+X 250000a 2500909 0
b3337f0 0 5c676d08 b33166c HDR+IX 400022 0 0
b56e094 0 5c676d08 b3337f0 HDR+X 800005 18b801 0
b572980 0 5c676d08 b7b33a4 HDR+X 700003 8fe90a 0
b7adef0 0 5c676d08 b56e094 HDR+IX 1a00001 0 0
b7b33a4 0 5c676d08 b56b8c4 HDR+IX 40001f 0 0
No remarkable in the onlinelog, buffer is 600000,
onstat -
IBM Informix Dynamic Server Version 7.31.UD8 -- On-Line -- Up 07:06:15 --3711120 Kbytes
oncheck -pc customer also shows something weird ...
... cut ....
TBLspace customer:informix.session_answer Index xpksess_ans fragment in
DBspace cust5idx
Physical Address 460000e
Creation date 02/24/2001 10:17:57
TBLspace Flags 802 Row Locking
TBLspace use 4 bit bit-maps
Maximum row size 332
Number of special columns 0
Number of keys 1
Number of extents 139
Current serial value 1
First extent size 4
Next extent size 5120
Number of pages allocated 266668
Number of pages used 261943
Number of data pages 0
Number of rows 0
Partition partnum 54525963
Partition lockid 35651608
Extents
... ....
251308 528c38e 5120
256428 52b5da5 5120
261548 52bd1f3 5120
Index information.
Number of indexes 1
Data record size 332
Index record size 2048
Number of records 765068
Unknown error message 0.
----- Original Message ----
From: Marcus Haarmann <marcus.haarmann@midoco.de>
To: ids@iiug.org
Sent: Sunday, December 2, 2007 4:09:47 PM
Subject: RE: Urgent help is needed... [10570]
Hi,
It is normally not a problem if you have bufwaits, but these should not
take long.
Is this a new problem, or has the database been stable in this setup for
a while ? Was any config param changed ?
Was the application changed ? If you have exclusive locks not being
released, you should check these first (onstat -k to find out the owner,
onstat -g sql to see what is the owner doing).Maybe you have a problem with wrong lock-levels ?
Pls send more details on your problem. Onstat -p output in the first
place, maybe you have remarkable output in online.log, what is your
current buffer size ?
A little more info would help ..
Marcus
-----Original Message-----
From: Kern Doe [mailto:kern_doe@yahoo.com]
Sent: Sunday, December 02, 2007 9:33 PM
To: ids@iiug.org
Subject: Urgent help is needed... [10569]
Hi Gurus,
We are having issues with the database (IDS: 7.31 UD8 on Solaris 8)
where we are seeing lots of waits on buffers "B" -- sometimes
applications fail with something like "cannot position within table" --
some processes that hold exclusive locks do not release causing other
ones to wait on and fail (time
out) -- I run oncheck -cc, this is what I'm getting
oncheck -cc customerValidating database customer
Validating systables for database customer Unknown error message 0.
IBM says it is because something holds on to systables which I'm not
convinced.
Can someone who has experience with this shreds some lights.
Thanks
Kern --
________________________________________________________________________
____________
Get easy, one-click access to your favorites.
Make Yahoo! your homepage.
http://www.yahoo.com/r/hs
************************************************************************
*******
Forum Note: Use "Reply" to post a response in the discussion forum.
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
________________________________________________________________________________
____
Never miss a thing. Make Yahoo your home page.
http://www.yahoo.com/r/hs
Kern Doe wrote:
Kern you are right that locks and buffer waits are independent. Also, if
all of the 'X' locks you see are the HDR+X and HDR+IX locks that's not
your problem. Every application takes an intent-exclusive (IX) lock on
the database's systables table and on the tablespace tablespace
(partition header) record of every table it has an open cursor against.
If your applications haven't changed and the server tuning hasn't
changed, then locks that were not causing errors before aren't suddenly
doing so.
On the other hand, if the database is growing, as it clearly is from
that one table's extents below, then there are several things that could
be suddenly causing wait delays long enough to cause what used to be
momentary locks that did not interfere with other apps/users to be
causing lock out errors like the 'cannot position' errors you are
seeing. Possible causes include slow disk, disk and/or channel
contention, swapping problems, insufficient buffers, too few LRU queues,
slow LRU flushing due to too few CLEANERS, poor application design
holding locks, excessive sequential scans, and others.
To track down what causes we can I'd suggest posting the following
output and we'll see what we can see:
Time since the server stats were last zero'd (restart or onstat -z run) and:
onstat -p
onstat -d
onstat -D
onstat -P
onstat -c
onstat -g iov
onstat -g iof
onstat -g glo
onstat -g seg
Art S. Kagel
> Thanks for your repply Marcus.
> Ok, I would think this is normal -- but I've seen output from oncheck -cc
> beore, locks (or xlocks) will cause a different error (ISAM error: record is
> locked.)
>
> The db had been stable for a while, no change to config or params. The big
> issues are xlocks, yes I do them -- they used to come and go, but now they
> stay
>
> onstat -k | grep X> aea8504 0 5c676d08 b7adef0 HDR+IX 1e00050 0 0
> b32e954 0 5c676d08 b572980 HDR+IX 1e00044 0 0
> b33166c 0 5c676d08 b32e954 HDR+X 250000a 2500909 0
> b3337f0 0 5c676d08 b33166c HDR+IX 400022 0 0
> b56e094 0 5c676d08 b3337f0 HDR+X 800005 18b801 0
> b572980 0 5c676d08 b7b33a4 HDR+X 700003 8fe90a 0
> b7adef0 0 5c676d08 b56e094 HDR+IX 1a00001 0 0
> b7b33a4 0 5c676d08 b56b8c4 HDR+IX 40001f 0 0
>
> No remarkable in the onlinelog, buffer is 600000,
> onstat -
> IBM Informix Dynamic Server Version 7.31.UD8 -- On-Line -- Up 07:06:15 --> 3711120 Kbytes
>
> oncheck -pc customer also shows something weird ...
> .... cut ....>
> TBLspace customer:informix.session_answer Index xpksess_ans fragment in
> DBspace cust5idx
>
> Physical Address 460000e
>
> Creation date 02/24/2001 10:17:57
>
> TBLspace Flags 802 Row Locking
>
> TBLspace use 4 bit bit-maps
>
> Maximum row size 332
>
> Number of special columns 0
>
> Number of keys 1
>
> Number of extents 139
>
> Current serial value 1
>
> First extent size 4
>
> Next extent size 5120
>
> Number of pages allocated 266668
>
> Number of pages used 261943
>
> Number of data pages 0
>
> Number of rows 0
>
> Partition partnum 54525963
>
> Partition lockid 35651608
>
> Extents
> .... ....
>
> 251308 528c38e 5120
>
> 256428 52b5da5 5120
>
> 261548 52bd1f3 5120
>
> Index information.
>
> Number of indexes 1
>
> Data record size 332
>
> Index record size 2048
>
> Number of records 765068
> Unknown error message 0.
>
> ----- Original Message ----
> From: Marcus Haarmann <marcus.haarmann@midoco.de>
> To: ids@iiug.org
> Sent: Sunday, December 2, 2007 4:09:47 PM
> Subject: RE: Urgent help is needed... [10570]
>
> Hi,
>
> It is normally not a problem if you have bufwaits, but these should not
> take long.
> Is this a new problem, or has the database been stable in this setup for
> a while ? Was any config param changed ?
> Was the application changed ? If you have exclusive locks not being
> released, you should check these first (onstat -k to find out the owner,
> onstat -g sql to see what is the owner doing).> Maybe you have a problem with wrong lock-levels ?
> Pls send more details on your problem. Onstat -p output in the first
> place, maybe you have remarkable output in online.log, what is your
> current buffer size ?
> A little more info would help ..
>
> Marcus
>
> -----Original Message-----
> From: Kern Doe [mailto:kern_doe@yahoo.com]
> Sent: Sunday, December 02, 2007 9:33 PM
> To: ids@iiug.org
> Subject: Urgent help is needed... [10569]
>
> Hi Gurus,
> We are having issues with the database (IDS: 7.31 UD8 on Solaris 8)
> where we are seeing lots of waits on buffers "B" -- sometimes
> applications fail with something like "cannot position within table" --
> some processes that hold exclusive locks do not release causing other
> ones to wait on and fail (time
> out) -- I run oncheck -cc, this is what I'm getting
>
> oncheck -cc customer> Validating database customer
>
> Validating systables for database customer Unknown error message 0.
>
> IBM says it is because something holds on to systables which I'm not
> convinced.
> Can someone who has experience with this shreds some lights.
> Thanks
> Kern --
>
> ________________________________________________________________________
> ____________
> Get easy, one-click access to your favorites.
> Make Yahoo! your homepage.
> http://www.yahoo.com/r/hs
>
> ************************************************************************
> *******
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
________________________________________________________________________________
____
> Never miss a thing. Make Yahoo your home page.
> http://www.yahoo.com/r/hs
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
Thank you Art.
Yesterday late afternoon, the application team closed down some "weblogic" to
reduce traffic. But still, one query kept getting hit frequently (it's the
nature of the app) and this one did not exit out causing xlocks situation,
waits on Locks and Buffers situation -- then more subsequent 'dominal effect'
-- we saw more and more xlocks. My guess that something was wrong with the
optimizer or something was wrong in the statistics causing the optimizer
confused and could not perform well with its execution plan (onstat -P showed
much higher percentage for data than btree).
My partner did a upstat high on one index column of a table that related to
the "bad" query. At the same time I frequently ran my quick and dirty script
to extract the columns or the index of other tables and ran similar upstat on
them too.
By the evening thing started to get better -- the "bad" I was talking about,
now runs and comes back in a fraction of a second -- everything else seems
normal (buffers, locks, xlocks, etc.) except for some high #mutexes with
waiters (we see this all frequently anyway -- haven't been able to fix it and
it is a separate issue).
Thanks
----- Original Message ----
From: Art S. Kagel (Oninit LLC) <art@oninit.com>
To: ids@iiug.org
Sent: Sunday, December 2, 2007 10:54:45 PM
Subject: Re: Urgent help is needed... [10572]
Kern Doe wrote:
Kern you are right that locks and buffer waits are independent. Also, if
all of the 'X' locks you see are the HDR+X and HDR+IX locks that's not
your problem. Every application takes an intent-exclusive (IX) lock on
the database's systables table and on the tablespace tablespace
(partition header) record of every table it has an open cursor against.
If your applications haven't changed and the server tuning hasn't
changed, then locks that were not causing errors before aren't suddenly
doing so.
On the other hand, if the database is growing, as it clearly is from
that one table's extents below, then there are several things that could
be suddenly causing wait delays long enough to cause what used to be
momentary locks that did not interfere with other apps/users to be
causing lock out errors like the 'cannot position' errors you are
seeing. Possible causes include slow disk, disk and/or channel
contention, swapping problems, insufficient buffers, too few LRU queues,
slow LRU flushing due to too few CLEANERS, poor application design
holding locks, excessive sequential scans, and others.
To track down what causes we can I'd suggest posting the following
output and we'll see what we can see:
Time since the server stats were last zero'd (restart or onstat -z run) and:
onstat -p
onstat -d
onstat -D
onstat -P
onstat -c
onstat -g iov
onstat -g iof
onstat -g glo
onstat -g seg
Art S. Kagel
> Thanks for your repply Marcus.
> Ok, I would think this is normal -- but I've seen output from oncheck -cc
> beore, locks (or xlocks) will cause a different error (ISAM error: record is
> locked.)
>
> The db had been stable for a while, no change to config or params. The big
> issues are xlocks, yes I do them -- they used to come and go, but now they
> stay
>
> onstat -k | grep X> aea8504 0 5c676d08 b7adef0 HDR+IX 1e00050 0 0
> b32e954 0 5c676d08 b572980 HDR+IX 1e00044 0 0
> b33166c 0 5c676d08 b32e954 HDR+X 250000a 2500909 0
> b3337f0 0 5c676d08 b33166c HDR+IX 400022 0 0
> b56e094 0 5c676d08 b3337f0 HDR+X 800005 18b801 0
> b572980 0 5c676d08 b7b33a4 HDR+X 700003 8fe90a 0
> b7adef0 0 5c676d08 b56e094 HDR+IX 1a00001 0 0
> b7b33a4 0 5c676d08 b56b8c4 HDR+IX 40001f 0 0
>
> No remarkable in the onlinelog, buffer is 600000,
> onstat -
> IBM Informix Dynamic Server Version 7.31.UD8 -- On-Line -- Up 07:06:15 --> 3711120 Kbytes
>
> oncheck -pc customer also shows something weird ...
> .... cut ....>
> TBLspace customer:informix.session_answer Index xpksess_ans fragment in
> DBspace cust5idx
>
> Physical Address 460000e
>
> Creation date 02/24/2001 10:17:57
>
> TBLspace Flags 802 Row Locking
>
> TBLspace use 4 bit bit-maps
>
> Maximum row size 332
>
> Number of special columns 0
>
> Number of keys 1
>
> Number of extents 139
>
> Current serial value 1
>
> First extent size 4
>
> Next extent size 5120
>
> Number of pages allocated 266668
>
> Number of pages used 261943
>
> Number of data pages 0
>
> Number of rows 0
>
> Partition partnum 54525963
>
> Partition lockid 35651608
>
> Extents
> .... ....
>
> 251308 528c38e 5120
>
> 256428 52b5da5 5120
>
> 261548 52bd1f3 5120
>
> Index information.
>
> Number of indexes 1
>
> Data record size 332
>
> Index record size 2048
>
> Number of records 765068
> Unknown error message 0.
>
> ----- Original Message ----
> From: Marcus Haarmann <marcus.haarmann@midoco.de>
> To: ids@iiug.org
> Sent: Sunday, December 2, 2007 4:09:47 PM
> Subject: RE: Urgent help is needed... [10570]
>
> Hi,
>
> It is normally not a problem if you have bufwaits, but these should not
> take long.
> Is this a new problem, or has the database been stable in this setup for
> a while ? Was any config param changed ?
> Was the application changed ? If you have exclusive locks not being
> released, you should check these first (onstat -k to find out the owner,
> onstat -g sql to see what is the owner doing).> Maybe you have a problem with wrong lock-levels ?
> Pls send more details on your problem. Onstat -p output in the first
> place, maybe you have remarkable output in online.log, what is your
> current buffer size ?
> A little more info would help ..
>
> Marcus
>
> -----Original Message-----
> From: Kern Doe [mailto:kern_doe@yahoo.com]
> Sent: Sunday, December 02, 2007 9:33 PM
> To: ids@iiug.org
> Subject: Urgent help is needed... [10569]
>
> Hi Gurus,
> We are having issues with the database (IDS: 7.31 UD8 on Solaris 8)
> where we are seeing lots of waits on buffers "B" -- sometimes
> applications fail with something like "cannot position within table" --
> some processes that hold exclusive locks do not release causing other
> ones to wait on and fail (time
> out) -- I run oncheck -cc, this is what I'm getting
>
> oncheck -cc customer> Validating database customer
>
> Validating systables for database customer Unknown error message 0.
>
> IBM says it is because something holds on to systables which I'm not
> convinced.
> Can someone who has experience with this shreds some lights.
> Thanks
> Kern --
>
> ________________________________________________________________________
> ____________
> Get easy, one-click access to your favorites.
> Make Yahoo! your homepage.
> http://www.yahoo.com/r/hs
>
> ************************************************************************
> *******
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
*******************************************************************************@@NL@