b-tree scanner tuning and expected behavior
Posted in 2008
A user on IDS 9.40FC9 (Solaris) saw btscanner threads suddenly peg a CPU for days; stopping them restored normal CPU. He had started 8 scanner threads manually with a very low threshold (2000) and no BTSCANNER onconfig entry. Respondents said that was far too many threads and too low a threshold, recommended 1-2 threads with threshold ~500000 and enabling range scans (rangesize 100), plus ALICE on later versions (with a warning about a crash defect before 10.00.FC9). Resolution: after stopping all threads and restarting with 2 threads, threshold 500000 and rangesize 100, cleaning times dropped to seconds and CPU normalised. A follow-up question about whether a constantly open index blocks cleaning went unanswered.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Server Administration, Triggers, Constraints & Referential Integrity, Platform-Specific Issues, Versions, Editions & End-of-Life
Hi,
we're having some trouble on production instances with the b-tree scanner
threads. We're using IDS 9.40FC6-9 on solaris 5.9 (sparc).
On some of these instances we observed that suddenly 'something' triggered an
increment of CPU usage: we had a guess on what was triggering it (an
application that was launched some db cleaning procedures), but it took quite
long to find out that is was the btscanner threads that suddenly got very busy
for a few days.
In this situation, we saw with onstat -g act that there was constantly a
btscanner thread running. When we stopped it (with onmode -C stop 1) the CPU
usage normalized.
We didn't do no b-tree scanner tuning/configuration so far (onconfig has no
btscanner entry), so I started to manually configure them to find out what the
right settings would be, and on one of those instances after a few days, whith
6 threads started, the situation normalized. But on a different instance, with
much less user activity, the CPU usage has still not normalized after more
than 5 days, the instance constantly using one of 2 CPUs at 100%.
So my question is, when the b-tree scanner tuning is done correctly, what is
the expected normal behavior for this threads? Are they supposed to start
suddenly running for long periods of time (say 2 - 3 days) or should their
activity, when properly configured, be almost unnoticeable?
And, what happens if you stop/kill this threads every time there's one
running, can this generate some damage to the database (inconsistencies or
index corruption)?
Thanks in advance,
Saludos.
GERARDO PADIERNA said:
> Hi,
> we're having some trouble on production instances with the b-tree scanner
> threads. We're using IDS 9.40FC6-9 on solaris 5.9 (sparc).
> On some of these instances we observed that suddenly 'something' triggered
> an
> increment of CPU usage: we had a guess on what was triggering it (an
> application that was launched some db cleaning procedures), but it took
> quite
> long to find out that is was the btscanner threads that suddenly got very
> busy
> for a few days.
> In this situation, we saw with onstat -g act that there was constantly a
> btscanner thread running. When we stopped it (with onmode -C stop 1) the
> CPU
> usage normalized.
> We didn't do no b-tree scanner tuning/configuration so far (onconfig has
> no
> btscanner entry), so I started to manually configure them to find out what
> the
> right settings would be, and on one of those instances after a few days,
> whith
> 6 threads started, the situation normalized. But on a different instance,
> with
> much less user activity, the CPU usage has still not normalized after more
> than 5 days, the instance constantly using one of 2 CPUs at 100%.
>
> So my question is, when the b-tree scanner tuning is done correctly, what
> is
> the expected normal behavior for this threads? Are they supposed to start
> suddenly running for long periods of time (say 2 - 3 days) or should their
> activity, when properly configured, be almost unnoticeable?
> And, what happens if you stop/kill this threads every time there's one
> running, can this generate some damage to the database (inconsistencies or
> index corruption)?
What, exactly, IS your btree cleaner configuration?
--
Bye now,
Obnoxio
http://obotheclown.blogspot.com/
On the instance that now has a 'normal' CPU usage, I stopped the single one
that was running and started 8 threads with:
onmode -C start 8
onmode -C threshold 2000That's it; now, the filtered output from onstat -g ath is:
juanclu1-[PROD]-(db1):~ $ onstat -g ath | grep btscann
72316 15c33b6c8 15b9ea980 2 cond wait btc sort 1cpu btscanner_0
72317 1712a6760 156499520 2 cond wait btc sort 1cpu btscanner_1
72318 170065028 15b9c5618 1 cond wait btc sort 1cpu btscanner_2
72319 16e79d468 15ed29068 1 cond wait btc sort 1cpu btscanner_3
72320 169394028 15b9b2850 2 sleeping secs: 7 4cpu btscanner_4
72321 15a8dd140 15649ad98 2 cond wait btc sort 3cpu btscanner_5
72322 1684371d0 15ed1b430 1 cond wait btc sort 1cpu btscanner_6
72323 15a883488 15ed172f0 1 cond wait btc sort 1cpu btscanner_7
On a different instance that didn't normalize after more than 5 days, I had 8
threads, too, started with:
onmode -C start 8 (after stopping the single one that was running, with onmode
-C stop)
onmode -C threshold 2000Today I added 2 more threads with onmode -C start 2, and now, shown in onstat
-g ath:
onstat -g ath | grep btscann11095 127a0aac8 12637e1b8 2 running 3cpu btscanner_0
11096 127659300 12637f208 2 cond wait btc sort 1cpu btscanner_1
11097 127659520 126383b70 2 sleeping secs: 114 10cpu btscanner_2
11098 127659740 12637e9e0 1 cond wait btc sort 1cpu btscanner_3
11099 127659960 126384bc0 1 cond wait btc sort 1cpu btscanner_4
11100 127659b80 126384398 1 cond wait btc sort 1cpu btscanner_5
11101 127659df0 126387488 1 cond wait btc sort 1cpu btscanner_6
11102 127936b78 126386438 1 cond wait btc sort 1cpu btscanner_7
11646 1280d14a8 126381ad0 1 cond wait btc sort 1cpu btscanner_8
11647 1288bd218 126388d00 1 cond wait btc sort 1cpu btscanner_9
and, active threads:
SISAL @ fenix @ /home/informix $ onstat -g act
IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 61 days 02:13:30
-- 852992 Kbytes
Running threads:
tid tcb rstcb prty status vp-class name
8 126d73b70 0 2 running 8tli tlitcppoll
9 126e12028 0 2 running 9tli tlitcppoll
11095 127a0aac8 12637e1b8 2 running 3cpu btscanner_0
11664 12766f028 126380a80 2 running 1cpu sqlexec
As I mentioned before, the onconfig files include no BTSCANNER entry.
Thanks.
GERARDO PADIERNA said:
> On the instance that now has a 'normal' CPU usage, I stopped the single
> one
> that was running and started 8 threads with:
> onmode -C start 8
Why so many? I don't think you need more than 1 or 2.
> onmode -C threshold 2000
Why so low? Try 500000.
--
Bye now,
Obnoxio
http://obotheclown.blogspot.com/
Since the configuration was not supplied, I will just make a simple
suggestion that you turn on range scanning if under 10.00.xC8
onmode -C rangesize 100
If you are version 10.00.xC8 or higher I would suggest turning on ALICE=
scanning to value 6 or higher.
onmode -C alice 6
This will greatly increase the performance of the index cleaning.
John F. Miller III
STSM, Support Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
=
"Obnoxio The =
Clown" =
<obnoxio@serendip =
To
ita.com> ids@iiug.org =
Sent by: =
cc
ids-bounces@iiug. =
org Subj=
ect
Re: b-tree scanner tuning and =
expected behavior [12818] =
07/21/2008 02:15 =
AM =
=
=
Please respond to =
ids@iiug.org =
=
=
GERARDO PADIERNA said:
> On the instance that now has a 'normal' CPU usage, I stopped the sing=
le
> one
> that was running and started 8 threads with:
> onmode -C start 8
Why so many? I don't think you need more than 1 or 2.
> onmode -C threshold 2000
Why so low? Try 500000.
--
Bye now,
Obnoxio
http://obotheclown.blogspot.com/
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
Just to add that ALICE will only work if the indexes are detached which they
should if they were created under v10, but if you migrated in place, this
may not be the case...
Regards.
On Mon, Jul 21, 2008 at 4:23 PM, John Miller iii <miller3@us.ibm.com> wrote:
> Since the configuration was not supplied, I will just make a simple
> suggestion that you turn on range scanning if under 10.00.xC8
>
> onmode -C rangesize 100>
> If you are version 10.00.xC8 or higher I would suggest turning on ALICE=
>
> scanning to value 6 or higher.
>
> onmode -C alice 6>
> This will greatly increase the performance of the index cleaning.
>
> John F. Miller III
> STSM, Support Architect
> miller3@us.ibm.com
> 503-578-5645
> IBM Informix Dynamic Server (IDS)
>
> =
>
> "Obnoxio The =
>
> Clown" =
>
> <obnoxio@serendip =
> To
>
> ita.com> ids@iiug.org =
>
> Sent by: =
> cc
>
> ids-bounces@iiug. =
>
> org Subj=
> ect
>
> Re: b-tree scanner tuning and =
>
> expected behavior [12818] =
>
> 07/21/2008 02:15 =
>
> AM =
>
> =
>
> =
>
> Please respond to =
>
> ids@iiug.org =
>
> =
>
> =
>
> GERARDO PADIERNA said:
> > On the instance that now has a 'normal' CPU usage, I stopped the sing=
> le
> > one
> > that was running and started 8 threads with:
> > onmode -C start 8>
> Why so many? I don't think you need more than 1 or 2.
>
> > onmode -C threshold 2000>
> Why so low? Try 500000.
>
> --
> Bye now,
> Obnoxio
>
> http://obotheclown.blogspot.com/
>
> ***********************************************************************=
> ********
>
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
> =
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
John Miller iii wrote:
> If you are version 10.00.xC8 or higher I would suggest turning on ALICE=
>
> scanning to value 6 or higher.
>
> onmode -C alice 6>
> This will greatly increase the performance of the index cleaning.
>
I suggest you turn on ALICE if your are on IDS 10.00.xC9 or higher.
We had a bad crash of IDS after turning on ALICE in IDS 10.00.FC8. It's
a defect fixed in 10.00.FC9 , hits if you clean large indexes in ALICE
mode.
Can you give the bug number?
On Mon, Jul 21, 2008 at 3:30 PM, tilleul17@web.de <tilleul17@web.de> wrote:
> John Miller iii wrote:
> > If you are version 10.00.xC8 or higher I would suggest turning on ALICE=
> >
> > scanning to value 6 or higher.
> >
> > onmode -C alice 6> >
> > This will greatly increase the performance of the index cleaning.
> >
> I suggest you turn on ALICE if your are on IDS 10.00.xC9 or higher.
> We had a bad crash of IDS after turning on ALICE in IDS 10.00.FC8. It's
> a defect fixed in 10.00.FC9 , hits if you clean large indexes in ALICE
> mode.
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Eric B. Rowell
Hi all,
I don't know exactly what configuration details you need, but I already posted
this:
1. We have IDS 9.40FCX (we have some instances with 10.00XC8, too, but no
problems with them)
2. No configuration at all in onconfig that makes any reference to BTSCANNER
3. We started manually up to 8 threads with the command:
onmode -C start #4. We played around with some threshold values, and the last one we had was:
onmode -C threshold 20005. As I didn' touch the priority, I guess it should be low (ok, I didn't
mention this before).
Nothing else.
Besides, I wanted to know if it can be considered a normal behavior that
suddenly one of those btscanner threads starts working, getting
uninterruptedly busy for several days consuming one CPU entirely, or if this
is an indication that something is not well configured or tuned (of course my
customers think it's not ok).
If you need the onconfig file or the output from some commands, please tell
me, but as far as I know, there's nothing more to configure (in IDS 9.40XXX at
least) in relation with the b-tree scanner than what I already told.
Thanks.
GERARDO PADIERNA said:
> Hi all,
> I don't know exactly what configuration details you need, but I already
> posted
> this:
> 1. We have IDS 9.40FCX (we have some instances with 10.00XC8, too, but no
> problems with them)
> 2. No configuration at all in onconfig that makes any reference to
> BTSCANNER> 3. We started manually up to 8 threads with the command:
> onmode -C start #> 4. We played around with some threshold values, and the last one we had
> was:
> onmode -C threshold 2000> 5. As I didn' touch the priority, I guess it should be low (ok, I didn't
> mention this before).
> Nothing else.
> Besides, I wanted to know if it can be considered a normal behavior that
> suddenly one of those btscanner threads starts working, getting
> uninterruptedly busy for several days consuming one CPU entirely, or if
> this
> is an indication that something is not well configured or tuned (of course
> my
> customers think it's not ok).
> If you need the onconfig file or the output from some commands, please
> tell
> me, but as far as I know, there's nothing more to configure (in IDS
> 9.40XXX at
> least) in relation with the b-tree scanner than what I already told.
> Thanks.
Post the output of:
onstat -C
onstat -C hot
onstat -C clean
You may also want to read this:
http://obotheclown.blogspot.com/2008/07/tech-tip-du-jour-quick-performance.html
Section 2c. Also, read the comments.
--
Bye now,
Obnoxio
http://obotheclown.blogspot.com/
here the outputs, and thanks for the references.
Regards.
juanclu1-[PROD]-(db1):~ $ onstat -C
IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days 21:33:27
-- 2015232 Kbytes
Btree Cleaner Info
BT scanner profile Information
==============================
Active Threads 8
Global Commands 20000 Building hot list
Number of partition scans 8433
Main Block 0x0000000156edad58
BTC Admin 0x000000015b9b2850
BTS info id Prio Partnum Key Cmd
0x1595ffbb0 0 High 0x00000000 0 20000 Building hot list
Number of leaves pages scanned 1717756000
Number of leaves with deleted items 2126059045
Time spent cleaning (sec) 71884
Number of index compresses 4440
Number of deleted items 523972
Number of index range scans 0
Number of index leaf scans 362
BTS info id Prio Partnum Key Cmd
0x15d7aebb0 1 High 0x00000000 0 20000 Building hot list
Number of leaves pages scanned -1003796269
Number of leaves with deleted items 1933096306
Time spent cleaning (sec) 56335
Number of index compresses 2702
Number of deleted items 402410
Number of index range scans 0
Number of index leaf scans 350
BTS info id Prio Partnum Key Cmd
0x1693518d8 2 Low 0x00000000 0 20000 Building hot list
Number of leaves pages scanned 125550
Number of leaves with deleted items 3689
Time spent cleaning (sec) 2
Number of index compresses 1848
Number of deleted items 187050
Number of index range scans 0
Number of index leaf scans 97
BTS info id Prio Partnum Key Cmd
0x1640809b0 3 Low 0x00000000 0 20000 Building hot list
Number of leaves pages scanned 116592
Number of leaves with deleted items 2425
Time spent cleaning (sec) 4
Number of index compresses 1682
Number of deleted items 167391
Number of index range scans 0
Number of index leaf scans 99
BTS info id Prio Partnum Key Cmd
0x169a26028 4 High 0x00000000 0 40 Yield N
Number of leaves pages scanned 208687
Number of leaves with deleted items 6698
Time spent cleaning (sec) 9
Number of index compresses 2811
Number of deleted items 298249
Number of index range scans 0
Number of index leaf scans 293
BTS info id Prio Partnum Key Cmd
0x15c9f38f0 5 High 0x006000CC 1 100000 Scan index
Number of leaves pages scanned 1885192740
Number of leaves with deleted items 1884963443
Time spent cleaning (sec) 8
Number of index compresses -525053193
Number of deleted items 393907
Number of index range scans 0
Number of index leaf scans 345
Scan Type Leaf
BTS info id Prio Partnum Key Cmd
0x1697e1ba8 6 Low 0x00000000 0 20000 Building hot list
Number of leaves pages scanned 122572
Number of leaves with deleted items 1540
Time spent cleaning (sec) 2
Number of index compresses 2110
Number of deleted items 201112
Number of index range scans 0
Number of index leaf scans 104
BTS info id Prio Partnum Key Cmd
0x1560165a8 7 Low 0x00000000 0 20000 Building hot list
Number of leaves pages scanned 107901
Number of leaves with deleted items 3626
Time spent cleaning (sec) 3
Number of index compresses 1729
Number of deleted items 174507
Number of index range scans 0
Number of index leaf scans 86
__________________________________________________________________
juanclu1-[PROD]-(db1):~ $ onstat -C hot
IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days 21:35:15
-- 2015232 Kbytes
Btree Cleaner Info
Index Hot List
==============
Current Item 2 List Created 12:00:54
List Size 1 List expires in 0 sec
Hit Threshold 100000 Range Scan Threshold -1
Partnum Key Hits
0x006000CB 1 8789 *
____________________________________________________________________
juanclu1-[PROD]-(db1):~ $ onstat -C clean
IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days 21:35:38
-- 2015232 Kbytes
Btree Cleaner Info
Index Cleaned Statistics
=========================
Partnum Key Dirty Hits Clean Time Pg Examined Items Del Pages/Sec
0x006000c3 1 1927 0 246 37638 246.00
0x006000c5 1 1508 2 2474 152556 1237.00
0x006000c8 1 457 0 0 11 0.00
0x006000c9 1 1192 15 371030 1388602 24735.33
0x006000ca 1 7470 92 466585 1597270 5071.58
0x006000cb 1 12682 161 569000 1733933 3534.16
0x006000cc 1 C 1690589 292236 -171845166 1725977 14108.88
0x00600031 1 18 0 0 0 0.00
0x00600071 1 1364 0 0 0 0.00
0x006000d1 1 1846 3 39103 709041 13034.33
0x006000f1 1 1364 0 0 0 0.00
0x00600032 1 18 0 0 0 0.00
0x006000d2 1 432 0 0 0 0.00
0x00600033 1 18 0 0 0 0.00
0x006000d3 1 320 0 0 0 0.00
0x006000b4 1 5 0 0 0 0.00
0x006000d4 1 320 0 0 0 0.00
0x00600035 1 18 0 0 0 0.00
0x00600055 1 2 0 0 0 0.00
0x00600055 2 2 0 0 0 0.00
0x00600095 1 16 0 0 0 0.00
0x00600096 1 750 15 353304 1363336 23553.60
0x006000d6 1 1183 0 692 48814 692.00
0x00600077 1 1 0 4 586 4.00
0x006000d7 1 1920 1 248 37637 248.00
0x006000f7 1 0 0 10 578 10.00
0x00600078 1 1 0 0 0 0.00
0x00600098 1 993 16 62999 311597 3937.44
0x006000d8 1 1068 0 723 14398 723.00
0x006000d9 1 1728 1 2194 153372 2194.00
0x006000da 1 21 5 7522 423596 1504.40
0x006000bb 1 1014 12 127990 710001 10665.83
0x006000db 1 1170 0 0 131 0.00
0x006000be 1 640 0 0 0 0.00
0x00800002 1 1 0 0 0 0.00
0x00800003 1 4 0 0 0 0.00
0x00800003 2 17 0 0 0 0.00
0x00800005 1 16 0 0 0 0.00
0x00800005 2 4 0 0 0 0.00
0x0080000d 1 9 0 0 0 0.00
0x00800010 1 2 0 0 0 0.00
0x00800011 1 27 0 0 0 0.00
0x00800011 2 6 0 0 0 0.00
0x00800011 3 1 0 0 0 0.00
0x00800013 1 0 0 14 1458 14.00
0x00800014 1 2 0 0 0 0.00
0x00800014 2 5 0 0 0 0.00
0x00800017 1 11 0 0 0 0.00
0x00800017 2 8 0 0 0 0.00
0x00800019 1 6 0 60 4340 60.00
0x00900002 1 14 0 0 0 0.00
0x00900002 2 9 0 0 0 0.00
0x00900003 1 239 0 0 0 0.00
0x00900003 2 17 0 0 0 0.00
0x00900004 1 809 0 0 0 0.00
0x00900004 2 48 0 0 0 0.00
0x00900005 1 18 0 0 0 0.00
0x00900005 2 24 0 0 0 0.00
0x009008a7 2 21 0 0 0 0.00
0x009008a8 1 20 0 0 0 0.00
0x009008a8 2 19 0 0 0 0.00
0x009008a9 1 32 0 0 0 0.00
0x009008a9 2 2 0 0 0 0.00
0x009008aa 1 10 0 0 0 0.00
0x009008aa 2 8 0 0 0 0.00
0x0090000c 1 37 0 0 0 0.00
0x0090000c 2 46 0 0 0 0.00
0x0090000c 3 808 0 0 0 0.00
0x0090000d 1 19 0 0 0 0.00
0x0090000d 2 6 0 0 0 0.00
0x0090000e 1 30 0 0 0 0.00
0x0090002e 1 51 0 0 0 0.00
0x0090002e 2 7 0 0 0 0.00
0x00900011 1 15 0 0 0 0.00
0x00900011 2 2 0 0 0 0.00
0x00900011 3 1 0 0 0 0.00
0x00900013 1 0 0 32 1726 32.00
0x009008d3 1 2 0 0 0 0.00
0x009008d3 2 2 0 0 0 0.00
0x00900014 1 16 0 0 0 0.00
0x00900014 2 15 0 0 0 0.00
0x00900017 1 6 0 0 0 0.00
0x00900017 2 5 0 0 0 0.00
0x009008b8 1 1395 0 0 0 0.00
0x00900019 1 0 0 643 16244 643.00
0x0090001a 1 1630 0 0 0 0.00
0x0090001c 2 1 0 0 0 0.00
0x0090001c 3 1 0 0 0 0.00
0x0090001c 4 12 0 0 0 0.00
0x009008be 1 1002 0 0 0 0.00
0x00a00020 1 5 0 0 0 0.00
0x00a
GERARDO PADIERNA said:
> here the outputs, and thanks for the references.
> Regards.
>
> juanclu1-[PROD]-(db1):~ $ onstat -C
>
> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
> 21:33:27
> -- 2015232 Kbytes>
> Btree Cleaner Info
> BT scanner profile Information
> ==============================
> Active Threads 8
> Global Commands 20000 Building hot list
> Number of partition scans 8433
> Main Block 0x0000000156edad58
> BTC Admin 0x000000015b9b2850
>
> BTS info id Prio Partnum Key Cmd
> 0x1595ffbb0 0 High 0x00000000 0 20000 Building hot list
>
> Number of leaves pages scanned 1717756000
Eeek!
> Number of leaves with deleted items 2126059045
>
> Time spent cleaning (sec) 71884
EEEEEEEEEEEEEEEEEEEEEEEEEEEEEK!
OK, only two of these cleaners are actually doing anything, but between
them they have been active for nearly 25% of your uptime. That is insane.
Increase your threshold to at least 500000, drop the other 6 cleaners and
then monitor the situation.
onmode -C threshold 500000
You may want to increase it even more.
> 0x15d7aebb0 1 High 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 56335
> 0x1693518d8 2 Low 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 2
> 0x1640809b0 3 Low 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 4
> 0x169a26028 4 High 0x00000000 0 40 Yield N
> Time spent cleaning (sec) 9
> 0x15c9f38f0 5 High 0x006000CC 1 100000 Scan index
> Time spent cleaning (sec) 8
> 0x1697e1ba8 6 Low 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 2
> 0x1560165a8 7 Low 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 3
> __________________________________________________________________
>
> juanclu1-[PROD]-(db1):~ $ onstat -C hot
>
> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
> 21:35:15
> -- 2015232 Kbytes>
> Btree Cleaner Info
>
> Index Hot List
> ==============
>
> Current Item 2 List Created 12:00:54
>
> List Size 1 List expires in 0 sec
>
> Hit Threshold 100000 Range Scan Threshold -1
Try setting range to 100: onmode -C range 100
> Partnum Key Hits
> 0x006000CB 1 8789 *
>
> ____________________________________________________________________
>
> juanclu1-[PROD]-(db1):~ $ onstat -C clean
>
> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
> 21:35:38
> -- 2015232 Kbytes>
> Btree Cleaner Info
>
> Index Cleaned Statistics
> =========================
> Partnum Key Dirty Hits Clean Time Pg Examined Items Del Pages/Sec
> 0x006000c9 1 1192 15 371030 1388602 24735.33
> 0x006000ca 1 7470 92 466585 1597270 5071.58
> 0x006000cb 1 12682 161 569000 1733933 3534.16
> 0x006000cc 1 C 1690589 292236 -171845166 1725977 14108.88
This is interesting ... why so many deletes?
> 0x00600096 1 750 15 353304 1363336 23553.60
> 0x006000bb 1 1014 12 127990 710001 10665.83
Check all these partitions. See if there are any indexes you can drop.
Have a look at your application, it seems to be doing a lot of index
deleting...
--
Bye now,
Obnoxio
http://obotheclown.blogspot.com/
I have had a similar situation with v9.40fc9 ... According to Tech
Support there are a couple of different bugs that can cause this
continuous cleaning. Of course there doesn't seem to be a pattern as to
what causes it ...
Tech Support had me add the following to the onconfig file ...
BTSCANNER NUM=1, PRIORITY=high, THRESHOLD=200000
That seems to have fixed the problem.
The systems I have seen this on were not real busy, so I haven't tried to
tune things better.
TS also did warn me about putting this in any system without testing it
first ....
Peter Logan
Senior Database Administrator
Phone: 616/878-8309
From:
"GERARDO PADIERNA" <g.padierna@gmail.com>
To:
ids@iiug.org
Date:
07/22/2008 05:15 AM
Subject:
Re: b-tree scanner tuning and expected behavior [12847]
Hi all,
I don't know exactly what configuration details you need, but I already
posted
this:
1. We have IDS 9.40FCX (we have some instances with 10.00XC8, too, but no
problems with them)
2. No configuration at all in onconfig that makes any reference to
BTSCANNER3. We started manually up to 8 threads with the command:
onmode -C start #4. We played around with some threshold values, and the last one we had
was:
onmode -C threshold 20005. As I didn' touch the priority, I guess it should be low (ok, I didn't
mention this before).
Nothing else.
Besides, I wanted to know if it can be considered a normal behavior that
suddenly one of those btscanner threads starts working, getting
uninterruptedly busy for several days consuming one CPU entirely, or if
this
is an indication that something is not well configured or tuned (of course
my
customers think it's not ok).
If you need the onconfig file or the output from some commands, please
tell
me, but as far as I know, there's nothing more to configure (in IDS
9.40XXX at
least) in relation with the b-tree scanner than what I already told.
Thanks.
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
I would first allow range scans to be run on your system. Currently
these are disabled. To do this you can run onmode -C rangesize 100
This means all tables with more than 100 pages will use range scaning.
To make this permanent you need to set the configuration parameter
BTSCANNERS. To include previously mentioned recommendation
your configuration parameter should look like this.
BTSCANNERS num=3D2,threshold=3D500000,rangsize=3D100
GERARDO PADIERNA said:
> here the outputs, and thanks for the references.
> Regards.
>
> juanclu1-[PROD]-(db1):~ $ onstat -C
>
> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
> 21:33:27
> -- 2015232 Kbytes>
> Btree Cleaner Info
> BT scanner profile Information
> =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D=3D=3D=3D=3D=3D
> Active Threads 8
> Global Commands 20000 Building hot list
> Number of partition scans 8433
> Main Block 0x0000000156edad58
> BTC Admin 0x000000015b9b2850
>
> BTS info id Prio Partnum Key Cmd
> 0x1595ffbb0 0 High 0x00000000 0 20000 Building hot list
>
> Number of leaves pages scanned 1717756000
Eeek!
> Number of leaves with deleted items 2126059045
>
> Time spent cleaning (sec) 71884
EEEEEEEEEEEEEEEEEEEEEEEEEEEEEK!
OK, only two of these cleaners are actually doing anything, but between=
them they have been active for nearly 25% of your uptime. That is insan=
e.
Increase your threshold to at least 500000, drop the other 6 cleaners a=
nd
then monitor the situation.
onmode -C threshold 500000
You may want to increase it even more.
> 0x15d7aebb0 1 High 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 56335
> 0x1693518d8 2 Low 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 2
> 0x1640809b0 3 Low 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 4
> 0x169a26028 4 High 0x00000000 0 40 Yield N
> Time spent cleaning (sec) 9
> 0x15c9f38f0 5 High 0x006000CC 1 100000 Scan index
> Time spent cleaning (sec) 8
> 0x1697e1ba8 6 Low 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 2
> 0x1560165a8 7 Low 0x00000000 0 20000 Building hot list
> Time spent cleaning (sec) 3
> __________________________________________________________________
>
> juanclu1-[PROD]-(db1):~ $ onstat -C hot
>
> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
> 21:35:15
> -- 2015232 Kbytes>
> Btree Cleaner Info
>
> Index Hot List
> =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D
>
> Current Item 2 List Created 12:00:54
>
> List Size 1 List expires in 0 sec
>
> Hit Threshold 100000 Range Scan Threshold -1
Try setting range to 100: onmode -C range 100
> Partnum Key Hits
> 0x006000CB 1 8789 *
>
> ____________________________________________________________________
>
> juanclu1-[PROD]-(db1):~ $ onstat -C clean
>
> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
> 21:35:38
> -- 2015232 Kbytes>
> Btree Cleaner Info
>
> Index Cleaned Statistics
> =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
=3D=3D
> Partnum Key Dirty Hits Clean Time Pg Examined Items Del Pages/Sec
> 0x006000c9 1 1192 15 371030 1388602 24735.33
> 0x006000ca 1 7470 92 466585 1597270 5071.58
> 0x006000cb 1 12682 161 569000 1733933 3534.16
> 0x006000cc 1 C 1690589 292236 -171845166 1725977 14108.88
This is interesting ... why so many deletes?
> 0x00600096 1 750 15 353304 1363336 23553.60
> 0x006000bb 1 1014 12 127990 710001 10665.83
Check all these partitions. See if there are any indexes you can drop.
Have a look at your application, it seems to be doing a lot of index
deleting...
--
Bye now,
Obnoxio
http://obotheclown.blogspot.com/
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
John Miller iii said:
> I would first allow range scans to be run on your system. Currently
> these are disabled. To do this you can run onmode -C rangesize 100
> This means all tables with more than 100 pages will use range scaning.
> To make this permanent you need to set the configuration parameter
> BTSCANNERS. To include previously mentioned recommendation
> your configuration parameter should look like this.
>
> BTSCANNERS num=3D2,threshold=3D500000,rangsize=3D100
Ah, gotta love that Bloatus Goats, don't we? Did you perchance mean:
BTSCANNERS num=2,threshold=500000,rangesize=100
:o)
> GERARDO PADIERNA said:
>> here the outputs, and thanks for the references.
>> Regards.
>>
>> juanclu1-[PROD]-(db1):~ $ onstat -C
>>
>> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
>> 21:33:27
>> -- 2015232 Kbytes>>
>> Btree Cleaner Info
>> BT scanner profile Information
>> =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
> =3D=3D=3D=3D=3D=3D=3D
>> Active Threads 8
>> Global Commands 20000 Building hot list
>> Number of partition scans 8433
>> Main Block 0x0000000156edad58
>> BTC Admin 0x000000015b9b2850
>>
>> BTS info id Prio Partnum Key Cmd
>> 0x1595ffbb0 0 High 0x00000000 0 20000 Building hot list
>>
>> Number of leaves pages scanned 1717756000
>
> Eeek!
>
>> Number of leaves with deleted items 2126059045
>>
>> Time spent cleaning (sec) 71884
>
> EEEEEEEEEEEEEEEEEEEEEEEEEEEEEK!
>
> OK, only two of these cleaners are actually doing anything, but between=
>
> them they have been active for nearly 25% of your uptime. That is insan=
> e.
> Increase your threshold to at least 500000, drop the other 6 cleaners a=
> nd
> then monitor the situation.
>
> onmode -C threshold 500000>
> You may want to increase it even more.
>
>> 0x15d7aebb0 1 High 0x00000000 0 20000 Building hot list
>> Time spent cleaning (sec) 56335
>
>> 0x1693518d8 2 Low 0x00000000 0 20000 Building hot list
>> Time spent cleaning (sec) 2
>
>> 0x1640809b0 3 Low 0x00000000 0 20000 Building hot list
>> Time spent cleaning (sec) 4
>
>> 0x169a26028 4 High 0x00000000 0 40 Yield N
>> Time spent cleaning (sec) 9
>
>> 0x15c9f38f0 5 High 0x006000CC 1 100000 Scan index
>> Time spent cleaning (sec) 8
>
>> 0x1697e1ba8 6 Low 0x00000000 0 20000 Building hot list
>> Time spent cleaning (sec) 2
>
>> 0x1560165a8 7 Low 0x00000000 0 20000 Building hot list
>> Time spent cleaning (sec) 3
>
>> __________________________________________________________________
>>
>> juanclu1-[PROD]-(db1):~ $ onstat -C hot
>>
>> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
>> 21:35:15
>> -- 2015232 Kbytes>>
>> Btree Cleaner Info
>>
>> Index Hot List
>> =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D
>>
>> Current Item 2 List Created 12:00:54
>>
>> List Size 1 List expires in 0 sec
>>
>> Hit Threshold 100000 Range Scan Threshold -1
>
> Try setting range to 100: onmode -C range 100
>
>> Partnum Key Hits
>> 0x006000CB 1 8789 *
>>
>> ____________________________________________________________________
>>
>> juanclu1-[PROD]-(db1):~ $ onstat -C clean
>>
>> IBM Informix Dynamic Server Version 9.40.FC9 -- On-Line -- Up 5 days
>> 21:35:38
>> -- 2015232 Kbytes>>
>> Btree Cleaner Info
>>
>> Index Cleaned Statistics
>> =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=
> =3D=3D
>> Partnum Key Dirty Hits Clean Time Pg Examined Items Del Pages/Sec
>> 0x006000c9 1 1192 15 371030 1388602 24735.33
>> 0x006000ca 1 7470 92 466585 1597270 5071.58
>> 0x006000cb 1 12682 161 569000 1733933 3534.16
>> 0x006000cc 1 C 1690589 292236 -171845166 1725977 14108.88
>
> This is interesting ... why so many deletes?
>
>> 0x00600096 1 750 15 353304 1363336 23553.60
>> 0x006000bb 1 1014 12 127990 710001 10665.83
>
> Check all these partitions. See if there are any indexes you can drop.
> Have a look at your application, it seems to be doing a lot of index
> deleting...
>
> --
> Bye now,
> Obnoxio
>
> http://obotheclown.blogspot.com/
>
> ***********************************************************************=
> ********
>
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
> =
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
>
--
Bye now,
Obnoxio
http://obotheclown.blogspot.com/
Eric Rowell wrote: > Can you give the bug number? > >> I suggest you turn on ALICE if your are on IDS 10.00.xC9 or higher. >> We had a bad crash of IDS after turning on ALICE in IDS 10.00.FC8. It's >> a defect fixed in 10.00.FC9 , hits if you clean large indexes in ALICE >> mode. No bug number , sorry. Defect was discovered internally - and initially considered more a cosmetic issue. However it has later been discovered it has the portntial to crash IDS server.
John,
Can you supply what the onconfig parameter BTSCANNER should look like in
11.1 or 11.5 ... Thanks ...
Peter Logan
Senior Database Administrator
Phone: 616/878-8309
From:
"John Miller iii" <miller3@us.ibm.com>
To:
ids@iiug.org
Date:
07/21/2008 11:24 AM
Subject:
Re: b-tree scanner tuning and expected behavior [12834]
Since the configuration was not supplied, I will just make a simple
suggestion that you turn on range scanning if under 10.00.xC8
onmode -C rangesize 100
If you are version 10.00.xC8 or higher I would suggest turning on ALICE=
scanning to value 6 or higher.
onmode -C alice 6
This will greatly increase the performance of the index cleaning.
John F. Miller III
STSM, Support Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
=
"Obnoxio The =
Clown" =
<obnoxio@serendip =
To
ita.com> ids@iiug.org =
Sent by: =
cc
ids-bounces@iiug. =
org Subj=
ect
Re: b-tree scanner tuning and =
expected behavior [12818] =
07/21/2008 02:15 =
AM =
=
=
Please respond to =
ids@iiug.org =
=
=
GERARDO PADIERNA said:
> On the instance that now has a 'normal' CPU usage, I stopped the sing=
le
> one
> that was running and started 8 threads with:
> onmode -C start 8
Why so many? I don't think you need more than 1 or 2.
> onmode -C threshold 2000
Why so low? Try 500000.
--
Bye now,
Obnoxio
http://obotheclown.blogspot.com/
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
Hi,
I guess it's solved now: but just changing the parameters didn't stop the
btscanner activity. First I had to stop all threads, then I started two
threads with:
onmode -C start 2
onmode -C threshold 500000
onmode -C rangesize 100and it's running fine now for about 5 days without these long high activity
periods that we observed earlier (we monitor the CPU activity with nagios and
get nice graphs that are really much better now).
The cleaning times now are between 2 and 32 seconds (started as I said about 5
days ago).
Now we will include the proper line in our onconfig files and I hope that's it.
One last question. As I monitored the cleaning activity with onstat -C clean,
I saw that one and the same index (partnum) was being cleaned for hours and
hours (or even days); I found via a 'select hex(partnum) from systabnames'
which index that was, and than I checked with onstat -g opn that this index
was constantly being kept open. Could this block the btscanner from actually
cleaning the index? Does the btscanner need exclusive access so that if some
session is keeping the table/index open the cleaning process is in some sort
of busy waiting status?
Thank you all.
Regards.
No the btree scanner does not need exclusive access to clean the
table.
John F. Miller III
STSM, Support Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
=
"GERARDO =
PADIERNA" =
<g.padierna@gmail =
To
.com> ids@iiug.org =
Sent by: =
cc
ids-bounces@iiug. =
org Subj=
ect
Re: b-tree scanner tuning and =
expected behavior [12936] =
07/29/2008 01:50 =
AM =
=
=
Please respond to =
ids@iiug.org =
=
=
Hi,
I guess it's solved now: but just changing the parameters didn't stop t=
he
btscanner activity. First I had to stop all threads, then I started two=
threads with:
onmode -C start 2
onmode -C threshold 500000
onmode -C rangesize 100and it's running fine now for about 5 days without these long high acti=
vity
periods that we observed earlier (we monitor the CPU activity with nagi=
os
and
get nice graphs that are really much better now).
The cleaning times now are between 2 and 32 seconds (started as I said
about 5
days ago).
Now we will include the proper line in our onconfig files and I hope th=
at's
it.
One last question. As I monitored the cleaning activity with onstat -C
clean,
I saw that one and the same index (partnum) was being cleaned for hours=
and
hours (or even days); I found via a 'select hex(partnum) from systabnam=
es'
which index that was, and than I checked with onstat -g opn that this i=
ndex
was constantly being kept open. Could this block the btscanner from
actually
cleaning the index? Does the btscanner need exclusive access so that if=
some
session is keeping the table/index open the cleaning process is in some=
sort
of busy waiting status?
Thank you all.
Regards.
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g