Re: IDS 10 erratic run times
Posted in 2005
Neil Truby reported wildly variable run times on IDS 10.00.FC3 under Solaris 9 (dbimports taking 9-14h, and a repeatable CREATE CLUSTER INDEX test taking 25-30 min normally but 50-120 min on roughly one run in two or three), with onstat showing threads sleeping and no I/O progress. Testing showed times were consistent when KAIO was disabled, and that 9.40FC7 behaved consistently with KAIO on or off, so he attributed it to KAIO on v10 and opened IBM case 439918 (plus 439945, since disabling KAIO made restores ~40x slower). Others suggested onmode -F, checking stacks/mutexes/spins, NOAGE and AIOVP counts, and raising it with the OS vendor too. No fix is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Logging & Checkpoints, Migration, Import/Export & Data Conversion, Platform-Specific Issues, Clustering, Grid & MACH11, Versions, Editions & End-of-Life
IDS 10.0FC3 on Solaris 2.9
>> This is a brand new server which we are preparing to migrate a v9 system
>> to.
Over the past few weeks I have run multiple dbimports from more-or-less the
same data source - certainly the same schema - but the import time varies
widely.
>> The fastest I've done it is 9h, the slowest 14. This weekend's took 13h
10m. There's never anyhting else going on on it. In the UPDATE STATISICS
phase the elapsed time was 3h: during the "fast" run it was just 1h. I had
a look during this period: often onstat -p onstat -m onstat -D would show
periods of perhaps 2 minutes with no activity at all: no bufreads or
physreads, no checkpointing for 20 minutes (the checkpoint interval is 5
mins), the relevant thread just "sleeping forever".
We've done a lot of extra work on this, and have now built a test case with
a smaller table - just repeatedly running drop index; add cluster index. It
normally takes 25-30 minutes. About 1 in every 2 or 3 runs it takens
significantly longer - anything from 50 to 120 mins. There seems to be no
discernible pattern - if the instance has been bounced for example. During a
slow run we see with an onstat -g iov - see below - absolutely nothing
increasing for seconds on end except the wake-up count.
But, if we disable kaio (this is Solaris, so it's substituting block for
character devices), we get a consistent performance every time. We have
repeated this 20 or more times with and without kaio so I am increasingly
confident that this is a pattern, not coincidenece, and that the problem is
with kaio on v10.
Our next stage is to repeat the tests with 9.40FC7. I'll be raising a tech
support call on Monday when they are open again. In the meantime if anyone
has any observations I'd welcome them
regards
Neil
$ onstat -g iov -r 1 | grep kio
kio 0 s 572.3 10581014 5492312 5088702 0 16255623 0.7 0
kio 0 s 572.2 10581014 5492312 5088702 0 16255678 0.7 0
kio 0 s 572.2 10581014 5492312 5088702 0 16255733 0.7 0
kio 0 s 572.2 10581014 5492312 5088702 0 16255786 0.7 0
kio 0 s 572.1 10581014 5492312 5088702 0 16255841 0.7 0
kio 0 s 572.1 10581014 5492312 5088702 0 16255896 0.7 0
kio 0 s 572.1 10581014 5492312 5088702 0 16255950 0.7 0
kio 0 s 572.0 10581014 5492312 5088702 0 16256003 0.7 0
kio 0 s 572.0 10581014 5492312 5088702 0 16256058 0.7 0
kio 0 s 572.0 10581014 5492312 5088702 0 16256112 0.7 0
kio 0 s 571.9 10581014 5492312 5088702 0 16256167 0.7 0
kio 0 s 571.9 10581014 5492312 5088702 0 16256222 0.7 0
kio 0 s 571.9 10581014 5492312 5088702 0 16256277 0.7 0
kio 0 s 571.9 10581014 5492312 5088702 0 16256333 0.7 0
kio 0 s 571.8 10581014 5492312 5088702 0 16256388 0.7 0
kio 0 s 571.8 10581014 5492312 5088702 0 16256438 0.7 0
kio 0 s 571.8 10581014 5492312 5088702 0 16256492 0.7 0
kio 0 s 571.7 10581014 5492312 5088702 0 16256544 0.7 0
kio 0 s 571.7 10581014 5492312 5088702 0 16256597 0.7 0
kio 0 s 571.7 10581014 5492312 5088702 0 16256650 0.7 0
kio 0 s 571.6 10581014 5492312 5088702 0 16256703 0.7 0
kio 0 s 571.6 10581014 5492312 5088702 0 16256757 0.7 0
kio 0 s 571.6 10581014 5492312 5088702 0 16256809 0.7 0
kio 0 s 571.5 10581014 5492312 5088702 0 16256862 0.7 0
kio 0 s 571.5 10581014 5492312 5088702 0 16256915 0.7 0
kio 0 s 571.5 10581014 5492312 5088702 0 16256968 0.7 0
kio 0 s 571.5 10581014 5492312 5088702 0 16257027 0.7 0
kio 0 s 571.4 10581014 5492312 5088702 0 16257079 0.7 0
kio 0 s 571.4 10581014 5492312 5088702 0 16257132 0.7 0
kio 0 s 571.4 10581014 5492312 5088702 0 16257189 0.7 0
kio 0 s 571.3 10581014 5492312 5088702 0 16257242 0.7 0
kio 0 s 571.3 10581014 5492312 5088702 0 16257295 0.7 0
kio 0 s 571.3 10581014 5492312 5088702 0 16257347 0.7 0
kio 0 s 571.2 10581014 5492312 5088702 0 16257400 0.7 0
kio 0 s 571.2 10581014 5492312 5088702 0 16257454 0.7 0
kio 0 s 571.2 10581014 5492312 5088702 0 16257510 0.7 0
kio 0 s 571.1 10581014 5492312 5088702 0 16257567 0.7 0
kio 0 s 571.1 10581014 5492312 5088702 0 16257621 0.7 0
kio 0 s 571.1 10581014 5492312 5088702 0 16257674 0.7 0
kio 0 s 571.1 10581014 5492312 5088702 0 16257726 0.7 0
kio 0 s 571.0 10581014 5492312 5088702 0 16257778 0.7 0
kio 0 s 571.0 10581014 5492312 5088702 0 16257835 0.7 0
kio 0 s 571.0 10581014 5492312 5088702 0 16257890 0.7 0
kio 0 s 570.9 10581014 5492312 5088702 0 16257947 0.7 0
kio 0 s 570.9 10581014 5492312 5088702 0 16258000 0.7 0
kio 0 s 570.9 10581014 5492312 5088702 0 16258056 0.7 0
kio 0 s 570.8 10581014 5492312 5088702 0 16258111 0.7 0
kio 0 s 570.8 10581014 5492312 5088702 0 16258167 0.7 0
kio 0 s 570.8 10581014 5492312 5088702 0 16258219 0.7 0
kio 0 s 570.7 10581014 5492312 5088702 0 16258270 0.7 0
kio 0 s 570.7 10581014 5492312 5088702 0 16258325 0.7 0
kio 0 s 570.7 10581014 5492312 5088702 0 16258378 0.7 0
kio 0 s 570.7 10581014 5492312 5088702 0 16258431 0.7 0
kio 0 s 570.6 10581014 5492312 5088702 0 16258486 0.7 0
kio 0 s 570.6 10581014 5492312 5088702 0 16258540 0.7 0
kio 0 s 570.6 10581014 5492312 5088702 0 16258595 0.7 0
kio 0 s 570.5 10581014 5492312 5088702 0 16258646 0.7 0
kio 0 s 570.5 10581014 5492312 5088702 0 16258700 0.7 0
kio 0 s 570.5 10581014 5492312 5088702 0 16258754 0.7 0
kio 0 s 570.4 10581014 5492312 5088702 0 16258810 0.7 0
kio 0 s 570.4 10581014 5492312 5088702 0 16258863 0.7 0
kio 0 s 570.4 10581014 5492312 5088702 0 16258917 0.7 0
kio 0 s 570.3 10581014 5492312 5088702 0 16258973 0.7 0
kio 0 s 570.3 10581014 5492312 5088702 0 16259027 0.7 0
kio 0 s 570.3 10581014 5492312 5088702 0 16259080 0.7 0
kio 0 s 570.3 10581014 5492312 5088702 0 16259138 0.7 0
kio 0 s 570.2 10581014 5492312 5088702 0 16259191 0.7 0
kio 0 s 570.2 10581014 5492312 5088702 0 16259243 0.7 0
kio 0 s 570.2 10581014 5492312 5088702 0 16259298 0.7 0
kio 0 s 570.1 10581014 5492312 5088702 0 16259353 0.7 0
kio 0 s 570.1 10581014 5492312 5088702 0 16259406 0.7 0
kio 0 s 570.1 10581014 5492312 5088702 0 16259461 0.7 0
kio 0 s 570.0 10581014 5492312
1. Try running onmode -F before each run (free shared memory
segments/blocks and general tidy up memory
structures).
2. What does
onstat -g stk <thread id>
give for the sqlexec (or sessions threads) and the kaio threads during
the freeze?
Repeat several times to see what functions it is getting stuck in.
3. What does onstat -g lmx and onstat -g con give during these times?
4/ How does onstat -g spi change on a bad run? Number of loops changes
a lot
somewhere on a bad run compared to a good run?
The threads must be waiting on something. Also seeing the stack for
each involved
thread would be useful.
David.
<david@smooth1.co.uk> wrote in message
news:1127598536.808774.302800@g14g2000cwa.googlegroups.com...
>
>
> 1. Try running onmode -F before each run (free shared memory
> segments/blocks and general tidy up memory
> structures).
Could do, but never do before a non-kaio run at it doesn't affcect them.
>
> 2. What does
>
> onstat -g stk <thread id>>
> give for the sqlexec (or sessions threads) and the kaio threads during
> the freeze?
>
> Repeat several times to see what functions it is getting stuck in.
>
> 3. What does onstat -g lmx and onstat -g con give during these times?
>
Just caught it during a bad run (almost 2 hours now on a run that at best
can take 22 minutes). Hope this is useful ...
$ onstat -g iov -r 1 | grep kio
kio 0 s 268.7 3476050 1884407 1591643 0 5785530 0.6 0
kio 0 s 268.7 3476050 1884407 1591643 0 5785588 0.6 0
kio 0 s 268.7 3476050 1884407 1591643 0 5785645 0.6 0
kio 0 s 268.6 3476050 1884407 1591643 0 5785704 0.6 0
kio 0 s 268.6 3476050 1884407 1591643 0 5785763 0.6 0
^C$ onstat -g sql
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:35:44 -- 2573312 Kbytes
Sess SQL Current Iso Lock SQL ISAM F.E.
Id Stmt type Database Lvl Mode ERR ERR Vers
Explain
21 CREATE INDEX caps CR Not Wait 0 0 9.03 Off
$ onstat -g stk 21
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:35:48 -- 2573312 Kbytes
Stack for thread: 21 aio vp 16
base: 0x0000000194407000
len: 69632
pc: 0x000000010090e344
tos: 0x0000000194417471state: sleeping
vp: 22
0x10090e344 oninit :: yield_processor_mvp + 0x678
sp=0x194417c70(0x10110a4b8, 0x1010dd8f0, 0x1010de1e8, 0x193014c48,
0x1943ec8b8, 0x193014bd0)
0x10091cdfc oninit :: iowork + 0xfc sp=0x194417da0 delta_sp=304(0x0,
0x1941dd850, 0x0, 0x1943ec850, 0x1941dd850, 0x9)
0x10090ead0 oninit :: startup + 0xfc sp=0x194417e50
delta_sp=176(0x10110a4b8, 0x1, 0x1010de1e8, 0x10110a4b8, 0x1010d8ae8, 0x2)
$ onstat -g stk 21
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:35:59 -- 2573312 Kbytes
Stack for thread: 21 aio vp 16
base: 0x0000000194407000
len: 69632
pc: 0x000000010090e344
tos: 0x0000000194417471state: sleeping
vp: 22
0x10090e344 oninit :: yield_processor_mvp + 0x678
sp=0x194417c70(0x10110a4b8, 0x1010dd8f0, 0x1010de1e8, 0x193014c48,
0x1943ec8b8, 0x193014bd0)
0x10091cdfc oninit :: iowork + 0xfc sp=0x194417da0 delta_sp=304(0x0,
0x1941dd850, 0x0, 0x1943ec850, 0x1941dd850, 0x9)
0x10090ead0 oninit :: startup + 0xfc sp=0x194417e50
delta_sp=176(0x10110a4b8, 0x1, 0x1010de1e8, 0x10110a4b8, 0x1010d8ae8, 0x2)
$ onstat -g stk 21
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:36:01 -- 2573312 Kbytes
Stack for thread: 21 aio vp 16
base: 0x0000000194407000
len: 69632
pc: 0x000000010090e344
tos: 0x0000000194417471state: sleeping
vp: 22
0x10090e344 oninit :: yield_processor_mvp + 0x678
sp=0x194417c70(0x10110a4b8, 0x1010dd8f0, 0x1010de1e8, 0x193014c48,
0x1943ec8b8, 0x193014bd0)
0x10091cdfc oninit :: iowork + 0xfc sp=0x194417da0 delta_sp=304(0x0,
0x1941dd850, 0x0, 0x1943ec850, 0x1941dd850, 0x9)
0x10090ead0 oninit :: startup + 0xfc sp=0x194417e50
delta_sp=176(0x10110a4b8, 0x1, 0x1010de1e8, 0x10110a4b8, 0x1010d8ae8, 0x2)
$ onstat -g stk 21
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:36:16 -- 2573312 Kbytes
Stack for thread: 21 aio vp 16
base: 0x0000000194407000
len: 69632
pc: 0x000000010090e344
tos: 0x0000000194417471state: sleeping
vp: 22
0x10090e344 oninit :: yield_processor_mvp + 0x678
sp=0x194417c70(0x10110a4b8, 0x1010dd8f0, 0x1010de1e8, 0x193014c48,
0x1943ec8b8, 0x193014bd0)
0x10091cdfc oninit :: iowork + 0xfc sp=0x194417da0 delta_sp=304(0x0,
0x1941dd850, 0x0, 0x1943ec850, 0x1941dd850, 0x9)
0x10090ead0 oninit :: startup + 0xfc sp=0x194417e50
delta_sp=176(0x10110a4b8, 0x1, 0x1010de1e8, 0x10110a4b8, 0x1010d8ae8, 0x2)
$ onstat -g lmx
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:36:51 -- 2573312 Kbytes
Locked mutexes:
mid addr name holder lkcnt waiter
waittime
Number of mutexes on VP free lists: 26
$ onstat -g lmx
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:36:54 -- 2573312 Kbytes
Locked mutexes:
mid addr name holder lkcnt waiter
waittime
Number of mutexes on VP free lists: 26
$ onstat -g lmx
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:36:55 -- 2573312 Kbytes
Locked mutexes:
mid addr name holder lkcnt waiter
waittime
Number of mutexes on VP free lists: 26
$ onstat -g con
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:37:01 -- 2573312 Kbytes
Conditions with waiters:
cid addr name waiter waittime
560 10a102408 bp_cond 115 213
$ onstat -g con
^[k
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:37:06 -- 2573312 Kbytes
Conditions with waiters:
cid addr name waiter waittime
560 10a102408 bp_cond 115 218
$ ^[k
ksh: ^[k: not found
$ onstat -g con
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:37:11 -- 2573312 Kbytes
Conditions with waiters:
cid addr name waiter waittime
560 10a102408 bp_cond 115 223
$ onstat -g con
IBM Informix Dynamic Server Version 10.00.FC3 -- On-Line -- Up
03:37:13 -- 2573312 Kbytes
Conditions with waiters:
cid addr name waiter waittime
560 10a102408 bp_cond 115 225
$ onstat -g iov -r 1 | grep kio
kio 0 s 266.8 3478334 1886691 1591643 0 5795567 0.6 0
kio 0 s 266.8 3478334 1886691 1591643 0 5795620 0.6 0
kio 0 s 266.7 3478334 1886691 1591643 0 5795673 0.6 0
kio 0 s 266.7 3478334 1886691 1591643 0 5795729 0.6 0
kio 0 i 266.7 3478334 1886691 1591643 0 5795784 0.6 0
kio 0 s 266.7 3478334 1886691 1591643 0 5795840 0.6 0
kio 0 s 266.7 3478334 1886691 1591643 0 5795895 0.6 0
kio 0 s 266.6 3478334 1886691 1591643 0 5795951 0.6 0
kio 0 s 266.6 3478334 1886691 1591643 0 5796008 0.6 0
kio 0 s 266.6 3478334 1886691 1591643 0 5796064 0.6 0
kio 0 s 266.6 3478334 1886691 1591643 0 5796119 0.6 0
"Neil Truby" <neil.truby@ardenta.com> wrote in message
news:3pliprFb403pU1@individual.net...
> IDS 10.0FC3 on Solaris 2.9
>
>>> This is a brand new server which we are preparing to migrate a v9 system
>>> to.
> Over the past few weeks I have run multiple dbimports from more-or-less
> the
> same data source - certainly the same schema - but the import time varies
> widely.
>
>>> The fastest I've done it is 9h, the slowest 14. This weekend's took 13h
> 10m. There's never anyhting else going on on it. In the UPDATE
> STATISICS
> phase the elapsed time was 3h: during the "fast" run it was just 1h. I
> had
> a look during this period: often onstat -p onstat -m onstat -D would show
> periods of perhaps 2 minutes with no activity at all: no bufreads or
> physreads, no checkpointing for 20 minutes (the checkpoint interval is 5
> mins), the relevant thread just "sleeping forever".
>
> We've done a lot of extra work on this, and have now built a test case
> with
> a smaller table - just repeatedly running drop index; add cluster index.
> It
> normally takes 25-30 minutes. About 1 in every 2 or 3 runs it takens
> significantly longer - anything from 50 to 120 mins. There seems to be no
> discernible pattern - if the instance has been bounced for example. During
> a
> slow run we see with an onstat -g iov - see below - absolutely nothing
> increasing for seconds on end except the wake-up count.
>
> But, if we disable kaio (this is Solaris, so it's substituting block for
> character devices), we get a consistent performance every time. We have
> repeated this 20 or more times with and without kaio so I am increasingly
> confident that this is a pattern, not coincidenece, and that the problem
> is
> with kaio on v10.
>
> Our next stage is to repeat the tests with 9.40FC7. I'll be raising a
> tech
> support call on Monday when they are open again. In the meantime if
> anyone
> has any observations I'd welcome them
>
> regards
> Neil
Although we haven't been able to do as much testing as I would have liked,
the problem does not seem to occur with 9.40FC7, in that the times for the
add clsuter index remain consistent whether or not kaio is enabled, unlike
v10 where the results with kaio enabled are highly variable, but consistent
with it disabled.
I conclude that our problem is with kaio and v10 at our installation.
"Neil Truby" <neil.truby@ardenta.com> wrote in message news:3polnvFbicgoU1@individual.net... > "Neil Truby" <neil.truby@ardenta.com> wrote in message > news:3pliprFb403pU1@individual.net... >> IDS 10.0FC3 on Solaris 2.9 > > I conclude that our problem is with kaio and v10 at our installation. IBM case 439918 refers. We have been sent some tuning suggestions for the actual test. I am a little wary of this since the problem is not really that the test (CREATE CLUSTER INDEX) runs slowly all the time, but that it is sporadically slow with kaio enabled *only*. But more testing and more information certainly can't hurt ... In the meantime we've disabled kaio on the databse servers and discovered another problem. With kaio disabled the backups are perhaps 75% longer (35 minutes instead of 20), but the *restore* runs 40 times more slowly (444 pages/second without kaio ; 10093 with kaio, that is 7.5 hours vs 20 mins for the restore). A separate case, 439945, refers here.
if it only occurs with KIO turned on, you may want to discussed KIO issues with OS vendor Neil Truby wrote: > "Neil Truby" <neil.truby@ardenta.com> wrote in message > news:3polnvFbicgoU1@individual.net... > > "Neil Truby" <neil.truby@ardenta.com> wrote in message > > news:3pliprFb403pU1@individual.net... > >> IDS 10.0FC3 on Solaris 2.9 > > > > I conclude that our problem is with kaio and v10 at our installation. > > IBM case 439918 refers. We have been sent some tuning suggestions for the > actual test. I am a little wary of this since the problem is not really > that the test (CREATE CLUSTER INDEX) runs slowly all the time, but that it > is sporadically slow with kaio enabled *only*. But more testing and more > information certainly can't hurt ... > > In the meantime we've disabled kaio on the databse servers and discovered > another problem. With kaio disabled the backups are perhaps 75% longer (35 > minutes instead of 20), but the *restore* runs 40 times more slowly (444 > pages/second without kaio ; 10093 with kaio, that is 7.5 hours vs 20 mins > for the restore). A separate case, 439945, refers here.
> Neil Truby wrote: >> "Neil Truby" <neil.truby@ardenta.com> wrote in message >> news:3polnvFbicgoU1@individual.net... >> > "Neil Truby" <neil.truby@ardenta.com> wrote in message >> > news:3pliprFb403pU1@individual.net... >> >> IDS 10.0FC3 on Solaris 2.9 >> > >> > I conclude that our problem is with kaio and v10 at our installation. >> >> IBM case 439918 refers. We have been sent some tuning suggestions for >> the >> actual test. I am a little wary of this since the problem is not really >> that the test (CREATE CLUSTER INDEX) runs slowly all the time, but that >> it >> is sporadically slow with kaio enabled *only*. But more testing and more >> information certainly can't hurt ... >> >> In the meantime we've disabled kaio on the databse servers and discovered >> another problem. With kaio disabled the backups are perhaps 75% longer >> (35 >> minutes instead of 20), but the *restore* runs 40 times more slowly (444 >> pages/second without kaio ; 10093 with kaio, that is 7.5 hours vs 20 mins >> for the restore). A separate case, 439945, refers here. "scottishpoet" <dryburghj@yahoo.com> wrote in message news:1127748148.419479.255360@g43g2000cwa.googlegroups.com... > if it only occurs with KIO turned on, you may want to discussed KIO > issues with OS vendor Hmm. Interesting point. If it's confirmed as a problem with IDS v10 and kaio (it doesn't seem to be a problem with 9.40FC7), how am I - the poor, ignorant customer - to know if I should get help from the database vendor or the OS vendor?
therein lines the dilema maybe reprot the problem to both organisations Neil Truby wrote: > > Neil Truby wrote: > >> "Neil Truby" <neil.truby@ardenta.com> wrote in message > >> news:3polnvFbicgoU1@individual.net... > >> > "Neil Truby" <neil.truby@ardenta.com> wrote in message > >> > news:3pliprFb403pU1@individual.net... > >> >> IDS 10.0FC3 on Solaris 2.9 > >> > > >> > I conclude that our problem is with kaio and v10 at our installation. > >> > >> IBM case 439918 refers. We have been sent some tuning suggestions for > >> the > >> actual test. I am a little wary of this since the problem is not really > >> that the test (CREATE CLUSTER INDEX) runs slowly all the time, but that > >> it > >> is sporadically slow with kaio enabled *only*. But more testing and more > >> information certainly can't hurt ... > >> > >> In the meantime we've disabled kaio on the databse servers and discovered > >> another problem. With kaio disabled the backups are perhaps 75% longer > >> (35 > >> minutes instead of 20), but the *restore* runs 40 times more slowly (444 > >> pages/second without kaio ; 10093 with kaio, that is 7.5 hours vs 20 mins > >> for the restore). A separate case, 439945, refers here. > > > "scottishpoet" <dryburghj@yahoo.com> wrote in message > news:1127748148.419479.255360@g43g2000cwa.googlegroups.com... > > if it only occurs with KIO turned on, you may want to discussed KIO > > issues with OS vendor > > Hmm. Interesting point. If it's confirmed as a problem with IDS v10 and > kaio (it doesn't seem to be a problem with 9.40FC7), how am I - the poor, > ignorant customer - to know if I should get help from the database vendor or > the OS vendor?
Neil Truby wrote:
>>Neil Truby wrote:
>>
>>>"Neil Truby" <neil.truby@ardenta.com> wrote in message
>>>news:3polnvFbicgoU1@individual.net...
>>>
>>>>"Neil Truby" <neil.truby@ardenta.com> wrote in message
>>>>news:3pliprFb403pU1@individual.net...
>>>>
>>>>>IDS 10.0FC3 on Solaris 2.9
>>>>
>>>>I conclude that our problem is with kaio and v10 at our installation.
>>>
>>>IBM case 439918 refers. We have been sent some tuning suggestions for
>>>the
>>>actual test. I am a little wary of this since the problem is not really
>>>that the test (CREATE CLUSTER INDEX) runs slowly all the time, but that
>>>it
>>>is sporadically slow with kaio enabled *only*. But more testing and more
>>>information certainly can't hurt ...
>>>
>>>In the meantime we've disabled kaio on the databse servers and discovered
>>>another problem. With kaio disabled the backups are perhaps 75% longer
>>>(35
>>>minutes instead of 20), but the *restore* runs 40 times more slowly (444
>>>pages/second without kaio ; 10093 with kaio, that is 7.5 hours vs 20 mins
>>>for the restore). A separate case, 439945, refers here.
>
>
>
> "scottishpoet" <dryburghj@yahoo.com> wrote in message
> news:1127748148.419479.255360@g43g2000cwa.googlegroups.com...
>
>>if it only occurs with KIO turned on, you may want to discussed KIO
>>issues with OS vendor
>
>
> Hmm. Interesting point. If it's confirmed as a problem with IDS v10 and
> kaio (it doesn't seem to be a problem with 9.40FC7), how am I - the poor,
> ignorant customer - to know if I should get help from the database vendor or
> the OS vendor?
>
>
What have you got for :
1. NOAGE?
i.e. when you have a "slow" run, what does ps -aldef show for the
oninit's priorities?
2. How many AIOVPs have you got?
i.e. Have you got enough when KAIO is disabled.
if there is no performance problem (ie its not eratic) then i don't
believechanging the AIO VPs will have any effect. Don't really want to
make AIO eratic!
TBP wrote:
> Neil Truby wrote:
> >>Neil Truby wrote:
> >>
> >>>"Neil Truby" <neil.truby@ardenta.com> wrote in message
> >>>news:3polnvFbicgoU1@individual.net...
> >>>
> >>>>"Neil Truby" <neil.truby@ardenta.com> wrote in message
> >>>>news:3pliprFb403pU1@individual.net...
> >>>>
> >>>>>IDS 10.0FC3 on Solaris 2.9
> >>>>
> >>>>I conclude that our problem is with kaio and v10 at our installation.
> >>>
> >>>IBM case 439918 refers. We have been sent some tuning suggestions for
> >>>the
> >>>actual test. I am a little wary of this since the problem is not really
> >>>that the test (CREATE CLUSTER INDEX) runs slowly all the time, but that
> >>>it
> >>>is sporadically slow with kaio enabled *only*. But more testing and more
> >>>information certainly can't hurt ...
> >>>
> >>>In the meantime we've disabled kaio on the databse servers and discovered
> >>>another problem. With kaio disabled the backups are perhaps 75% longer
> >>>(35
> >>>minutes instead of 20), but the *restore* runs 40 times more slowly (444
> >>>pages/second without kaio ; 10093 with kaio, that is 7.5 hours vs 20 mins
> >>>for the restore). A separate case, 439945, refers here.
> >
> >
> >
> > "scottishpoet" <dryburghj@yahoo.com> wrote in message
> > news:1127748148.419479.255360@g43g2000cwa.googlegroups.com...
> >
> >>if it only occurs with KIO turned on, you may want to discussed KIO
> >>issues with OS vendor
> >
> >
> > Hmm. Interesting point. If it's confirmed as a problem with IDS v10 and
> > kaio (it doesn't seem to be a problem with 9.40FC7), how am I - the poor,
> > ignorant customer - to know if I should get help from the database vendor or
> > the OS vendor?
> >
> >
>
> What have you got for :
>
> 1. NOAGE?
>
> i.e. when you have a "slow" run, what does ps -aldef show for the
> oninit's priorities?
>
> 2. How many AIOVPs have you got?
> i.e. Have you got enough when KAIO is disabled.
oops should have read Neil's post more carefully!
"TBP" <TheBigPotato@NotHere.Co.Uk> wrote in message
news:XOVZe.18852$wm3.1269@newsfe6-win.ntli.net...
> Neil Truby wrote:
>>>Neil Truby wrote:
>>>
>>>>"Neil Truby" <neil.truby@ardenta.com> wrote in message
>>>>news:3polnvFbicgoU1@individual.net...
>>>>
>>>>>"Neil Truby" <neil.truby@ardenta.com> wrote in message
>>>>>news:3pliprFb403pU1@individual.net...
>>>>>
>>>>>>IDS 10.0FC3 on Solaris 2.9
>>>>>
>>>>>I conclude that our problem is with kaio and v10 at our installation.
>>>>
>>>>IBM case 439918 refers. We have been sent some tuning suggestions for
>>>>the
>>>>actual test. I am a little wary of this since the problem is not really
>>>>that the test (CREATE CLUSTER INDEX) runs slowly all the time, but that
>>>>it
>>>>is sporadically slow with kaio enabled *only*. But more testing and
>>>>more
>>>>information certainly can't hurt ...
>>>>
>>>>In the meantime we've disabled kaio on the databse servers and
>>>>discovered
>>>>another problem. With kaio disabled the backups are perhaps 75% longer
>>>>(35
>>>>minutes instead of 20), but the *restore* runs 40 times more slowly (444
>>>>pages/second without kaio ; 10093 with kaio, that is 7.5 hours vs 20
>>>>mins
>>>>for the restore). A separate case, 439945, refers here.
>>
>>
>>
>> "scottishpoet" <dryburghj@yahoo.com> wrote in message
>> news:1127748148.419479.255360@g43g2000cwa.googlegroups.com...
>>
>>>if it only occurs with KIO turned on, you may want to discussed KIO
>>>issues with OS vendor
>>
>>
>> Hmm. Interesting point. If it's confirmed as a problem with IDS v10 and
>> kaio (it doesn't seem to be a problem with 9.40FC7), how am I - the poor,
>> ignorant customer - to know if I should get help from the database vendor
>> or
>> the OS vendor?
>>
>>
>
> What have you got for :
>
> 1. NOAGE?
1
> i.e. when you have a "slow" run, what does ps -aldef show for the
> oninit's priorities?
See below - all oninits have the same priorities although the problem is
happening right now.
> 2. How many AIOVPs have you got?
Sometimes 4, sometimes 32, sometimes 65.
> i.e. Have you got enough when KAIO is disabled.
Even ther 32nd is quite heavily used so I could increase them some more.
$ ps -adelf | grep init
8 S root 1 0 0 40 20 ? 163 ? Sep 24 ?
0:02 /etc/init -
8 S root 27266 27257 0 40 20 ? 811670 ? 16:18:47 ?
0:00 oninit -v
8 S root 27293 27257 0 40 20 ? 811670 ? 16:18:50 ?
0:00 oninit -v
8 S root 10981 10963 0 40 20 ? 76310 ? 23:43:13 ?
0:00 oninit -v
8 S root 27285 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 27296 27257 0 40 20 ? 811670 ? 16:18:51 ?
0:00 oninit -v
8 S root 19941 19937 0 40 20 ? 76310 ? 07:57:52 ?
0:00 oninit -v
8 S root 19945 19937 0 40 20 ? 76310 ? 07:57:55 ?
0:00 oninit -v
8 S root 19939 19937 0 40 20 ? 76310 ? 07:57:50 ?
0:00 oninit -v
8 S root 27261 27257 0 40 20 ? 811670 ? 16:18:45 ?
0:00 oninit -v
8 S root 27264 27257 0 40 20 ? 811670 ? 16:18:47 ?
0:00 oninit -v
8 S informix 19936 1 0 40 10 ? 76310 ? 07:57:49 ?
0:03 oninit -v
8 S root 27269 27257 0 40 20 ? 811670 ? 16:18:47 ?
0:00 oninit -v
8 S informix 19938 19937 0 40 10 ? 76310 ? 07:57:50 ?
0:06 oninit -v
8 S root 19943 19937 0 40 20 ? 76310 ? 07:57:54 ?
0:00 oninit -v
8 S root 27276 27257 0 40 20 ? 811670 ? 16:18:48 ?
0:00 oninit -v
8 S root 27284 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 27257 27256 0 40 20 ? 811670 ? 16:18:43 ?
0:01 oninit -v
8 S root 27275 27257 0 40 20 ? 811670 ? 16:18:48 ?
0:00 oninit -v
8 S root 27277 27257 0 40 20 ? 811670 ? 16:18:48 ?
0:00 oninit -v
8 S root 27286 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 27280 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 27265 27257 0 40 20 ? 811670 ? 16:18:47 ?
0:00 oninit -v
8 S root 27273 27257 0 40 20 ? 811670 ? 16:18:48 ?
0:00 oninit -v
8 S root 19946 19937 0 40 20 ? 76310 ? 07:57:55 ?
0:00 oninit -v
8 S root 19940 19937 0 40 20 ? 76310 ? 07:57:51 ?
0:00 oninit -v
8 S root 27279 27257 0 40 20 ? 811670 ? 16:18:48 ?
0:00 oninit -v
8 O informix 27256 1 2 47 10 ? 811687 16:18:30 ?
10:24 oninit -v
8 S root 27274 27257 0 40 20 ? 811670 ? 16:18:48 ?
0:00 oninit -v
8 S root 27272 27257 0 40 20 ? 811670 ? 16:18:48 ?
0:00 oninit -v
8 S root 27282 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 19944 19937 0 40 20 ? 76310 ? 07:57:54 ?
0:00 oninit -v
8 S root 27288 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 27268 27257 0 40 20 ? 811670 ? 16:18:47 ?
0:00 oninit -v
8 S root 19942 19937 0 40 20 ? 76316 ? 07:57:53 ?
0:00 oninit -v
8 S root 27262 27257 0 40 20 ? 811676 ? 16:18:46 ?
0:00 oninit -v
8 S root 19937 19936 0 40 20 ? 76310 ? 07:57:50 ?
0:00 oninit -v
8 S root 27283 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 27278 27257 0 40 20 ? 811670 ? 16:18:48 ?
0:00 oninit -v
8 S root 27281 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 27295 27257 0 40 20 ? 811670 ? 16:18:51 ?
0:00 oninit -v
8 S root 27291 27257 0 40 20 ? 811670 ? 16:18:50 ?
0:00 oninit -v
8 S root 27270 27257 0 40 20 ? 811670 ? 16:18:47 ?
0:00 oninit -v
8 S root 27290 27257 0 40 20 ? 811670 ? 16:18:50 ?
0:00 oninit -v
8 S root 27259 27257 0 40 20 ? 811670 ? 16:18:43 ?
0:00 oninit -v
8 S root 27289 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S informix 27258 27257 2 40 10 ? 811687 ? 16:18:43 ?
20:17 oninit -v
8 S root 27260 27257 0 40 20 ? 811670 ? 16:18:44 ?
0:00 oninit -v
8 S root 27294 27257 0 40 20 ? 811670 ? 16:18:50 ?
0:00 oninit -v
8 S root 27287 27257 0 40 20 ? 811670 ? 16:18:49 ?
0:00 oninit -v
8 S root 27292 27257 0 40 20 ? 811670 ? 16:18:50 ?
0:00 oninit -v
8 S root 10963 10962 0 40 20 ? 76310 ? 23:43:08 ?
0:00 oninit -v
8 S informix 10039 10013 0 50 20 ? 126 ? 19:56:54 pts/30:00 grep init
8 S root 10967 10963 0 40 20 ? 76316 ? 23:43:11 ?
0:00 oninit -v
8 S root 10965 10963 0 40 20 ? 76310 ? 23:43:09 ?@
1.ok... what does condition bp_cond mean?
does anyone have an internals manuals?
2. onstat -g stk is not for the thread I wanted
a) onstat -g ses <session id> for the session what is the name of the
threads that are running.
b) onstat -g stk <thread id> for each sessions thread and
onstat -g stk for each kaio thread.
3. How does onstat -g spi change over 10 seconds during a bad run.
Either the engine is waiting internally so internal mutexs/ spin locks
will show this. Also
strange function names on the stack might show something.