IDS 9.20 with a single cpu vp
Posted in 2000
Neil Truby reported that on IDS 9.20 configured with a single CPU VP, one heavy-duty session froze all other sessions for 5-20 seconds, whereas the same test on 7.24 (or on 9.20 with a second CPU VP added) showed sub-second waits. Madison Pruet asked for diagnostics while the hang occurs (onstat -g ath, -g lmx, repeated -g act) plus stack traces via 'kill -7' on the CPU VP or a debugger, warning this can crash some platforms, and discussed why onstat -g stk can't reliably capture a running thread's stack. Others urged escalating the support case above priority 4. No root cause or fix is recorded; the practical workaround was adding a second CPU VP.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Server Administration, Versions, Editions & End-of-Life
Continuing my story about the whole of an IDS 9.20 instance being frozen out of 10-20 seconds on end by a single heavy-duty process: I have repeated my test (heavy duty program running in one session; dbaccess->table info on the other) on a v7.24 instance. I can confirm that the symptom we are experiencing on 9.20 (second session frozen out for a long period when only one cpu vp is configured) does not occur on> the v7.24 one. The wait times are sub-second on v7.24 (and on v9.20 when another cpu vp is added), but anything from 5-20 seconds on 9.20 with a single cpu vp. Perhaps someone with detailed knowledge can explain why IDS9.20 behaves so differently, and disastrously so, to IDS 7.24? thanks Neil
Neil,
We'll probably need a bit of your help to isolate the problem. I don't know if
you can easily reproduce this in a controlled situation or not, but here goes.
First of all try to get the following onstats when the problem is occuring.....
1) onstat -g ath
2) onstat -g lmx
3) several onstat -g act.
Now for the fun part....
We need to get several stack traces when the hang is happening. This will give
some idea as to what the thread is doing that is hanging the instance. On most
platforms you can do this by issueing a "kill -7 <pid_of_cpuvp>. BUT WATCH
OUT... On some platforms this will cause a crash. I know that DEC will crash.
I'm not sure if HP will give a good stack or not. (Of course, if we get a
crash, then we might be able to get the stack as well... ;-) but let's try to
not be so drastic.)
Another alternative would be to attach with a debugger. I don't know if you are
comfortable with gdb, dbx, and/or adb. But before you attempt any of these, you
should contact tech support, explain the problem, and see if they want you to
persue with the stacks or not.
Neil Truby wrote:
> Continuing my story about the whole of an IDS 9.20 instance being frozen out
> of 10-20 seconds on end by a single heavy-duty process:
>
> I have repeated my test (heavy duty program running in one session;
> dbaccess->table info on the other) on a v7.24 instance. I can confirm that
> the symptom we are experiencing on 9.20 (second session frozen out for a
> long period when only one cpu vp is configured) does not occur on> the v7.24
> one.
>
> The wait times are sub-second on v7.24 (and on v9.20 when another cpu vp is
> added), but anything from 5-20 seconds on 9.20 with a single cpu vp.
>
> Perhaps someone with detailed knowledge can explain why IDS9.20 behaves so
> differently, and disastrously so, to IDS 7.24?
>
> thanks
> Neil
>> But before you attempt any of these, you should contact tech support, explain the problem, and see if they want you to persue with the stacks or not. This leads me on to a supplementary question! I've raised the matter with Tech Support. The problem's been assigned a Priority 4 - the lowest priority a problem can be given in the UK. I've debated with myself whether this is valid:- the problem has caused me a great deal of grief in production, but now I know about it I can assign an extra cpu vp. Do you and others feel it's worth getting to the bottom of? Anyway, I can collect the stats if you feel it's worthwhile. Thanks for the interest. Neil
Madison Pruet wrote in message <399F4FA2.75F504DB@home.com>...
>Neil,
>
>We need to get several stack traces when the hang is happening. This will
give
>some idea as to what the thread is doing that is hanging the instance. On
most
>platforms you can do this by issueing a "kill -7 <pid_of_cpuvp>. BUT WATCH
>OUT... On some platforms this will cause a crash. I know that DEC will
crash.
What about onstat -g stk <thread-id> ?
Neil Truby wrote in message <8no0ko$6al$1@lyonesse.netcom.net.uk>... >>> But before you attempt any of these, you should contact tech support, >explain the problem, and see if they want you to persue with the stacks or >not. > >This leads me on to a supplementary question! I've raised the matter with >Tech Support. The problem's been assigned a Priority 4 - the lowest >priority a problem can be given in the UK. I've debated with myself whether >this is valid:- the problem has caused me a great deal of grief in >production, but now I know about it I can assign an extra cpu vp. Do you >and others feel it's worth getting to the bottom of? > Yes! Raise it to priority 2. You control the priority of the case. >Anyway, I can collect the stats if you feel it's worthwhile. > Yep! >Thanks for the interest. > >Neil > > > >
You have the right to insist that the problem be esclated. I would prefer for the root of this problem to be analyzed, understood, and corrected. This is true with any problem. If as you indicate the problem is causing you grief, then I would contact tech support and request esclation. However, as I indicated, getting to the root problem may require a bit of help from you. Neil Truby wrote: > >> But before you attempt any of these, you should contact tech support, > explain the problem, and see if they want you to persue with the stacks or > not. > > This leads me on to a supplementary question! I've raised the matter with > Tech Support. The problem's been assigned a Priority 4 - the lowest > priority a problem can be given in the UK. I've debated with myself whether > this is valid:- the problem has caused me a great deal of grief in > production, but now I know about it I can assign an extra cpu vp. Do you > and others feel it's worth getting to the bottom of? > > Anyway, I can collect the stats if you feel it's worthwhile. > > Thanks for the interest. > > Neil
It is rather difficult to get the stack for a running thread.
smooth1 wrote:
> Madison Pruet wrote in message <399F4FA2.75F504DB@home.com>...
> >Neil,
> >
>
> >We need to get several stack traces when the hang is happening. This will
> give
> >some idea as to what the thread is doing that is hanging the instance. On
> most
> >platforms you can do this by issueing a "kill -7 <pid_of_cpuvp>. BUT WATCH
> >OUT... On some platforms this will cause a crash. I know that DEC will
> crash.
>
> What about onstat -g stk <thread-id> ?
I guess a little more explaination is in order....
The only way that you can create a stack trace is via the return address from a
call function. This may or may not yet be in the physical stack memory. For
instance, with Solaris, there is a so called circular register stack which
remains within the cpu register set. This regiseter stack is not actually moved
to the memory portion of the stack unless there is an overflow of the cpu
register stack memory. Otherwise, the register stack remains only in the
processor registers. Even on those platforms which utilize the stack memory
directly, the 'current' function will not appear until that function makes a call
to another function. This is because it's the return address that appears in the
stack and identifies the functions within the stack.. Also, the currently
running thread will be constantly changing it's stack, especially the memory near
the top of the stack. (Bottom on those platforms whose stack grows down.)
Because of all of this, it becomes rather difficult to get a stack trace of the
currently running function except by the "kill -7" technique or by attaching with
some form of a debugger.
Madison Pruet wrote:
> It is rather difficult to get the stack for a running thread.
>
> smooth1 wrote:
>
> > Madison Pruet wrote in message <399F4FA2.75F504DB@home.com>...
> > >Neil,
> > >
> >
> > >We need to get several stack traces when the hang is happening. This will
> > give
> > >some idea as to what the thread is doing that is hanging the instance. On
> > most
> > >platforms you can do this by issueing a "kill -7 <pid_of_cpuvp>. BUT WATCH
> > >OUT... On some platforms this will cause a crash. I know that DEC will
> > crash.
> >
> > What about onstat -g stk <thread-id> ?
Madison Pruet wrote in message <39A02422.21A0A67D@home.com>...
>It is rather difficult to get the stack for a running thread.
>
>smooth1 wrote:
>
True, I was just thinking of using a supported method
(onstat -g stk) rather than an undocumented one
(kill -7 of cpu vp)!
Especially as the next question is which cpu vp? By the time you
issue the kill, the thread might have migrated to a different
CPU VP!
>> Madison Pruet wrote in message <399F4FA2.75F504DB@home.com>...
>> >Neil,
>> >
>>
>> >We need to get several stack traces when the hang is happening. This
will
>> give
>> >some idea as to what the thread is doing that is hanging the instance.
On
>> most
>> >platforms you can do this by issueing a "kill -7 <pid_of_cpuvp>. BUT
WATCH
>> >OUT... On some platforms this will cause a crash. I know that DEC will
>> crash.
>>
>> What about onstat -g stk <thread-id> ?
>
Valid point. But then this thread is "with a single cpu vp" ;-)
smooth1 wrote:
> Madison Pruet wrote in message <39A02422.21A0A67D@home.com>...
> >It is rather difficult to get the stack for a running thread.
> >
> >smooth1 wrote:
> >
>
> True, I was just thinking of using a supported method
> (onstat -g stk) rather than an undocumented one
> (kill -7 of cpu vp)!
>
> Especially as the next question is which cpu vp? By the time you
> issue the kill, the thread might have migrated to a different
> CPU VP!
>
> >> Madison Pruet wrote in message <399F4FA2.75F504DB@home.com>...
> >> >Neil,
> >> >
> >>
> >> >We need to get several stack traces when the hang is happening. This
> will
> >> give
> >> >some idea as to what the thread is doing that is hanging the instance.
> On
> >> most
> >> >platforms you can do this by issueing a "kill -7 <pid_of_cpuvp>. BUT
> WATCH
> >> >OUT... On some platforms this will cause a crash. I know that DEC will
> >> crash.
> >>
> >> What about onstat -g stk <thread-id> ?
> >
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g