large number of shared memory segments
Posted in 1999
Topics: Server Administration, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
IDS 7.24.UC5, HP9000 running HP-UX 11.0
We had an interesting situation last night where the server
un-surprisingly ) appeared to not accept any user requests from either
soctcp or shmem connections even though the engine was up and running and
was issuing checkpoints. We couldn't manually issue checkpoints either, the
terminal would freeze, and no onmode commands would work. Having ensured
that no processes were in critical sections, I eventually had to kill the
master oninit process using kill -11 ( -9 is a no no I know ).
The onstat commands appeared to be working fine providing whatever stats we
wanted, one of which was the onstat -g seg which gave us about 32 additional
V segments each 8MB in size. ( at this point alarm bells started to ring ).
My problem, however, was that I didn't know of a way to trace back what
process ( or processes ) caused so many segments to be added in such a short
time - over a period of about 1 hour.
Can any of you guys/galls point me in a few directions that I may need to
investigate so that if/when it happens again, I can narrow the search down
to a particular user/process AQAP.
I have already upped the SHMVIRTSIZE to 750MB ( from 500MB) and SHMADD to
64MB ( from 8MB) to reduce the number of segments the next time it needs to
add them, and also implemented a script to raise the alarm if > 3 segments
get allocated.
I must point out that this happened when there were relatively few ( ~300 )
users on the system not running anything massive.
I wait with baited breath.
TIA
Sean
Sean Kelsey wrote:
>
> IDS 7.24.UC5, HP9000 running HP-UX 11.0
>
> We had an interesting situation last night where the server
> un-surprisingly ) appeared to not accept any user requests from either
> soctcp or shmem connections even though the engine was up and running and
> was issuing checkpoints. We couldn't manually issue checkpoints either, the
> terminal would freeze, and no onmode commands would work. Having ensured
> that no processes were in critical sections, I eventually had to kill the
> master oninit process using kill -11 ( -9 is a no no I know ).
>
> The onstat commands appeared to be working fine providing whatever stats we
> wanted, one of which was the onstat -g seg which gave us about 32 additional
> V segments each 8MB in size. ( at this point alarm bells started to ring ).
>
> My problem, however, was that I didn't know of a way to trace back what
> process ( or processes ) caused so many segments to be added in such a short
> time - over a period of about 1 hour.
>
> Can any of you guys/galls point me in a few directions that I may need to
> investigate so that if/when it happens again, I can narrow the search down
> to a particular user/process AQAP.
>
Your online log will give the time when the segments were added. I have
a 'sniffer' script that I run in cron every five minutes to give me a
snapshot of what 4ge and gnt (Lawson) processes are running. It's not
perfect, but it gets me in the area.
> I have already upped the SHMVIRTSIZE to 750MB ( from 500MB) and SHMADD to
> 64MB ( from 8MB) to reduce the number of segments the next time it needs to
> add them, and also implemented a script to raise the alarm if > 3 segments
> get allocated.
>
Anyone dropping / recreating indices or running PDQ?
John Carlson
Informix DBA
WHSmith USA
Sean,
We were running into the same problem. Found that one of the reports that
had recently been created wasn't tested as well as it should have been. We
created a cron script to monitor the shared memory segs. and display a mail
message to the informix log in when they started increasing. Once the mail
message appeared we ran the onstat commands to determine the culprit.
HTH,
Brian
Sean Kelsey <chilliinc@hotmail.com> wrote in article
<36cc0478.0@145.227.194.253>...
> IDS 7.24.UC5, HP9000 running HP-UX 11.0
>
> We had an interesting situation last night where the server
> un-surprisingly ) appeared to not accept any user requests from either
> soctcp or shmem connections even though the engine was up and running and
> was issuing checkpoints. We couldn't manually issue checkpoints either,
the
> terminal would freeze, and no onmode commands would work. Having ensured
> that no processes were in critical sections, I eventually had to kill the
> master oninit process using kill -11 ( -9 is a no no I know ).
>
> The onstat commands appeared to be working fine providing whatever stats
we
> wanted, one of which was the onstat -g seg which gave us about 32
additional
> V segments each 8MB in size. ( at this point alarm bells started to ring
).
>
> My problem, however, was that I didn't know of a way to trace back what
> process ( or processes ) caused so many segments to be added in such a
short
> time - over a period of about 1 hour.
>
> Can any of you guys/galls point me in a few directions that I may need to
> investigate so that if/when it happens again, I can narrow the search
down
> to a particular user/process AQAP.
>
> I have already upped the SHMVIRTSIZE to 750MB ( from 500MB) and SHMADD to
> 64MB ( from 8MB) to reduce the number of segments the next time it needs
to
> add them, and also implemented a script to raise the alarm if > 3
segments
> get allocated.
>
> I must point out that this happened when there were relatively few ( ~300
)
> users on the system not running anything massive.
>
> I wait with baited breath.
>
> TIA
>
> Sean
>
>
>
>
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g