Re: trouble
Posted in 2000
Thanks for the many helpful suggestions re: our troubles. The shared memory
stuff I think I have sorted, at least re: what to do if it happens again
(use icprm or icps).
Re: the number of "oninit -s" processes running, I'm counting approximately
120. Truth be known, this is not something I've looked at regularly, but
it seems more than normal. Dumb question -- oninit -s from the command line
means 'Initialize shared memory, leave in quiescent mode.' Does this mean I
have 120 processes trying to take Informix to quiescent mode???? That
certainly doesn't sound good. Or am I completely misinterpreting what those
processes are doing? If I am misinterpreting the results of the grep, and
seeing all those processes is normal, that would be extremely comforting to
know!
I do onstat -g segs on a pretty regular basis, and our problems started when
I made a relatively minor change in SHMVIRTSIZE from 220000 to 250000. At
the time, Informix was taking up 4 additional segments (in addition to the
first two).
Right at the moment, onstat -g seg looks like this:
Segment Summary:
(resident segments are not locked)
id key addr size ovhd class blkused blkfree
29 1382303745 30000000 22405120 1168 R 2731 4
30 1382303746 31580000 122880000 2468 V 5273 9727
Total: - - 145285120 - - 8004 9731
As you can see, I pulled the SHMVIRTSIZE parameter down considerably, and
for whatever reason it seems to be happy at that figure for now, probably
because we have cancelled some automated jobs that usually run overnight.
I am getting some warnings now and then that the size of my physical logs is
too low (5000) and that it should be 20 times maximum user threads. Am I
correct in assuming that onstat -u | wc -l will give me an approximation of
active user threads at any given time? By that yardstick, our physical log
is currently 22 times active user threads, but if those automated jobs were
running, we'd be seeing at least twice as many threads, meaning that our
physical log size is indeed inadequate.
We are currently scheduled to do some more work on things over the weekend,
and one other thing I'm planning to do is set up a DBSPACETEMP. We
originally set this at the default (root), and I'm wondering if some of our
problems could be related to temporary tables filling up the rootdbs, which
is currently sized at 10000 2K pages with about 3000 pages free. We have
already moved the physical logs to a separate dbspace, and if we tear down
and rebuild again will make this ever so much larger (am I correct in
assuming that it's important to have one large contiguous chunk for physical
logs?).
Thanks again,
Mark McDonough
Knowledge Systems Inc.
Mark McDonough wrote:
> Hi folks,
>
> We've been having some strange problems this week. We run an
> ever-growing number of small databases on Informix 7.2n, running on IRIX
> 6.n.
>
> For ages, life has been good, but suddenly we're seeing the following in
> the online log:
>
> 15:08:02 semget: errno = 28
> 15:08:02 create_vp: cannot allocate memory>
> When I do a ps -ef | grep oninit, I'm seeing tons of oninit -s processes
> out there.
>
> I'm not a trained dba but the famous "guy who read the manual."
>
> We have had a couple of nasty crashes in the last 48 hours, one of which
> ultimately involved going to backup data and reinitializing Informix.
> When we attempted to bring Informix online again after the crashes, the
> error messages were as follows:
>
> shmget: [EEXIST][17]: key 52644801: shared memory already exists
> 09:18:15 mt_shm_init: can't create resident segment>
> My interpretation of this was that something in the engine believed that
> it already had shared memory, even though it did not, and therefore
> would no create shared memory. This problem was fixed only by
> reinitializing.
>
> What else should I be looking at?
>
> Mark McDonough
>
> --
--
=> Experience EXPOVenture at http://www.expoventure.com <=
"What a virtual tradeshow should be."