How do onit processes work?
Posted in 2004
A sysadmin on IDS 7.23.UC6/HP-UX 10.20 saw two disjoint sets of oninit processes (12 from September, 33 from October 8), with onstat only showing the newer ones, plus creeping non-shared memory use, and asked whether oninits respawn by themselves. Respondents said Informix never restarts itself and suggested either a second instance or someone silently restarting after a hang; they also advised checking the online log, ipcs/ps parentage, and routine monitoring. The poster found the answer: there were separate test and production instances, and only production had been restarted. He concluded there was a slow memory leak and planned periodic restarts; others urged upgrading the old IDS/HP-UX versions.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Networking & sqlhosts Configuration
Hi,
I've been searching in the archive and can't quite find the detail I'm
interested in.
The situation is that we've an Informix database running version
7.23.UC6 on HPUX 10.20. To my knowledge this was started on Sept 17th
and by the middle of last week was using a fair amount of memory
(ordinary memory not shared memory). Then on Oct 8 around 400mb of
memory got freed up and some new oninit's got fired up.
There's 12 oninit processes running from Sept 17th and 33 that were
started last week. Onstat now seems to say the database was started last
week and only lists virtual processors with PID's that were started last
week. It looks to me like these processes are totally detached from each
other.
And yet, apparently no one restarted the database, we had no reports of
outages etc.
How does it work? Will oninits periodically respawn? For that matter is
it normal for memory usage to creep up? The "old" onit's look active to
me, they've still got network connections, files open etc (using lsof).
Would it be a good idea to periodically restart informix to free up this
memory?
Ian
On Tue, 12 Oct 2004 05:45:38 -0400, Ian Spare wrote:
Is it possible that someone started second instance, managing independent disk
space/databases/etc., on the 8th? This would explain this somewhat, though if
your INFORMIXSERVER setting *should* point to the original instance and onstat
-g glo shows the new oninits, it is more likely that something happened to
hang the original oninit processes and someone destroyed their shared memory
segments and restarted the engine without killing the original processes.
Informix will NOT automatically restart itself. If noone owns up to having
restarted it, it may be possible that someone programmed your ALARMPROGRAM
script/task (or some other periodic monitoring script) to detect the engine
failure, clean up, and restart the instance. I've seen Oracle installations
set up this way (mostly by Oracle consultants who don't seem to trust their
own product to remain online ;-} ). In any case, unless this is an
intentional second instance with independent disks (then there would be 2 sets
of entries in sqlhosts for this machine and a second ONCONFIG file as evidence
to look for), I would bring both instances down (onmode -ky) clean up shared
memory (ipcrm) and finally kill any oninit processes that are left (standard
procedure is to do kill -TERM followed by kill -PIPE then kill -KILL if the
first two don't do the job).
Art S. Kagel
> Hi,
>
> I've been searching in the archive and can't quite find the detail I'm
> interested in.
>
> The situation is that we've an Informix database running version 7.23.UC6 on
> HPUX 10.20. To my knowledge this was started on Sept 17th and by the middle
> of last week was using a fair amount of memory (ordinary memory not shared
> memory). Then on Oct 8 around 400mb of memory got freed up and some new
> oninit's got fired up.
>
> There's 12 oninit processes running from Sept 17th and 33 that were started
> last week. Onstat now seems to say the database was started last week and
> only lists virtual processors with PID's that were started last week. It
> looks to me like these processes are totally detached from each other.
>
> And yet, apparently no one restarted the database, we had no reports of
> outages etc.
>
> How does it work? Will oninits periodically respawn? For that matter is it
> normal for memory usage to creep up? The "old" onit's look active to me,
> they've still got network connections, files open etc (using lsof).
>
> Would it be a good idea to periodically restart informix to free up this
> memory?
>
> Ian
Ian Spare wrote:
> I've been searching in the archive and can't quite find the detail I'm
> interested in.
>
> The situation is that we've an Informix database running version
> 7.23.UC6 on HPUX 10.20.
Ouch! IDS 7.23 is not Y2K-compliant and should not still be in use.
I suspect HP-UX 10.20 is close to obsolete too. You should be
considering upgrades.
> To my knowledge this was started on Sept 17th
> and by the middle of last week was using a fair amount of memory
> (ordinary memory not shared memory).
Sounds like there might be a slow memory leak. Do you have any really
long-running sessions? How much help do you get from monitoring with
'onstat'? (Which options are available in 7.23 to monitor memory
usage -- I've forgotten, assuming I ever knew.)
> Then on Oct 8 around 400mb of
> memory got freed up and some new oninit's got fired up.
If you had a really long-running session (daemon process?) that
finally expired (possibly out of memory itself), then the server might
have been able to release the resources consumed by the long running
session.
> There's 12 oninit processes running from Sept 17th and 33 that were
> started last week.
Intriguing. How many oninit processes do you normally run with? Are
the 12 and the 33 both running out of the same (only?) INFORMIXDIR?
Are the 12 still busy? Can you track the process hierarchy (ps
listing of PID and PPID - there should be one process in the 45 owned
by PPID = 1 (init - the Unix init process), and everything else should
be descended from the PID that has PPID = 1. If you have two PIDs
with PPID = 1, then you have two instances of IDS running - or more.
> Onstat now seems to say the database was started last
> week and only lists virtual processors with PID's that were started last
> week.
OK - so the other lot should be moribund. Are they still busy at all?
> It looks to me like these processes are totally detached from each
> other.
Worrying.
> And yet, apparently no one restarted the database, we had no reports of
> outages etc.
Well, this is where the routine system monitoring that Lester Knutsen
advocated in one of his talks at the IBM DM Tech Conference in Las
Vegas, and which I advocated in one of my talks at the DMTC, and which
Dan Wood advocated in one of hist talks at the DMTC, would come in handy.
The sort of routine monitoring I have in mind is probably triggered by
cron, and simply logs key statistics (onstat, ps, ...) on an hourly
basis, with daily, weekly, monthly cleanups and analysis. Lester was
concerned with also recording system performance information with
iostat, vmstat, sar, etc. He also suggested, perfectly sensibly,
storing the data in a database. My version stopped short of both
aspects - but if you're concerned about overall system health,
monitoring the system is helpful.
The big advantages of this sort of monitoring is that you (a)
establish the 'normal operating condition' of your system, which helps
you diagnose when the current behaviour is out of whack, and (b) it
usually helps you diagnose when the problem occurred.
Obviously, if you don't have this monitoring in place, it is too late
to find out all that happened. However, it might help for next time.
Have you tracked through the online log file? You should find records
about what happened a week ago in there.
> How does it work?
That's a big question.
> Will oninits periodically respawn?
Not in the ordinary course of events. If you add CPU VPs, or if you
enable PDQ, you might get more oninits
> For that matter is it normal for memory usage to creep up?
It should only happen if there's a process (thread, session) which is
forcing the server to use more memory because it is carelessly written
and is not freeing the resources it is asking the server to consume.
Of course, it could be a creepy-crawly bug in the version of IDS you
have - again, you'd know better if you had the monitoring in place.
> The "old" onit's look active to
> me, they've still got network connections, files open etc (using lsof).
Do they show up on 'top' or an equivalent? Can you detect any CPU
usage? If, for sake of argument, these were pseudo-zombie AIO VPs
that for some reason had not died, then they might have a number of
files open without actively writing to them or reading from them.
How much shared memory is in use?
> Would it be a good idea to periodically restart informix to free up this
> memory?
It's not a bad idea to restart every year or so - you could do it more
frequently if you found that you did have problematic memory growth.
But it would be worth tracking why the memory use grows, especially if
there are any long-running sessions. OTOH, there's no formal reason
to restart it unless something goes wrong.
Gut feel: you had a weird crash about a week ago, and someone isn't
'fessing up to having restarted the server. The two disjoint sets of
oninit processes is weird - it should mean that you have two
independent servers running (multiple residency). If the first server
still had shared memory, it should prevent the second instance of the
same server number from running again. Have a good look at what ipcs
tells you - how many lots of shared memoy does it show? I'm inclined
to think that someone did something fairly dramatic a week ago, and
you have finally spotted it. Look at all the available o/s logs to
see whether there is anything you can see that might help - login
records, etc. System log. Anywhere.
Good luck.
--
Jonathan Leffler #include <disclaimer.h>
Email: jleffler@earthlink.net, jleffler@us.ibm.com
Guardian of DBD::Informix v2003.04 -- http://dbi.perl.org/
what does the online log say happened last week when the number of
processes changed (probably on October 8th) do you see the engine
beeing shut down or at least restarted?
It is possible to manually add oninits but I don't think I met an
instance that magically spawned them
Or did someome start a second IDS instance on the same machine?
Ian Spare <Ian_Spare@hotmail.com> wrote in message news:<2t1ne2F1p714oU1@uni-berlin.de>...
> Hi,
>
> I've been searching in the archive and can't quite find the detail I'm
> interested in.
>
> The situation is that we've an Informix database running version
> 7.23.UC6 on HPUX 10.20. To my knowledge this was started on Sept 17th
> and by the middle of last week was using a fair amount of memory
> (ordinary memory not shared memory). Then on Oct 8 around 400mb of
> memory got freed up and some new oninit's got fired up.
>
> There's 12 oninit processes running from Sept 17th and 33 that were
> started last week. Onstat now seems to say the database was started last
> week and only lists virtual processors with PID's that were started last
> week. It looks to me like these processes are totally detached from each
> other.
>
> And yet, apparently no one restarted the database, we had no reports of
> outages etc.
>
> How does it work? Will oninits periodically respawn? For that matter is
> it normal for memory usage to creep up? The "old" onit's look active to
> me, they've still got network connections, files open etc (using lsof).
>
> Would it be a good idea to periodically restart informix to free up this
> memory?
>
> Ian
Ian,
I'm guessing that you already know, but the version of Informix and
HP-UX are pretty old. For your own sanity you should upgrade.
That said, there were a lot of problems with HP-UX 10.20 and shared
memory usage. I remember a summer when HP-UX 10.20 was a current
version when HP-UX released over 100 patches in one month and many of
these patches concerned shared memory allocation. There was an
Informix bug (sorry, can't remember the number) concerning shared
memory allocation that wasn't fixed on HP-UX 10.20 because HP-UX kept
changing 10.20 on Informix.
Those two points said, yes, it is normal for Informix to allocate
extra memory over the life of the instance, and to spawn and kill
oninit processes. The "onstat -g seg" command will help you keep track
of which oninit processes are allocated for a particular duty, and how
much memory each is using. The onconfig parameter "SHMVIRTSIZE" is the
initial memory allocation for the instance, and "SHMADD" is the
onconfig parameter that is the size of extra shared memory
allocations. If you see in your online.log that the instance has added
three extra shared memory segments, then stop your instance and
increase SHMVIRTSIZE accordingly:
new SHMVIRTSIZE = old SHMVIRTSIZE + (3 * SHMADD)
While the instance will work with fragmented shared memory, for
performance reasons it's best to have a single large shared memory
allocation.
It is a good idea to periodically restart the Informix instance to
free unused shared memory, or you can use the "onmode -F" command.
Hope this information helps.
Brice Avila
Ian Spare <Ian_Spare@hotmail.com> wrote in message news:<2t1ne2F1p714oU1@uni-berlin.de>...
> Hi,
>
> I've been searching in the archive and can't quite find the detail I'm
> interested in.
>
> The situation is that we've an Informix database running version
> 7.23.UC6 on HPUX 10.20. To my knowledge this was started on Sept 17th
> and by the middle of last week was using a fair amount of memory
> (ordinary memory not shared memory). Then on Oct 8 around 400mb of
> memory got freed up and some new oninit's got fired up.
>
> There's 12 oninit processes running from Sept 17th and 33 that were
> started last week. Onstat now seems to say the database was started last
> week and only lists virtual processors with PID's that were started last
> week. It looks to me like these processes are totally detached from each
> other.
>
> And yet, apparently no one restarted the database, we had no reports of
> outages etc.
>
> How does it work? Will oninits periodically respawn? For that matter is
> it normal for memory usage to creep up? The "old" onit's look active to
> me, they've still got network connections, files open etc (using lsof).
>
> Would it be a good idea to periodically restart informix to free up this
> memory?
>
> Ian
Art S. Kagel wrote:
> On Tue, 12 Oct 2004 05:45:38 -0400, Ian Spare wrote:
>
> Is it possible that someone started second instance, managing independent disk
> space/databases/etc., on the 8th? This would explain this somewhat, though if
> your INFORMIXSERVER setting *should* point to the original instance and onstat
> -g glo shows the new oninits, it is more likely that something happened to
> hang the original oninit processes and someone destroyed their shared memory
> segments and restarted the engine without killing the original processes.
>
That explains it, reading this and the other helpful posts here helped
me understand what happened. I'm not a DBA for Informix or any other
database so it's hard to know where to look without these hints. I can
now see we've got a test instance as well as a prod instance. The prod
instance must have been restarted but the test one wasn't, obviously the
most simple solution of course but it hadn't occured to me. That's why
the old processes looked active to me, they were :-)
I'm still digesting the other comments but it's clear to me we do have a
slow memory leak, it's normal memory that's leaking, the shared memory
usage is static. A periodic restart looks the best way to deal with this.
Many thanks guys !
Ian
"Ian Spare" <Ian_Spare@hotmail.com> wrote in message news:2t2p2fF1qvb4rU1@uni-berlin.de... > Art S. Kagel wrote: > > On Tue, 12 Oct 2004 05:45:38 -0400, Ian Spare wrote: > > > I'm still digesting the other comments but it's clear to me we do have a > slow memory leak, it's normal memory that's leaking, the shared memory > usage is static. A periodic restart looks the best way to deal with this. Migrating to supported versions would be the best way to deal with it, as a previous poster suggested. There were a lot of memory problems in 10.20 that in my expereince at several 24x7 sites, are fixed now in 11.0 and 11.11.
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g