Assert Failed: semop: errno = 43 (EIDRM)
An Informix® engine on Linux aborts with an assertion failure from semop(),
often repeatedly and at the same time of day. The engine is not at fault: something outside it
has deleted its System V semaphore set while it was running. This page explains how to
recognise the problem, find what is doing it, and stop it. The error-code view of the same
number is on -43.
Symptom
Assert Failed: semop: errno = 43 Who: Thread(0, idle, 0, 20) File: mt_fn.c Line: 1687 Stack: afstack <- afhandler <- afcrash_interface <- P <- idle_processor <- startup
- A thread (here an idle thread on a VP) was blocked in
P(), a semaphore wait, whensemop()failed. - The engine aborts and has to be restarted.
- The distinguishing feature is often regularity: the same time, to the second, every day.
What errno 43 means
On Linux, errno 43 is EIDRM, Identifier removed. semop() returns it when the
semaphore set is deleted while a process is waiting on it. On BSD the same number is
EPROTONOSUPPORT, which is why this condition is easily overlooked from the number alone.
The conclusion is firm: the engine did not fail internally. Something external removed its IPC.
Who can delete SysV IPC from outside
systemd-logindwithRemoveIPC=yes. When a user's last session ends, logind removes every SysV object whose owner uid or gid matches that user. The compiled-in default isyeseven whenlogind.confshows#RemoveIPC=yescommented out.ipcrm, run by hand or from a script.onclean, or a secondoniniton the sameSERVERNUM, clearing segments it believes are stale.- A container sharing the host IPC namespace (
--ipc=host) running cleanup tools.
Narrowing it down
| Check | What it tells you |
|---|---|
id informix | A UID of 1000 or above is not a system user, so logind will clean up its IPC. |
grep RemoveIPC /etc/systemd/logind.conf | Commented out means the default, yes. |
journalctl -u systemd-logind around the crash time | No session ending at that moment points away from logind. |
last informix | Shows whether the engine owner ever logs in interactively. |
Crash times in online.log | The same second every day means a scheduled or automated trigger, not a random fault. |
| Full journal at the crash second | Look for a burst of short SSH sessions, cron jobs or timers starting in the same second. |
user@UID.service stop lines in the journal | Absent means that user never fully logged out, so logind's RemoveIPC did not fire for them. |
The usual outcome: logind is ruled out when the owning user always has a session open, and the
trigger turns out to be a command run by an automated job at exactly the crash second
(onclean, ipcrm, or onmode/oninit run with the wrong environment).
Catch the culprit: auditd
The decisive step. Add these rules before the next crash window:
auditctl -a always,exit -F arch=b64 -S semctl -F a2=0 -k ipc_rm # IPC_RMID on semaphores auditctl -a always,exit -F arch=b64 -S shmctl -F a1=0 -k ipc_rm # IPC_RMID on shared memory
After the crash:
ausearch -k ipc_rm -i | grep -E 'exe=|comm=|auid=|uid=' auditctl -D -k ipc_rm # remove the rules afterwards
exe= names the process that deleted the IPC and auid= the login user behind it.
Informix's own clean shutdown also shows up, so read the entry from the crash second.
Other checks
id <automation-user> # is the engine's group its primary group?
ipcs -s -c ; ipcs -m -c # who owns the semaphores and shared memory
systemctl list-timers --all # timers firing at the crash time
crontab -l ; crontab -l -u informix # and any other user's crontab
docker inspect --format '{{.Name}} {{.HostConfig.IpcMode}}' $(docker ps -aq)
Also check whether a key used for automation has a command= restriction in
authorized_keys, and what changed on the system on or just before the first crash date.
If the automated logins are not yours, treat that as a security incident first.
Hardening
- Set
RemoveIPC=noin/etc/systemd/logind.conf, thensystemctl restart systemd-logind. Existing sessions are not ended. - Start the engine from a systemd unit (
User=informix) rather than by hand from a login session, so it does not depend on anyone's sessions. - If another user's primary group is
informix, make it secondary (usermod -g user -aG informix user) so that user's logout cannot match the engine's IPC by gid.
Recovery after a crash
ipcs -s ; ipcs -m # look for leftover informix-owned sets onclean -ky # or ipcrm -s / ipcrm -m oninit # fast recovery runs on startup
See also
- Error -43 and -36, the same
EIDRMcondition by errno number. - Working with Assert Failures
- Inter-Process Communication