ontape backup failed
Posted in 2010
On Tru64 with IDS 9.40.FC2, ontape failed every 30-40 days with "Server is in an incompatible state or user authentication failed", plus -952 password errors and a shmat "not enough core" error; only a reboot or engine restart fixed it. Art Kagel diagnosed it as shared-memory exhaustion: the engine had added extra virtual segments and vmstat showed almost no free memory, so the archive couldn't attach another segment. Running onmode -F freed/consolidated memory and let ontape run without a reboot; adding RAM and watching for an engine/application memory leak was advised, with a caution that frequent onmode -F on that vintage could crash the engine.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Backup & Restore
OS - Compaq Tru64 UNIX V5.1A
Informix IDS Version 9.40.FC2
When running ontape backup receive the following message.
Server is in an incompatible state or user authentication failed.
Program over.
Nothing have been chnage to IDS or UNIX. It was working fine. We notice that
if we reboot the system or restart the IDS, the ontape backup works fine.
Welcome any suggestions for fixing the problem.
Thanks in advance.
Kanti
On 14 June 2010 14:15, KANTI BAKARANIA <kbakarania@pascoclerk.com> wrote:
> OS - Compaq Tru64 UNIX V5.1A
> Informix IDS Version 9.40.FC2
>
> When running ontape backup receive the following message.
>
> Server is in an incompatible state or user authentication failed.
>
> Program over.
>
> Nothing have been chnage to IDS or UNIX. It was working fine. We notice that
> if we reboot the system or restart the IDS, the ontape backup works fine.
>
> Welcome any suggestions for fixing the problem.
>
> Thanks in advance.
>
> Kanti
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
How long does it take, after a reboot (either time or number of
archive), before you get this error ?
Anything in the ids log file ? Is ontape always run by hte same user ?
What does an onstat -p show immediately before/after a failed attempt
?
Keith
Thank you Keith, Here is the information you asked for.
How long does it take, after a reboot (either time or number of archive),
before you get this error?
It takes about 30 to 40 days.
ontape is starting from cron as Informix user.
The online.log has following entry.
10:30:16 Maximum server connections 550
10:35:16 Fuzzy Checkpoint Completed: duration was 0 seconds, 256 buffers not
flushed,
timestamp: 745890119.
10:35:16 Checkpoint loguniq 794586, logpos 0x3bc038, timestamp: 745890120
10:35:16 Maximum server connections 550
10:35:52 Logical Log 794586 Complete, timestamp:745916202.
10:37:24 Password Validation for user [ USER
. ] failed!
10:37:24 Check for password aging/account lock-out.
10:37:24 listener-thread: err = -952: oserr = 0: errstr = USER
: User (USER
)
's password is not correct for the database server.
10:39:00 Logical Log 794587 Complete, timestamp:746123113.
10:40:16 Fuzzy Checkpoint Completed: duration was 0 seconds, 263 buffers not f
lushed,
timestamp: 746198287.
10:40:16 Checkpoint loguniq 794588, logpos 0xb4720, timestamp: 746198287
10:40:16 Maximum server connections 550
Onstat p before the backup.
# onstat -p
IBM Informix Dynamic Server Version 9.40.FC2 -- On-Line -- Up 32 days 12:55:
52 -- 1024000 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
145667959 168312516 3189867634 95.43 2233709 7701131 60734228 96.32
isamtot open start read write rewrite delete commit rollbk
2074493020 25107323 63662020 1705594378 13841386 366577 121182 335265 10
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
20 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 44115.75 6201.78 671 1729
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
9493367 2 1782291415 0 0 85 1002557 408128
ixda-RA idx-RA da-RA RA-pgsused lchwaits
6987288 3203309 100243651 110237144 139696
Onstat p after the error.
# onstat -p
IBM Informix Dynamic Server Version 9.40.FC2 -- On-Line -- Up 32 days 12:57:
35 -- 1024000 Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
145900144 168546818 3196466118 95.44 2233968 7702393 60769440 96.32
isamtot open start read write rewrite delete commit rollbk
2079380129 25184823 63917951 1709470204 13843566 366826 121215 335529 10
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
21 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 44209.47 6221.65 671 1729
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
9517495 2 1786494763 0 0 85 1004196 408917
ixda-RA idx-RA da-RA RA-pgsused lchwaits
7009843 3203354 100415356 110430584 140035
I am getting following error.
10:36:32 shmat: [12]: operating system error
-12 Not enough core.
An operating-system error code with the meaning shown was unexpectedly
returned to the database server. Core probably refers to data space in
memory that an operating-system function needed. Look for other
operating-system error messages that might give more information.
Thank you
Kanti
Run an onstat -g seg. Do you have one or a few virtual segments ('V' in the
class columns) or do you have many of them? If many it may be that you've
run out of physical memory.
It sounds to me like your engine is allocating additional virtual segments
over time until at the point before it gets the error it has allocated a GB
(see the '1024000' in the onstat headings below). At this point, the archive
is requiring the engine to allocate yet another shared memory segment and
there is no more memory available on your system.
Does vmstat report any free memory available?
One thing to keep in mind is the IDS 9.40xC2 is VERY old circa 2002 and was
rather buggy. You may be running into an internal memory leak in the engine
that's causing the engine to "lose" segments within the virtual memory
segments and so require more memory as time goes on. You can also try
running "onmode -F" to see if that will free up and consolidate enough
contiguous memory to permit the archive to run without requiring an
additional shared memory segment.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Mon, Jun 14, 2010 at 10:51 AM, KANTI BAKARANIA <kbakarania@pascoclerk.com
> wrote:
> Thank you Keith, Here is the information you asked for.
>
> How long does it take, after a reboot (either time or number of archive),
> before you get this error?
>
> It takes about 30 to 40 days.
>
> ontape is starting from cron as Informix user.
>
> The online.log has following entry.
>
> 10:30:16 Maximum server connections 550
> 10:35:16 Fuzzy Checkpoint Completed: duration was 0 seconds, 256 buffers
> not
> flushed,
> timestamp: 745890119.
> 10:35:16 Checkpoint loguniq 794586, logpos 0x3bc038, timestamp: 745890120
>
> 10:35:16 Maximum server connections 550
> 10:35:52 Logical Log 794586 Complete, timestamp:745916202.
> 10:37:24 Password Validation for user [ USER
. ] failed!
> 10:37:24 Check for password aging/account lock-out.
> 10:37:24 listener-thread: err = -952: oserr = 0: errstr = USER
: User
> (USER
)
> 's password is not correct for the database server.
>
> 10:39:00 Logical Log 794587 Complete, timestamp:746123113.
> 10:40:16 Fuzzy Checkpoint Completed: duration was 0 seconds, 263 buffers
> not f
> lushed,
> timestamp: 746198287.
> 10:40:16 Checkpoint loguniq 794588, logpos 0xb4720, timestamp: 746198287
>
> 10:40:16 Maximum server connections 550
>
> Onstat p before the backup.
>
> # onstat -p>
> IBM Informix Dynamic Server Version 9.40.FC2 -- On-Line -- Up 32 days
> 12:55:
> 52 -- 1024000 Kbytes>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 145667959 168312516 3189867634 95.43 2233709 7701131 60734228 96.32
>
> isamtot open start read write rewrite delete commit rollbk
> 2074493020 25107323 63662020 1705594378 13841386 366577 121182 335265 10
>
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 20 0 0 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 0 44115.75 6201.78 671 1729
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 9493367 2 1782291415 0 0 85 1002557 408128
>
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 6987288 3203309 100243651 110237144 139696
>
> Onstat p after the error.
>
> # onstat -p>
> IBM Informix Dynamic Server Version 9.40.FC2 -- On-Line -- Up 32 days
> 12:57:
> 35 -- 1024000 Kbytes>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 145900144 168546818 3196466118 95.44 2233968 7702393 60769440 96.32
>
> isamtot open start read write rewrite delete commit rollbk
> 2079380129 25184823 63917951 1709470204 13843566 366826 121215 335529 10
>
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 21 0 0 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 0 44209.47 6221.65 671 1729
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 9517495 2 1786494763 0 0 85 1004196 408917
>
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 7009843 3203354 100415356 110430584 140035
>
> I am getting following error.
>
> 10:36:32 shmat: [12]: operating system error
>
> -12 Not enough core.
>
> An operating-system error code with the meaning shown was unexpectedly
> returned to the database server. Core probably refers to data space in
> memory that an operating-system function needed. Look for other
> operating-system error messages that might give more information.
>
> Thank you
> Kanti
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--000e0cd75516eb418f0488ff82ff
Art:
Here is the info
# onstat -g seg
IBM Informix Dynamic Server Version 9.40.FC2 -- On-Line -- Up 32 days 14:20:
15 -- 1024000 Kbytes
Segment Summary:
id key addr size ovhd class blkused bl
kfree
0 1381386241 200000000 595591168 437760 R* 144776 63
2
1 1381386242 223800000 360710144 11520 V 69380 18
684
2 1381386243 239000000 8388608 768 M 1231 81
7
20483 1381386244 239800000 41943040 1792 V 1 10
239
4 1381386245 23c000000 41943040 1792 V 1 10
239
Total: - - 1048576000 - - 215389 40
611
(* segment locked in memory)
# vmstat
Virtual Memory Statistics: (pagesize = 8192)
procs memory pages intr cpu
r w u act free wire fault cow zero react pin pout in sy cs us sy id
7 538 38 109K 165K 108K 137G 9M 22M 803 9M 0 1K 4K 12K 17 6 78
I am running onmode -F in backup cron job after the backup is done.
Thank you
Kanti
So, you are showing two additional virtual segments, beyond the initial
segment, of ~40MB each. If the engine is trying to allocate an additional
segment or 40MB that would be a problem since vmstat shows that you only
have 165KB Free. Looks like you may have to add memory to your system.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Mon, Jun 14, 2010 at 12:34 PM, KANTI BAKARANIA <kbakarania@pascoclerk.com
> wrote:
> Art:
>
> Here is the info
>
> # onstat -g seg>
> IBM Informix Dynamic Server Version 9.40.FC2 -- On-Line -- Up 32 days
> 14:20:
> 15 -- 1024000 Kbytes>
> Segment Summary:
> id key addr size ovhd class blkused bl
> kfree
> 0 1381386241 200000000 595591168 437760 R* 144776 63
> 2
> 1 1381386242 223800000 360710144 11520 V 69380 18
> 684
> 2 1381386243 239000000 8388608 768 M 1231 81
> 7
> 20483 1381386244 239800000 41943040 1792 V 1 10
> 239
> 4 1381386245 23c000000 41943040 1792 V 1 10
> 239
> Total: - - 1048576000 - - 215389 40
> 611
>
> (* segment locked in memory)
>
> # vmstat
> Virtual Memory Statistics: (pagesize = 8192)
> procs memory pages intr cpu
> r w u act free wire fault cow zero react pin pout in sy cs us sy id
> 7 538 38 109K 165K 108K 137G 9M 22M 803 9M 0 1K 4K 12K 17 6 78
>
> I am running onmode -F in backup cron job after the backup is done.
>
> Thank you
>
> Kanti
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--000e0cd755169310eb0489003390
Thank you Art,
After I ran the command onmode -F, the ontape worked without rebooting the
machine.
I have question can we run onmode -F while users are on the system or after
the hours. I ran the command after hours.
How can we find out who (which sesson) using the more memory? I think one of
our program may have memory leak, I don't know how can I find out that is it
really a memory leak or not?
Thank you again.
Kanti
onstat -g ses <sessid>
Will provide some info about a session. There is lots more available in the
SMI (sysmaster) database. See tables:
syssessions
syssrsprof
sysrstcb
sysscblst - this last has columns curheap, memused and memtotal.
The linking key is 'sid' the session id column.
Here's the thing about onmode -F: It should be safe, and in most versions
of IDS it IS safe, however, there have been versions, and I'm not certain
that 9.40.xC2 isn't one of those, where running onmode -F more than
occassionally will crash the engine. During the period when there were
memory leaks in the engine (around the time IBM acquired Informix which is
when 9.40 was released) some folk tried to get around the problem by running
onmode -F hourly from a cron. Unfortunately after several hours the enginewould crash with a segmentation fault. So, caviat emptor!
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
IIUG Board of Directors (art@iiug.org)
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Tue, Jun 15, 2010 at 9:27 AM, KANTI BAKARANIA
<kbakarania@pascoclerk.com>wrote:
> Thank you Art,
>
> After I ran the command onmode -F, the ontape worked without rebooting the
> machine.
>
> I have question can we run onmode -F while users are on the system or after
> the hours. I ran the command after hours.
>
> How can we find out who (which sesson) using the more memory? I think one
> of
> our program may have memory leak, I don't know how can I find out that is
> it
> really a memory leak or not?
>
> Thank you again.
>
> Kanti
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--000e0cd7098ccc9eb3048911dde7
Art: Thanks you for the info. Thanks again. Kanti