Re: IDS crashes on Windows
Posted in 2005
> Strange that it would run for four years
> before it's become a problem, though.
Agreed. If it is a handle count leak, could some change in your
applications have triggered it like:
1. stored procedure system command failing (I vaguely remember that
caused a resource problem in some versions).
2. Set debug statement being used in some stored procedures.
3. Additional user activity.
3. ???
If this is a resource leak you might get a more accurate error message
(instead or errno=0) with 9.40.TC5 due to the fix for PTS 166079 CALLS
TO WRITEFILE ON WINDOWS DO NOT TRAP FOR WORKING SET ERRORS and others.
Also consider using diskmon from sysinternals.com to trap for the
underlying errors from WriteFile(). It should give you output like this:
147950 8:21:16 AM oninit.exe:1812 WRITE
c:\\IFMXDATA\\mysvr\\mysvr_dat.000 * 0xC00000A1 Offset: 1478434816
Length: 4096
147951 8:21:17 AM oninit.exe:1812 WRITE
c:\\IFMXDATA\\mysvr\\mysvr_dat.000 * 0xC00000A1 Offset: 1478434816
Length: 4096
147952 8:21:17 AM oninit.exe:1812 WRITE
c:\\IFMXDATA\\mysvr\\rootdbs_dat.000 * 0xC00000A1 Offset: 12566528
Length: 4096
In this example, 0xC00000A1 represents STATUS_WORKING_SET_QUOTA.
Guy
Everett Mills wrote:
> Thanks Guy-
> I'll check into that. Strange that it would run for four years
> before it's become a problem, though.
>
>
>
> --EEM
>
>
>
>
>>-----Original Message-----
>>From: Guy Bowerman [mailto:gbowerman@yahoo.com]
>>Sent: Tuesday, November 29, 2005 10:22 PM
>>To: informix-list@iiug.org
>>Subject: Re: IDS crashes on Windows
>>
>>Everett,
>>
>>The symptoms you describe indicate a resource leak. There may well be
>>bugs in the IDS version you have that leak handles (like PTS 169797 -
>>SET DEBUG FILE STATEMENT CAUSES MEMORY LEAK ON WINDOWS - HANDLE COUNT
>>GROWS - fixed in 9.40.TC6).
>>
>>To confirm a handle count leak, open the Windows Task Manager and
>>highlight the oninit.exe process. Then pick View->Select Columns and
>>check "Handle Count".
>>
>>If the handle count for the oninit process is going up with time then
>
> it
>
>> is probably what is causing your problem.
>>
>>Be aware that in addition to the bug mentioned above there is also one
>>new resource leak that was introduced in 9.40.TC5:
>>PTS 174339 CONSOLE MESSAGES ARE NOT WRITTEN ON WINDOWS AND FILE HANDLE
>>IS NOT CLOSED CAUSING RESOURCE LEAK.
>>
>>This one is fixed in 9.40.TC8.
>>
>>Regards
>>Guy
>>
>>Everett Mills wrote:
>>
>>>Guys-
>>> I've been seeing these error messages show up in my online log.
>>>Several occur in a row, and either clear themselves up or crash the
>>>engine. I am thinking that the RAID1 Array running the database is
>>>getting flakey, but how can you tell which drive, or if it's the
>
> RAID
>
>>>card? I am running IDS 9.30.TC2 (old, I know. I'll be installing
>>>9.40.TC5 over Christmas) on Windows 2000 server. We have enabled
>
> the
>
>>>Windows hardware monitoring, but nothing ever shows up in the event
>
> log.
>
>>>I have run defrags and thorough scandisks. I have also blown away
>
> all
>
>>>of the chunks, recreated the files and restored a backup. The
>
> effect
>
>>>remains the same, every four or five days, it starts to generate
>
> errors
>
>>>and eventually dies. Here is an extract from my online.log:
>>>
>>>
>>>19:41:08 KAIO: error in kaio_WRITE, kaiocbp = 0x14ffe1c8, errno = 0
>>>19:41:08 fildes = 480 (gfd 6), buf = 0x147A7000, nbytes = 4096,>
> offset
>
>>>= 126353408
>>>19:41:09 KAIO: error in kaio_WRITE, kaiocbp = 0x14feee94, errno = 0
>>>19:41:09 fildes = 480 (gfd 6), buf = 0x156BC000, nbytes = 4096,>
> offset
>
>>>= 190603264
>>>19:41:09 KAIO: error in kaio_WRITE, kaiocbp = 0x1528506c, errno = 0
>>>19:41:09 fildes = 480 (gfd 6), buf = 0x15287000, nbytes = 32768,>
> offset
>
>>>= 225316864
>>>19:41:09 KAIO: error in kaio_WRITE, kaiocbp = 0x14ffe280, errno = 0
>>>19:41:09 fildes = 480 (gfd 6), buf = 0x0EE73000, nbytes = 4096,>
> offset
>
>>>= 560599040
>>>19:41:09 KAIO: error in kaio_WRITE, kaiocbp = 0x14ffe110, errno = 0
>>>19:41:09 fildes = 480 (gfd 6), buf = 0x13EF3000, nbytes = 4096,>
> offset
>
>>>= 593719296
>>>19:41:10 Checkpoint Completed: duration was 1 seconds.
>>>19:41:10 Checkpoint loguniq 140098, logpos 0x1a3018
>>>
>>>19:41:10 Maximum server connections 13
>>>19:45:39 KAIO: error in kaio_WRITE, kaiocbp = 0x14ff382c, errno = 0
>>>19:45:39 fildes = 448 (gfd 14), buf = 0x0D8EB000, nbytes = 4096,>
> offset
>
>>>= 169680896
>>>19:46:09 Checkpoint Completed: duration was 0 seconds.
>>>19:46:10 Checkpoint loguniq 140098, logpos 0x1a9018
>>>
>>>19:46:10 Maximum server connections 13>>>
>>> <<Informix Dynamic Server>>> Logical Log 140098 Complete.
>>>19:48:03 Logical Log 140098 Complete.
>>>19:51:14 Checkpoint Completed: duration was 0 seconds.
>>>19:51:14 Checkpoint loguniq 140099, logpos 0xcb018
>>>
>>>19:51:15 Maximum server connections 13
>>>19:52:08 KAIO: error in kaio_WRITE, kaiocbp = 0x14fed49c, errno = 0
>>>19:52:08 fildes = 468 (gfd 4), buf = 0x0D8DA000, nbytes = 36864,>
> offset
>
>>>= 3375104
>>>19:52:09 KAIO: error in kaio_WRITE, kaiocbp = 0x14fed49c, errno = 0
>>>19:52:09 fildes = 468 (gfd 4), buf = 0x0D8DA000, nbytes = 36864,>
> offset
>
>>>= 3375104
>>>19:52:09 Assert Failed: I/O error, Mirror Chunk
>>>'C:\\IFMXDATA\\DDGPLC\\rootdbs_dat.000' -- Offline
>>>19:52:09 Informix Dynamic Server Version 9.30.TC2
>>>19:52:09 Who: Session(34969,
>>>plcdata@nbpdc-cs1.corpnbp.nationalbeef.com, 118028, 0)
>>> Thread(34986, sqlexec, 0, 1)
>>> File: rsbuff.c Line: 4086
>>>19:52:09 Results: Chunk is now unusable
>>>19:52:10 Action: Repair and restore from mirror or archive
>>>19:52:10 stack trace for pid 10112 written to C:\\tmp\\af.8c92b448
>>>19:52:17 See Also: C:\\tmp\\af.8c92b448
>>>19:52:17 I/O error, Mirror Chunk>
> 'C:\\IFMXDATA\\DDGPLC\\rootdbs_dat.000'
>
>>>-- Offline
>>>19:52:17 I/O error, Mirror Chunk>
> 'C:\\IFMXDATA\\DDGPLC\\rootdbs_dat.000'
>
>>>-- Offline
>>>19:52:19 Assert Failed: Chunk 1 is being taken OFFLINE.
>>>19:52:19 Informix Dynamic Server Version 9.30.TC2
>>>19:52:19 Who: Session(34969,
>>>plcdata@nbpdc-cs1.corpnbp.nationalbeef.com, 118028, 0)
>>> Thread(34986, sqlexec, 0, 1)
>>> File: rsmirror.c Line: 1766
>>>19:52:19 Results: Dynamic Server must abort
>>>19:52:19 Action: Reinitialize shared memory
>>>19:52:19 stack trace for pid 10112 written to C:\\tmp\\af.8c92b448
>>>19:52:25 See Also: C:\\tmp\\af.8c92b448
>>>19:52:25 Chunk 1 is being taken OFFLINE.
>>>19:52:25 Chunk 1 is being taken OFFLINE.
>>>19:52:25 Assert Failed: INFORMIX-OnLine Must ABORT>>> Critical media failure.
>>>
>>>
>>>The af files contain the same info. I'll be calling tech support,
>
> as
>
>>>well.
>>>
>>>
>>> Thanks in advance, if someone
>>>has any ideas.
>