Error 104 - File descriptor...
Posted in 2013
A 4GL batch job on IDS 11.50 (AIX 6.1) intermittently failed with SQL -261 / ISAM -104 'too many files open', and the poster questioned whether -104 really relates to OS file descriptors. IBM support (Jacques Renaut) confirmed it does not: the -104 comes from an internal per-thread limit of 32767 open table file descriptors, so monitoring should count onstat -g opn lines per thread id, not overall. No sysmaster equivalent exists, though sysrstcb.nopens gives a rough high-water mark. No fix for the application itself was recorded; the poster planned to open a PMR requesting better monitoring.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Storage & Space Management, Error Codes & Troubleshooting, Connectivity: ESQL/C, 4GL & Embedded SQL, Networking & sqlhosts Configuration, Platform-Specific Issues, Versions, Editions & End-of-Life
Hi ,
ifx 11.50 FC9x6 , AIX 6.1
Yesterday, one of our interface batch jobs (4gl) become stopping with
error 261/104 few times during the day.
Sql : -261 Cannot create file for table ().
Isam : -104 ISAM error: too many files open.
I notice this occur because the amount of data grew .
And then, later (after rush hour) it just run without problem (the same
data).
We already got into this problem before, where we open a PMR and the
solution was increase the OS ulimit/open file parameter.
Where we change to 10.000 .
I always had the doubt if this error (-104) have a real relation with
O.S. limits...
After doing a better research, at this point my conclusion is : Doesn't !!
---->>>> Appreciate the opinion of the developers/gurus here to my/our
clarification how this works.
What I found and convince me the error 104 doesn't have direct relation
with OS limits is :
1) This message from Jacques (17/may/2011) on informix-list (usenet).
Saying about the internal concept of "file descriptor"
>> I'd suggest contacting support. I seriously doubt the 104 error on
>> the start violations would have anything to do with OS level file
>> descriptors. The sqlexec thread on the server would not be opening
>> anything additionally on that command, as a table in the server is not
>> a unix file construct. All the oninit processes would already have
>> all the chunks open, so the only OS file descriptors involved would be
>> the ones for chunks or for the network connections. We have an
>> internal concept of "file descriptor" that this is likely referring to
>> that relates to tables, but the limit on the number of these should be
>> very large, like 32k if I remember right. So I would think that that
>> error is either getting generated when it shouldn't, or possibly an
>> error is occurring but we're maybe reporting the wrong error or a
>> misleading error.
>> Jacques Renaut
>> IBM Informix Advanced Support
>> APD Team
2) At the moment we got the error -104 the onstat -g iog return the
values bellow, what is very low number and I know they represents the #
of chunks and other files:
$ og iog
IBM Informix Dynamic Server Version 11.50.FC9X6 -- On-Line -- Up 5 days
17:57:56 -- 183239776 Kbytes
AIO global info:
9 aio classes
152 open files
192 max global files
3) Today , monitoring a session with is using a lot of memory
(considering our AVG), and there no relation with the problem yesterday
I notice it with a log of "File Descriptors" (isfd) over the same table.
I believe this is some programing flaw , something like open or prepare
the same cursor repeatedly.
But the focus here is , the same table, appear lot of times, for the
same thread with a lot of ISFD.
$ og opn 7287878 | egrep "0x0250001d|part"tid rstcb isfd op_mode op_flags partnum ucount
ocount lockmode
7287878 0x0700001daa3cca68 103 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 107 0x00000400 0x00000403 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 125 0x00000402 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 129 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 141 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 170 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 206 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 223 0x00000402 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 225 0x00000402 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 230 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 232 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 238 0x00000402 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 298 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 308 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 312 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 322 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 332 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 336 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 342 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 357 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 369 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 371 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 375 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 379 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 381 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 383 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 387 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 391 0x00000402 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 399 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 401 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 407 0x00000402 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 411 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 413 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 453 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 469 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 473 0x00000400 0x00000403 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 499 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 542 0x00000400 0x00000403 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 546 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 548 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 570 0x00000400 0x00000403 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 572 0x00000402 0x00000403 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 617 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 678 0x00000400 0x00000407 0x0250001d 45
0 9
7287878 0x0700001daa3cca68 697 0x00000400 0x00000407 0x0250001d 45
0 9
Comments please!!
And if possible, a tip where monitor the amount of the same "file
descriptor" showed by "onstat -g opn"
Regards
Cesar
Complementing...
The way I will use to monitor the amount of this FDs , without sure if
is correct or not....
# of open Files Descriptors.
$ onstat -g opn | awk '{print $3}' | sort | uniq -c | wc -l
721
Check what FDs is + shared.....
$ onstat -g opn | awk '{print $3}' | sort | uniq -c | sort -nk2
...
On 24/1/2013 12:47, Cesar Inacio Martins wrote:
> Hi ,
>
> ifx 11.50 FC9x6 , AIX 6.1
>
> Yesterday, one of our interface batch jobs (4gl) become stopping with
> error 261/104 few times during the day.
>
> Sql : -261 Cannot create file for table ().
>
> Isam : -104 ISAM error: too many files open.
>
> I notice this occur because the amount of data grew .
> And then, later (after rush hour) it just run without problem (the same
> data).
>
> We already got into this problem before, where we open a PMR and the
> solution was increase the OS ulimit/open file parameter.
> Where we change to 10.000 .
>
> I always had the doubt if this error (-104) have a real relation with
> O.S. limits...
> After doing a better research, at this point my conclusion is : Doesn't !!
>
> ---->>>> Appreciate the opinion of the developers/gurus here to my/our
> clarification how this works.
>
> What I found and convince me the error 104 doesn't have direct relation
> with OS limits is :
>
> 1) This message from Jacques (17/may/2011) on informix-list (usenet).
> Saying about the internal concept of "file descriptor"
>
>>> I'd suggest contacting support. I seriously doubt the 104 error on
>>> the start violations would have anything to do with OS level file
>>> descriptors. The sqlexec thread on the server would not be opening
>>> anything additionally on that command, as a table in the server is not
>>> a unix file construct. All the oninit processes would already have
>>> all the chunks open, so the only OS file descriptors involved would be
>>> the ones for chunks or for the network connections. We have an
>>> internal concept of "file descriptor" that this is likely referring to
>>> that relates to tables, but the limit on the number of these should be
>>> very large, like 32k if I remember right. So I would think that that
>>> error is either getting generated when it shouldn't, or possibly an
>>> error is occurring but we're maybe reporting the wrong error or a
>>> misleading error.
>>> Jacques Renaut
>>> IBM Informix Advanced Support
>>> APD Team
> 2) At the moment we got the error -104 the onstat -g iog return the
> values bellow, what is very low number and I know they represents the #
> of chunks and other files:
> $ og iog>
> IBM Informix Dynamic Server Version 11.50.FC9X6 -- On-Line -- Up 5 days
> 17:57:56 -- 183239776 Kbytes>
> AIO global info:
>
> 9 aio classes
> 152 open files
> 192 max global files
>
> 3) Today , monitoring a session with is using a lot of memory
> (considering our AVG), and there no relation with the problem yesterday
> I notice it with a log of "File Descriptors" (isfd) over the same table.
> I believe this is some programing flaw , something like open or prepare
> the same cursor repeatedly.
> But the focus here is , the same table, appear lot of times, for the
> same thread with a lot of ISFD.
>
> $ og opn 7287878 | egrep "0x0250001d|part"> tid rstcb isfd op_mode op_flags partnum ucount
> ocount lockmode
> 7287878 0x0700001daa3cca68 103 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 107 0x00000400 0x00000403 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 125 0x00000402 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 129 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 141 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 170 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 206 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 223 0x00000402 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 225 0x00000402 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 230 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 232 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 238 0x00000402 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 298 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 308 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 312 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 322 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 332 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 336 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 342 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 357 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 369 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 371 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 375 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 379 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 381 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 383 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 387 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 391 0x00000402 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 399 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 401 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 407 0x00000402 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 411 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 413 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 453 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 469 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 473 0x00000400 0x00000403 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 499 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 542 0x00000400 0x00000403 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 546 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 548 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 570 0x00000400 0x00000403 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 572 0x00000402 0x00000403 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 617 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 678 0x00000400 0x00000407 0x0250001d 45
> 0 9
> 7287878 0x0700001daa3cca68 697 0x00000400 0x00000407 0x0250001d 45
> 0 9
>
> Comments please!!
>
> And if possible, a tip where monitor the amount of the same "file
> descriptor" showed by "onstat -g opn"
>
> Regards
> Cesar
>
>
>
*******************************************************************************
>
Original Post:
Complementing...
The way I will use to monitor the amount of this FDs , without sure if
is correct or not....
# of open Files Descriptors.
$ onstat -g opn | awk '{print $3}' | sort | uniq -c | wc -l
721
Check what FDs is + shared.....
$ onstat -g opn | awk '{print $3}' | sort | uniq -c | sort -nk2
...
Response:
You need to look at the FD's per thread id. I think the 1st command you have
isn't giving you what you think, which I think you're trying to get is the
total number of FD's (and by FD I mean the internal IDS FD not the OS
definition of FD). However, the limit that causes the 104 error to be returned
is more then 32767 FD for 1 thread. So onstat -g opn by it self gives output
for all threads. You need something to then go and sum up each grouping of
lines by thread id, so you can tell how many FD's each thread has. So I think
it would be more something like this you'd want to run:
onstat -g opn | awk '{print $1}' | uniq -c | sort -nk2
That should get you a count of FD's per thread id. The 1st column is the
number of fd's 2nd column is Thread id. (You'd still want to clean that up
some as it has stuff for the blank lines separating each thread id output and
stuff). But I think it's closer to what you are looking for in terms of trying
to monitor if any thread is getting close to the 32767 open table limit in the
serer that generates 104 errors.
Jacques Renaut
IBM Informix Advanced Support
APD Team
Hi Jacques,
Thanks for your comments.
So , I'm considering your answer as confirmation about my understanding
of 104 errors.
Do you know if exists a technotes or someplace with IBM official doc
where explain this limits?
And how monitoring it when the error 104 occurs ?
I looking at sysmaster trying to write a SQL where give me similar
result of onstat -g opn... without success for now.
Do you have something to share?
Regards
Cesar
On 24/1/2013 15:14, JACQUES RENAUT wrote:
> Original Post:
>
> Complementing...
>
> The way I will use to monitor the amount of this FDs , without sure if
> is correct or not....
>
> # of open Files Descriptors.
> $ onstat -g opn | awk '{print $3}' | sort | uniq -c | wc -l>
> 721
>
> Check what FDs is + shared.....
> $ onstat -g opn | awk '{print $3}' | sort | uniq -c | sort -nk2
> ....>
> Response:
>
> You need to look at the FD's per thread id. I think the 1st command you have
> isn't giving you what you think, which I think you're trying to get is the
> total number of FD's (and by FD I mean the internal IDS FD not the OS
> definition of FD). However, the limit that causes the 104 error to be
returned
> is more then 32767 FD for 1 thread. So onstat -g opn by it self gives output
> for all threads. You need something to then go and sum up each grouping of
> lines by thread id, so you can tell how many FD's each thread has. So I think
> it would be more something like this you'd want to run:
>
> onstat -g opn | awk '{print $1}' | uniq -c | sort -nk2>
> That should get you a count of FD's per thread id. The 1st column is the
> number of fd's 2nd column is Thread id. (You'd still want to clean that up
> some as it has stuff for the blank lines separating each thread id output and
> stuff). But I think it's closer to what you are looking for in terms of
trying
> to monitor if any thread is getting close to the 32767 open table limit in
the
> serer that generates 104 errors.
>
> Jacques Renaut
> IBM Informix Advanced Support
> APD Team
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Jacques , I forgot.
If possible, a documentation with explanation why the same table could
appear many times open for the same thread.
Cursor open and not closed/freerepeatedly , multiple cursors open ,
prepared constantly , any other reason!??
I don't know how suggest offically to IBM, I think this subject give us
a useful technotes.
Regards
Cesar
On 24/1/2013 15:14, JACQUES RENAUT wrote:
> Original Post:
>
> Complementing...
>
> The way I will use to monitor the amount of this FDs , without sure if
> is correct or not....
>
> # of open Files Descriptors.
> $ onstat -g opn | awk '{print $3}' | sort | uniq -c | wc -l>
> 721
>
> Check what FDs is + shared.....
> $ onstat -g opn | awk '{print $3}' | sort | uniq -c | sort -nk2
> ....>
> Response:
>
> You need to look at the FD's per thread id. I think the 1st command you have
> isn't giving you what you think, which I think you're trying to get is the
> total number of FD's (and by FD I mean the internal IDS FD not the OS
> definition of FD). However, the limit that causes the 104 error to be
returned
> is more then 32767 FD for 1 thread. So onstat -g opn by it self gives output
> for all threads. You need something to then go and sum up each grouping of
> lines by thread id, so you can tell how many FD's each thread has. So I think
> it would be more something like this you'd want to run:
>
> onstat -g opn | awk '{print $1}' | uniq -c | sort -nk2>
> That should get you a count of FD's per thread id. The 1st column is the
> number of fd's 2nd column is Thread id. (You'd still want to clean that up
> some as it has stuff for the blank lines separating each thread id output and
> stuff). But I think it's closer to what you are looking for in terms of
trying
> to monitor if any thread is getting close to the 32767 open table limit in
the
> serer that generates 104 errors.
>
> Jacques Renaut
> IBM Informix Advanced Support
> APD Team
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Original post:
Hi Jacques,
Thanks for your comments.
So , I'm considering your answer as confirmation about my understanding
of 104 errors.
Do you know if exists a technotes or someplace with IBM official doc
where explain this limits?
And how monitoring it when the error 104 occurs ?
I looking at sysmaster trying to write a SQL where give me similar
result of onstat -g opn... without success for now.
Do you have something to share?
Regards
Cesar
Response:
I found this doc
http://publib.boulder.ibm.com/infocenter/idshelp/v115/index.jsp?topic=%2Fcom.ibm
.adref.doc%2Fids_adr_0722.htm
but it actually looks like that information is misleading/incorrect? I'm
referring to the "maximum number of open tables per user and join" row. It
states it's "dynamic allocation". Which is correct at the lowest level, but
then there's a higher level where the 32767 limit kicks in (again 32767 open
tables per thread).
I did take a look at sysmaster and there does not appear to be a way to get
the same output as onstat -g opn command. A less accurate but possibly close
way to gauge would be to look at the sysrstcb table. A query like "select tid,
nopens from sysrstcb where tid != 0" will give you the max size of the open
file fd table (the open file fd table is what onstat -g opn is printing). So
nopens gives you the max size that it got to (not what's currently open) and
if that nopens got to 32767 that thread could possibly hit the 104 error. Also
that number I believe will jump in increments of 64...so it'd go 64 to 128 to
192 etc...
Jacques Renaut
IBM Informix Advanced Support
APD Team
Jacques ,
Thank you for you answer , although not the answer I wanted , help me to
understand this situation.
I will open a PMR requesting a feature (onstat option) where able to
monitor the real usage of this limits and avoid miss understand the
relation of it with O.S. file descriptors.
Regards
Cesar
On 25/1/2013 14:14, JACQUES RENAUT wrote:
>
> Response:
>
> I found this doc
>
>
>
http://publib.boulder.ibm.com/infocenter/idshelp/v115/index.jsp?topic=%2Fcom.ibm
.adref.doc%2Fids_adr_0722.htm
>
> but it actually looks like that information is misleading/incorrect? I'm
> referring to the "maximum number of open tables per user and join" row. It
> states it's "dynamic allocation". Which is correct at the lowest level, but
> then there's a higher level where the 32767 limit kicks in (again 32767 open
> tables per thread).
>
> I did take a look at sysmaster and there does not appear to be a way to get
> the same output as onstat -g opn command. A less accurate but possibly close
> way to gauge would be to look at the sysrstcb table. A query like "select
tid,
> nopens from sysrstcb where tid != 0" will give you the max size of the open
> file fd table (the open file fd table is what onstat -g opn is printing). So
> nopens gives you the max size that it got to (not what's currently open) and
> if that nopens got to 32767 that thread could possibly hit the 104 error.
Also
> that number I believe will jump in increments of 64...so it'd go 64 to 128 to
> 192 etc...
>
> Jacques Renaut
> IBM Informix Advanced Support
> APD Team
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g