Re: HPL problem on 9.40
Posted in 2004
--0__=08BBE4F0DFBF86638f9e8a93df938690918c08BBE4F0DFBF8663
Content-type: text/plain; charset=US-ASCII
Hi Andrew,
But why do we want to force all the header conversion all at once.
They will anyway occur when you are going to those tables/pages.
Also providing links to /dev/null can result in a security hole.
So we actually check for the file type in the code.
Now comming to the issue about hang and producing an AF...(Can you send the
FE AF files?)
I wanted to know exactly which version are you on.
Also please open a case with IBM Informix tech support to report this
problem with AF files.
It looks like some corruption might have occured to the pages that pload is
accessing.
Thanks,
Pravin.
owner-informix-list@iiug.org wrote on 03/25/2004 11:16:13 AM:
> Some more information on this problem and its symptoms:
>
> We're running a large number (172) of HPL unload jobs one after the
> other from the command line using onpladm.
>
> The unloads are being done to /dev/null as we only want the pages to
> be read into memory to force the header conversion that is performed
> on each page following an "onmode -BC 2" in 9.40.
>
> We found that HPL would not allow us to create a device array that
> used /dev/null (either directly or through a symbolic link) so we got
> round this by creating an array which used pipes that cat the output
> to /dev/null.
>
> After a number of the unloads have completed onpladm will successfully
> submit a job:
>
> Connecting to onpload, Please wait...
> Successful connection to onpload established
> Thu Mar 25 15:06:22 2004
>
> SHMBASE 0xc00000003492f000
> CLIENTNUM 0x0000000049010000
> Session ID 90>
> Unload Database -> system
> Query Name -> AUTO.90
> Device Array -> dummy_arr
> Query Mapping -> AUTO.90
> Query -> select * from fplbc for read only
> Convert Reject -> /tmp/fplbc_unl.rej
>
> and then hang indefinitely with the session reporting something like:
>
> tid name rstcb flags curstk status
> 5966 sqlexec c0000000175ee258 Y--P--- 131520 cond
> wait(netnorm)
>
> At this time a number of af files are produced but the engine will not
> crash. There is nothing in the related /tmp/onploaderr file. The af
> files have the following message repeated in them:
>
> 15:06:23 Found during mt_shm_malloc_segid 5
> 15:06:23 Pool 'afpool' (0xc000000034a31040)
> 15:06:23 Bad free block 0xc000000034c915f8
> blk-64
> c000000034c915b8: 00000000 00000000 00000000 00000000 ........
> ........
> c000000034c915c8: 00000020 00000000 00000005 00000000 ... ....
> ........
> c000000034c915d8: 00000000 00000000 00000000 00000000 ........
> ........
> c000000034c915e8: 00000000 00000000 00000020 00000000 ........ ...
> ....>
> If the hanging session is cancelled (either interrupted or killed with
> onmode -z) and the job re-submitted it fails with a 255 error,
> reporting:>
> Connecting to onpload, Please wait...
> Error connecting to Socket
>
> But no af files are created this time.
>
> The unload jobs have been specified to use a sockets connection. When
> the "Error connecting to Socket" is returned the output from netstat
> -an |grep 5000 (where 5000 is the port used by this service) looks
> like this:
>
> tcp 0 0 193.118.114.33.5000 *.*
> LISTEN
> tcp 0 0 193.118.114.33.62469 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62470 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62472 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62476 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62477 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62478 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62479 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62480 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62481 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62482 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62484 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62485 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62487 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.62488 193.118.114.33.5000
> TIME_WAIT
> tcp 0 0 193.118.114.33.5000 193.118.114.33.62483
> TIME_WAIT
> tcp 0 0 193.118.114.33.62497 193.118.114.33.5000
> TIME_WAIT
>
> If I wait until all/most of the TIME_WAIT's have gone (maybe 30
> seconds) and re-submit the job it connects fine and runs to completion
> without problem.
--0__=08BBE4F0DFBF86638f9e8a93df938690918c08BBE4F0DFBF8663
Content-type: text/html; charset=US-ASCII
Content-Disposition: inline
<html><body>
<p>Hi Andrew,<br>
<br>
But why do we want to force all the<tt> header conversion all at once.</tt><br>
<tt>They will anyway occur when you are going to those tables/pages.</tt><br>
<br>
<tt>Also providing links to /dev/null can result in a security hole.</tt><br>
<tt>So we actually check for the file type in the code.</tt><br>
<br>
<tt>Now comming to the issue about hang and producing an AF...(Can you send the FE AF files?)</tt><br>
<tt>I wanted to know exactly which version are you on.</tt><br>
<br>
<tt>Also please open a case with IBM Informix tech support to report this problem with AF files.</tt><br>
<tt>It looks like some corruption might have occured to the pages that pload is accessing.</tt><br>
<br>
Thanks,<br>
Pravin.<br>
<br>
<tt>owner-informix-list@iiug.org wrote on 03/25/2004 11:16:13 AM:<br>
<br>
> Some more information on this problem and its symptoms:<br>
> <br>
> We're running a large number (172) of HPL unload jobs one after the<br>
> other from the command line using onpladm.<br>
> <br>
> The unloads are being done to /dev/null as we only want the pages to<br>
> be read into memory to force the header conversion that is performed<br>
> on each page following an "onmode -BC 2" in 9.40.<br>
> <br>
> We found that HPL would not allow us to create a device array that<br>
> used /dev/null (either directly or through a symbolic link) so we got<br>
> round this by creating an array which used pipes that cat the output<br>
> to /dev/null.<br>
> <br>
> After a number of the unloads have completed onpladm will successfully<br>
> submit a job:<br>
> <br>
> Connecting to onpload, Please wait...<br>
> Successful connection to onpload established<br>
> Thu Mar 25 15:06:22 2004<br>
> <br>
> SHMBASE 0xc00000003492f000<br>
> CLIENTNUM