HPL problem on 9.40
Posted in 2004
Topics: Installation, Setup & Upgrades, Platform-Specific Issues
HHP-UX 11.11.
We have a problem with HPL. On selected tables, this is about 10 of 100-odd
we are working on:
We can run onpload against these tables in 9.21 FC4. When we try with
9.40FC3W2 at level 2 it fails. Usually it hangs, as in the second case
below, but occasionally we get the "sockets" error as in the first case.
informix@hamilton[/exports/andy/forcepageconv] onpladm run job povendor_unl
>
Connecting to onpload, Please wait...
Error connecting to Socket
informix@hamilton[/exports/andy/forcepageconv] onpladm run job povendor_unl
>
Connecting to onpload, Please wait...
Successful connection to onpload established
informix@hamilton[/exports/andy/forcepageconv] ^?^[
informix@hamilton[/exports/andy/forcepageconv]
informix@hamilton[/exports/andy/forcepageconv] ^[
informix@hamilton[/exports/andy/forcepageconv] onpladm run job povendor_unl
>
Connecting to onpload, Please wait...
Successful connection to onpload established
...
Points to note:
- the jobs are identical between databases (the onpload db is converted as
part of the v9.21 -> v9.40 upgrade)
- it's always the same tables
- they never work, ever, but all the other tables work, always.
- it's not a table structure problem as, even if we re-build the tables with
ALTER FRAGMENT ... INIT, HPL still fails.
Any ideas or thoughts? It's raised with IBM but I haven't been given the
case number yet.
thanks
Neil Truby t:01932 724027
Director m:07798 811708
Ardenta Limited e:neil.truby@ardenta.com
Some more information on this problem and its symptoms:
We're running a large number (172) of HPL unload jobs one after the
other from the command line using onpladm.
The unloads are being done to /dev/null as we only want the pages to
be read into memory to force the header conversion that is performed
on each page following an "onmode -BC 2" in 9.40.
We found that HPL would not allow us to create a device array that
used /dev/null (either directly or through a symbolic link) so we got
round this by creating an array which used pipes that cat the output
to /dev/null.
After a number of the unloads have completed onpladm will successfully
submit a job:
Connecting to onpload, Please wait...
Successful connection to onpload established
Thu Mar 25 15:06:22 2004
SHMBASE 0xc00000003492f000
CLIENTNUM 0x0000000049010000
Session ID 90
Unload Database -> system
Query Name -> AUTO.90
Device Array -> dummy_arr
Query Mapping -> AUTO.90
Query -> select * from fplbc for read only
Convert Reject -> /tmp/fplbc_unl.rej
and then hang indefinitely with the session reporting something like:
tid name rstcb flags curstk status
5966 sqlexec c0000000175ee258 Y--P--- 131520 cond
wait(netnorm)
At this time a number of af files are produced but the engine will not
crash. There is nothing in the related /tmp/onploaderr file. The af
files have the following message repeated in them:
15:06:23 Found during mt_shm_malloc_segid 5
15:06:23 Pool 'afpool' (0xc000000034a31040)
15:06:23 Bad free block 0xc000000034c915f8
blk-64
c000000034c915b8: 00000000 00000000 00000000 00000000 ........
........
c000000034c915c8: 00000020 00000000 00000005 00000000 ... ....
........
c000000034c915d8: 00000000 00000000 00000000 00000000 ........
........
c000000034c915e8: 00000000 00000000 00000020 00000000 ........ ...
....
If the hanging session is cancelled (either interrupted or killed with
onmode -z) and the job re-submitted it fails with a 255 error,
reporting:
Connecting to onpload, Please wait...
Error connecting to Socket
But no af files are created this time.
The unload jobs have been specified to use a sockets connection. When
the "Error connecting to Socket" is returned the output from netstat
-an |grep 5000 (where 5000 is the port used by this service) looks
like this:
tcp 0 0 193.118.114.33.5000 *.*
LISTEN
tcp 0 0 193.118.114.33.62469 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62470 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62472 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62476 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62477 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62478 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62479 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62480 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62481 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62482 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62484 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62485 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62487 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.62488 193.118.114.33.5000
TIME_WAIT
tcp 0 0 193.118.114.33.5000 193.118.114.33.62483
TIME_WAIT
tcp 0 0 193.118.114.33.62497 193.118.114.33.5000
TIME_WAIT
If I wait until all/most of the TIME_WAIT's have gone (maybe 30
seconds) and re-submit the job it connects fine and runs to completion
without problem.