Handling files over 2Gbytes with the HPL
Posted in 1999
Hi,
We have a small, but growing Data Warehouse (running on a HP K460, HP-UX
10.20, IDS 7.30). Some of the tables require loading with more than 2Gbytes
of data at a time. Currently, we perform some very inelegant file splitting
on these files as the HPL is limited to 2Gbytes per load file. This is based
on:
1. Getting the logical record length from the formats table of the onpload
database
2. Getting the file size from the file system
3. Calculating the number of records in the file
4. Using the unix split command to split the file into as few parts as
possible that ensure each part is under 2 Gbytes.
5. Generating and then loading device array data into the onpload database
for these files
6. Running the load using the files as a device array.
I guess the better solution would be for the programs to generate multiple
output files so we don't need to split them. However, this is not simple as
we use a rather inflexible COBOL code generator. Also, the programs don't
know how big the files are going to be. We would have to always (say),
generate 5 output files - until a program processes more than 10 Gbytes of
data.
I tried writing an intelligent file splitter which I'd hoped would run in
parallel with the hpl, feeding the data via named pipes (thus using more cpu
time, but keeping elapsed time down). However, this fails as the hpl only
supports 'real' files or tapes. The load works fine if the data is written
to files and is then loaded. The pipe support mentioned for the hpl is the
more usual '|' type of pipe.
OK. So does anyone have any better ideas for handling big file loads, where
the data comes from a single source?
Thanks
Martyn Hodgson (martyn.hodgson@eaglestar.co.uk)
Zurich Insurance
Chletenham
UK