Re: Handling files over 2Gbytes with the HPL
Posted in 1999
Topics: Stored Procedures & SPL, Platform-Specific Issues, Versions, Editions & End-of-Life
From: "Martyn Hodgson" <martyn.hodgson@eaglestar.co.uk>
>
>We have a small, but growing Data Warehouse (running on a HP K460, HP-UX
>10.20, IDS 7.30). Some of the tables require loading with more than 2Gbytes
>of data at a time. Currently, we perform some very inelegant file splitting
>on these files as the HPL is limited to 2Gbytes per load file. This is
>based
>on:
Excuse me if this shows my considerable ignorance, but how do you get HPUX
10.20 to accept files bigger than 2GB? I tried this the other day and was
rudely refused on a K220.
>1. Getting the logical record length from the formats table of the onpload
>database
>2. Getting the file size from the file system
>3. Calculating the number of records in the file
>4. Using the unix split command to split the file into as few parts as
>possible that ensure each part is under 2 Gbytes.
[split on HPUX10.20 will allow you to specify a target file size. This may
or may not help you]
>5. Generating and then loading device array data into the onpload database
>for these files
>6. Running the load using the files as a device array.
>
>I guess the better solution would be for the programs to generate multiple
>output files so we don't need to split them. However, this is not simple as
>we use a rather inflexible COBOL code generator. Also, the programs don't
>know how big the files are going to be. We would have to always (say),
>generate 5 output files - until a program processes more than 10 Gbytes of
>data.
Can't you determine up front how much data will be generated and then split
based on that?
>OK. So does anyone have any better ideas for handling big file loads, where
>the data comes from a single source?
What is the source system? Some kind of IBM M/F?
Anyway, I'd go with the way you're doing it now, but I've always been a bit
simple-minded. :-)
>Chletenham
.^^^^^^^^^^
I looked on the map, but I couldn't find this place anywhere... :-)
______________________________________________________
Get Your Private, Free Email at http://www.hotmail.com
Obnoxio The Clown wrote: > From: "Martyn Hodgson" <martyn.hodgson@eaglestar.co.uk> > > > >We have a small, but growing Data Warehouse (running on a HP K460, HP-UX > >10.20, IDS 7.30). Some of the tables require loading with more than 2Gbytes > >of data at a time. Currently, we perform some very inelegant file splitting > >on these files as the HPL is limited to 2Gbytes per load file. This is > >based > >on: > > Excuse me if this shows my considerable ignorance, but how do you get HPUX > 10.20 to accept files bigger than 2GB? I tried this the other day and was > rudely refused on a K220. Hmmmm, I haven't played with 10.20 in a while, so let me get back to you on that one. I think you can specify it in the kernel parameters, but thats a different point. One thing you can do is on your outbound Cobol program, there may be several ways you can do this. 1) In JCL, you should be able to check the file size and split it based on the number of records. (You know the record length since its fixed right?) 2) Do a record count and after so many records, you can then split the file. Since its been a long, long time since I played with the Big Iron, you should be able to close the first output file and open a second one within the cobol program. If not, then here's another kludge. You track your position in the output file. After x number of records. You save your position and abend. (You need to track the number of records in the file and your last position.) If your last position isn't the same as the number of records, you re-run the program passing in the last record count so you know how many records you need to ignore before writing to the file. (I think if you index the file, you can jump just to the last position.) [I'm basicly saying that you use a simple process recovery method and just force the abend when the output file gets too large.] You will of course have to get a timestamp in the name of the file so you don't overwrite the file. Now another neat idea, sort your output to match the primary index of your Data Mart. Won't that help speed up the loading process? -Just some food for thought from your Uncle Mikey