Using High Perf Loader to unload entire database?
Posted in 2008
Topics: Performance & Tuning, Data Types & Schema Design, Migration, Import/Export & Data Conversion
Good morning,
Well, now that we can begun to setup tables using Advanced Data-Types,
specifically LVARCHAR and CLOB types, we can no longer use Informix native
utilities ONUNLOAD and ONLOAD because both of those utilities only support
legacy data-types. So, I have been fuddling around with Informix high
performance loader with onpload and ipload (GUI) and I noticed that I can
only do one table at a time :o( I=B9m totally bummed. I recall something
something from one of the IIUG conference sessions earlier this year that i=
t
was possible to unload and entire DB to file. My question is, has anyone
done that with HPL? Reason why I am asking is that we have a test and
training environment for my programmers with Informix that I refresh the
databases with onload, and now I would like to use HPL to support advanced
data-types.
Thank-you very much in advance for any insight, wisdom, or experience with
this utility.
Jonathan Smaby
Pomona College
---
My favorite quote about the handling of the Iraq war: If you try to fail,
and succeed, which have you done? ~George Carlin (1937-2008)
-------------------------------------------------------------
This message has been scanned by Postini anti-virus software.
=0D
You wrote:
Good morning,
Well, now that we can begun to setup tables using Advanced Data-Types,
specifically LVARCHAR and CLOB types, we can no longer use Informix native
utilities ONUNLOAD and ONLOAD because both of those utilities only support
legacy data-types. So, I have been fuddling around with Informix high
performance loader with onpload and ipload (GUI) and I noticed that I can
only do one table at a time :o( I=B9m totally bummed. I recall something
something from one of the IIUG conference sessions earlier this year that i=
t
was possible to unload and entire DB to file. My question is, has anyone
done that with HPL? Reason why I am asking is that we have a test and
training environment for my programmers with Informix that I refresh the
databases with onload, and now I would like to use HPL to support advanced
data-types.
-----
Consider:
IF only a few tables use the "Advanced Data-Types" then unload those tables
using HPL and then drop the table(s). Then, do your usual migrations thing. On
the target machine, do your usual data import thing and then manually create
the tables of interest and use HPL to load the "special" data.
Haven't do it myself.
HPL is Great! I have used it for single tables before but recently I needed
to do a migration to new hardware and with help from this list I was able to
do a lot of tables at the same time. On my old hardware there was a point
where to many tables slowed things down but you CAN do more than one at a
time using scripts to drive the onpload'er rather than using the ipload gui.
The IBM doc's are weak on the subject but someone sent me a sample script
that would take a list of tables and build the jobs which I could run in any
order or in any quantity.
Attached is an email about the subject and an email with a custom script a
guy sent me that does a real nice job of automating the building of the jobs
to run in parallel.
Tim Ertl
413-442-9000 x6211
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Jonathan Smaby
Sent: Wednesday, August 20, 2008 1:49 PM
To: ids@iiug.org
Subject: Using High Perf Loader to unload entire database? [13164]
Good morning,
Well, now that we can begun to setup tables using Advanced Data-Types,
specifically LVARCHAR and CLOB types, we can no longer use Informix native
utilities ONUNLOAD and ONLOAD because both of those utilities only support
legacy data-types. So, I have been fuddling around with Informix high
performance loader with onpload and ipload (GUI) and I noticed that I can
only do one table at a time :o( I=B9m totally bummed. I recall something
something from one of the IIUG conference sessions earlier this year that i=
t
was possible to unload and entire DB to file. My question is, has anyone
done that with HPL? Reason why I am asking is that we have a test and
training environment for my programmers with Informix that I refresh the
databases with onload, and now I would like to use HPL to support advanced
data-types.
Thank-you very much in advance for any insight, wisdom, or experience with
this utility.
Jonathan Smaby
Pomona College
---
My favorite quote about the handling of the Iraq war: If you try to fail,
and succeed, which have you done? ~George Carlin (1937-2008)
-------------------------------------------------------------
This message has been scanned by Postini anti-virus software.
=0D
****************************************************************************
***
Forum Note: Use "Reply" to post a response in the discussion forum.
Tim,
Glad to hear it worked for you. We are still running 9.40 and 7.31 so the
multiple buffer sizes have never come into play for us. By the way, there
was a series of three articles on HPL in an IBM on-line magazine. You might
find them interesting. Here are the links:
http://www.ibmdatabasemag.com/blog/main/archives/2008/05/the_informix_hi.htm
l
http://www.ibmdatabasemag.com/blog/main/archives/2008/06/the_informix_hi_1.h
tml
http://www.ibmdatabasemag.com/blog/main/archives/2008/07/the_informix_hi_2.h
tml
Rob Schmitz
Embarq Data Management
913-534-3474
rob.b.schmitz@embarq.com
www.embarq.com
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Tim
Ertl
Sent: Sunday, July 27, 2008 3:41 PM
To: ids@iiug.org
Subject: RE: ONPLADM , IPLOAD & ONPLOAD scripts [12915]
Bob, I spent the day today using your utility to perform our conversion in 5
hours TOTAL without removing any data before hand. My previous attempt was
estimated to be over 32hours. This was GREAT! Many Many thanks for your
help!
Tim Ertl
413-442-9000 x6211
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Schmitz, Rob B [EQ]
Sent: Friday, July 25, 2008 12:20 PM
To: ids@iiug.org
Subject: RE: ONPLADM , IPLOAD & ONPLOAD scripts [12904]
Tim,
I have such a script. It is not intended to cover all situations, just those
our group has run into. It serves our needs. You might be able to use it as
is
or you might be able to start with it and then modify going forward. The
script is named hpl_create.ksh and it uses a text file named
hpl_create.list.
The hpl_create.list file has one line per table. Each line contains the name
of the table, the number of devices (unload files), the letter e or d
indicating express or deluxe mode, and the root name of the job and
device(s).
The job name will be the root name with "_job" tacked on the end and the
device name will be the root name with "_dvc" tacked on the end.
For example,
mytab 12 e mytab
Will create both a load and an unload job named mytab_job with 12 unload
files
defined by device name mytab_dvc. The jobs will use express mode.
There are restrictions on using express mode (row size, extended data types,
etc.) and the script performs a few rudimentary checks and may change an
express specification to deluxe if needed.
The script expects you to pass the instance name, the database name, and the
path for the location of your unload files (device path). With all of our
scripts, we source a script that sets our environment for a particular
instance. Also, what we call an instance is essentially the extent name of
the
onconfig file and we use it for many things to identify an instance. It is
different (usually shorter) than the INFORMIXSERVER value.
The first time a job is created for a table with this script, you will see
several errors related to the "delete job" lines in the script. As I say,
this
is a somewhat crude script that meets our needs and it has never been worth
our time to make it squeaky clean.
This script was not written to meet anyone's needs but our own; however, you
may find some useful bits in it. (caveat lector)
Here is the text of the script:
----------------------------------------------------------------------------
-------------
#!/usr/bin/ksh
#---------------------------------------------------------------------------
----
# hpl_create.ksh
#
# Notes on job creation options:
# -flu means create both an unload job and a load job
# adding an "a" means treat data source as a device array
# adding a "c" means deluxe mode
# adding an "N" means it is a deluxe mode without replication
# adding an "n" means no-conversion (this is fastest, must be express)
# Without the "c" option, the mode will default to express
#---------------------------------------------------------------------------
----
if [ "$#" -lt 3 ]
then
echo "USAGE: `basename $0` instance database device_path"
exit 1
else
export INST=$1
export DB=$2
export DEVICE_PATH=$3
fi
... ~informix/infx_env $INST
#---------------------------------------------------------------------------
----
# Create the jobs
#---------------------------------------------------------------------------
----
cat hpl_create.list | tr '{A-Z}' '{a-z}' \\\\
| while read TAB NUM_DEVICES MODE JOBNAME_ROOT
do
echo " unload to hpl_create.tabinfo delimiter ' '
Jonathan,
If you can get away with it, I suggest using the onpladm utility as opposed to
the GUI. Onpladm is so much easier to use. Below are some HPL functions from a
korn shell script I wrote. It is not a load and go script - you'll need to add
some details, and there are some error checking routines and info messages I
print that will not apply for you - but the nuts and bolts are here. If you
are decent with KSH it should be pretty easy. These functions will create an
hpl load job for a table. In my environment I have a load_database script that
builds a list of tables that I want to load and then passes the table names to
the load_table script. The functions below are from the load_table script. In
the load_database script I also manage process load as a single hpl load can
fork multiple processes and on a busy server it could cause a load issue. I
try to keep the # of concurrent HPL loads under 9. But you may be able to get
away with much more. We are somewhat hardware challenged. Anyways if you have
any questions email me (sorry for the format - I dunno how to make it look any
better):
################################################################################
####
# Initial hpl function. This first checks to see if we are loading from
another #
# database server, it then checks to see if the onpload db exists. If the db
does #
# not exist we call setup_hpl to setup all job details. When the job details
are #
# created the onpload db is automatically created. If the onpload db does
exist we #
# check to see if there is a load job defined for this table, if there is not
we #
# do a second check to see if there are old objects left in the onpload db. If
#
# orphans are found they are deleted. After we determine that no objects exist
we #
# call setup_hpl to create a load job. If a job is found with all objects
defined #
# we call run_hpl to start the job. #
################################################################################
####
do_hpl_load()
{
################################################
# Account for loads from other servers. #
################################################
if [[ ${BKUPSERVER} == ${SERVER} ]]
then
JOB_NAME=${DBNAME}:${TABLE}_load
DEV_NAME=${DBNAME}:${TABLE}.load
else
JOB_NAME=${DBNAME}:${TABLE}@${BKUPSERVER}_load
DEV_NAME=${DBNAME}:${TABLE}@${BKUPSERVER}.load
fi
echo "Checking to see if onpload database exists..."
dbaccess sysmaster@${SERVER} - <<EOT!
output to '${WORKDIR}/dbcheck.out' without headings
select count(*) from sysdatabases
where name = "onpload"EOT!
RET=$?
if [[ ${RET} -gt 0 ]]
then
systrack.sh $0 ERROR Cannot connect to sysmaster@${SERVER} to setup HPL job
echo "Unable to connect to sysmaster@${SERVER}" >>
${WORKDIR}/$$_load_${TABLE}.txt
send_mail
fi
DBCOUNT=`cat ${WORKDIR}/dbcheck.out`
if [[ ${DBCOUNT} -ne 1 ]]
then
echo "onpload database does not exist - no jobs defined."
echo "Calling setup_hpl to create load job, onpload db will be automatically
created."
setup_hpl
else
echo "onpload database exists."
echo "Checking to see if ${DBNAME}:${TABLE} has an hpl job defined."
onpladm describe job ${JOB_NAME} -fl > /dev/null 2>&1
RET=$?
if [[ ${RET} -ne 0 ]]
then
echo "No job defined for ${DBNAME}:${TABLE}."
echo "Checking for orphaned hpl devices..."
onpladm describe device ${DEV_NAME} > /dev/null 2>&1
RET=$?
if [[ ${RET} -eq 0 ]]
then
echo "Device found - deleting device..."
cleanup_hpl
else
echo "No devices found. Proceeding to setup hpl job..."
setup_hpl
fi
else
echo "HPL job located - beginning load process..."
run_hpl
fi
fi
}
##################################################################
# Function to create the hpl load job. We first cat strings to #
# a .spec file needed initially for hpl to create the load job. #
# this file will be removed the next time this table is loaded. #
# After the .spec file is created we use onpladm to create the #
# necessary details for the table's load job. Finally we create #
# the actual job with the onpladm utility. Once a job is created #
# for a given table this function will not be called again. Once #
# the details and the job have been created we call run_hpl to #
# start the load process. #
##################################################################
setup_hpl()
{
###################################################################
# Check to see if we need to load from another database's DATADIR #
###################################################################
if [[ ${BKUPSERVER} == ${SERVER} ]]
then
echo "Local database load"
DEV_FILE=${WORKDIR}/${TABLE}.spec.load
touch $DEV_FILE
echo "BEGIN OBJECT DEVICEARRAY ${DEV_NAME} " > ${DEV_FILE}
echo "BEGIN SEQUENCE" >> ${DEV_FILE}
echo "TYPE PIPE" >> ${DEV_FILE}
echo "FILE" >> ${DEV_FILE}
echo "TAPEBLOCKSIZE 0" >> ${DEV_FILE}
echo "TAPEDEVICESIZE 0" >> ${DEV_FILE}
echo 'PIPECOMMAND "gunzip < '${OUTDIR}/${TABLE}'.unl.gz"' >> ${DEV_FILE}
echo "END SEQUENCE" >> ${DEV_FILE}
echo "END OBJECT" >> ${DEV_FILE}
else
echo "Loading from other database - calling vcs_config to set env properly."
DEV_FILE=${WORKDIR}/$TABLE".spec.remote.load"
touch $DEV_FILE
. ${INEXEC}/vcs_config.sh ${BKUPSERVER}
if [ ${ERROR_FLAG} = "Y" ]
then
echo "Unable to set environment to ${BKUPSERVER}" >
${WORKDIR}/$$_load_${TABLE}.txt
send_mail
fi
INDIR=${DATA_ENV_DIR}/informix/${BKUPDBNAME}
INFILE=${INDIR}/${TABLE}.unl.gz
echo "BEGIN OBJECT DEVICEARRAY ${DEV_NAME} " > ${DEV_FILE}
echo "BEGIN SEQUENCE" >> ${DEV_FILE}
echo "TYPE PIPE" >> ${DEV_FILE}
echo "FILE" >> ${DEV_FILE}
echo "TAPEBLOCKSIZE 0" >> ${DEV_FILE}
echo "TAPEDEVICESIZE 0" >> ${DEV_FILE}
echo 'PIPECOMMAND "gunzip < '${INFILE}'"' >> ${DEV_FILE}
echo "END SEQUENCE" >> ${DEV_FILE}
echo "END OBJECT" >> ${DEV_FILE}
. ${INEXEC}/vcs_config.sh ${SERVER}
if [ ${ERROR_FLAG} = "Y" ]
then
echo "Unable to set environment ${SERVER}" > ${WORKDIR}/$$_load_${TABLE}.txt
send_mail
fi
fi
onpladm create object -F ${DEV_FILE}
RET=$?
if [[ ${RET} -ne 0 ]]
then
echo "Problem creating device array. Will attempt to clean up job and try
again."
cleanup_hpl
fi
onpladm create job ${JOB_NAME} -d ${DEV_NAME} -D ${DBNAME} -t ${TABLE} -fla
RET=$?
echo "Return code was ${RET}@@DQ@