Re: Loading ASCII data
Posted in 1994
>From: Ting Barrow <tingb@delphi.com> >Subject: Re: Loading ASCII data >Date: Thu, 24 Nov 94 23:35:29 -0500 >X-Informix-List-Id: <news.9991> >Just a thought: we get same message frequently when the load file has >an incorrect number of pipes - usually because someone typed one in. >Informix also bombs on trying to load the backslash character (octal 134 >in ASCII). Could this be your problem? One way of checking this out, on Unix, is to use the tr and uniq commands: tr -cd '|\\012' < data.file | uniq -c Given the data: aaa|bbbb|ccc|dddd| aab|bbbb|ccc|dddd| aac|bbbb-ccc|dddd| aad|bbbb|ccc|dddd| aae|bbbb|ccc|dddd| It produces the output: 2 |||| 1 ||| 2 |||| The odd line doesn't show up very well here because of the small data file; I've used it where the repetitive counts are numbers in the many thousands of rows, and the odd flaky line shows up with a count of one. You can work out which line is wrong by adding up the numbers as you go. You often get a pair of 1-counts, representing two halves of a single line that was split somehow. Limitations -- it certainly doesn't deal with pipe characters embedded in strings (would anyone care to provide a regular expression (for sed, preferably) which recognises that \\| and \\\\\\| ... are OK, | and \\\\| ... aren't), and it doesn't handle text blobs. You can add an awk script to the end of the pipeline to do the counting for you. awk '{ cum+= $1; printf("%6d %6d %s\\n", cum, $1, $2); }' I find this pipeline, converted to script form, excellent for assessing whether random external data will actually load or whether I need to clean it up first. Yours, Jonathan Leffler (johnl@informix.com) #include <disclaimer.h>