Text files
Posted in 2004
Topics: General Discussion
Hi, I'm looking for a way to identify and remove extra line feeds or page eject codes within an input file. Any suggestions would be appreciated. Tony
>>> "Demeis, Tony" <Tony.Demeis@moh.gov.on.ca> 09/30/2004 12:05:03 >>> Hi, I'm looking for a way to identify and remove extra line feeds or page eject codes within an input file. <<< Tony, I once had occasion to remove imbedded CRLF's from an extract file. I used a perl script, which follows. Two things: 1) it was designed around the fact that the file was supposed to contain records with 10 fields separated by 9 pipes, and only one of the fields ("description" - see imbedded comments) might contain extraneous CRLF's, and 2) it was a quick-and-dirty, so I'm not excessively proud of it. That said, it may give you some ideas, so here you are, warts and all: #!/usr/bin/perl $count = @ARGV; if (!count) { print "Usage: perl crstrip.pl inpfile\\ "; exit; } else { $d = $ARGV[0]; } open (FILEIN, $d) || die "Couldn't open $d\\ "; open (FILEOUT, ">bldg_permit_correct.txt") || die "Couldn't open output\\ "; open (FILEBUST, ">bldg_permit_busted.txt") || die "Couldn't open busted\\ "; $lc = 0; $pipepos = 0; $pipeoffset = 0; $busted = 0; $cumcount = 0; @inparray; @outarray; while (<FILEIN>) { chomp $_; $lc += 1; if (($lc % 100000) == 0) { # write occasional stats on stdout print "$lc\\ "; } $pipecount = 0; # index function doesn't work if char occurs at position 0, $_ =~ s/^\\\\|/ \\\\|/; # so insert space $pipepos = index($_, "|"); # count number of pipe characters if ($pipepos > 0 ) { $pipecount = 1; } while ($pipepos > 0) { $pipeoffset = $pipepos + 1; $pipepos = index($_, "|", $pipeoffset); if ($pipepos > 0) { $pipecount += 1; } } if ($pipecount < 9) { # this is a fragmented row $busted += 1; $_ =~ s/\\\\\\\\/ /; # strip spurious escape characters $_ =~ s/\\ //g; # . and crlf pairs $_ =~ s/\\\\r//g; # . and solo cr's @inparray = split('\\\\|', $_); if ($busted > 1) { # assume first field is continuation of description $lastval = pop(@outarray); if ($pipecount == 0) { # sole contents of line? $lastval .= $_; # append to last array item push(@outarray, $lastval); # and resave description } else { $lastval .= $inparray[0]; # append to last array item push(@outarray, $lastval); # resave description shift(@inparray); # description fragment saved, push(@outarray, @inparray); # now save the rest of row } } else { @outarray = @inparray; # start new row in fresh output array } $cumcount += $pipecount; if ($cumcount >= 9) { # 9 pipes means row complete $fullrec = join('|', @outarray); print FILEBUST "$fullrec|\\ "; # write delimited row $cumcount = 0; # reset counters $busted = 0; } } else { # undamaged row print FILEOUT "$_\\ "; # write as is } } HTH, Dick Cunningham Thurston County, WA, USA
You didn't say what OS, but for unix try "dos2unix" or "unix2dos". When I worked on solaris we had such a utility, but not sure where it came from. Take a look at: http://www.iconv.com/dos2unix.htm "Demeis, Tony" <Tony.Demeis@moh.gov.on.ca> Sent by: forum.subscriber@iiug.org 09/30/2004 02:05 PM To: ids@iiug.org cc: Subject: Text files [3507] Hi, I'm looking for a way to identify and remove extra line feeds or page eject codes within an input file. Any suggestions would be appreciated. Tony
Tony, I once had occasion to remove imbedded CRLF's from an extract file. I used a perl script, which follows. Two things: 1) it was designed around the fact that the file was supposed to contain records with 10 fields separated by 9 pipes, and only one of the fields ("description" - see imbedded comments) might contain extraneous CRLF's, and 2) it was a quick-and-dirty, so I'm not excessively proud of it. That said, it may give you some ideas, so here you are, warts and all: #!/usr/bin/perl $count = @ARGV; if (!count) { print "Usage: perl crstrip.pl inpfile\\ "; exit; } else { $d = $ARGV[0]; } open (FILEIN, $d) || die "Couldn't open $d\\ "; open (FILEOUT, ">bldg_permit_correct.txt") || die "Couldn't open output\\ "; open (FILEBUST, ">bldg_permit_busted.txt") || die "Couldn't open busted\\ "; $lc = 0; $pipepos = 0; $pipeoffset = 0; $busted = 0; $cumcount = 0; @inparray; @outarray; while (<FILEIN>) { chomp $_; $lc += 1; if (($lc % 100000) == 0) { # write occasional stats on stdout print "$lc\\ "; } $pipecount = 0; # index function doesn't work if char occurs at position 0, $_ =~ s/^\\\\|/ \\\\|/; # so insert space $pipepos = index($_, "|"); # count number of pipe characters if ($pipepos > 0 ) { $pipecount = 1; } while ($pipepos > 0) { $pipeoffset = $pipepos + 1; $pipepos = index($_, "|", $pipeoffset); if ($pipepos > 0) { $pipecount += 1; } } if ($pipecount < 9) { # this is a fragmented row $busted += 1; $_ =~ s/\\\\\\\\/ /; # strip spurious escape characters $_ =~ s/\\ //g; # . and crlf pairs $_ =~ s/\\\\r//g; # . and solo cr's @inparray = split('\\\\|', $_); if ($busted > 1) { # assume first field is continuation of description $lastval = pop(@outarray); if ($pipecount == 0) { # sole contents of line? $lastval .= $_; # append to last array item push(@outarray, $lastval); # and resave description } else { $lastval .= $inparray[0]; # append to last array item push(@outarray, $lastval); # resave description shift(@inparray); # description fragment saved, push(@outarray, @inparray); # now save the rest of row } } else { @outarray = @inparray; # start new row in fresh output array } $cumcount += $pipecount; if ($cumcount >= 9) { # 9 pipes means row complete $fullrec = join('|', @outarray); print FILEBUST "$fullrec|\\ "; # write delimited row $cumcount = 0; # reset counters $busted = 0; } } else { # undamaged row print FILEOUT "$_\\ "; # write as is } } HTH, Dick Cunningham Thurston County, WA, USA
?? sed/awk? -----Original Message----- From: "Demeis, Tony" <Tony.Demeis@moh.gov.on.ca> To: ids@iiug.org Date: Thu, 30 Sep 2004 15:05:03 -0400 (EDT) Subject: Text files [3507] Hi, I'm looking for a way to identify and remove extra line feeds or page eject codes within an input file. Any suggestions would be appreciated. Tony Jean Sagi jeansagi@myrealbox.com jeansagi@yahoo.com