sorting dates in ksh
Posted in 2004
Topics: General Discussion
Greetings, I need help figuring out a way to sort out the data sample below. I want to sort it by the second field first, then the thierd field. The goal is to present the data by date then time in ascending order. I'm working in HPUX ksh. I tried sort -k against field 2 and it won't sort. I suspect the forward slashes need to be defined in some way... or somthing. I can't find the answer in the man page or on the Internet. Can anyone provide a sort command that will sort by $2 then $3? Thanks, -Chris Harry 05/04/04 10:45:00 Bill 05/03/04 19:00:00 Sally 05/04/04 06:30:00 Ben 05/03/04 23:30:00 Jimmy 05/04/04 00:30:00 Benny 05/03/04 18:30:00 Mike 05/03/04 20:30:00 Susan 05/03/04 21:00:00 Larry 05/03/04 21:30:00 Joe 05/03/04 22:00:00 Nancy 05/09/04 09:30:00 George 05/09/04 11:00:00 sending to informix-list
chris.staubin@reebok.com wrote: > Greetings, > > I need help figuring out a way to sort out the data sample below. I want > to sort it by the second field first, then the thierd field. The goal is > to present the data by date then time in ascending order. I'm working in > HPUX ksh. > > I tried sort -k against field 2 and it won't sort. I suspect the forward > slashes need to be defined in some way... or somthing. I can't find the > answer in the man page or on the Internet. > > Can anyone provide a sort command that will sort by $2 then $3? > > Thanks, > -Chris > > Harry 05/04/04 10:45:00 > Bill 05/03/04 19:00:00 > Sally 05/04/04 06:30:00 > Ben 05/03/04 23:30:00 > Jimmy 05/04/04 00:30:00 > Benny 05/03/04 18:30:00 > Mike 05/03/04 20:30:00 > Susan 05/03/04 21:00:00 > Larry 05/03/04 21:30:00 > Joe 05/03/04 22:00:00 > Nancy 05/09/04 09:30:00 > George 05/09/04 11:00:00 > sending to informix-list and the relevance to informix is exactly what....
chris.staubin@reebok.com wrote: > > Greetings, > > I need help figuring out a way to sort out the data sample below. I want > to sort it by the second field first, then the thierd field. The goal is > to present the data by date then time in ascending order. I'm working in > HPUX ksh. > > I tried sort -k against field 2 and it won't sort. I suspect the forward > slashes need to be defined in some way... or somthing. I can't find the > answer in the man page or on the Internet. > > Can anyone provide a sort command that will sort by $2 then $3? > > Thanks, > -Chris > > Harry 05/04/04 10:45:00 > Bill 05/03/04 19:00:00 > Sally 05/04/04 06:30:00 > Ben 05/03/04 23:30:00 > Jimmy 05/04/04 00:30:00 > Benny 05/03/04 18:30:00 > Mike 05/03/04 20:30:00 > Susan 05/03/04 21:00:00 > Larry 05/03/04 21:30:00 > Joe 05/03/04 22:00:00 > Nancy 05/09/04 09:30:00 > George 05/09/04 11:00:00 > sending to informix-list untested, but give it a try: sort -t' ' +1 -2 +2 -3 my_input >my_output (this only works if HPUX sort does support the archaic form of what today is -k, but hey, I am old, very old :) dic_k -- Richard Kofler SOLID STATE EDV Dienstleistungen GmbH Vienna/Austria/Europe
chris.staubin@reebok.com wrote: > Greetings, > [SNIP] sort -b +1
chris.staubin@reebok.com wrote: > I need help figuring out a way to sort out the data sample below. I want > to sort it by the second field first, then the thierd field. The goal is > to present the data by date then time in ascending order. I'm working in > HPUX ksh. > > I tried sort -k against field 2 and it won't sort. I suspect the forward > slashes need to be defined in some way... or somthing. I can't find the > answer in the man page or on the Internet. > > Can anyone provide a sort command that will sort by $2 then $3? > > Thanks, > -Chris > > Harry 05/04/04 10:45:00 > Bill 05/03/04 19:00:00 > Sally 05/04/04 06:30:00 > Ben 05/03/04 23:30:00 > Jimmy 05/04/04 00:30:00 > Benny 05/03/04 18:30:00 > Mike 05/03/04 20:30:00 > Susan 05/03/04 21:00:00 > Larry 05/03/04 21:30:00 > Joe 05/03/04 22:00:00 > Nancy 05/09/04 09:30:00 > George 05/09/04 11:00:00 So, some people spent the whole of the run-up to Y2K in a coma? Use 4 digits for the year. Please! :-) It isn't clear which order the fields within the date are ordered; unless the order is year, month, day, it is going to be difficult to sort that data without pre-processing it. Sort assumes that all the fields in a sort are separated by the same delimiter. If you use white space as the delimiter, the second field has to be self-sorting into date order. If, as I suspect, the order is month, day, year within the field, then you have big problems - you can't sort by sub-fields. So, given that any preprocessing is required, you may as well do the job thoroughly. I have to assume you don't have anybody who goes by a non-hyphenated double word first name - no "C J" nor "Nancy Sue" nor ... So, I'd probably do it in Perl - but you want it in ksh, so I'll use ksh to run Perl as well as sort and sed. perl -n -w -e ' chomp; my(@F) = split /\\s+/, $_; my(@D) = split /\\//, $F[1]; $D[2] += 2000; print "$D[2]-$D[0]-$D[1]T$F[2]|$_\\n"; ' "$@" | sort | sed -e 's/[^|]*|//' Chances are that on big data sets, this will work faster than a complex sort criterion -- there was a study done on the behaviour of /bin/sort (at AT&T when AT&T still owned Unix) which showed that simplifying the sort criteria, even using awk as the pre-processor and post-processor, speeded things up a lot. So, what does that gibberish in Perl do? The options tell it to give warnings (-w), to read each line automatically but not print it (-n; -p would print the line, or $_, automatically), and the -e says the Perl script follows. The chomp removes the newline; the first split breaks the data at strings of one or more blanks; the second splits the second field (counting from zero) on slashes; the addition puts 2000 in front of the date (beware strings such as 05/02/99 - they are not handled, so actually the addition is mostly unnecessary), then the print formats the date and time as an ISO 8601 string (which sorts based on alphanumeric criteria) followed by a pipe symbol and the original input. The sort does its stuff; if two people have the same date and time set, then the names go in alphabetic order. The sed strips off the generated sort key. Note that by numbering the input records and sorting them with the record number as the final part of the sort, you achieve a stable sort - which is otherwise extremely hard to do with a quicksort based sort tool. (A stable sort keeps records that compare equal in the order in which they were presented in the original data). -- Jonathan Leffler #include <disclaimer.h> Email: jleffler@earthlink.net, jleffler@us.ibm.com Guardian of DBD::Informix v2003.04 -- http://dbi.perl.org/