AWK question
Posted in 2011
The poster wanted to do the equivalent of "SELECT col1, COUNT(*) ... GROUP BY 1" on a flat Unix text file without loading it into Informix. Resolved: use an awk associative array, e.g. arr[$1]++ for each line, then loop "for (k in arr) print k, arr[k]" in an END block (sample scripts from Jack Parker and Art Kagel). Walt Hultgren offered a simpler alternative, "cut | sort | uniq -c". A follow-up confirmed multiple awk steps can be piped, but combining them into one awk script is cheaper; Walt warned older awk versions have fixed input-record length limits.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: SQL Development & Query Writing, Migration, Import/Export & Data Conversion
If I have a text file in Unix, how can I do the following on the text file
using awk ?
select column1, count(*)
from tablename
group by 1
Sorry, I did not know how else to ask the question, but I am trying to do this
on a text file, without having to load the file into Informix, running the
select and then unloading the result again.
Ie. I want to do this straight in Unix without using Informix. My knowledge of
awk is very basic, ie:
Cat filename | awk '{print $1}' ..... print column 1
Cat filename | awk 's+=$1{print s}' ...... sum of column 1
but I do not know how to do a count, group by in awk.
NOTE: This e-mail message is subject to the MTN Group disclaimer see
http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
Load the values into an array and then count the elements of the array.
{
arr[$1]+=3D1 # Add one to arr[$1} for every occurence
}
END {
for x in (arr) {
print x, arr[x]
}
}
On Oct 11, 2011, at 12:44 PM, Dirk Cornel.... wrote:
> If I have a text file in Unix, how can I do the following on the text =
file=20
> using awk ?=20
>=20
> select column1, count(*)=20
> from tablename=20
> group by 1=20>=20
> Sorry, I did not know how else to ask the question, but I am trying to =
do this=20
> on a text file, without having to load the file into Informix, running =
the=20
> select and then unloading the result again.=20
> Ie. I want to do this straight in Unix without using Informix. My =
knowledge of=20
> awk is very basic, ie:=20
>=20
> Cat filename | awk '{print $1}' ..... print column 1=20
>=20
> Cat filename | awk 's+=3D$1{print s}' ...... sum of column 1=20
>=20
> but I do not know how to do a count, group by in awk.=20
>=20
> NOTE: This e-mail message is subject to the MTN Group disclaimer see=20=
> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx=20
>=20
>=20
> =
**************************************************************************=
*****=20
> Forum Note: Use "Reply" to post a response in the discussion forum.=20=
>=20
Something like this:
awk '
BEGIN {idx=0;}
{
if (name[$1] == $1) {
count[$1] += 1;
} else {
name[$1] = $1;
count[$1] = 1;
}
}
END {
for (aname in name) {
print aname, count[aname];
}
}' filepath
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel.... <moolma_dc@mtn.co.za>wrote:
> If I have a text file in Unix, how can I do the following on the text file
> using awk ?
>
> select column1, count(*)
> from tablename
> group by 1>
> Sorry, I did not know how else to ask the question, but I am trying to do
> this
> on a text file, without having to load the file into Informix, running the
> select and then unloading the result again.
> Ie. I want to do this straight in Unix without using Informix. My knowledge
> of
> awk is very basic, ie:
>
> Cat filename | awk '{print $1}' ..... print column 1
>
> Cat filename | awk 's+=$1{print s}' ...... sum of column 1
>
> but I do not know how to do a count, group by in awk.
>
> NOTE: This e-mail message is subject to the MTN Group disclaimer see
> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--90e6ba6e827a3d918704af08cf7e
My mailer screwed that line up. I'm going to re-type it with a space =
(where there should be none) between each character:
a r r [ $ 1 ] + =3D $ 1
j.
On Oct 11, 2011, at 12:51 PM, Jack Parker wrote:
> Load the values into an array and then count the elements of the =
array.=20
>=20
> {=20
> arr[$1]+=3D3D1 # Add one to arr[$1} for every occurence=20
> }=20
>=20
> END {=20
> for x in (arr) {=20
>=20
> print x, arr[x]=20
> }=20
> }=20
>=20
> On Oct 11, 2011, at 12:44 PM, Dirk Cornel.... wrote:=20
>=20
>> If I have a text file in Unix, how can I do the following on the text =
=3D=20
> file=3D20=20
>> using awk ?=3D20=20
>> =3D20=20
>> select column1, count(*)=3D20=20
>> from tablename=3D20=20
>> group by 1=3D20=20>> =3D20=20
>> Sorry, I did not know how else to ask the question, but I am trying =
to =3D=20
> do this=3D20=20
>> on a text file, without having to load the file into Informix, =
running =3D=20
> the=3D20=20
>> select and then unloading the result again.=3D20=20
>> Ie. I want to do this straight in Unix without using Informix. My =3D=20=
> knowledge of=3D20=20
>> awk is very basic, ie:=3D20=20
>> =3D20=20
>> Cat filename | awk '{print $1}' ..... print column 1=3D20=20
>> =3D20=20
>> Cat filename | awk 's+=3D3D$1{print s}' ...... sum of column 1=3D20=20=
>> =3D20=20
>> but I do not know how to do a count, group by in awk.=3D20=20
>> =3D20=20
>> NOTE: This e-mail message is subject to the MTN Group disclaimer =
see=3D20=3D=20
>=20
>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx=3D20=20
>> =3D20=20
>> =3D20=20
>> =3D=20
> =
**************************************************************************=
=3D=20
> *****=3D20=20
>> Forum Note: Use "Reply" to post a response in the discussion =
forum.=3D20=3D=20
>=20
>> =3D20=20
>=20
>=20
> =
**************************************************************************=
*****=20
> Forum Note: Use "Reply" to post a response in the discussion forum.=20=
>=20
If you are running in dbaccess then do the following:
output to pipe "awk '{print $1}' "
select * from systables
John F. Miller III
STSM, Embedability Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
ids-bounces@iiug.org wrote on 10/11/2011 09:44:23 AM:
> From: "Dirk Cornel...." <moolma_dc@mtn.co.za>
> To: ids@iiug.org
> Date: 10/11/2011 09:55 AM
> Subject: AWK question [25165]
> Sent by: ids-bounces@iiug.org
>
> If I have a text file in Unix, how can I do the following on the text
file
> using awk ?
>
> select column1, count(*)
> from tablename
> group by 1>
> Sorry, I did not know how else to ask the question, but I am trying
> to do this
> on a text file, without having to load the file into Informix, running
the
> select and then unloading the result again.
> Ie. I want to do this straight in Unix without using Informix. My
> knowledge of
> awk is very basic, ie:
>
> Cat filename | awk '{print $1}' ..... print column 1
>
> Cat filename | awk 's+=$1{print s}' ...... sum of column 1
>
> but I do not know how to do a count, group by in awk.
>
> NOTE: This e-mail message is subject to the MTN Group disclaimer see
> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Does it have to be awk? How about a "cut -f1-dx ...| sort | uniq -c" sequence.
Just a thought.
Walt.
On Oct 11, 2011, at 1:11 PM, John Miller iii wrote:
> If you are running in dbaccess then do the following:
>
> output to pipe "awk '{print $1}' "
> select * from systables>
> John F. Miller III
> STSM, Embedability Architect
> miller3@us.ibm.com
> 503-578-5645
> IBM Informix Dynamic Server (IDS)
>
> ids-bounces@iiug.org wrote on 10/11/2011 09:44:23 AM:
>
>> From: "Dirk Cornel...." <moolma_dc@mtn.co.za>
>> To: ids@iiug.org
>> Date: 10/11/2011 09:55 AM
>> Subject: AWK question [25165]
>> Sent by: ids-bounces@iiug.org
>>
>> If I have a text file in Unix, how can I do the following on the text
> file
>> using awk ?
>>
>> select column1, count(*)
>> from tablename
>> group by 1>>
>> Sorry, I did not know how else to ask the question, but I am trying
>> to do this
>> on a text file, without having to load the file into Informix, running
> the
>> select and then unloading the result again.
>> Ie. I want to do this straight in Unix without using Informix. My
>> knowledge of
>> awk is very basic, ie:
>>
>> Cat filename | awk '{print $1}' ..... print column 1
>>
>> Cat filename | awk 's+=$1{print s}' ...... sum of column 1
>>
>> but I do not know how to do a count, group by in awk.
>>
>> NOTE: This e-mail message is subject to the MTN Group disclaimer see
>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
>>
>>
>>
>
>
*******************************************************************************
>
>> Forum Note: Use "Reply" to post a response in the discussion forum.
>>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
--
Walt Hultgren
walt@altasenta.com
No, it doesn't have to be awk, any method will do :-)
Thanks very much for all the replies.
> -----Original Message-----
> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
> Walt Hultgren
> Sent: Tuesday, 11 October 2011 08:19 PM
> To: ids@iiug.org
> Subject: Re: AWK question [25170]
>
> Does it have to be awk? How about a "cut -f1-dx ...| sort | uniq -c"
> sequence.
> Just a thought.
>
> Walt.
>
> On Oct 11, 2011, at 1:11 PM, John Miller iii wrote:
>
> > If you are running in dbaccess then do the following:
> >
> > output to pipe "awk '{print $1}' "
> > select * from systables> >
> > John F. Miller III
> > STSM, Embedability Architect
> > miller3@us.ibm.com
> > 503-578-5645
> > IBM Informix Dynamic Server (IDS)
> >
> > ids-bounces@iiug.org wrote on 10/11/2011 09:44:23 AM:
> >
> >> From: "Dirk Cornel...." <moolma_dc@mtn.co.za>
> >> To: ids@iiug.org
> >> Date: 10/11/2011 09:55 AM
> >> Subject: AWK question [25165]
> >> Sent by: ids-bounces@iiug.org
> >>
> >> If I have a text file in Unix, how can I do the following on the
> text
> > file
> >> using awk ?
> >>
> >> select column1, count(*)
> >> from tablename
> >> group by 1> >>
> >> Sorry, I did not know how else to ask the question, but I am trying
> >> to do this
> >> on a text file, without having to load the file into Informix,
> running
> > the
> >> select and then unloading the result again.
> >> Ie. I want to do this straight in Unix without using Informix. My
> >> knowledge of
> >> awk is very basic, ie:
> >>
> >> Cat filename | awk '{print $1}' ..... print column 1
> >>
> >> Cat filename | awk 's+=$1{print s}' ...... sum of column 1
> >>
> >> but I do not know how to do a count, group by in awk.
> >>
> >> NOTE: This e-mail message is subject to the MTN Group disclaimer see
> >> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> >>
> >>
> >>
> >
> >
> ***********************************************************************
> ********
> >
> >> Forum Note: Use "Reply" to post a response in the discussion forum.
> >>
> >
> >
> >
> ***********************************************************************
> ********
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
>
> --
> Walt Hultgren
> walt@altasenta.com
>
>
> ***********************************************************************
> ********
> Forum Note: Use "Reply" to post a response in the discussion forum.
NOTE: This e-mail message is subject to the MTN Group disclaimer see
http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
Last question ... PS. I love Unix :-)
I assume if I have multiple awk "functions" like the one below, I can pipe one
awk to the next ?
Eg.
Awk ' ..... do my stuff ..... '
|
Awk ' ..... do some more stuff ..... '
> -----Original Message-----
> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
> Art Kagel
> Sent: Tuesday, 11 October 2011 06:58 PM
> To: ids@iiug.org
> Subject: Re: AWK question [25167]
>
> Something like this:
>
> awk '
> BEGIN {idx=0;}
> {
>
> if (name[$1] == $1) {
>
> count[$1] += 1;
>
> } else {
>
> name[$1] = $1;
>
> count[$1] = 1;
>
> }
> }
> END {
>
> for (aname in name) {
>
> print aname, count[aname];
>
> }
> }' filepath
>
> Art
>
> Art S. Kagel
> Advanced DataTools (www.advancedatatools.com)
> Blog: http://informix-myview.blogspot.com/
>
> Disclaimer: Please keep in mind that my own opinions are my own
> opinions and
> do not reflect on my employer, Advanced DataTools, the IIUG, nor any
> other
> organization with which I am associated either explicitly, implicitly,
> or by
> inference. Neither do those opinions reflect those of other individuals
> affiliated with any entity with which I am affiliated nor those of the
> entities themselves.
>
> On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel....
> <moolma_dc@mtn.co.za>wrote:
>
> > If I have a text file in Unix, how can I do the following on the text
> file
> > using awk ?
> >
> > select column1, count(*)
> > from tablename
> > group by 1> >
> > Sorry, I did not know how else to ask the question, but I am trying
> to do
> > this
> > on a text file, without having to load the file into Informix,
> running the
> > select and then unloading the result again.
> > Ie. I want to do this straight in Unix without using Informix. My
> knowledge
> > of
> > awk is very basic, ie:
> >
> > Cat filename | awk '{print $1}' ..... print column 1
> >
> > Cat filename | awk 's+=$1{print s}' ...... sum of column 1
> >
> > but I do not know how to do a count, group by in awk.
> >
> > NOTE: This e-mail message is subject to the MTN Group disclaimer see
> > http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> >
> >
> >
> >
> ***********************************************************************
> ********
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
> --90e6ba6e827a3d918704af08cf7e
>
>
> ***********************************************************************
> ********
> Forum Note: Use "Reply" to post a response in the discussion forum.
NOTE: This e-mail message is subject to the MTN Group disclaimer see
http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
Totally, but you can wrap it all into one awk program if you like and =
do:
whatever | awk -f progrm.awk
j.=20
On Oct 12, 2011, at 5:01 AM, Dirk Cornel.... wrote:
> Last question ... PS. I love Unix :-)=20
>=20
> I assume if I have multiple awk "functions" like the one below, I can =
pipe one=20
> awk to the next ?=20
>=20
> Eg.=20
>=20
> Awk ' ..... do my stuff ..... '=20
>=20
> |=20
>=20
> Awk ' ..... do some more stuff ..... '=20
>=20
>> -----Original Message-----=20
>> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of=20=
>> Art Kagel=20
>> Sent: Tuesday, 11 October 2011 06:58 PM=20
>> To: ids@iiug.org=20
>> Subject: Re: AWK question [25167]=20
>>=20
>> Something like this:=20
>>=20
>> awk '=20
>> BEGIN {idx=3D0;}=20
>> {=20
>>=20
>> if (name[$1] =3D=3D $1) {=20
>>=20
>> count[$1] +=3D 1;=20
>>=20
>> } else {=20
>>=20
>> name[$1] =3D $1;=20
>>=20
>> count[$1] =3D 1;=20
>>=20
>> }=20
>> }=20
>> END {=20
>>=20
>> for (aname in name) {=20
>>=20
>> print aname, count[aname];=20
>>=20
>> }=20
>> }' filepath=20
>>=20
>> Art=20
>>=20
>> Art S. Kagel=20
>> Advanced DataTools (www.advancedatatools.com)=20
>> Blog: http://informix-myview.blogspot.com/=20
>>=20
>> Disclaimer: Please keep in mind that my own opinions are my own=20
>> opinions and=20
>> do not reflect on my employer, Advanced DataTools, the IIUG, nor any=20=
>> other=20
>> organization with which I am associated either explicitly, =
implicitly,=20
>> or by=20
>> inference. Neither do those opinions reflect those of other =
individuals=20
>> affiliated with any entity with which I am affiliated nor those of =
the=20
>> entities themselves.=20
>>=20
>> On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel....=20
>> <moolma_dc@mtn.co.za>wrote:=20
>>=20
>>> If I have a text file in Unix, how can I do the following on the =
text=20
>> file=20
>>> using awk ?=20
>>>=20
>>> select column1, count(*)=20
>>> from tablename=20
>>> group by 1=20>>>=20
>>> Sorry, I did not know how else to ask the question, but I am trying=20=
>> to do=20
>>> this=20
>>> on a text file, without having to load the file into Informix,=20
>> running the=20
>>> select and then unloading the result again.=20
>>> Ie. I want to do this straight in Unix without using Informix. My=20
>> knowledge=20
>>> of=20
>>> awk is very basic, ie:=20
>>>=20
>>> Cat filename | awk '{print $1}' ..... print column 1=20
>>>=20
>>> Cat filename | awk 's+=3D$1{print s}' ...... sum of column 1=20
>>>=20
>>> but I do not know how to do a count, group by in awk.=20
>>>=20
>>> NOTE: This e-mail message is subject to the MTN Group disclaimer see=20=
>>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx=20
>>>=20
>>>=20
>>>=20
>>>=20
>> =
***********************************************************************=20=
>> ********=20
>>> Forum Note: Use "Reply" to post a response in the discussion forum.=20=
>>>=20
>>>=20
>>=20
>> --90e6ba6e827a3d918704af08cf7e=20
>>=20
>>=20
>> =
***********************************************************************=20=
>> ********=20
>> Forum Note: Use "Reply" to post a response in the discussion forum.=20=
>=20
> NOTE: This e-mail message is subject to the MTN Group disclaimer see=20=
> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx=20
>=20
>=20
> =
**************************************************************************=
*****=20
> Forum Note: Use "Reply" to post a response in the discussion forum.=20=
>=20
Yes. Or just combine the functionality into one script. Any text that
passes one braced section falls through to the next one unless the section
ends with a 'next;' statement. So:
cat file | awk ' {some stuff}' | awk '{other stuff}'
is equivalent to:
cat file | awk '{some stuff'} {other stuff}'
with less overhead because only one copy of awk is running.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Wed, Oct 12, 2011 at 5:01 AM, Dirk Cornel.... <moolma_dc@mtn.co.za>wrote:
> Last question ... PS. I love Unix :-)
>
> I assume if I have multiple awk "functions" like the one below, I can pipe
> one
> awk to the next ?
>
> Eg.
>
> Awk ' ..... do my stuff ..... '
>
> |
>
> Awk ' ..... do some more stuff ..... '
>
> > -----Original Message-----
> > From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
> > Art Kagel
> > Sent: Tuesday, 11 October 2011 06:58 PM
> > To: ids@iiug.org
> > Subject: Re: AWK question [25167]
> >
> > Something like this:
> >
> > awk '
> > BEGIN {idx=0;}
> > {
> >
> > if (name[$1] == $1) {
> >
> > count[$1] += 1;
> >
> > } else {
> >
> > name[$1] = $1;
> >
> > count[$1] = 1;
> >
> > }
> > }
> > END {
> >
> > for (aname in name) {
> >
> > print aname, count[aname];
> >
> > }
> > }' filepath
> >
> > Art
> >
> > Art S. Kagel
> > Advanced DataTools (www.advancedatatools.com)
> > Blog: http://informix-myview.blogspot.com/
> >
> > Disclaimer: Please keep in mind that my own opinions are my own
> > opinions and
> > do not reflect on my employer, Advanced DataTools, the IIUG, nor any
> > other
> > organization with which I am associated either explicitly, implicitly,
> > or by
> > inference. Neither do those opinions reflect those of other individuals
> > affiliated with any entity with which I am affiliated nor those of the
> > entities themselves.
> >
> > On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel....
> > <moolma_dc@mtn.co.za>wrote:
> >
> > > If I have a text file in Unix, how can I do the following on the text
> > file
> > > using awk ?
> > >
> > > select column1, count(*)
> > > from tablename
> > > group by 1> > >
> > > Sorry, I did not know how else to ask the question, but I am trying
> > to do
> > > this
> > > on a text file, without having to load the file into Informix,
> > running the
> > > select and then unloading the result again.
> > > Ie. I want to do this straight in Unix without using Informix. My
> > knowledge
> > > of
> > > awk is very basic, ie:
> > >
> > > Cat filename | awk '{print $1}' ..... print column 1
> > >
> > > Cat filename | awk 's+=$1{print s}' ...... sum of column 1
> > >
> > > but I do not know how to do a count, group by in awk.
> > >
> > > NOTE: This e-mail message is subject to the MTN Group disclaimer see
> > > http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> > >
> > >
> > >
> > >
> > ***********************************************************************
> > ********
> > > Forum Note: Use "Reply" to post a response in the discussion forum.
> > >
> > >
> >
> > --90e6ba6e827a3d918704af08cf7e
> >
> >
> > ***********************************************************************
> > ********
> > Forum Note: Use "Reply" to post a response in the discussion forum.
>
> NOTE: This e-mail message is subject to the MTN Group disclaimer see
> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--90e6ba21219b49245304af182e4f
Great, thank you. I knew some very basic one liners in awk, but awk is a whole
language in itself. Very cool. Did my 1st multiple line script yesterday.
Dirk
> -----Original Message-----
> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
> Art Kagel
> Sent: Wednesday, 12 October 2011 01:19 PM
> To: ids@iiug.org
> Subject: Re: AWK question [25177]
>
> Yes. Or just combine the functionality into one script. Any text that
> passes one braced section falls through to the next one unless the
> section
> ends with a 'next;' statement. So:
>
> cat file | awk ' {some stuff}' | awk '{other stuff}'
>
> is equivalent to:
>
> cat file | awk '{some stuff'} {other stuff}'
>
> with less overhead because only one copy of awk is running.
>
> Art
>
> Art S. Kagel
> Advanced DataTools (www.advancedatatools.com)
> Blog: http://informix-myview.blogspot.com/
>
> Disclaimer: Please keep in mind that my own opinions are my own
> opinions and
> do not reflect on my employer, Advanced DataTools, the IIUG, nor any
> other
> organization with which I am associated either explicitly, implicitly,
> or by
> inference. Neither do those opinions reflect those of other individuals
> affiliated with any entity with which I am affiliated nor those of the
> entities themselves.
>
> On Wed, Oct 12, 2011 at 5:01 AM, Dirk Cornel....
> <moolma_dc@mtn.co.za>wrote:
>
> > Last question ... PS. I love Unix :-)
> >
> > I assume if I have multiple awk "functions" like the one below, I can
> pipe
> > one
> > awk to the next ?
> >
> > Eg.
> >
> > Awk ' ..... do my stuff ..... '
> >
> > |
> >
> > Awk ' ..... do some more stuff ..... '
> >
> > > -----Original Message-----
> > > From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf
> Of
> > > Art Kagel
> > > Sent: Tuesday, 11 October 2011 06:58 PM
> > > To: ids@iiug.org
> > > Subject: Re: AWK question [25167]
> > >
> > > Something like this:
> > >
> > > awk '
> > > BEGIN {idx=0;}
> > > {
> > >
> > > if (name[$1] == $1) {
> > >
> > > count[$1] += 1;
> > >
> > > } else {
> > >
> > > name[$1] = $1;
> > >
> > > count[$1] = 1;
> > >
> > > }
> > > }
> > > END {
> > >
> > > for (aname in name) {
> > >
> > > print aname, count[aname];
> > >
> > > }
> > > }' filepath
> > >
> > > Art
> > >
> > > Art S. Kagel
> > > Advanced DataTools (www.advancedatatools.com)
> > > Blog: http://informix-myview.blogspot.com/
> > >
> > > Disclaimer: Please keep in mind that my own opinions are my own
> > > opinions and
> > > do not reflect on my employer, Advanced DataTools, the IIUG, nor
> any
> > > other
> > > organization with which I am associated either explicitly,
> implicitly,
> > > or by
> > > inference. Neither do those opinions reflect those of other
> individuals
> > > affiliated with any entity with which I am affiliated nor those of
> the
> > > entities themselves.
> > >
> > > On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel....
> > > <moolma_dc@mtn.co.za>wrote:
> > >
> > > > If I have a text file in Unix, how can I do the following on the
> text
> > > file
> > > > using awk ?
> > > >
> > > > select column1, count(*)
> > > > from tablename
> > > > group by 1> > > >
> > > > Sorry, I did not know how else to ask the question, but I am
> trying
> > > to do
> > > > this
> > > > on a text file, without having to load the file into Informix,
> > > running the
> > > > select and then unloading the result again.
> > > > Ie. I want to do this straight in Unix without using Informix. My
> > > knowledge
> > > > of
> > > > awk is very basic, ie:
> > > >
> > > > Cat filename | awk '{print $1}' ..... print column 1
> > > >
> > > > Cat filename | awk 's+=$1{print s}' ...... sum of column 1
> > > >
> > > > but I do not know how to do a count, group by in awk.
> > > >
> > > > NOTE: This e-mail message is subject to the MTN Group disclaimer
> see
> > > > http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> > > >
> > > >
> > > >
> > > >
> > >
> ***********************************************************************
> > > ********
> > > > Forum Note: Use "Reply" to post a response in the discussion
> forum.
> > > >
> > > >
> > >
> > > --90e6ba6e827a3d918704af08cf7e
> > >
> > >
> > >
> ***********************************************************************
> > > ********
> > > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> > NOTE: This e-mail message is subject to the MTN Group disclaimer see
> > http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> >
> >
> >
> >
> ***********************************************************************
> ********
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
> --90e6ba21219b49245304af182e4f
>
>
> ***********************************************************************
> ********
> Forum Note: Use "Reply" to post a response in the discussion forum.
NOTE: This e-mail message is subject to the MTN Group disclaimer see
http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
On Oct 12, 2011, at 8:10 AM, Dirk Cornel.... wrote:
> Great, thank you. I knew some very basic one liners in awk, but awk is a
whole
> language in itself. Very cool. Did my 1st multiple line script yesterday.
>
> Dirk
Congratulations! You are now an official "developer" (if you weren't one
already). :-)
One more point about awk: I'm not familiar with all of the various
implementations of awk, but at least the older ones tended to have what we
would now consider fairly low limitations on the maximum allowed length of any
one input record. Those limits would not be dynamically resized, either. You
got what you got. This may not apply to the awk in in your environment (awk,
nawk, gawk, ...), but it's something to be aware of.
I don't recall seeing what your overall application is. If what you're
planning will be mission critical and/or subject to very long input records, I
suggest you try some stress testing with what you consider a worst case set of
input.
HTH,
Walt.
>
>> -----Original Message-----
>> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
>> Art Kagel
>> Sent: Wednesday, 12 October 2011 01:19 PM
>> To: ids@iiug.org
>> Subject: Re: AWK question [25177]
>>
>> Yes. Or just combine the functionality into one script. Any text that
>> passes one braced section falls through to the next one unless the
>> section
>> ends with a 'next;' statement. So:
>>
>> cat file | awk ' {some stuff}' | awk '{other stuff}'
>>
>> is equivalent to:
>>
>> cat file | awk '{some stuff'} {other stuff}'
>>
>> with less overhead because only one copy of awk is running.
>>
>> Art
>>
>> Art S. Kagel
>> Advanced DataTools (www.advancedatatools.com)
>> Blog: http://informix-myview.blogspot.com/
>>
>> Disclaimer: Please keep in mind that my own opinions are my own
>> opinions and
>> do not reflect on my employer, Advanced DataTools, the IIUG, nor any
>> other
>> organization with which I am associated either explicitly, implicitly,
>> or by
>> inference. Neither do those opinions reflect those of other individuals
>> affiliated with any entity with which I am affiliated nor those of the
>> entities themselves.
>>
>> On Wed, Oct 12, 2011 at 5:01 AM, Dirk Cornel....
>> <moolma_dc@mtn.co.za>wrote:
>>
>>> Last question ... PS. I love Unix :-)
>>>
>>> I assume if I have multiple awk "functions" like the one below, I can
>> pipe
>>> one
>>> awk to the next ?
>>>
>>> Eg.
>>>
>>> Awk ' ..... do my stuff ..... '
>>>
>>> |
>>>
>>> Awk ' ..... do some more stuff ..... '
>>>
>>>> -----Original Message-----
>>>> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf
>> Of
>>>> Art Kagel
>>>> Sent: Tuesday, 11 October 2011 06:58 PM
>>>> To: ids@iiug.org
>>>> Subject: Re: AWK question [25167]
>>>>
>>>> Something like this:
>>>>
>>>> awk '
>>>> BEGIN {idx=0;}
>>>> {
>>>>
>>>> if (name[$1] == $1) {
>>>>
>>>> count[$1] += 1;
>>>>
>>>> } else {
>>>>
>>>> name[$1] = $1;
>>>>
>>>> count[$1] = 1;
>>>>
>>>> }
>>>> }
>>>> END {
>>>>
>>>> for (aname in name) {
>>>>
>>>> print aname, count[aname];
>>>>
>>>> }
>>>> }' filepath
>>>>
>>>> Art
>>>>
>>>> Art S. Kagel
>>>> Advanced DataTools (www.advancedatatools.com)
>>>> Blog: http://informix-myview.blogspot.com/
>>>>
>>>> Disclaimer: Please keep in mind that my own opinions are my own
>>>> opinions and
>>>> do not reflect on my employer, Advanced DataTools, the IIUG, nor
>> any
>>>> other
>>>> organization with which I am associated either explicitly,
>> implicitly,
>>>> or by
>>>> inference. Neither do those opinions reflect those of other
>> individuals
>>>> affiliated with any entity with which I am affiliated nor those of
>> the
>>>> entities themselves.
>>>>
>>>> On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel....
>>>> <moolma_dc@mtn.co.za>wrote:
>>>>
>>>>> If I have a text file in Unix, how can I do the following on the
>> text
>>>> file
>>>>> using awk ?
>>>>>
>>>>> select column1, count(*)
>>>>> from tablename
>>>>> group by 1>>>>>
>>>>> Sorry, I did not know how else to ask the question, but I am
>> trying
>>>> to do
>>>>> this
>>>>> on a text file, without having to load the file into Informix,
>>>> running the
>>>>> select and then unloading the result again.
>>>>> Ie. I want to do this straight in Unix without using Informix. My
>>>> knowledge
>>>>> of
>>>>> awk is very basic, ie:
>>>>>
>>>>> Cat filename | awk '{print $1}' ..... print column 1
>>>>>
>>>>> Cat filename | awk 's+=$1{print s}' ...... sum of column 1
>>>>>
>>>>> but I do not know how to do a count, group by in awk.
>>>>>
>>>>> NOTE: This e-mail message is subject to the MTN Group disclaimer
>> see
>>>>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
>>>>>
>>>>>
>>>>>
>>>>>
>>>>
>> ***********************************************************************
>>>> ********
>>>>> Forum Note: Use "Reply" to post a response in the discussion
>> forum.
>>>>>
>>>>>
>>>>
>>>> --90e6ba6e827a3d918704af08cf7e
>>>>
>>>>
>>>>
>> ***********************************************************************
>>>> ********
>>>> Forum Note: Use "Reply" to post a response in the discussion forum.
>>>
>>> NOTE: This e-mail message is subject to the MTN Group disclaimer see
>>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
>>>
>>>
>>>
>>>
>> ***********************************************************************
>> ********
>>> Forum Note: Use "Reply" to post a response in the discussion forum.
>>>
>>>
>>
>> --90e6ba21219b49245304af182e4f
>>
>>
>> ***********************************************************************
>> ********
>> Forum Note: Use "Reply" to post a response in the discussion forum.
>
> NOTE: This e-mail message is subject to the MTN Group disclaimer see
> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
--
Walt Hultgren
walt@altasenta.com
Good point Walt. The worst that I've seen are the awk's on Solaris and AIX
which also have some old bugs like not properly supporting input variables
(-v varname=value). On Solaris you usually also have nawk which works fine
and some installations of Solaris have GNU awk (gawk) which is available for
AIX as well. Nawk and gawk have many extensions to the basic awk language
as well.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Wed, Oct 12, 2011 at 10:08 AM, Walt Hultgren <walt@altasenta.com> wrote:
> On Oct 12, 2011, at 8:10 AM, Dirk Cornel.... wrote:
>
> > Great, thank you. I knew some very basic one liners in awk, but awk is a
> whole
> > language in itself. Very cool. Did my 1st multiple line script yesterday.
> >
> > Dirk
>
> Congratulations! You are now an official "developer" (if you weren't one
> already). :-)
>
> One more point about awk: I'm not familiar with all of the various
> implementations of awk, but at least the older ones tended to have what we
> would now consider fairly low limitations on the maximum allowed length of
> any
> one input record. Those limits would not be dynamically resized, either.
> You
> got what you got. This may not apply to the awk in in your environment
> (awk,
> nawk, gawk, ...), but it's something to be aware of.
>
> I don't recall seeing what your overall application is. If what you're
> planning will be mission critical and/or subject to very long input
> records, I
> suggest you try some stress testing with what you consider a worst case set
> of
> input.
>
> HTH,
>
> Walt.
>
> >
> >> -----Original Message-----
> >> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
> >> Art Kagel
> >> Sent: Wednesday, 12 October 2011 01:19 PM
> >> To: ids@iiug.org
> >> Subject: Re: AWK question [25177]
> >>
> >> Yes. Or just combine the functionality into one script. Any text that
> >> passes one braced section falls through to the next one unless the
> >> section
> >> ends with a 'next;' statement. So:
> >>
> >> cat file | awk ' {some stuff}' | awk '{other stuff}'
> >>
> >> is equivalent to:
> >>
> >> cat file | awk '{some stuff'} {other stuff}'
> >>
> >> with less overhead because only one copy of awk is running.
> >>
> >> Art
> >>
> >> Art S. Kagel
> >> Advanced DataTools (www.advancedatatools.com)
> >> Blog: http://informix-myview.blogspot.com/
> >>
> >> Disclaimer: Please keep in mind that my own opinions are my own
> >> opinions and
> >> do not reflect on my employer, Advanced DataTools, the IIUG, nor any
> >> other
> >> organization with which I am associated either explicitly, implicitly,
> >> or by
> >> inference. Neither do those opinions reflect those of other individuals
> >> affiliated with any entity with which I am affiliated nor those of the
> >> entities themselves.
> >>
> >> On Wed, Oct 12, 2011 at 5:01 AM, Dirk Cornel....
> >> <moolma_dc@mtn.co.za>wrote:
> >>
> >>> Last question ... PS. I love Unix :-)
> >>>
> >>> I assume if I have multiple awk "functions" like the one below, I can
> >> pipe
> >>> one
> >>> awk to the next ?
> >>>
> >>> Eg.
> >>>
> >>> Awk ' ..... do my stuff ..... '
> >>>
> >>> |
> >>>
> >>> Awk ' ..... do some more stuff ..... '
> >>>
> >>>> -----Original Message-----
> >>>> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf
> >> Of
> >>>> Art Kagel
> >>>> Sent: Tuesday, 11 October 2011 06:58 PM
> >>>> To: ids@iiug.org
> >>>> Subject: Re: AWK question [25167]
> >>>>
> >>>> Something like this:
> >>>>
> >>>> awk '
> >>>> BEGIN {idx=0;}
> >>>> {
> >>>>
> >>>> if (name[$1] == $1) {
> >>>>
> >>>> count[$1] += 1;
> >>>>
> >>>> } else {
> >>>>
> >>>> name[$1] = $1;
> >>>>
> >>>> count[$1] = 1;
> >>>>
> >>>> }
> >>>> }
> >>>> END {
> >>>>
> >>>> for (aname in name) {
> >>>>
> >>>> print aname, count[aname];
> >>>>
> >>>> }
> >>>> }' filepath
> >>>>
> >>>> Art
> >>>>
> >>>> Art S. Kagel
> >>>> Advanced DataTools (www.advancedatatools.com)
> >>>> Blog: http://informix-myview.blogspot.com/
> >>>>
> >>>> Disclaimer: Please keep in mind that my own opinions are my own
> >>>> opinions and
> >>>> do not reflect on my employer, Advanced DataTools, the IIUG, nor
> >> any
> >>>> other
> >>>> organization with which I am associated either explicitly,
> >> implicitly,
> >>>> or by
> >>>> inference. Neither do those opinions reflect those of other
> >> individuals
> >>>> affiliated with any entity with which I am affiliated nor those of
> >> the
> >>>> entities themselves.
> >>>>
> >>>> On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel....
> >>>> <moolma_dc@mtn.co.za>wrote:
> >>>>
> >>>>> If I have a text file in Unix, how can I do the following on the
> >> text
> >>>> file
> >>>>> using awk ?
> >>>>>
> >>>>> select column1, count(*)
> >>>>> from tablename
> >>>>> group by 1> >>>>>
> >>>>> Sorry, I did not know how else to ask the question, but I am
> >> trying
> >>>> to do
> >>>>> this
> >>>>> on a text file, without having to load the file into Informix,
> >>>> running the
> >>>>> select and then unloading the result again.
> >>>>> Ie. I want to do this straight in Unix without using Informix. My
> >>>> knowledge
> >>>>> of
> >>>>> awk is very basic, ie:
> >>>>>
> >>>>> Cat filename | awk '{print $1}' ..... print column 1
> >>>>>
> >>>>> Cat filename | awk 's+=$1{print s}' ...... sum of column 1
> >>>>>
> >>>>> but I do not know how to do a count, group by in awk.
> >>>>>
> >>>>> NOTE: This e-mail message is subject to the MTN Group disclaimer
> >> see
> >>>>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> >>>>>
> >>>>>
> >>>>>
> >>>>>
> >>>>
> >> ***********************************************************************
> >>>> ********
> >>>>> Forum Note: Use "Reply" to post a response in the discussion
> >> forum.
> >>>>>
> >>>>>
> >>>>
> >>>> --90e6ba6e827a3d918704af08cf7e
> >>>>
> >>>>
> >>>>
> >> ***********************************************************************
> >>>> ********
> >>>> Forum Note: Use "Reply" to post a response in the discussion forum.
> >>>
> >>> NOTE: This e-mail message is subject to the MTN Group disclaimer see
> >>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> >>>
> >>>
> >>>
> >>>
> >> ***********************************************************************
> >> ********
> >>> Forum Note: Use "Reply"
awk , 5 minutes to learn, months to master. What's nice is that you can =
productive with it over that time, and the mastery is only months, not =
years. Not as powerful as perl, but then the doc for awk is a few pages =
long instead of a phone book.
j.
On Oct 12, 2011, at 10:42 AM, Art Kagel wrote:
> Good point Walt. The worst that I've seen are the awk's on Solaris and =
AIX=20
> which also have some old bugs like not properly supporting input =
variables=20
> (-v varname=3Dvalue). On Solaris you usually also have nawk which =
works fine=20
> and some installations of Solaris have GNU awk (gawk) which is =
available for=20
> AIX as well. Nawk and gawk have many extensions to the basic awk =
language=20
> as well.=20
>=20
> Art=20
>=20
> Art S. Kagel=20
> Advanced DataTools (www.advancedatatools.com)=20
> Blog: http://informix-myview.blogspot.com/=20
>=20
> Disclaimer: Please keep in mind that my own opinions are my own =
opinions and=20
> do not reflect on my employer, Advanced DataTools, the IIUG, nor any =
other=20
> organization with which I am associated either explicitly, implicitly, =
or by=20
> inference. Neither do those opinions reflect those of other =
individuals=20
> affiliated with any entity with which I am affiliated nor those of the=20=
> entities themselves.=20
>=20
> On Wed, Oct 12, 2011 at 10:08 AM, Walt Hultgren <walt@altasenta.com> =
wrote:=20
>=20
>> On Oct 12, 2011, at 8:10 AM, Dirk Cornel.... wrote:=20
>>=20
>>> Great, thank you. I knew some very basic one liners in awk, but awk =
is a=20
>> whole=20
>>> language in itself. Very cool. Did my 1st multiple line script =
yesterday.=20
>>>=20
>>> Dirk=20
>>=20
>> Congratulations! You are now an official "developer" (if you weren't =
one=20
>> already). :-)=20
>>=20
>> One more point about awk: I'm not familiar with all of the various=20
>> implementations of awk, but at least the older ones tended to have =
what we=20
>> would now consider fairly low limitations on the maximum allowed =
length of=20
>> any=20
>> one input record. Those limits would not be dynamically resized, =
either.=20
>> You=20
>> got what you got. This may not apply to the awk in in your =
environment=20
>> (awk,=20
>> nawk, gawk, ...), but it's something to be aware of.=20
>>=20
>> I don't recall seeing what your overall application is. If what =
you're=20
>> planning will be mission critical and/or subject to very long input=20=
>> records, I=20
>> suggest you try some stress testing with what you consider a worst =
case set=20
>> of=20
>> input.=20
>>=20
>> HTH,=20
>>=20
>> Walt.=20
>>=20
>>>=20
>>>> -----Original Message-----=20
>>>> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf =
Of=20
>>>> Art Kagel=20
>>>> Sent: Wednesday, 12 October 2011 01:19 PM=20
>>>> To: ids@iiug.org=20
>>>> Subject: Re: AWK question [25177]=20
>>>>=20
>>>> Yes. Or just combine the functionality into one script. Any text =
that=20
>>>> passes one braced section falls through to the next one unless the=20=
>>>> section=20
>>>> ends with a 'next;' statement. So:=20
>>>>=20
>>>> cat file | awk ' {some stuff}' | awk '{other stuff}'=20
>>>>=20
>>>> is equivalent to:=20
>>>>=20
>>>> cat file | awk '{some stuff'} {other stuff}'=20
>>>>=20
>>>> with less overhead because only one copy of awk is running.=20
>>>>=20
>>>> Art=20
>>>>=20
>>>> Art S. Kagel=20
>>>> Advanced DataTools (www.advancedatatools.com)=20
>>>> Blog: http://informix-myview.blogspot.com/=20
>>>>=20
>>>> Disclaimer: Please keep in mind that my own opinions are my own=20
>>>> opinions and=20
>>>> do not reflect on my employer, Advanced DataTools, the IIUG, nor =
any=20
>>>> other=20
>>>> organization with which I am associated either explicitly, =
implicitly,=20
>>>> or by=20
>>>> inference. Neither do those opinions reflect those of other =
individuals=20
>>>> affiliated with any entity with which I am affiliated nor those of =
the=20
>>>> entities themselves.=20
>>>>=20
>>>> On Wed, Oct 12, 2011 at 5:01 AM, Dirk Cornel....=20
>>>> <moolma_dc@mtn.co.za>wrote:=20
>>>>=20
>>>>> Last question ... PS. I love Unix :-)=20
>>>>>=20
>>>>> I assume if I have multiple awk "functions" like the one below, I =
can=20
>>>> pipe=20
>>>>> one=20
>>>>> awk to the next ?=20
>>>>>=20
>>>>> Eg.=20
>>>>>=20
>>>>> Awk ' ..... do my stuff ..... '=20
>>>>>=20
>>>>> |=20
>>>>>=20
>>>>> Awk ' ..... do some more stuff ..... '=20
>>>>>=20
>>>>>> -----Original Message-----=20
>>>>>> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On =
Behalf=20
>>>> Of=20
>>>>>> Art Kagel=20
>>>>>> Sent: Tuesday, 11 October 2011 06:58 PM=20
>>>>>> To: ids@iiug.org=20
>>>>>> Subject: Re: AWK question [25167]=20
>>>>>>=20
>>>>>> Something like this:=20
>>>>>>=20
>>>>>> awk '=20
>>>>>> BEGIN {idx=3D0;}=20
>>>>>> {=20
>>>>>>=20
>>>>>> if (name[$1] =3D=3D $1) {=20
>>>>>>=20
>>>>>> count[$1] +=3D 1;=20
>>>>>>=20
>>>>>> } else {=20
>>>>>>=20
>>>>>> name[$1] =3D $1;=20
>>>>>>=20
>>>>>> count[$1] =3D 1;=20
>>>>>>=20
>>>>>> }=20
>>>>>> }=20
>>>>>> END {=20
>>>>>>=20
>>>>>> for (aname in name) {=20
>>>>>>=20
>>>>>> print aname, count[aname];=20
>>>>>>=20
>>>>>> }=20
>>>>>> }' filepath=20
>>>>>>=20
>>>>>> Art=20
>>>>>>=20
>>>>>> Art S. Kagel=20
>>>>>> Advanced DataTools (www.advancedatatools.com)=20
>>>>>> Blog: http://informix-myview.blogspot.com/=20
>>>>>>=20
>>>>>> Disclaimer: Please keep in mind that my own opinions are my own=20=
>>>>>> opinions and=20
>>>>>> do not reflect on my employer, Advanced DataTools, the IIUG, nor=20=
>>>> any=20
>>>>>> other=20
>>>>>> organization with which I am associated either explicitly,=20
>>>> implicitly,=20
>>>>>> or by=20
>>>>>> inference. Neither do those opinions reflect those of other=20
>>>> individuals=20
>>>>>> affiliated with any entity with which I am affiliated nor those =
of=20
>>>> the=20
>>>>>> entities themselves.=20
>>>>>>=20
>>>>>> On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel....=20
>>>>>> <moolma_dc@mtn.co.za>wrote:=20
>>>>>>=20
>>>>>>> If I have a text file in Unix, how can I do the following on the=20=
>>>> text=20
>>>>>> file=20
>>>>>>> using awk ?=20
>>>>>>>=20
>>>>>>> select column1, count(*)=20
>>>>>>> from tablename=20
>>>>>>> group by 1=20>>>>>>>=20
>>>>>>> Sorry, I did not know how else to ask the question, but I am=20
>>>> trying=20
>>>>>> to do=20
>>>>>>> this=20
>>>>>>> on a text file, without having to load the file into Informix,=20=
>>>>>> running the=20
>>>>>>> select and then unloading the result again.=20
>>>>>>> Ie. I want to do this straight in Unix without using Informix. =
My=20
>>>>>> knowledge=20
>>>>>>> of=20
>>>>>>> awk is very basic, ie:=20
>>>>>>>=20
>>>>>>> Cat filename | awk '{print $1}' ..... print column 1=20
>>>>>>>=20
>>>>>>>
Thanks Walt :-)
Regarding the developer part, I actually started as a developer, and then
later moved into DBA and SysAdmin work.
> -----Original Message-----
> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
> Walt Hultgren
> Sent: Wednesday, 12 October 2011 04:09 PM
> To: ids@iiug.org
> Subject: Re: AWK question [25181]
>
> On Oct 12, 2011, at 8:10 AM, Dirk Cornel.... wrote:
>
> > Great, thank you. I knew some very basic one liners in awk, but awk
> is a
> whole
> > language in itself. Very cool. Did my 1st multiple line script
> yesterday.
> >
> > Dirk
>
> Congratulations! You are now an official "developer" (if you weren't
> one
> already). :-)
>
> One more point about awk: I'm not familiar with all of the various
> implementations of awk, but at least the older ones tended to have what
> we
> would now consider fairly low limitations on the maximum allowed length
> of any
> one input record. Those limits would not be dynamically resized,
> either. You
> got what you got. This may not apply to the awk in in your environment
> (awk,
> nawk, gawk, ...), but it's something to be aware of.
>
> I don't recall seeing what your overall application is. If what you're
> planning will be mission critical and/or subject to very long input
> records, I
> suggest you try some stress testing with what you consider a worst case
> set of
> input.
>
> HTH,
>
> Walt.
>
> >
> >> -----Original Message-----
> >> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf
> Of
> >> Art Kagel
> >> Sent: Wednesday, 12 October 2011 01:19 PM
> >> To: ids@iiug.org
> >> Subject: Re: AWK question [25177]
> >>
> >> Yes. Or just combine the functionality into one script. Any text
> that
> >> passes one braced section falls through to the next one unless the
> >> section
> >> ends with a 'next;' statement. So:
> >>
> >> cat file | awk ' {some stuff}' | awk '{other stuff}'
> >>
> >> is equivalent to:
> >>
> >> cat file | awk '{some stuff'} {other stuff}'
> >>
> >> with less overhead because only one copy of awk is running.
> >>
> >> Art
> >>
> >> Art S. Kagel
> >> Advanced DataTools (www.advancedatatools.com)
> >> Blog: http://informix-myview.blogspot.com/
> >>
> >> Disclaimer: Please keep in mind that my own opinions are my own
> >> opinions and
> >> do not reflect on my employer, Advanced DataTools, the IIUG, nor any
> >> other
> >> organization with which I am associated either explicitly,
> implicitly,
> >> or by
> >> inference. Neither do those opinions reflect those of other
> individuals
> >> affiliated with any entity with which I am affiliated nor those of
> the
> >> entities themselves.
> >>
> >> On Wed, Oct 12, 2011 at 5:01 AM, Dirk Cornel....
> >> <moolma_dc@mtn.co.za>wrote:
> >>
> >>> Last question ... PS. I love Unix :-)
> >>>
> >>> I assume if I have multiple awk "functions" like the one below, I
> can
> >> pipe
> >>> one
> >>> awk to the next ?
> >>>
> >>> Eg.
> >>>
> >>> Awk ' ..... do my stuff ..... '
> >>>
> >>> |
> >>>
> >>> Awk ' ..... do some more stuff ..... '
> >>>
> >>>> -----Original Message-----
> >>>> From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf
> >> Of
> >>>> Art Kagel
> >>>> Sent: Tuesday, 11 October 2011 06:58 PM
> >>>> To: ids@iiug.org
> >>>> Subject: Re: AWK question [25167]
> >>>>
> >>>> Something like this:
> >>>>
> >>>> awk '
> >>>> BEGIN {idx=0;}
> >>>> {
> >>>>
> >>>> if (name[$1] == $1) {
> >>>>
> >>>> count[$1] += 1;
> >>>>
> >>>> } else {
> >>>>
> >>>> name[$1] = $1;
> >>>>
> >>>> count[$1] = 1;
> >>>>
> >>>> }
> >>>> }
> >>>> END {
> >>>>
> >>>> for (aname in name) {
> >>>>
> >>>> print aname, count[aname];
> >>>>
> >>>> }
> >>>> }' filepath
> >>>>
> >>>> Art
> >>>>
> >>>> Art S. Kagel
> >>>> Advanced DataTools (www.advancedatatools.com)
> >>>> Blog: http://informix-myview.blogspot.com/
> >>>>
> >>>> Disclaimer: Please keep in mind that my own opinions are my own
> >>>> opinions and
> >>>> do not reflect on my employer, Advanced DataTools, the IIUG, nor
> >> any
> >>>> other
> >>>> organization with which I am associated either explicitly,
> >> implicitly,
> >>>> or by
> >>>> inference. Neither do those opinions reflect those of other
> >> individuals
> >>>> affiliated with any entity with which I am affiliated nor those of
> >> the
> >>>> entities themselves.
> >>>>
> >>>> On Tue, Oct 11, 2011 at 12:44 PM, Dirk Cornel....
> >>>> <moolma_dc@mtn.co.za>wrote:
> >>>>
> >>>>> If I have a text file in Unix, how can I do the following on the
> >> text
> >>>> file
> >>>>> using awk ?
> >>>>>
> >>>>> select column1, count(*)
> >>>>> from tablename
> >>>>> group by 1> >>>>>
> >>>>> Sorry, I did not know how else to ask the question, but I am
> >> trying
> >>>> to do
> >>>>> this
> >>>>> on a text file, without having to load the file into Informix,
> >>>> running the
> >>>>> select and then unloading the result again.
> >>>>> Ie. I want to do this straight in Unix without using Informix. My
> >>>> knowledge
> >>>>> of
> >>>>> awk is very basic, ie:
> >>>>>
> >>>>> Cat filename | awk '{print $1}' ..... print column 1
> >>>>>
> >>>>> Cat filename | awk 's+=$1{print s}' ...... sum of column 1
> >>>>>
> >>>>> but I do not know how to do a count, group by in awk.
> >>>>>
> >>>>> NOTE: This e-mail message is subject to the MTN Group disclaimer
> >> see
> >>>>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> >>>>>
> >>>>>
> >>>>>
> >>>>>
> >>>>
> >>
> ***********************************************************************
> >>>> ********
> >>>>> Forum Note: Use "Reply" to post a response in the discussion
> >> forum.
> >>>>>
> >>>>>
> >>>>
> >>>> --90e6ba6e827a3d918704af08cf7e
> >>>>
> >>>>
> >>>>
> >>
> ***********************************************************************
> >>>> ********
> >>>> Forum Note: Use "Reply" to post a response in the discussion
> forum.
> >>>
> >>> NOTE: This e-mail message is subject to the MTN Group disclaimer
> see
> >>> http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> >>>
> >>>
> >>>
> >>>
> >>
> ***********************************************************************
> >> ********
> >>> Forum Note: Use "Reply" to post a response in the discussion forum.
> >>>
> >>>
> >>
> >> --90e6ba21219b49245304af182e4f
> >>
> >>
> >>
> ***********************************************************************
> >> ********
> >> Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> > NOTE: This e-mail message is subject to the MTN Group disclaimer see
> > http://www.mtn.co.za/SUPPORT/LEGAL/Pages/EmailDisclaimer.aspx
> >
> >
> >
> ***********************************