Re: Names matching
Posted in 1994
In article <32urqr$c7s@emory.mathcs.emory.edu> tilh@sin-co.sin-ro.DHL.COM (Ti Lian Hwang) writes: >We have a problem of matching names (esp. Company names) from one table >to another. Fun, isn't it? >One of the table is actually a consignee database. > >The other is keyed in probably half way around the world and represents >the recipient of the item. > >We need to match both this records. As we all know, data entry of proper >nouns always result in spelling errors, acronyns, missing or extra "the"s,"of"s >etc. > >Does anyone knows what's a good way of doing this, short of using an expert >system (we only have a 4GL and a database !) ? Would soundex help ? No reason you can't implement an expert system with nothing but 4GL and a database (and objects commonly found around the home)! I have been attempting something like this, so I'll ramble on a bit about my approach. I don't trust pattern matching enough to do anything automatically based on apparent "matches", so I just print out reports and say "Hey, the following companies have similar names - you might want to decide whether they are really the same or not". I basically take the list of names and calculate "signatures" for each name. To do this I drop all the "THE"s, "OF"s, "INC"s, "CORP"s, etc. I drop all the vowels, replace all the double letters by single instances of the same letter, etc. I guess I am doing a sort of poor man's SOUNDEX here. I don't think it is too hard to do a real soundex - the code gets posted here now and then (I forget by whom). Then I have my report list all the different companies that map to the same signature. Then I rely on humans to investigate whether they truly are the same company or not. I find that the easy way to drop the "THE"s etc is to unload the names to an ascii file and pipe that file through a sed command such as: sed 's/ THE / /g' Then pipe it though "sed 's/ OF / /g' " etc. Later I load the ascii file back into a table and proceed from there. I am sure you could do this in "pure 4GL" instead, but those unix utilities are so darned fast: why not use them (especially if you are searching for similar names in a set of 50,000 company names)? I could probably expand on some of my "etc"s above if anyone is interested, (i.e. just which are the content-free words that should be dropped?) Enough rambling .... Paul