The method on this page finds duplicate owner records in three passes: exact matches on a field that should never repeat, exact matches again once the formats are standardized, and near matches that a person reviews before anything merges. The U.S. Census Bureau built the mailing list for its 1992 Census of Agriculture on the same principle, with software designating links and clerks settling the uncertain pairs, and together they marked 20.9 percent of six million records as duplicates (Winkler, U.S. Census Bureau). Each merge keeps the original records, so it can be undone.
Six million records from twelve sources
The Census Bureau built the mailing lists for its 1987 and 1992 Censuses of Agriculture from six million records taken from 12 different sources, and had to find every farm listed in them more than once. William Winkler's 1993 account compares the two rounds on that common base (Winkler, U.S. Census Bureau).
In 1987 a computer rule compared records that shared a ZIP code, using the surname, the first letter of the first name and the numbers in the address. It designated 6.6 percent of the records as duplicates and 28.9 percent as possible duplicates. Reviewing those took 14,000 person hours, as many as 75 clerks for three months, and turned up another 7.5 percent. Winkler notes that many duplicates were never found, and that estimates built on the list may have suffered for it.
In 1992 the Bureau used new matching software built on the Fellegi-Sunter model, written to tolerate typing errors. The software designated 12.8 percent as duplicates and sent 19.7 percent to clerical review, where 6,500 person hours found another 8.1 percent. The computer run took 22 days. Building the software had cost $110,000. Computer and clerks together found 14.1 percent duplicates in 1987 and 20.9 percent in 1992.
What a duplicate costs a small list
An invented example: a list of 10,000 owner records holds 1,000 duplicates. It describes 9,000 owners. A reply rate computed over 10,000 comes out a tenth too low, and so does any other rate divided by the record count. A skip trace priced per record pays for 1,000 owners twice. For the owner, one copy shows a call on Tuesday and the other shows nothing, and a second caller dials the same person.
What one record stands for
A landowner with four parcels is one person and four properties. The list's purpose decides what happens next. In a list of properties, those four rows are correct, and merging them drops three parcels. In a list of people, they are one owner entered four times. A list that mixes both gets split into an owners tab and a properties tab before any matching starts.
The duplicate rate before cleanup
Measured before the cleanup and again after it, the rate shows whether the cleanup worked.
- Export every record with its record number, its source and the date it was created.
- Pull 200 records at random. A column of random numbers, sorted, is enough.
- For each of the 200, search the full file by last name and by phone number, and mark whether another record is the same owner.
- Divide the records that have a twin by 200.
An invented example: 20 of the 200 have a twin, a rate of 10 percent. At 95 percent confidence the Wilson method puts the list's true rate between about 7 and 15 percent; the guide on how fast records go stale shows that arithmetic with its source. At 7 percent, about 700 of 10,000 records have a twin somewhere in the file; at 15 percent, about 1,500.
Pass one, exact matches
An exact match on a field that should never repeat for two owners, such as a full email address or a phone number stored one way on every record, is nearly always the same person. A shared office line or a family email breaks that rule, so any value that matches three or more records goes to a person before it merges.
Pass two, standardized formats
Many duplicates differ only in formatting: (406) 555-0147 on one record, 406.555.0147 on the other. Each field used for matching is standardized, and pass one runs again.
- Phone numbers. The International Telecommunication Union's Recommendation E.164 gives every international number a country code of one to three digits followed by the national number, at most 15 digits in all (ITU-T E.164). Written in that structure with a leading plus sign and digits only, both become +14065550147.
- Email addresses. RFC 5321, the standard for sending mail, says servers treat the part before the @ as case sensitive, and the same section discourages relying on that (RFC 5321, section 2.4). A lowercased copy in its own column does the matching, and the address stays as the owner typed it.
- Mailing addresses. The Postal Service's Publication 28, dated October 2024, lists the standard abbreviations for street suffixes and for unit designators such as apartment and suite (Publication 28). Address software certified under the Postal Service's CASS program has to assign ZIP + 4 and carrier route codes with at least 98.5 percent accuracy in testing, and delivery point codes with 100 percent (CASS certification).
- Names. First and last names sit in separate fields, titles and suffixes come out, and company, trust and estate names get a field of their own. "Smith Family Trust" and "John Smith" are related, and they are two records.
Pass three, near matches a person reviews
The pairs left over differ by typing: Katherine and Kathryn, Mc Donald and McDonald, a house number with two digits swapped. Three measures score how alike two values are.
- Sound codes. Soundex codes a surname as its first letter and three digits for how it sounds, so names spelled several ways file together. The National Archives uses it to index census records, and it codes Washington as W-252 (National Archives).
- Edit distance. NIST's Dictionary of Algorithms and Data Structures defines the Levenshtein distance as the smallest number of insertions, deletions and substitutions that turns one string into another (NIST). Jon and John are one insertion apart, and so are Jon and Joan.
- Name scoring from census work. William Winkler extended Jaro's string comparator to handle typing variation in first names and surnames, in matching work for the 1990 census (Winkler, 1990).
Jon and Joan show why a score alone does not merge anything. The Census Bureau's decision rule sorts each pair into links, nonlinks and possible links, and clerks decide the possible links (Winkler, U.S. Census Bureau). A list kept by one person works with two cutoffs on the score. Above the high one, a person gives each pair a quick yes. Between the two, the records are compared side by side, and pairs below the low one stay apart.
Comparing only records that share a value
The number of pairs grows fast: 10,000 records make 49,995,000 of them (10,000 times 9,999, divided by 2). Record linkage cuts the work with blocking, which compares only records that already share a value. Winkler's example takes two lists of 300,000 business records for one state, 90 billion possible pairs, and compares only the 30 million pairs that share a ZIP code. A duplicate whose blocking field is itself wrong is never compared, so a second round blocks on a different field, such as the phone number.
Two ways a merge goes wrong
A matching rule can join two different owners into one record, or leave one owner as two. Precision is the share of flagged pairs that really are duplicates; recall is the share of real duplicates that were found. Fifty merged pairs pulled at random and read by a person give a precision estimate. On owner records the two errors cost different amounts: a missed duplicate costs a second call, and a wrong merge hands one owner's history, and possibly that owner's request to stop contact, to someone else.
A merge that can be undone
- Every original record stays, marked as merged into the surviving record, with the date and the rule that matched it.
- Each field notes which value won and why: the phone with the latest confirmation date, the ownership from the county record, the name as the owner wrote it.
- A value a person entered after talking to the owner outranks a value a data provider appended, unless that person marked it out of date.
- A request to stop contact on any merged record carries to the survivor.
Households and lists that outgrow a sheet
- Households and companies need rules of their own. Two people at one mailing address may be spouses, tenants or strangers, and the matching fields cannot tell which.
- A list that several people edit, or that outgrows one spreadsheet, moves into a database that keeps merge history before pass three runs.
Sources
- William E. Winkler, "Matching and Record Linkage," U.S. Census Bureau, 1993: census.gov
- William E. Winkler, "String Comparator Metrics and Enhanced Decision Rules in the Fellegi-Sunter Model of Record Linkage," 1990: eric.ed.gov
- International Telecommunication Union, Recommendation ITU-T E.164, "The international public telecommunication numbering plan," February 2026 edition: itu.int
- Internet Engineering Task Force, RFC 5321, "Simple Mail Transfer Protocol," October 2008, section 2.4: rfc-editor.org
- United States Postal Service, Publication 28, "Postal Addressing Standards," October 2024: pe.usps.com
- United States Postal Service, PostalPro, "CASS Certification": postalpro.usps.com
- National Archives, "The Soundex Indexing System": archives.gov
- National Institute of Standards and Technology, Dictionary of Algorithms and Data Structures, "Levenshtein distance," modified May 15, 2019: nist.gov