Probabilistic Name and Address Cleaning and Standardisation.

Peter Christen, Tim Churches, Justin Zhu

2002 · 12 citations · 0 references

Abstract

In the absence of a shared unique key, an ensemble of nonunique personal attributes such as names and addresses is often used to link data from disparate sources. Such data matching is widely used when assembling data warehouses and business mailing lists, and is a foundation of many longitudinal epidemiological and other health related studies. Unfortunately,