Estimating the success of re-identifications in incomplete datasets using generative models

Luc Rocher, Julien M. Hendrickx, Yves-Alexandre de Montjoye

Nature Communications · 2019 · 769 citations · 46 references

DOIFull text

Open access

TL;DR

Rich medical, behavioral, and socio‑demographic data are essential for data‑driven research, but their collection raises privacy concerns that are traditionally addressed by de‑identification and sampling. The study proposes a generative copula‑based method to estimate the likelihood of correctly re‑identifying a specific individual in highly incomplete datasets. The method uses a generative copula model to estimate re‑identification likelihoods in incomplete datasets. The method achieves AUCs of 0.84–0.97 for predicting individual uniqueness, shows that 99.98% of Americans would be re‑identified with 15 demographic attributes, and indicates that heavily sampled anonymized datasets fail to meet GDPR anonymization standards.

Abstract

Abstract While rich medical, behavioral, and socio-demographic data are key to modern data-driven research, their collection and use raise legitimate privacy concerns. Anonymizing datasets through de-identification and sampling before sharing them has been the main tool used to address those concerns. We here propose a generative copula-based method that can accurately estimate the likelihood of a specific person to be correctly re-identified, even in a heavily incomplete dataset. On 210 populations, our method obtains AUC scores for predicting individual uniqueness ranging from 0.84 to 0.97, with low false-discovery rate. Using our model, we find that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes. Our results suggest that even heavily sampled anonymized datasets are unlikely to satisfy the modern standards for anonymization set forth by GDPR and seriously challenge the technical and legal adequacy of the de-identification release-and-forget model.

References

46