Privacy-preserving <i>k</i> -means clustering over vertically partitioned data

Jaideep Vaidya, Chris Clifton

2003 · 628 citations · 23 references

Concepts

TL;DR

Privacy concerns can prevent data sharing, yet distributed knowledge discovery can yield valid results while safeguarding data disclosure. The study proposes a k‑means clustering method for vertically partitioned data where each site holds different attributes of the same entities. Each site learns the cluster assignment of each entity but learns nothing about the attributes held by other sites.

Abstract

Privacy and security concerns can prevent sharing of data, derailing data mining projects. Distributed knowledge discovery, if done correctly, can alleviate this problem. The key is to obtain valid results, while providing guarantees on the (non)disclosure of data. We present a method for k-means clustering when different sites contain different attributes for a common set of entities. Each site learns the cluster of each entity, but learns nothing about the attributes at other sites.

References

23