2006 · 204 citations · 27 references
Privacy ProtectionEngineeringMachine LearningNonlinear KernelsInformation SecuritySupport Vector MachineData ScienceData MiningPattern RecognitionPrivacy SystemData ManagementPrivacy ServiceKnowledge DiscoveryData PrivacyComputer ScienceDifferential PrivacyPrivacyData SecurityCryptographyTraditional Data MiningBig Data
Traditional Data Mining and Knowledge Discovery algorithms assume free access to data, either at a centralized location or in federated form. Increasingly, privacy and security concerns restrict this access, thus derailing data mining projects. What we need is distributed knowledge discovery that is sensitive to this problem. The key is to obtain valid results, while providing guarantees on the non-disclosure of data. Support vector machine classification is one of the most widely used classification methodologies in data mining and machine learning. It is based on solid theoretical foundations and has wide practical application. This paper proposes a privacy-preserving solution for support vector machine (SVM) classification, PP-SVM for short. Our solution constructs the global SVM classification model from the data distributed at multiple parties, without disclosing the data of each party to others. We assume that data is horizontally partitioned -- each party collects the same features of information for different data objects. We quantify the security and efficiency of the proposed method, and highlight future challenges.
27
Yuhai Wu, Vladimir Vapnik · Technometrics · 1999 · 26.9K citations
Handbook of applied cryptography
Choice Reviews Online · 1997 · 10.4K citations
Valuable Reference, Cryptographic Primitive, Rapid Access +19
Privacy-preserving data mining
Rakesh Agrawal, Ramakrishnan Srikant · ACM SIGMOD Record · 2000 · 3K citations
Privacy-preserving Data Mining, Engineering, Machine Learning +23