2018 · 31 citations · 34 references
Artificial IntelligenceEngineeringMachine LearningCausal UnlearningMachine Learning ToolInformation ForensicsCausal InferenceData ScienceData MiningPattern RecognitionAdversarial Machine LearningData Pollution AttackLeakage (Machine Learning)Causal ModelTraining SetPredictive AnalyticsKnowledge DiscoveryData PollutionComputer ScienceSynthetic DataModel MaintenanceData TreatmentBig Data
Machine learning systems, though being successful in many real-world applications, are known to remain prone to errors and attacks. A major attack, called data pollution, injects maliciously crafted training data samples into the training set, causing the system to learn an incorrect model and subsequently misclassify testing samples. A natural solution to a data pollution attack is to remove the polluted data from the training set and relearn a clean model. Unfortunately, the training set of a real-world machine learning system can contain millions of samples; it is thus hopeless for an administrator to manually inspect all of them to weed out the polluted ones.
34
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky et al. · 2014 · 34.2K citations
LIBLINEAR: A Library for Large Linear Classification
Rong-En Fan, Kai‐Wei Chang, Cho‐Jui Hsieh et al. · 2008 · 6.6K citations
Active Learning Literature Survey
Burr Settles · Minds at UW (University of Wisconsin) · 2009 · 4.8K citations · Full text
Mining high-speed data streams
Pedro Domingos, Geoff Hulten · 2000 · 2.2K citations · Full text