2013 · 11 citations · 13 references
Cluster ComputingEngineeringMachine LearningStreaming AlgorithmData StreamConcept DriftData ScienceData MiningPattern RecognitionManagementData IntegrationParallel ComputingBig DataData ManagementHigh-performance Data AnalyticsPredictive AnalyticsKnowledge DiscoveryComputer ScienceData Stream ManagementScalability IssueEvolving Data StreamsData Stream MiningTraditional Data MiningEnsemble Algorithm
Unlike traditional data mining where data is static, mining algorithms for data streams must process the data "on the fly" and update the class decision boundaries as the stream progresses to address the challenges of concept drift and feature evolution. In our current work, we have proposed a multi-tiered ensemble based fast and robust method, which rapidly learns the concepts in a data stream, predicts labels for new data with strong accuracy, and agilely tracks the dynamic changes in the evolving concepts and feature space. Bottleneck of our current work is, it needs to build ADABOOST ensemble for each numeric feature. This can face scalability issue as number of features can be very large at times in data stream. In this paper we propose a method to parallelize the independent parts of that work using a MapReduce framework. This increases scalability and achieves a significant speedup without compromising classification accuracy. We demonstrate the performance of our approach in terms of speedup, scale up and classification accuracy.
13
Sanjay Ghemawat · Communications of the ACM · 2008 · 18.4K citations · Full text
Improved boosting algorithms using confidence-rated predictions
Robert E. Schapire, Yoram Singer · 1998 · 2.6K citations
Improved Boosting Algorithms Using Confidence-rated Predictions
Robert E. Schapire, Yoram Singer · Machine Learning · 1999 · 1.9K citations · Full text
What's In A Name? Malay Seals As Onomastic Sources
Annabel Teh Gallop · Malay Literature · 2018 · 1.5K citations · Full text
Language Documentation, Southeast Asia, East Asian Studies +8