2013 · 209 citations · 29 references
Software MaintenanceEngineeringMachine LearningBusiness IntelligencePeters FilterQuality PredictionFault ForecastingBusiness AnalyticsData ScienceData MiningManagementStatisticsQuantitative ManagementComputational Learning TheoryFeature EngineeringPredictive AnalyticsKnowledge DiscoveryPredictive ModelingComputer ScienceInformation ManagementReliability PredictionStatistical Learning TheorySoftware TestingBusinessLife CycleFailure Prediction
How can we find data for quality prediction? Early in the life cycle, projects may lack the data needed to build such predictors. Prior work assumed that relevant training data was found nearest to the local project. But is this the best approach? This paper introduces the Peters filter which is based on the following conjecture: When local data is scarce, more information exists in other projects. Accordingly, this filter selects training data via the structure of other projects. To assess the performance of the Peters filter, we compare it with two other approaches for quality prediction. Within-company learning and cross-company learning with the Burak filter (the state-of-the-art relevancy filter). This paper finds that: 1) within-company predictors are weak for small data-sets; 2) the Peters filter+cross-company builds better predictors than both within-company and the Burak filter+cross-company; and 3) the Peters filter builds 64% more useful predictors than both within-company and the Burak filter+cross-company approaches. Hence, we recommend the Peters filter for cross-company learning.
29
Leo Breiman · Machine Learning · 2001 · 119.3K citations · Full text
Mark Hall, Eibe Frank, Geoffrey Holmes et al. · ACM SIGKDD Explorations Newsletter · 2009 · 17.8K citations
Data clustering: 50 years beyond K-means
Anil K. Jain · Pattern Recognition Letters · 2009 · 8.9K citations