Electronic Journal of Statistics · 2017 · 63 citations · 33 references
Artificial IntelligenceEngineeringMachine LearningModel TuningHyperparameter EstimationBayesian OptimizationData ScienceData MiningLarge DatasetsRobot LearningMachine Learning ModelKnowledge DiscoverySmall SubsetsComputer ScienceDeep LearningModel OptimizationBayesian Optimization ProcedureParameter TuningStatistical Inference
Bayesian optimization has become a successful tool for optimizing the hyperparameters of machine learning algorithms, such as support vector machines or deep neural networks. Despite its success, for large datasets, training and validating a single configuration often takes hours, days, or even weeks, which limits the achievable performance. To accelerate hyperparameter optimization, we propose a generative model for the validation error as a function of training set size, which is learned during the optimization process and allows exploration of preliminary configurations on small subsets, by extrapolating to the full dataset. We construct a Bayesian optimization procedure, dubbed FABOLAS, which models loss and training time as a function of dataset size and automatically trades off high information gain about the global optimum against computational cost. Experiments optimizing support vector machines and deep neural networks show that FABOLAS often finds high-quality solutions 10 to 100 times faster than other state-of-the-art Bayesian optimization methods or the recently proposed bandit strategy Hyperband.
33
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio et al. · Proceedings of the IEEE · 1998 · 56.5K citations · Full text
Engineering, Machine Learning, Multilayer Neural Networks +17
<tt>emcee</tt>: The MCMC Hammer
Publications of the Astronomical Society of the Pacific · 2013 · 11.3K citations · Full text
Random search for hyper-parameter optimization
James Bergstra, Yoshua Bengio · 2012 · 7.9K citations