2020 · 89 citations · 26 references
Cluster ComputingEngineeringMachine LearningAdvanced ComputingMachine Learning ToolModel Batch SizeComputer ArchitectureProduction Machine LearningData ScienceHigh-performance ArchitectureParallel ComputingPerformance PredictionMl Prediction PipelinesPredictive AnalyticsComputer EngineeringComputer ScienceDeep LearningHardware AccelerationParallel Programming
Serving ML prediction pipelines spanning multiple models and hardware accelerators is a key challenge in production machine learning. Optimally configuring these pipelines to meet tight end-to-end latency goals is complicated by the interaction between model batch size, the choice of hardware accelerator, and variation in the query arrival process.
26
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala et al. · 2017 · 11.1K citations
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford et al. · 2010 · 2.4K citations · Full text
The Nonstochastic Multiarmed Bandit Problem
Peter Auer, Nicolò Cesa‐Bianchi, Yoav Freund et al. · SIAM Journal on Computing · 2002 · 2.2K citations
Matt Welsh, David Culler · 2001 · 840 citations
Event-driven Architecture, Cluster Computing, New Design +11