2023 · 15 citations · 34 references
Performance variations caused by anomalies in modern High Performance Computing (HPC) systems lead to decreased efficiency, impaired application performance, and increased operational costs. While machine learning (ML)-based frameworks for automated anomaly detection (often based on time series telemetry data) are gaining popularity in the literature, practical deployment challenges are often overlooked. Some ML-based frameworks require extensive customization, while others need a rich set of labeled samples, none of which are feasible for a production HPC system.
34
Marco Túlio Ribeiro, Sameer Singh, Carlos Guestrin · 2016 · 14K citations
Karl Pearson · The London Edinburgh and Dublin Philosophical Magazine and Journal of Science · 1900 · 3.9K citations