Concurrency and Computation Practice and Experience · 2015 · 545 citations · 34 references
EngineeringBig Data AnalyticsSoftware EngineeringHardware SystemsDynamic Workflow SystemWorkflow GraphData ScienceComputing SystemsSystems EngineeringDynamic WorkflowsParallel ComputingHigh-throughput ComputingData ManagementWorkflow TechnologyDistributed SystemsComputer ScienceWorkflow Management SystemData-intensive ComputingWorkflow ExecutionScientific Workflow SystemParallel ProgrammingWorkflow SoftwareSystem Software
The paper introduces FireWorks, a workflow system for high‑throughput calculations on supercomputing centers, and outlines its performance, limitations, and future development plans. FireWorks is built on Python and MongoDB, employing job packing, failure detection, provenance tracking, duplicate detection, and dynamic workflow graph modification to support concurrent, fault‑tolerant, high‑throughput computations. The system has processed over 50 million CPU‑hours of computational chemistry and materials science work, demonstrating the effectiveness of its high‑throughput features and providing performance data and identified limitations. © 2015 John Wiley & Sons, Ltd.
Summary This paper introduces FireWorks, a workflow software for running high‐throughput calculation workflows at supercomputing centers. FireWorks has been used to complete over 50 million CPU‐hours worth of computational chemistry and materials science calculations at the National Energy Research Supercomputing Center. It has been designed to serve the demanding high‐throughput computing needs of these applications, with extensive support for (i) concurrent execution through job packing, (ii) failure detection and correction, (iii) provenance and reporting for long‐running projects, (iv) automated duplicate detection, and (v) dynamic workflows (i.e., modifying the workflow graph during runtime). We have found that these features are highly relevant to enabling modern data‐driven and high‐throughput science applications, and we discuss our implementation strategy that rests on Python and NoSQL databases (MongoDB). Finally, we present performance data and limitations of our approach along with planned future work. Copyright © 2015 John Wiley & Sons, Ltd.
34
Commentary: The Materials Project: A materials genome approach to accelerating materials innovation
Anubhav Jain, Shyue Ping Ong, Geoffroy Hautier et al. · APL Materials · 2013 · 11.9K citations · Full text