Source-linked AI summary
Interactive Supercomputing on 40,000 Cores for Machine Learning and Data Analysis
Albert Reuther, Jeremy Kepner, Chansup Byun, Siddharth Samsi, William Arcand, David Bestor, Bill Bergeron, Vijay Gadepally, Michael Houle, Matthew Hubbell, Michael Jones, Anna Klein, Lauren Milechin, Julia Mullen, Andrew Prout, Antonio Rosa, Charles Yee, Peter Michaleas
TL;DR
Scaling interactive machine-learning and data-analysis launches to tens of thousands of cores is difficult because schedulers and application dependencies can create long startup times. This paper engineers immediate launches on a 40,000-core supercomputer, achieving under-5-second launches for 32,000+ TensorFlow cores and under-40-second launches for 260,000+ MATLAB/Octave processes.
Problem
Scaling interactive launches to 40,000-core jobs remained challenging, with initial Slurm launches of interactive MATLAB/Octave jobs taking 30–60 minutes.
Method
The system uses immediate launches with user resource limits and one scheduler-issued launcher per node to spawn and background application processes.
Results
Under 5 seconds launches 32,000+ cores for TensorFlow, while under 40 seconds launches exceed 260,000 MATLAB/Octave processes.
Takeaways & Limitations
These launch capabilities support rapid interactive exploration of neural-network trade-offs and rapid prototyping, algorithm development, and data analysis.
Takeaways & Limitations
At high node and process counts, launch times rise because serving files to many processes creates backpressure from the central Lustre file system.
Abstract
from arXiv · showhide
Interactive massively parallel computations are critical for machine learning and data analysis. These computations are a staple of the MIT Lincoln Laboratory Supercomputing Center (LLSC) and has required the LLSC to develop unique interactive supercomputing capabilities. Scaling interactive machine learning frameworks, such as TensorFlow, and data analysis environments, such as MATLAB/Octave, to tens of thousands of cores presents many technical challenges - in particular, rapidly dispatching many tasks through a scheduler, such as Slurm, and starting many instances of applications with thousands of dependencies. Careful tuning of launches and prepositioning of applications overcome these challenges and allow the launching of thousands of tasks in seconds on a 40,000-core supercomputer. Specifically, this work demonstrates launching 32,000 TensorFlow processes in 4 seconds and launching 262,000 Octave processes in 40 seconds. These capabilities allow researchers to rapidly explore novel machine learning architecture and data analysis algorithms.