Source-linked AI summary
Compute Trends Across Three Eras of Machine Learning
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, Pablo Villalobos
TL;DR
The paper asks how training compute for milestone ML systems has changed over time, motivated by compute’s role as a quantifiable proxy for ML progress. It curates a dataset of more than 100 systems and analyzes compute trends across proposed eras. Training compute doubled every 18 months from 1952 to 2010, every 6 months from 2010 to 2022, and a large-scale trend emerged in late 2015 at 2–3 orders of magnitude above the previous trend.
Problem
Predicting ML progress is difficult because it depends on many factors, while compute is a regular factor and quantifiable proxy for ML research progress.
Method
The paper curates a dataset of training compute for more than 100 milestone ML systems and analyzes how compute trends grow over time.
Results
18-month doubling from 1952 to 2010, 6-month doubling from 2010 to 2022, and a large-scale trend from late 2015 at 2 to 3 orders of magnitude above the previous trend with 10-month doubling.
Takeaways & Limitations
The authors frame compute history in three eras: Pre Deep Learning, Deep Learning, and Large-Scale.
Takeaways & Limitations
The authors are not confident in the distinction between large-scale and regular-scale models.
Abstract
from arXiv · showhide
Compute, data, and algorithmic advances are the three fundamental factors that guide the progress of modern Machine Learning (ML). In this paper we study trends in the most readily quantified factor - compute. We show that before 2010 training compute grew in line with Moore's law, doubling roughly every 20 months. Since the advent of Deep Learning in the early 2010s, the scaling of training compute has accelerated, doubling approximately every 6 months. In late 2015, a new trend emerged as firms developed large-scale ML models with 10 to 100-fold larger requirements in training compute. Based on these observations we split the history of compute in ML into three eras: the Pre Deep Learning Era, the Deep Learning Era and the Large-Scale Era. Overall, our work highlights the fast-growing compute requirements for training advanced ML systems.
1 Introduction
The paper investigates training-compute demand in milestone ML models because compute is a regular, quantifiable proxy for ML progress. It curates and analyzes a dataset spanning three proposed compute eras.
- Compute is presented as an unusually regular factor influencing questions about future ML capabilities, automation, and societal change.
- Compute can serve as a quantifiable proxy for the progress of ML research because model scaling is related to AI capabilities.
- The paper conducts a detailed investigation into the compute demand of milestone ML models over time.
- 123 milestone ML systems are curated and annotated with the compute required to train them.
- The authors tentatively frame compute trends as three eras and estimate doubling times for each.
- The dataset, figures, and interactive visualization are publicly available, and the authors check their results through appendices examining alternative interpretations and prior-work differences.
2 Related work
Prior studies reported differing compute and parameter-growth trends using limited or differently scoped datasets. This paper broadens the data collection and incorporates information from related public model-tracking initiatives.
- Amodei and Hernandez reported a 3.4-month training-compute doubling time from 2012 to 2018 using 15 ML systems.
- Sastry and colleagues estimated approximately two-year training-compute doubling between 1959 and 2012 after adding 10 pre-2012 papers.
- Lyzhov argued growth stalled after 2018, finding GPT-3 required only 1.5× the training compute of AlphaGo Zero.
- Parameter-count studies found 18–24-month doubling across application domains, with language-model doubling accelerating to 4–8 months between 2016 and 2018.
- Other work linked rising compute requirements to increasingly infeasible progress and identified cost, hardware, and engineering constraints on continued scaling.
- The paper incorporates work from Akronomicon, Computer Progress, and AI Tracker into its dataset, alongside estimates informed by inference-compute research.
- Compared with prior work, the dataset contains three times more ML models and extends coverage through 2022.
3 Trends
The analysis identifies faster compute growth after the advent of Deep Learning and a separate large-scale-model trend emerging around 2015–2016. Regular-scale growth continues across that later transition, while the era boundary and large-scale distinction remain qualified.
- 3.1 The transition to Deep Learning: Before Deep Learning, training compute doubled every 17–29 months; afterward, the overall trend doubled every 4–9 months.
- 3.1 The transition to Deep Learning: The Pre Deep Learning trend roughly matches Moore’s law, with computational performance commonly simplified as doubling every two years.
- 3.1 The transition to Deep Learning: The Deep Learning Era’s start is uncertain because the transition shows no noticeable discontinuity, and results barely change when dated to 2010 or 2012.
- 3.2 Trends in the Large-Scale era: Around 2015–2016, a new large-scale trend emerged with AlphaGo and continued through the present, involving systems trained by large corporations.
- 3.2 Trends in the Large-Scale era: The regular-scale trend remains continuous before and after 2016, doubling every 5–6 months.
- 3.2 Trends in the Large-Scale era: Large-scale models show an apparent 9–10-month doubling time, although limited data means this possible slowdown may reflect noise.
- 3.2 Trends in the Large-Scale era: Separating large-scale and regular-scale models helps reconcile prior estimates based on limited samples and a single assumed trend.
4 Conclusion
The paper curates training-compute data for more than 100 milestone ML systems and identifies three eras with distinct scaling patterns. It argues that compute growth accelerated around 2010, while large-scale models emerged in late 2015 and exceeded the prior trend by orders of magnitude.
- 18 months, 6 months, and 10 months are the reported doubling times for 1952–2010, 2010–2022, and the large-scale trend, respectively.The large-scale trend began 2 to 3 orders of magnitude above the previous trend.
- Around 2010, compute growth accelerated as the field transitioned into the Deep Learning Era.
- In late 2015, companies began releasing large-scale models such as AlphaGo that surpassed the previous trend.The paper marks this as the beginning of the Large-Scale Era.
- Framing compute trends as three eras helps explain observed discontinuities, although the paper is not confident in distinguishing large-scale from regular-scale models.
- The growing training-compute trend highlights the strategic importance of hardware infrastructure, computing clusters, and engineers with expertise to use them.
- The study does not analyze data trends, leaving dataset size and its relationship to compute for future work.
A Methods
The study constructs a curated dataset of milestone ML systems, estimates missing training-compute values, and applies statistical procedures to fit trends. Its selection and measurement choices include subjective judgments, uncertainty adjustments, outlier filtering, and known dataset limitations.
- Data selection: Milestone models are selected primarily from papers showing learning, experimental results, state-of-the-art advances, and at least one notability criterion.For models from 2020 onward, assessing these criteria is harder, so selection falls back to subjective judgment.
- Limitations: The dataset is biased and may contain mistakes, so the authors caution against drawing strong conclusions from it.The investigation’s limitations are discussed in Appendix H.
- Compute estimation: Training compute is estimated from forward-pass compute or GPU time when papers do not report it.The estimation reasoning is annotated in the corresponding dataset cells.
- Outlier handling: Five of 123 systems are excluded as low-compute outliers using a local log-compute Z-score threshold two standard deviations below the mean.The comparison window is 1.5 years around each model’s publication date.
- Outlier handling: High-compute outliers after 2016 are selected using a Z > 0.76 threshold determined after visual inspection.The thresholds were chosen to automatically reproduce visually identified outliers.
B Analyzing record-setting models
The record-setting-model analysis provides a corroborating view of compute growth, showing a slow era before 2010, faster growth from 2010 to 2015, and a discontinuity around September 2015. Results after 2015 depend materially on whether AlphaGo Zero and AlphaGo Master are included.
- Record-setting models are models that set a compute-demand record by outcompeting all previously released models.
- Record-setting models support the paper’s broader conclusions, but their compute budgets are likely dominated by outliers and expensive efforts to push the state of the art.
- 1957–2010 shows slow compute growth, while 2010–2015 shows fast growth in the record-setting models.
- Around September 2015, the record-setting-model trend exhibits a discontinuity.
- Including AlphaGo Zero and AlphaGo Master produces a one-year doubling time for 2015–2022.Excluding them yields a trend with a doubling time similar to 2010–2015.
C Trends in different domains
The paper analyzes compute trends separately across vision, language, games, and other domains because architectures may differ. Vision and language follow the overall growth pattern, whereas games show no consistent trend and other grouped domains broadly track the aggregate.
- Domain analysis: The analysis separates vision, language, games, and other domains because different architectures may follow different scaling laws.
- Vision and language: Vision and language show fairly consistent growth over time, following the same doubling pattern as the overall dataset.
- Games: Games show no consistent compute trend.The authors suggest sparse data, heterogeneous games, or less systematic progress as possible explanations.
- Other domains: Other domains grouped together appear to follow the overall dataset’s trend.The group includes domains with fewer than 10 systems each, such as speech, robotics, and recommender systems.
- Era boundary: The paper uses 2010 as the default start of the Deep Learning Era, while noting that results do not change when using 2012.The choice is supported by GPU use, competitive deep neural networks, and speech-recognition adoption.
E Comparison to Amodei & Hernandez’s analysis
The paper attributes its different doubling-time estimate from Amodei and Hernandez to extended data and, especially, separating regular-scale from large-scale trends. It favors the interpretation that these are distinct trends.
- 5.7 months is the paper’s estimated compute-doubling time from 2012 to 2022, compared with Amodei and Hernandez’s 3.4 months from 2012 to 2018.The analyses use different sample sizes, time periods, and trend identification.
- Between 2015 and 2017, a distinct Large-Scale Era trend emerged, and the paper evaluates both separate-trend and single-trend interpretations.The single-trend interpretation resembles Amodei and Hernandez’s analysis.
- Separating large-scale models leaves a regular-scale trend with a similar doubling time before and after 2017.The paper argues that mixing regular-scale and large-scale models produces Amodei and Hernandez’s different result.
- The authors favor separating the trends because the large-scale account better predicts post-2017 developments and reflects a drastic departure in funding.They note that Lyzhov found the single-trend account did not extend past 2017.
F Are large-scale models a different category?
The paper hypothesizes that extraordinarily compute-intensive projects form a distinct flagship-model category, while acknowledging uncertainty about their economic and categorical status. The boundary remains partly judgment-based.
- From 2016 onward, some companies spent substantially more compute and money than previous trends predicted, including AlphaGo Zero and AlphaStar.AlphaGo Zero’s estimated cost was $35M, while AlphaStar’s was estimated at $12M.
- Without inside knowledge, the paper cannot determine whether these projects continued an existing trend or represented categorically different projects.The authors specifically question whether expected economic returns were significantly larger for some models.
- The dividing line for large-scale models is uncertain because several named systems could reasonably be placed on either side.The paper gives NASv3, Libratus, Megatron-LM, T5-3B, and others as borderline examples.
- Different Z-value thresholds produce only small differences in the selected Large-Scale models.Figure 8 illustrates an alternate reasonable selection using threshold Z = 0.54.
G Causes of the possible slowdown of training budget in large-scale models between 2016 and 2022
Large-scale-model compute increased more slowly than the overall trend between 2016 and 2022. The paper lists chip shortages, infrastructure constraints, budget caps, and undisclosed models as possible explanations.
- 10 months was the large-scale-model compute doubling time, versus 6 months for the overall trend.The comparison covers increasing compute between 2016 and 2022.
- A 2020–2022 global chip shortage may have slowed growth by raising GPU prices and limiting hardware availability.The shortage followed strong equipment demand, supply shocks, and trade frictions.
- HPC constraints may limit massive training runs through memory and communication-bandwidth bottlenecks.These constraints require models to be partitioned across groups of layers and trained in parallel, making efficient engineering difficult.
- Budget caps may constrain further scaling because compute-intensive training runs can cost millions of dollars.The paper cites an estimate of $10M in cloud-computing costs for Google’s T5 project.
- Undisclosed large models may distort the observed trend because many compute-intensive systems originate in corporate AI labs and are not published.
H Limitations
The analysis has uncertainties from dataset selection, compute estimation, non-sampling errors, and limited verification. The authors also caution that the dataset is biased and may contain mistakes.
- Compute estimates are generally expected to be accurate within about a factor of two because inputs such as utilization rate and FLOP/s are uncertain.The authors introduce noise during bootstrapping to account for this uncertainty.
- Non-sampling errors remain possible, including incorrect calculations, although the calculations are intended to be verifiable through dataset annotations.
- The dataset overrepresents academic, English-language, and subjectively notable ML systems.Closed-source commercial systems and papers omitting training-time information are less represented; notable models tend to be larger and newer.
- A higher proportion of large and recent models increases the estimated doubling times.