Source-linked AI summary
Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model
Alexandra Sasha Luccioni, Sylvain Viguier, Anne-Laure Ligozat
TL;DR
Training and deploying large ML models consume substantial resources, while their broader life-cycle carbon footprint remains difficult to quantify. The paper estimates BLOOM’s emissions across equipment manufacturing, energy use, training, and real-time API deployment, reporting 24.69 tonnes of CO2eq from dynamic power consumption alone and broader estimates that include additional processes. It concludes that carbon accounting requires more precise data and improved reporting methods.
Problem
ML models require substantial computational resources, energy, and materials, but the field lacks precise, systematic estimates covering their broader carbon footprint.
Method
The study connects emissions from BLOOM’s equipment manufacturing, training energy, infrastructure operation, and real-time API deployment using life-cycle assessment where information is available.
Results
24.69 tonnes of CO2eq were attributed to BLOOM’s dynamic training power consumption, while the study estimates emissions across additional life-cycle processes and deployment.
Takeaways & Limitations
The paper provides a broader framework for connecting carbon-emission sources in training and deploying a 176B-parameter language model and identifies deployment measurement as a starting point.
Takeaways & Limitations
The estimates remain constrained by unavailable information, including exact hardware-manufacturing data and the representativeness of one real-time deployment configuration.
Abstract
from arXiv · showhide
Progress in machine learning (ML) comes with a cost to the environment, given that training ML models requires significant computational resources, energy and materials. In the present article, we aim to quantify the carbon footprint of BLOOM, a 176-billion parameter language model, across its life cycle. We estimate that BLOOM's final training emitted approximately 24.7 tonnes of~\carboneq~if we consider only the dynamic power consumption, and 50.5 tonnes if we account for all processes ranging from equipment manufacturing to energy-based operational consumption. We also study the energy requirements and carbon emissions of its deployment for inference via an API endpoint receiving user queries in real-time. We conclude with a discussion regarding the difficulty of precisely estimating the carbon footprint of ML models and future research directions that can contribute towards improving carbon emissions reporting.
1 Introduction
The paper motivates systematic carbon-footprint tracking for ML because computing infrastructure emissions are substantial but difficult to estimate precisely. It studies BLOOM as an initial broader estimate spanning equipment manufacturing, training, and deployment.
- ML infrastructure contributes to ICT-sector emissions, but the extent of its contribution remains unclear because global computing infrastructure is distributed.
- BLOOM is a 176-billion-parameter LLM whose training requires millions of GPU hours and emits carbon.
- The study estimates BLOOM’s broader carbon footprint, including computing-equipment manufacturing and API-based model deployment.
- Rather than target an exact emissions number, the study estimates relative contributions from stages of the deployment process.
2 Related Work
Prior work has estimated training emissions, developed carbon-accounting tools, and examined additional life-cycle factors. However, estimates remain incomplete, tools are seldom reported in ML publications, and broader evidence across models and use cases is needed.
- Empirical Studies on ML CO2 Emissions: Most empirical studies estimate CO2 emissions from model training, while broader studies examine trends in ML energy requirements and emissions.
- Empirical Studies on ML CO2 Emissions: Existing studies disagree on whether ML-model emissions will grow or shrink, motivating estimates across a broader variety of models and use cases.
- Tools for Estimating Carbon Impact: Carbon-estimation tools either track energy and emissions during training or produce higher-level post-training estimates.
- Tools for Estimating Carbon Impact: These tools remain seldom used in ML publications, and their reported estimates vary significantly.
- Additional Factors: Complementary research considers conference attendance, hardware manufacturing, full ML life cycles, and certification of social and environmental impacts.
3 Background and Methodology
BLOOM is a 176-billion-parameter multilingual model developed through the year-long BigScience workshop and trained on the Jean Zay cluster. The methodology adapts life-cycle assessment to available information, covering equipment manufacturing through deployment and expressing greenhouse gases as CO2 equivalents.
- 3.1 The BLOOM Model: BLOOM has 176 billion parameters and was trained on 1.6 terabytes of data spanning 46 natural and 13 programming languages.
- 3.1 The BLOOM Model: The model was developed during a year-long BigScience workshop involving over a thousand researchers and trained on CNRS’s Jean Zay computer cluster.
- 3.1 The BLOOM Model: The project included data processing, tokenization, architecture engineering, evaluation, and smaller experiments that informed the final 176B-parameter architecture.
- 3.2 Methodology: Because complete cradle-to-grave information was unavailable, the study applies LCA to stages from training-equipment manufacturing through model deployment.
- 3.2 Methodology: CO2 equivalents convert different greenhouse gases into a common unit using their global-warming potential relative to carbon dioxide.
4 Results: Carbon Emissions of the BLOOM Model
The study estimates BLOOM’s carbon footprint across training and deployment by including dynamic energy, embodied emissions, idle infrastructure, and API inference. Training emissions include 24.69 tonnes of CO2eq from dynamic energy consumption, while deployment measurements show substantial energy use even without incoming requests.
- Life-cycle scope: The study compares BLOOM’s total life-cycle contributors, including training energy, equipment manufacturing, infrastructure overhead, and deployment inference.Its scope includes embodied emissions, energy-based operational emissions, and API deployment alongside model training.
- Embodied Emissions: 11.2 tonnes of CO2eq came from embodied emissions associated with the servers and GPUs used during BLOOM training.The estimate includes approximately 7.57 tonnes for servers and 3.64 tonnes for GPUs, excluding other infrastructure such as network switches and cooling equipment.
- Dynamic Power Consumption: 24.69 tonnes of CO2eq were attributed to BLOOM’s dynamic training energy consumption from 433,195 kWh of electricity.The estimate uses 1.08 million GPU hours, 400W GPU TDP, and a grid carbon intensity of approximately 57 gCO2eq/kWh.
- Idle Power Consumption: Idle-consumption accounting broadens the estimate beyond GPU dynamic power to include networking, storage, cooling, and other infrastructure needed for training.Infrastructure mode measures networking, datacenter maintenance, and cooling with servers off; idle mode measures powered-on but unused servers; dynamic mode measures active training.
- Deployment and Inference: 914 kWh powered the BLOOM API instance over approximately 18 days, with 75.3% consumed by GPUs, 22.7% by RAM, and 2% by the CPU.The deployment handled 230,768 real-time requests without batching, while energy consumption remained approximately 0.28 kWh during a 10-minute interval with almost no requests.
- Deployment and Inference: The deployment case is only one of many possible configurations because hardware, inference batch size, and deployment region can vary.The authors present the case as a starting point for estimating deployment emissions rather than a universal estimate.
5 Discussion and Future Work
The paper compares BLOOM’s carbon footprint with other LLMs and broadens accounting beyond final training to experimentation, embodied emissions, deployment, and research activities. It identifies substantial uncertainty in these estimates and calls for more granular reporting and broader environmental assessment.
- 5.1 Comparisons with other LLMs: BLOOM emitted 25 tonnes during training, less than half of OPT’s 70 tonnes and 20 times less than GPT-3’s 502 tonnes.BLOOM nevertheless consumed 433 MWh, slightly more than OPT’s 324 MWh, because carbon intensity differed across energy sources.
- 5.1 Comparisons with other LLMs: Carbon comparisons remain difficult because accounting methods differ and model hardware, token counts, architecture, and datacenter PUE affect reported values.The paper notes that PUE values are relatively similar across efficient datacenters, contributing relatively little to overall training footprints.
- 5.2 Carbon Footprint of the BigScience Workshop: 3.46 million GPU hours across BigScience experiments consumed 1,163,032 kWh and emitted approximately 66.29 tonnes of CO2eq through dynamic power.The total includes 2.2 million V100 GPU hours and 1.24 million A100 GPU hours, with final BLOOM training representing only part of the activity.
- 5.2 Carbon Footprint of the BigScience Workshop: 35.8 tonnes of CO2eq came from intermediate-model experimentation, exceeding emissions from final-model training.Model evaluation emitted 2.46 tonnes, while benchmarking, data processing, and tokenization emitted 3.3 tonnes.
- 5.2 Carbon Footprint of the BigScience Workshop: Including embodied emissions and idle equipment consumption raised the workshop estimate to 123.82 tonnes of CO2eq.The paper also notes that emissions from deployment and usage of additional shared models could not be accounted for.
- 5.3 Future Work: The authors call for more precise hardware-specific embodied-emissions data, empirical inference studies, granular carbon reporting, and broader environmental-impact assessment.They emphasize reporting energy consumption, carbon intensity, PUE, research and development, evaluation, and benchmarking rather than a single aggregate figure.