Source-linked AI summary
Holistically Evaluating the Environmental Impact of Creating Language Models
Jacob Morrison, Clara Na, Jared Fernandez, Tim Dettmers, Emma Strubell, Jesse Dodge
TL;DR
AI model developers often disclose final-training energy or carbon but provide less information about development, hardware manufacturing, and water use. This paper measures those impacts across the OLMo lifecycle using detailed power and environmental accounting. Across the model series, the authors report 493 tCO2eq and 2,769 kL of water, while identifying development costs and measurement transparency as major concerns.
Problem
Environmental reporting for AI models often omits development, hardware manufacturing, and water use despite the growing environmental costs of model creation and deployment.
Method
The study estimates energy, carbon, and water impacts across model development, final training, inference, and embodied hardware for OLMo models spanning 20 million to 13 billion active parameters.
Results
493 tCO2eq and 2,769 kL of water were attributed to the model series, including development, hardware manufacturing, and final training.
Takeaways & Limitations
Development costs are substantial and often unreported, while inference can eventually outweigh training costs for widely deployed models.
Takeaways & Limitations
Important sources, including transportation and hardware end-of-life disposal, remain difficult to measure without proprietary industry information.
Abstract
from arXiv · showhide
As the performance of artificial intelligence systems has dramatically increased, so too has the environmental impact of creating these systems. While many model developers release estimates of the power consumption and carbon emissions from the final training runs for their latest models, there is comparatively little transparency into the impact of model development, hardware manufacturing, and total water usage throughout. In this work, we estimate the real-world environmental impact of developing a series of language models, ranging from 20 million to 13 billion active parameters, trained on up to 5.6 trillion tokens each. When accounting for hardware manufacturing, model development, and our final training runs, we find that our series of models released 493 metric tons of carbon emissions, equivalent to powering about 98 homes in the United States for one year, and consumed 2.769 million liters of water, equivalent to about 24.5 years of water usage by a person in the United States, even though our data center is extremely water-efficient. We measure and report the environmental impact of our model development; to the best of our knowledge we are the first to do so for LLMs, and we find that model development, the impact of which is generally not disclosed by most model developers, amounted to ~50% of that of training. By looking at detailed time series data for power consumption, we also find that power usage throughout training is not consistent, fluctuating between ~15% and ~85% of our hardware's maximum power draw, with negative implications for grid-scale planning as demand continues to grow. We close with a discussion on the continued difficulty of estimating the environmental impact of AI systems, and key takeaways for model developers and the public at large.
1 INTRODUCTION
AI's rapid growth has made the environmental cost of developing and deploying language models increasingly important to measure. This paper estimates impacts across model development, training, inference, hardware, energy, carbon, and water, while emphasizing incomplete industry transparency.
- AI's rapid progress has increased the environmental costs associated with developing and deploying large language and multimodal models.Training requires substantial computational resources, energy, carbon emissions, and water consumption.
- The study estimates energy use, carbon emissions, and water consumption for OLMo models ranging from 20 million to 13 billion active parameters and 1.7 to 5.6 trillion tokens.It also includes upstream embodied impacts and downstream inference estimates.
- Model development, including hyperparameter tuning and pre-training experiments, is measured alongside final training and inference.The authors identify this as the first such reporting for large language model development to their knowledge.
- Power consumption is measured at sub-second intervals rather than assumed to equal GPUs' theoretical maximum, while deployment impacts are estimated across model sizes and scenarios.This approach is intended to improve reporting of both training variability and inference costs.
- The paper concludes that greater industry transparency is needed because much larger systems deployed globally may produce emissions tens or hundreds of times larger than those reported here.It assigns primary reporting and reduction responsibility to developers of the largest models.
2 RELATED WORK
Prior environmental assessments of AI models provide partial coverage, while water use, embodied carbon, and development costs remain comparatively underreported. Existing studies measure selected operational or manufacturing impacts but leave important parts of the lifecycle uncharacterized.
- Most publicly available models report no climate impact, although some recent studies estimate emissions from manufacturing, training electricity, or idle cluster electricity.Coverage remains uneven across environmental dimensions.
- Prior work on language and vision models has measured electricity and carbon emissions with granular or region-specific data, but generally omits development costs, water consumption, or inference.The cited studies therefore provide only partial lifecycle coverage.
- Water-consumption estimates for AI systems often rely on speculative training locations and energy use, while embodied-carbon estimates remain limited by opaque hardware manufacturing.These gaps are especially pronounced for closed models and state-of-the-art computational hardware.
3 METHODOLOGY
The methodology measures environmental impacts across development, training, inference, hardware, energy, carbon, and water rather than reporting only final-training cost. It combines detailed power monitoring with data-center efficiency, grid, cooling, hardware, and deployment assumptions.
- The study expands conventional final-training accounting to include development, training, inference, embodied hardware impacts, water use, and operational greenhouse-gas emissions.This supports a more comprehensive model-lifecycle assessment.
- Operational impacts are calculated from power use, data-center efficiency, and local grid carbon intensity, with carbon emissions represented as CO2e = P · PUE · CI.The study uses cluster-specific assumptions for Texas and Iowa facilities.
- Development costs include hyperparameter experiments and other runs commonly excluded from final-training estimates.The paper frames the cost of a scientific result as depending on per-example cost, dataset size, and the number of experiments.
- Water consumption combines onsite and offsite water-use effectiveness, with closed-loop cooling assigned 0 liters per kWh onsite and cluster-specific offsite values.The assumed offsite values are 1.29 L/kWh for Jupiter and 3.10 L/kWh for Augusta.
- Power during development and training is logged at sub-second intervals for one node and extrapolated across nodes, making the reported GPU-only estimates lower bounds on total power.This monitoring avoids assuming constant maximum GPU draw.
- Embodied emissions and water are amortized over hardware lifetime and multiplied by GPU hours used for development and training.The analysis covers models from 20 million to 13 billion active parameters trained on standard multi-GPU servers.
- Deployment costs are estimated rather than observed because the models are not deployed, using chat-style scenarios and single-H100 inference measurements.Measured inference excludes overhead from holding models in memory or listening for requests.
4 RESULTS
The results quantify environmental costs across hardware manufacturing, model development, final training, and simulated inference for the OLMo model series. Development and training together account for substantial emissions and water use, while deployment costs and power fluctuations introduce additional considerations.
- Hardware manufacturing: 22 tCO2eq and 4.8 kL of water were attributed to hardware manufacturing under a four-year GPU lifespan assumption.The estimate amortized embodied impacts over 1.65 million GPU hours.
- Development: 159 tCO2eq and 843 kL of water were consumed by development runs, with about 70% of development costs incurred at the 7B and 13B scales.Development included controlled experiments, initialization studies, mid-training recipes, hyperparameter selection, and data-mixture scaling experiments.
- Putting it in perspective: 493 tCO2eq were emitted and at least 2,769 kL of water were consumed across the model series.These totals include hardware manufacturing, development, and final training.
- Inference: Most models required hundreds of millions to tens of billions of inferences to outweigh training costs, except the most over-trained models.The paper notes that deployment-focused models can reach this tipping point as usage grows.
- Power fluctuations during training: GPU power remained steady during training but dropped quickly while saving checkpoints, producing inconsistent power consumption over time.The authors connect similar fluctuations in other systems to checkpointing and synchronization between nodes.
5 DISCUSSION
The discussion highlights that environmental impacts arise across development, training, deployment, water use, and power consumption, while reporting remains incomplete. It also connects smaller, cheaper models and checkpointing practices to growing downstream resource demands and grid-management challenges.
- 5.1 MORE TRANSPARENCY IS (STILL) NEEDED: GPU manufacturing impacts remain essentially unknown because manufacturers do not provide reliable embodied-emissions data.The paper identifies this as a persistent obstacle to estimating total model-development impact.
- 5.1 MORE TRANSPARENCY IS (STILL) NEEDED: Development costs from failed runs, hyperparameter searches, and architecture testing represent a substantial, often unreported share of environmental impact.The paper emphasizes that this matters especially for AutoML and scaling-law experiments, where many trained models may be discarded.
- 5.1 MORE TRANSPARENCY IS (STILL) NEEDED: Water consumption is substantial even for comparatively small models and varies sharply with data-center cooling and electricity-generation methods.The authors argue that missing information about when, where, and how models are trained limits quantification of the issue.
- 5.2 SMALL CHOICES DURING TRAINING CAN HAVE LARGE IMPACTS: Smaller deployment-optimized models can reduce training and inference costs while expanding usage across APIs and devices.The discussion notes that lower costs may increase total downstream impact, especially as models are deployed in new scenarios.
- 5.2 SMALL CHOICES DURING TRAINING CAN HAVE LARGE IMPACTS: Transparency is needed across the full deployment pipeline because smaller models are increasingly used in scenarios that can greatly increase total inference.Many immediate-use requests cannot be batched to exploit cheaper or cleaner energy.
- 5.2 SMALL CHOICES DURING TRAINING CAN HAVE LARGE IMPACTS: Checkpointing causes rapid power drops during training, creating repeated supply-and-demand fluctuations that challenge large-scale power-grid control.The figure reports active-training power above 600W and checkpointing power just above 100W for a single node.
A.1 ADDITIONAL INFERENCE SIMULATION DETAILS
The additional simulations benchmark inference using ShareGPT prompts in an online chat setting and provide supplementary results for a larger model set. The authors note that these controlled assumptions may favor OLMo in some comparisons, although they do not consider the effect significant.
- A.1 ADDITIONAL INFERENCE SIMULATION DETAILS: Inference simulations use ShareGPT prompts in an online chat setting, with additional results reported for a larger model set.The appendix points readers to Table 4 for the expanded simulations.
- A.1 ADDITIONAL INFERENCE SIMULATION DETAILS: Shorter training context lengths may give OLMo models an advantage on longer inference examples, although the authors do not consider this factor significant.They also observe that Llama 3.1 8B is faster and less energy intensive than OLMo 7B models in their measurements.
A.2 LIMITATIONS
The limitations concern simplified inference simulations, uncertain embodied-impact estimates, and restricted generalizability of observed training-cost trends. The reported resource figures therefore apply most directly to the studied settings and assumptions.
- A.2 LIMITATIONS: Inference and deployment estimates rely on controlled, limited simulations rather than real serving data across models.The authors also lack information about most other models’ actual usage.
- A.2 LIMITATIONS: The inference experiments omit variations such as quantization, decoding algorithms, likelihood evaluation, and edge-device deployment.They simulate default SGLang token-ingestion and generation settings rather than the full range of practitioner configurations.
- A.2 LIMITATIONS: Training-cost trends that are linear with parameter count across four orders of magnitude may not hold across all scales or settings.Decentralized or multi-data-center training could introduce substantially greater communication overhead.
- A.2 LIMITATIONS: Inference resource measurements account for active GPU processes but exclude CPU, RAM, and server-overhead usage, making them lower bounds in similar settings.The table also uses a specific cluster and fixed WUE, PUE, and carbon-intensity coefficients.