Source-linked AI summary

Life-Cycle Emissions of AI Hardware: A Cradle-To-Grave Approach and Generational Trends

Ian Schneider, Hui Xu, Stephan Benecke, David Patterson, Keguo Huang, Parthasarathy Ranganathan, Cooper Elsworth

arXiv:2502.01671v1cs.ARcs.AI

TL;DR

AI hardware emissions lack comprehensive, standardized assessment across the hardware lifecycle, limiting consistent quantification. This paper applies a first-party cradle-to-grave LCA to five TPUs and introduces CCI to compare emissions per computation. It reports a 3x CCI improvement from TPU v4i to TPU v6e, while identifying scope boundaries and metric limitations.

  • Problem

    Comprehensive and consistent assessments of AI hardware greenhouse-gas emissions remain limited despite growing AI compute requirements.

  • Method

    The study applies a cradle-to-grave LCA to five TPUs using first-party data, direct fleet measurements, a defined hardware functional unit, and manufacturing inventories.

  • Results

    3x CCI improvement is measured from TPU v4i to TPU v6e across Google’s TPU fleet.

  • Takeaways & Limitations

    CCI provides a way to evaluate AI hardware emissions per computation and estimate training or inference carbon footprints, while hardware and software efficiency gains can compound.

  • Takeaways & Limitations

    The functional unit excludes peripheral components beyond accelerator and host trays, auxiliary compute and storage resources, and edge or end-user devices.

Abstract

from arXiv · show

Specialized hardware accelerators aid the rapid advancement of artificial intelligence (AI), and their efficiency impacts AI's environmental sustainability. This study presents the first publication of a comprehensive AI accelerator life-cycle assessment (LCA) of greenhouse gas emissions, including the first publication of manufacturing emissions of an AI accelerator. Our analysis of five Tensor Processing Units (TPUs) encompasses all stages of the hardware lifespan - from raw material extraction, manufacturing, and disposal, to energy consumption during development, deployment, and serving of AI models. Using first-party data, it offers the most comprehensive evaluation to date of AI hardware's environmental impact. We include detailed descriptions of our LCA to act as a tutorial, road map, and inspiration for other computer engineers to perform similar LCAs to help us all understand the environmental impacts of our chips and of AI. A byproduct of this study is the new metric compute carbon intensity (CCI) that is helpful in evaluating AI hardware sustainability and in estimating the carbon footprint of training and inference. This study shows that CCI improves 3x from TPU v4i to TPU v6e. Moreover, while this paper's focus is on hardware, software advancements leverage and amplify these gains.

1 Introduction

This study addresses limited comprehensive accounting of AI hardware emissions by applying a cradle-to-grave LCA to five TPUs and introducing CCI for comparing carbon emissions per computation. It also measures operational emissions directly and examines generational changes in hardware efficiency.

  • Scope and motivation: Comprehensive cradle-to-grave LCA covers AI hardware emissions from raw-material extraction through manufacturing, operation, transportation, and retirement.The inventory includes data-center construction, development, serving, and end-of-life stages, while excluding edge devices and auxiliary compute and storage under the defined functional unit.
  • Scope and motivation: Five Google TPUs are evaluated using first-party data, targeting AI accelerators and their attached host computers.The paper presents the analysis as a comprehensive assessment intended to inform quantification and reduction of data-center service emissions.
  • Metrics and contributions: CCI is a normalized CO2e/performance metric for comparing AI machines with different performance and carbon footprints; lower CCI is better.Operational and embodied CCI variants support narrower evaluations.
  • Metrics and contributions: Direct fleet-wide measurement improves operational-emissions estimates over TDP proxies by capturing actual utilization, software, and power heterogeneity.The paper notes that deployed TPU v4 utilization can vary by 60%.
  • Novelty and generational analysis: The study is presented as the first published AI-accelerator LCA to estimate manufacturing emissions and document a repeatable process for future LCAs.Earlier AI-product studies generally covered only one or two life-cycle stages.
  • Novelty and generational analysis: CCI improves 3x from TPU v4i to TPU v6e across directly measured production accelerators.The comparison provides a longer-term view of operational and embodied efficiency across successive hardware generations.

2 Definition of concepts and assumptions

The study defines an AI computer functional unit, lifecycle assumptions, emissions accounting choices, and a manufacturing method for cradle-to-gate assessment. It combines standardized accounting with proprietary chip parameters while explicitly excluding peripheral and auxiliary resources.

  • Functional unit and boundaries: The functional unit is one data-center AI computer containing one or more accelerator trays connected to one host tray.Peripheral components beyond the trays and auxiliary computing and storage resources are excluded, while cooling overheads are included in operational emissions.
  • Functional unit and boundaries: A six-year AI-machine lifespan is used as a reasonable assumption for the lifecycle assessment.The lifecycle categories follow the GHG Protocol accounting and reporting standard.
  • Emissions accounting: Operational emissions use IPCC AR5 emission factors to provide a standardized and comparable assessment of AI hardware climate impact.Emissions are reported in grams or kilograms of CO2e.
  • Emissions accounting: Scope 2 emissions are calculated by multiplying operational electricity consumption by an electricity emission factor, with location-based and market-based methods considered.Location-based accounting excludes carbon-free-energy procurement, while market-based accounting reflects purchased electricity sources.
  • Manufacturing assessment: The LCA follows ISO 14040/14044 and combines IMEC virtual-fab inventories with proprietary chip-manufacturing parameters for advanced TPU modeling.Parameters include technology node, die size, yield rates, and fluorinated-GHG abatement ratios; the result is a published estimate of accelerator manufacturing emissions.

3 CCI: A novel CO2e/performance metric

The paper introduces compute carbon intensity (CCI), a CO2e-per-utilized-FLOP metric grounded in first-party workload measurements, to compare AI hardware across generations and account for embodied carbon. It also examines utilization, data-format, and reliability considerations that affect how CCI should be interpreted.

  • CCI measures CO2e per utilized floating-point operation, using a fixed computation amount rather than a FLOPs-per-second rate.Its unit is CO2e/FLOP.
  • CCI approximates CO2e/Goodput by using actual workload power and performance data while naturally reflecting underutilization.The approach does not include the additional goodput adjustment for wasted work caused by unreliability.
  • CCI can estimate training emissions from approximate FLOP counts; TPU v4 yields about 107 tonnes of CO2e for GPT-3, versus 89 tonnes on TPU v5p.The TPU v4 estimate combines 25 tonnes of embodied and 82 tonnes of market-based operational CO2e.
  • Operational CCI uses measured utilized FLOPs and whole-system energy, enabling comparisons that include embodied carbon and account for lower-carbon electricity sources.FLOPs are aggregated across chips and five-minute intervals for the period of interest.
  • CCI comparisons adjust for utilization differences with duty-cycle measurement and propensity-score weighting; this reduces the v6e-over-v4i improvement from 74% to 66%.The same adjustment reduces v5p over v4 from 24% to 16%, while v5e remains at 14% over v4i.
  • CCI treats measured FLOPs equally across numerical formats, making the effects of narrower data types difficult to identify.The authors identify evaluating quality and CO2e trade-offs across numerical formats as future work.
  • CCI is anchored to the algorithms and workloads represented by the current fleet, so changes that reduce both computation and energy may not change the metric proportionally.The paper notes that a 10% reduction in required FLOPs would suggest a 10% carbon reduction, while fleet-wide reductions in both energy and effective FLOPs could leave CCI unchanged.

4 Results and discussion

Across TPU generations, performance-normalized lifecycle carbon intensity improves substantially even as newer machines can have higher absolute manufacturing emissions. Operational electricity remains the dominant lifecycle source, while cleaner electricity further improves CCI.

  • Generational trends: 3x: TPU v6e improves compute carbon intensity versus TPU v4i.The quadrupling of v6e’s systolic array size partly explains the gain.
  • Life-cycle emissions: Operational emissions account for approximately 70% with Google’s CFE procurement and approximately 90% without it.Google’s CFE purchases reduce operational emissions by more than half under market-based accounting.
  • Life-cycle emissions: Under hourly 24/7 accounting, v6e has 3,305 kg lifetime electricity emissions and 182 gCO2e/ExaFLOP operational CCI.Operational emissions represent 82% of total v6e emissions under this accounting standard.
  • Cleaner electricity scenarios: 90% CFE for operations and electricity-related manufacturing would improve v6e’s total CCI by 4.6x.The scenario assumes 90% CFE globally for both use and manufacturing electricity.
  • Manufacturing trends: Embodied CCI decreases 66% from v5e to v6e because peak performance improves 4.7x while manufacturing emissions increase 1.8x.Newer TPUs may be more carbon intensive per machine but deliver substantially greater workload throughput.
  • Manufacturing trends: Manufacturing emissions increase 1.8x from v4i to v6e, driven mainly by larger TPU dies, HBM modules, and host DRAM capacity.The v6e host has 1,536 GB of DRAM versus 512 GB in v5e.

5 LCA versus Corporate Inventories

The paper’s lifecycle accounting differs from Google’s annual corporate inventory because the two approaches use different amortization rules and inventory boundaries. Consequently, annual hardware emissions do not directly represent the lifecycle emissions of deployed machines.

  • Lifecycle comparison: Using market-based operational electricity, approximately 25% of the studied AI hardware lifecycle emissions come from manufacturing.The study finds operational electricity emissions dominate the lifecycle total.
  • Accounting methodology: LCA amortizes hardware emissions over its lifetime, whereas Google’s corporate inventory reports capital-good emissions fully in the first year.Only approximately 1/6 of a hardware lifetime and 1/20 of a data-center lifetime occur in the first year.
  • Inventory boundaries: The LCA covers an AI accelerator and host machine, while the corporate inventory includes wider non-AI sources such as consumer-device manufacturing.The differing functional-unit boundaries make the inventories non-equivalent.
  • Interpretation: Annual corporate hardware emissions are driven by first-year purchases and data-center growth rather than directly tracking the lifecycle of deployed machines.This difference follows from both the accounting methodology and Google’s continuing data-center expansion.

6 Related work

Prior work has assessed selected AI workloads and hardware components, but the literature lacks a comprehensive and consistent cradle-to-grave view. This paper situates its LCA and CCI metric within those fragmented assessments.

  • Scope of prior work: Existing studies often focus on training or serving rather than the full AI hardware lifecycle.Retirement emissions and comprehensive accelerator manufacturing assessments remain limited.
  • Hardware assessments: Prior hardware estimates use placeholders or extrapolations, whereas the TPU results report 208–585 kg for TPU manufacturing and 2300–4672 kg per TPU machine.These comparisons illustrate substantial differences between proxy estimates and the paper’s TPU measurements.
  • Hardware assessments: Published embodied-emissions estimates vary widely, including 423–15,593 kgCO2e across six server-computer LCAs.The reviewed estimates average approximately 3900 ± approximately 5900 kg.
  • Operational assessments: Training-emissions estimates are highly sensitive to assumptions such as electricity emission factors and the energy-sampling window.A cited meta-analysis reports training emissions per model increased approximately 100x from 2012 to 2023.
  • Carbon-intensity metrics: CCI is proposed as CO2e/ExaFLOP to normalize emissions by computation when comparing AI hardware and workloads.The metric is intended to support comparisons across differing performance and carbon footprints.
  • Measurement assumptions: TDP can overestimate TPU energy use because its ratio to actual average power varies from 2X–6X.The paper therefore distinguishes rated power from measured average power.

7 Conclusion

The study finds that operational emissions dominate AI hardware’s lifetime footprint, while CCI and workload emissions decline across TPU generations. It also shows that hardware and software efficiency improvements can compound carbon reductions.

  • 3x hardware and 3x software-efficiency improvements would yield a 9x reduction in carbon emissions for the same model quality.CCI captures hardware emissions per ExaFLOP, while software efficiency reduces the FLOPs required for a given model quality.
  • More advanced TPU chips and increasing memory demand drive embodied emissions, but manufacturing CCI declines each generation.Memory alone represents more than a third of embodied emissions, while improved hardware design offsets rising manufacturing emissions on a performance-normalized basis.
  • 70% to 90% of total emissions come from operations, while manufacturing contributes under 25% and data center construction under 5%.
  • 27% to 39% lower gCO2e per training step was observed over one TPU generation when training the same model.
  • 3x improvement in CCI was delivered from TPU v4i to TPU v6e across Google’s fleet.
  • The LCA combines cradle-to-gate manufacturing assessment, direct fleet-wide power measurements, and a system boundary covering manufacturing, operations, and end-of-life.The functional unit is one AI computer containing accelerator trays and a host tray; peripheral infrastructure and auxiliary resources are excluded, while cooling overheads are included.

A.2.1 Life-cycle inventory data.

The life-cycle inventory combines primary and secondary data with bottom-up component accounting, then estimates emissions across hardware production, transport, operation, data-center infrastructure, and end-of-life.

  • Inventory construction: LCI data combine bills of materials, supplier-provided primary data, industry datasets, and detailed teardowns to fill component data gaps.The repositories include Sphera Professional Databases and ecoinvent, with foreground models configured for relevant primary data.
  • Inventory construction: The product-system scope accounts for physical tray components including processors, memory, PCBAs, thermal systems, mechanical parts, and electromechanical components.
  • Transportation: Transportation emissions cover ground and international shipping from manufacturing sites through hubs to data centers, with shipment weight including packaging.Components use ocean or air transport models depending on the hardware category.
  • Operational emissions: Operational energy is measured at tray PSUs, aggregated across machine trays, averaged over representative machine-days, and projected across a six-year lifetime with PUE overhead.Measurements are collected from deployed machines, with missing sensor observations excluded.
  • End of life: End-of-life recovery could offset embodied emissions by up to 4%, but the analysis conservatively claims no avoided-burden credit because outcomes depend on reuse and treatment routes.

B.1 Manufacturing and transportation emissions

The manufacturing methodology builds detailed tray inventories using technology-node-specific semiconductor models and component-level data, yielding substantial manufacturing and transportation emissions for TPU v5e.

  • Method: The TPU v5e LCA begins with bills of materials for accelerator and host trays, followed by detailed LCI construction and LCA modeling.IMEC.netzero supplies technology-node-specific data for large integrated circuits.
  • Method: The manufacturing model incorporates advanced nodes, advanced packaging, supplier-specific direct-emissions abatement, and country-specific electricity mixes.Examples include 5 nm, 3 nm, and 2 nm nodes, TSVs in HBM, and 2.5D silicon-interposer stacking.
  • Results: 2,277 kgCO2e/machine is estimated for TPU v5e manufacturing, comprising two accelerator trays and one host tray.The estimate is reported as GWP100 at the machine level.
  • Results: 471 kgCO2e is added by transportation to TPU v5e machine-level scope 3 emissions.
  • Results: 62% of TPU v5e tray manufacturing emissions come from TPU chips and thermal-management solutions.For the host tray, DRAM DIMMs and the mainboard dominate emissions.
  • Operational accounting: Measured fleet power is preferred to TDP because TDP can be 2x to 3x higher than actual measured power.The study therefore uses average hourly power rather than TDP or a fixed TDP fraction.

C Introduction to 24/7 CFE

The paper presents 24/7 Carbon-free Energy accounting as a more temporally and geographically precise way to assess electricity emissions than conventional annual matching.

  • Motivation: Annual CFE matching can leave companies reliant on carbon-emitting grid electricity for over 50% of demand despite matching annual consumption.The paper links this mismatch to rapidly increasing AI electricity demand in regions with substantial fossil generation.
  • Accounting principle: 24/7 CFE requires each purchased MWh to match a consumed MWh within the same local grid boundary and hour.
  • Accounting principle: The 24/7 emissions metric is higher than current market-based emissions while still reflecting reductions relative to the location-based metric.Geographically or temporally mismatched purchases cannot reduce reported electricity emissions under this method.
  • Implementation: Shortfalls count against the goal when grid supply and clean-energy contracts cannot provide 100% CFE in a given hour.CFE percentages are reported ex post after load, production, and grid-mix data are settled.
  • Implication: The paper recommends adopting more accurate accounting standards when calculating Compute Carbon Intensity and addressing AI electricity emissions.

D Initiatives and Process Flow of Manufacturing Emissions Methodology

The paper describes industry efforts to standardize ICT and data-center-equipment LCAs through shared category rules, configurable models, and aligned data practices.

  • Process flow: Figure 7 presents the process flow used in the paper.
  • Industry standardization: Google is collaborating on Product Category Rules to streamline supplier data collection and improve comparability across ICT product carbon assessments.The initiative seeks alignment with ISO 14040/44/67 and the GHG Protocol.
  • Industry standardization: An Open Compute Project workstream is developing data-center-equipment LCA practices and configurable, parameterized model building blocks.The workstream also addresses data accuracy and coverage for hotspot identification and carbon-aware design.

E On-Duty Machine Power for Benchmark Workloads

The study estimates on-duty machine power from workload telemetry, filtering out inactive periods and validating incomplete runs before calculating per-run averages.

  • Power and duty-cycle data are collected at five-minute intervals throughout each workload run.
  • Only timestamps where every assigned machine reaches at least 80% duty cycle are included, excluding inactive periods.
  • Figure 8 presents duty cycle and power over time for representative RLHF on v5e and SFT on v6e runs.
  • For each run, average on-duty machine power is computed across all machines and included timestamps.
  • Incomplete runs can contribute energy and Scope 2 carbon per step when completed steps provide meaningful metrics, subject to manual power validation.

F Propensity Score Weighting

The paper uses propensity score weighting to compare TPU generations at balanced utilization levels, reducing distortion from differing duty-cycle patterns and producing more modest CCI improvement estimates.

  • Newer platforms combine inherent efficiency gains with higher, more energy-efficient utilization rates, complicating direct generational comparisons.
  • Propensity score weighting balances utilization levels across generations and reduces utilization-related confounding.
  • The method compares powerful TPUs v4 and v5p separately from cost-efficient TPUs v4i, v5e, and v6e.
  • Propensity scores are calculated from generation proportions within grouped duty-cycle levels, then used for inverse probability weighting.
  • Weights are inversely related to assignment propensity, giving less-represented accelerators greater weight in the balanced comparison.
  • Table 4 reports results before and after weighting, with equalized duty cycles and a more accurate, more modest estimate of CCI improvement.
Loading 2502.01671v1…