Source-linked AI summary

PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation

Krishna Rao, Andrew Dumit, Shaena Ulissi, Jacob Feintzeig, P. James Joyce, Daniel Frank, Steven Watson, Jonathan Glidden, Gizem Ilayda Dinc, Travis M. Kwee

arXiv:2608.27716v1cs.AI

TL;DR

PCFBench addresses the lack of evaluations that test both final PCF estimates and the intermediate steps producing them. It decomposes cradle-to-gate PCF estimation into expert-annotated tasks and finds substantial weaknesses in compositional assembly and document extraction, supporting diagnostic evaluation of LLMs.

  • Problem

    Existing PCF evaluations measure aggregate emissions or isolated subtasks, leaving compositional intermediate-step correctness insufficiently evaluated.

  • Method

    PCFBench decomposes cradle-to-gate PCF estimation into six typed, independently evaluated tasks using expert-annotated data.

  • Results

    Compositional pipelines reach only 37–58% within 2× of verified kgCO2e truth, while models also show low extraction claim-F1 and under-specification failures.

  • Takeaways & Limitations

    PCFBench provides a multi-axis diagnostic for assessing whether AI-generated emissions estimates expose their compositional steps and support trustworthy use.

  • Takeaways & Limitations

    The benchmark is limited to cradle-to-gate impacts, excluding use-phase and end-of-life emissions, and its 614-item dataset has limited fine-grained statistical power.

Abstract

from arXiv · show

AI systems are being deployed on high-stakes, domain-specific workflows that demand correctness not just in the final output, but at every intermediate step. One such workflow is estimating a product carbon footprint (PCF), the greenhouse-gas emissions attributable to a physical product. AI agents are increasingly being used to generate PCFs, but existing evaluations score either total emissions (hiding error sources and cancelling mistakes) or sub-tasks in isolation (missing compositional interactions). We introduce PCFBench, the first benchmark to carve PCF modeling into independently-evaluable tasks that require decomposition, retrieval, ontology matching, and numerical extraction. It comprises 614 expert-labelled items across six tasks. Together they probe reasoning under under-specification, conflicting context, and numerical constraints. Across eight frontier LLMs from four providers, no single model dominates. Although the strongest models estimate total product emissions within 2 times of declared totals on 77% of products, this rate drops to 37-58% when the PCF is generated step by step, with only 45-75% obeying mass conservation. These failures undermine the transparency practitioners need to compare products and drive decarbonization. We release the dataset and evaluation harness to support targeted progress.

1 Introduction

PCF estimation is a context-dependent, multi-step workflow whose aggregate evaluation can hide intermediate failures. PCFBENCH decomposes cradle-to-gate estimation into typed, independently evaluated tasks with expert annotations.

  • A PCF is the kgCO2e attributable to one declared product unit, calculated through LCA and used in EPDs, procurement, and climate disclosures.
  • Context-dependent LCA judgments vary with study goals, geography, time, and reporting conventions, creating under-specification challenges for LLMs.
  • Intermediate PCF stages are needed to attribute impacts, compare products consistently, quantify decarbonization levers, and support verification and hotspot analysis.
  • PCFBENCH decomposes cradle-to-gate estimation into six LLM-evaluable tasks with typed schemas, task-specific metrics, and 614 expert-annotated items.
  • The pipeline recursively decomposes products, triages mapping decisions, maps materials, extracts material and energy rates, and predicts total emissions before deterministic aggregation.
  • PCFBENCH contributes the first PCF benchmark, expert-labelled datasets, and per-task profiles for eight frontier models plus a specialized mapping baseline.

2 Related Work

Existing LLM sustainability evaluations largely measure aggregate emissions or declarative knowledge rather than the operational, compositional capabilities required for PCF workflows. PCFBENCH addresses this gap through task-resolved evaluation aligned with documented failure modes.

  • Existing LLM-based PCF systems are evaluated only on aggregate kgCO2e, hiding which sub-task caused an error.
  • PCFBENCH provides task-relevant public ground truth across its benchmark coverage, enabling comparisons with prior datasets.
  • Prior sustainability benchmarks cover emission factors, sustainability knowledge, SDG monitoring, or declarative LCA knowledge rather than operational PCF capability.
  • PCFBENCH applies compositional evaluation across domain-decomposed tasks and targets failure modes including conflicting context, numerical reasoning, ontology alignment, and under-specification.

3 Benchmark Design

PCFBENCH models cradle-to-gate PCF estimation as a recursive pipeline whose expert-dependent stages are evaluated separately, while aggregation remains deterministic. Direct and compositional modes are scored against the same EPD ground truths.

  • The PCF sum multiplies each physical input rate q_i by its emission factor EF_i and adds the resulting contributions.
  • Six expert-judgment points identify inputs, triage decomposition, select emission factors, estimate material and energy quantities, and predict total kgCO2e.
  • Direct evaluation asks one LLM call to return a single kgCO2e estimate from EPD product context without intermediate decomposition.
  • Compositional evaluation chains Tasks 1–5 and deterministically sums per-component contributions, making every intermediate step inspectable.
  • Both evaluation modes produce one kgCO2e value per product and use a 2× acceptability threshold against the same EPD ground truths.
  • Task 6, multiply and sum, is deterministic and therefore not evaluated as an LLM task.

4 Datasets

PCFBENCH contains 614 expert-labelled items spanning six evaluable tasks, seven product categories, and broad emissions ranges. Its datasets cover decomposition, triage, mapping, document extraction, and EPD-based total-emissions prediction.

  • 614 items span six evaluable tasks, seven product categories, and five orders of magnitude in declared kgCO2e.
  • Task 1 contains 94 EPD-derived BOM items, with a median of four components and a maximum of eight per product.
  • Task 2 contains 200 balanced map-versus-decompose decisions, with 100 items in each label class.
  • Task 3 contains 109 material-mapping cases spanning edge cases, specialized challenges, and no-good-match scenarios.
  • Tasks 4–5 use 36 technical PDFs yielding 89 ground-truth claims supported by 227 evidence quotes.
  • Figure 2 encodes category counts with number, shading, and log(count + 1) colour, while its histogram shows the log-scale kgCO2e-per-kg distribution.
  • Task 7 contains 175 third-party-verified EPDs with declared kgCO2e values from 0.062 to 34.1.

5 Results and Discussion

PCFBENCH reveals that compositional PCF estimation is substantially harder than direct emission prediction, with errors arising across decomposition, mapping, numerical constraints, context use, and extraction. Performance varies by task, model, cost, and access to tools or context, while intermediate failures limit transparency.

  • Compositional versus direct estimation: 37–58% of products fall within 2× of verified kgCO2e in compositional evaluation, compared with 60–77% for direct prediction.Opus 4.6 and Gemini 3.1 Pro lead compositional performance at 58%, while DeepSeek reaches 37%.
  • Decomposition and mapping: 0.64–0.75 decomposition recall leaves a quarter to a third of expert-listed components missing, despite 0.85–0.92 precision.On aluminium foil, Opus 4.6 returns 2 of 5 expert-listed components, with additives, finishes, and alloying elements commonly omitted.
  • Numerical constraints: 25–55% of compositional BOMs violate mass conservation, while 47–86% contain at least one zero-mass ghost component.Both biases push final kgCO2e estimates downward.
  • Context effects: About 10 points of average Task 7 within-2× accuracy come from adding material composition, with little further movement after adding region.For Aluminium ingot – Semi-Primary, Opus 4.6 shifts from 4.5 kgCO2e/kg using the name alone to 0.82 after disclosure of 90.8% post-consumer scrap, against a declared 0.94.
  • Model and tool comparisons: Tool access improves Task 3 mapping unevenly but yields only marginal gains for top Task 2 triage performance.Five of eight models gain on Task 3, while Gemini 3 Flash loses 0.13; best triage accuracy rises from 0.705 non-agentic to 0.725 agentic.
  • Scientific document extraction: 0.27–0.51 material claim-F1 and 0.29–0.53 energy claim-F1 show that extraction remains weak because models over-extract scope-incompatible claims.In one query, Opus 4.6 returned six expert values plus thirty extras, including a value that included upstream polymer production despite the exclusion request.
  • Model and tool comparisons: No single model dominates: six different systems take the per-task crowns, and the cost–accuracy frontier is non-monotone.Gemini 3 Flash tops the mean normalized accuracy ranking and accuracy per dollar.

6 Limitations and Future Work

PCFBENCH is limited by dataset scale, cradle-to-gate scope, omitted document retrieval, and independent rather than fully paired task datasets. Its dual-use risks include greenwashing, erroneous emissions propagation, and regulatory misuse.

  • Limitations: 614 expert-annotated items are small relative to synthetic or crowd-sourced benchmarks, and fine-grained strata with fewer than 10 items lack statistical power.The annotation depth supports stratified analyses but increases collection cost.
  • Limitations: All tasks target cradle-to-gate impacts, excluding use-phase and end-of-life emissions that dominate some product categories.This defines the benchmark’s principal scope boundary.
  • Limitations: Document retrieval is out of scope because Tasks 4–5 evaluate extraction given the relevant technical document.Upstream locating of literature, supplier datasheets, handbooks, or trade databases is not evaluated.
  • Limitations: The datasets are not chained per product, so the compositional pipeline cannot measure fully matched real-product error propagation.Extraction PDFs are not paired with EPDs, and BOM, triage, and mapping labels are not attached to individual EPDs.
  • Future Work and Risks: AI-generated emission factors can propagate unnoticed into procurement, supplier scorecards, disclosures, and carbon-credit issuance without source and activity-ID provenance.Practitioners should preserve provenance so auditors can independently re-verify inputs.

7 Conclusion

PCFBENCH introduces a process-aware benchmark for evaluating PCF estimation sub-tasks and overall footprint prediction. Its conclusion emphasizes diagnostic transparency while showing that compositional estimation remains substantially harder than direct total prediction.

  • 7 Conclusion: PCFBENCH is the first decomposed benchmark for AI-generated PCF estimation, using task-specific metrics and expert-annotated data for each expert-workflow sub-task.The benchmark spans reasoning under under-specification, long-document numerical extraction, and order-of-magnitude estimation.
  • 7 Conclusion: PCFBENCH provides a shared yardstick for exposing compositional steps and assessing whether AI-produced emissions numbers are transparent enough for decarbonization use.The benchmark is intended as a multi-axis diagnostic for frontier LLMs.
  • 7 Conclusion: Total-kgCO2e prediction complements isolated sub-task evaluation by testing whether a model can produce a defensible overall footprint estimate.The evaluation uses 175 third-party-verified EPDs under four progressively informative context settings.

B.1 Task 1: Product Decomposition Dataset

Task 1 evaluates whether models can infer an ordered bill of materials from product information without explicit composition percentages. The dataset and scoring accommodate ambiguity in fabrication level, while the 94-product reuse creates a documented limitation.

  • Task 1: Product Decomposition: Models must list up to 8 input materials at practitioner-relevant granularity, ordered from highest to lowest mass contribution.Input rates and percentages are excluded to isolate compositional knowledge.
  • Task 1: Product Decomposition: Consumer packaging counts as a BOM component, whereas distribution packaging used only for inter-facility transport is out of scope.This boundary follows the cradle-to-gate framing.
  • Task 1: Product Decomposition: The aluminium example requires recognizing recycled scrap, primary aluminium, remelted ingots, and alloying elements from the product description without percentages.Components are ordered by mass contribution.
  • Task 1: Product Decomposition: A Gemini 2.5 Flash judge aligns predicted and expected components across 1-to-1 and 1-to-N compositional matches before computing precision, recall, F1, Kendall τ, and exact-set match.This accommodates synonyms and legitimate differences in fabrication level.
  • Limitations: The dataset reuses the same 94 products as the composition-bearing Task 7 subset, so cross-task analyses must account for overlap.The paper also notes that 94 items are small relative to real-world product diversity and that judge scoring has some non-determinism.

B.4 Task 3: Background Database Mapping Dataset

Task 3 tests context-sensitive selection of defensible ecoinvent mappings from raw material names and optional organizational context. Its deliberately difficult, ambiguous dataset is constrained by narrow coverage and sparse context fields.

  • Task 3: Background Database Mapping: The 109-item mapping dataset covers typical mappings, abbreviations, foreign-language names, catalysts, packaging materials, and no-good-match cases curated by sustainability experts.Inputs reflect raw, unstructured material names encountered in industrial BOMs.
  • Task 3: Background Database Mapping: Each item supplies a raw material name, optional context fields, and an unordered list of equally valid ecoinvent reference products.A prediction is correct when it matches any acceptable option; 38 items also specify banned substrings.
  • Task 3: Background Database Mapping: The abbreviation “Pca” is disambiguated toward PET variants rather than polycarbonate by the receiving-supplier context “water bottle maker.”All three listed products are acceptable options, but context identifies the relevant PET variants.
  • Task 3: Background Database Mapping: 60% of items have vagueness severity 4 or 5, while only 17% have severity 1, concentrating the dataset on difficult inputs.Challenge sets include vague inputs, abbreviations, foreign-language names, catalysts, and graceful-degradation cases.
  • Task 3: Background Database Mapping: 58% of items have exactly one acceptable option, 29% have two, and the remainder have 3–5 options, reflecting ambiguity in the mapping space.The multi-option structure treats all listed mappings as valid at score time.
  • Limitations: The dataset is small relative to ecoinvent’s approximately 10,000-activity vocabulary, with uneven category coverage and only 24 items containing contextual fields.The context-ablation evaluation is therefore limited to a small subset.
  • Candidate Set: The curated picklist filters ecoinvent v3.11 to practitioner-relevant market activities, retaining three energy markets by explicit exception because they lack mass-based units.The candidate set is designed for direct-match decisions rather than the full vocabulary.

F.6 Extraction (Tasks 4–5)

Tasks 4–5 evaluate document-grounded extraction and prior-knowledge estimation under controlled context settings, using strict claim-level matching. The broader pipeline also chains these outputs into compositional PCF estimates, where missing or zero-mass components create measurable failures.

  • Extraction evaluation: The headline extraction comparison varies only the system prompt and whether the source document is attached.Query-only prompts invite best-guess estimates, whereas query-plus-document prompts prohibit estimation and hallucination.
  • Extraction evaluation: Claim F1 is computed from exact value–unit tuple matches, with precision and recall aggregated across all items.Unit aliases are normalized through a fixed table, while unsupported synonyms or conversions are not credited.
  • Extraction evaluation: Strict matching credits a claim only when its canonicalized value agrees through six decimal places with the ground truth.The rule is intended to measure recovery of the annotator’s number rather than approximate agreement within a unit family.
  • Compositional pipeline: The compositional pipeline chains decomposition, triage, mapping, material-rate estimation, and energy-rate estimation before deterministic aggregation.Each component contribution is combined with an emission factor to produce one kgCO2e value.
  • Compositional pipeline: 25–55% is the headline mass-conservation violation rate, defined as products with material sums below 100% or above 300%.The upper threshold is intended to distinguish gross over-listing from ordinary process losses, scrap, or recycled-content double-counting.
  • Compositional pipeline: 47–86% of runs contain at least one ghost component, a listed BOM entry assigned zero mass and therefore excluded from the emissions sum.This failure mode is identified as the largest mass-balance problem and a likely contributor to under-prediction.

H.2 Task 1 judge: human-agreement audit

The Task 1 compositional-match judge was audited against human judgments, while the benchmark reports broader task-specific performance and calibration diagnostics. The audit found high agreement and only a small expected effect on the published F1 estimate.

  • Judge audit: 44 of 50 (88%) judge verdicts agreed with the human audit, with a Wilson 95% CI of [0.76, 0.94].Five disagreements were missed valid matches and one was an invalid accepted match, indicating a small conservative bias.
  • Judge audit: The expected mean F1 correction is approximately +0.010, raising Gemini 3.1 Pro’s published F1 from 0.740 to approximately 0.750.The shift is small enough that model ordering is preserved.
  • Compositional diagnostics: The compositional diagnostic separates below-100% and above-300% mass sums from ghost components, which are listed with zero mass.Ghosts are shown separately because they always push estimates downward and occur much more often than the other two failure modes.
  • Calibration: Task 3 mapping is approximately calibrated with ECE = 0.086 over 871 pooled predictions, whereas Task 2 triage is severely overconfident with ECE = 0.258 over 1,600.The pooled reliability curves show under-confidence in mapping and over-confidence in triage across models.
  • Calibration: Every model compresses confidence to [0.5, 1.0], so abstention thresholds at or below 0.5 filter almost no items.The lower half of the confidence scale is essentially unused.

I Annotation Guidelines

PCFBENCH’s annotation process combines domain expertise, structured tooling, task-specific instructions, and deterministic reconciliation. The resulting benchmark includes curated difficulty and ambiguity cases, but its extraction labels and judge-based components retain documented limitations.

  • Annotator profile: Six sustainability scientists created the expert-curated annotations, including three PhDs in LCA or adjacent sustainability disciplines.All annotators encounter mapping and technical-document screening tasks in their professional work.
  • Annotation process: The curation philosophy targeted realistic LCA situations, a trivial-to-hard difficulty range, and multiple product categories.Desk research surfaced diagnostic cases exercising specific failure modes, including ambiguity and trace-content traps.
  • Annotation process: Annotators used a structured application exposing candidate picklists or PDFs and persisting each person’s decisions separately.Separate records enabled deterministic cross-annotator reconciliation during dataset construction.
  • Reconciliation: Tasks 4–5 use a unanimous-on-include rule because objective truth is unavailable for some difficult scope and granularity judgments.The paper avoids reducing this uncertainty to a single kappa-style agreement number.
  • Dataset structure: PCFBENCH contains 614 task items across six evaluation tasks, represented as per-task JSONL records with inputs, expected outputs, and metadata.The benchmark is a curated slice rather than an exhaustive enumeration.
  • Limitations: Known residual noise includes consensus bias toward prominent claims, an LLM-based Task 1 judge with 88% human-audit agreement, and a region-agnostic Tasks 2–3 candidate set.These properties bound interpretation of the benchmark’s labels and transfer to other settings.

J.5 Uses

The dataset is released for evaluating and extending PCF and general LLM-agent capabilities, including compositional reasoning, extraction, matching, and calibrated decision-making. Its use remains bounded by geographic and source-domain coverage, and it is not a substitute for audited regulatory LCA reporting.

  • Supported uses: PCFBENCH supports sustainability-agent fine-tuning, per-step reward shaping, and studies comparing compositional with monolithic PCF pipelines.It also enables generic probes of decomposition, error attribution, evidence-grounded extraction, numerical reasoning, ontology matching, and tool use.
  • Supported uses: Tasks 2–3 support comparisons of calibrated abstention, semantic matching, and tool-use versus in-context retrieval on shared items.Both tasks include single-shot and agentic variants.
  • Scope boundaries: The candidate set for Tasks 2–3 is region-agnostic, and geographic stratification is identified as future work.This limits direct coverage of regional variation in those tasks.
  • Scope boundaries: Tasks 4–5 sources skew toward industrial-process literature, so generalization to other process classes is untested at this scale.The paper also characterizes the Task 1 judge as a strong but imperfect signal.
  • Scope boundaries: PCFBENCH metrics do not verify a system for production deployment or justify marketing it as a verified LCA tool.Regulatory PCF reporting still requires an audited LCA.
  • Distribution and maintenance: The dataset is publicly distributed through Hugging Face with companion code, licensing terms, versioned releases, and tracked errata.Users are advised to pin reported numbers to a release tag because point releases may correct annotations or expand coverage.

K Broader Impact

PCFBENCH is intended to diagnose process-level reliability in AI-generated PCF estimation and support comparison, training, and validation. Its broader-impact discussion emphasizes human oversight, provenance, methodological boundaries, contamination risks, and uneven coverage.

  • Intended use: PCFBENCH is intended to identify reliable and unreliable PCF sub-tasks, compare systems on an expert-annotated reference, and support training and validation.The benchmark is framed as a diagnostic measurement instrument rather than a general certification.
  • Deployment boundary: Human-in-the-loop deployment is recommended because median absolute relative error on Task 7 remains above 25% across all models and settings.The error exceeds 40% in the name-only setting, and the authors discourage treating leaderboard performance as a substitute for human LCA review.
  • Risk controls: Incorrect extraction or mapping can propagate an apparently plausible emission factor into procurement, supplier scorecards, disclosures, and carbon-credit issuance.The paper recommends preserving provenance to source documents and checking generated factors before consequential use.
  • Regulatory scope: PCFBENCH performance cannot establish compliance with mandatory climate-disclosure regimes because the benchmark does not encode their specific methodological requirements.The authors specifically discourage citing PCFBENCH scores in regulatory filings or disclosure attestations.
  • Benchmark integrity: The public EPDs and technical documents may occur in pre-training data, so memorization-driven gains cannot be ruled out.The authors also note that widespread adoption could generate training data as providers optimize against the leaderboard.
  • Coverage and bias: Scores are not equally informative across products or regions because the mapping picklist is region-agnostic and product-category coverage is uneven.The benchmark therefore may not expose country- or grid-specific activity-selection failures.

L Dataset Availability, Licensing, and Data Rights

PCFBENCH provides public dataset, harness, outputs, and reproduction resources, with access and licensing conditions varying by task and use. Full compositional reproduction additionally depends on an ecoinvent licence.

  • Availability: The dataset and evaluation resources are hosted on Hugging Face and GitHub, with unrestricted access for non-commercial academic research.A Croissant metadata document accompanies each dataset.
  • EPD rights: Tasks 1 and 7 redistribute selected EPD summary fields under written permission from EPD International for academic research.The permission covers fields including product information, declared unit, declared cradle-to-gate kgCO2e, and key supporting data.
  • Extraction data: Tasks 4 and 5 use publicly available open-access technical PDFs and release source URLs, extraction queries, OCR text, and adjudicated ground-truth claims.Ground truth includes values, units, and verbatim evidence quotes from the sources.
  • Mapping rights: Task 3 mapping labels and challenge annotations are author-created and released under CC BY-NC-SA-4.0 for academic and non-commercial use.ecoinvent reference-product names are included for identification only and remain ecoinvent property.
  • Triage data: Task 2 triage items contain expert-labelled map-versus-decompose decisions, anonymised supply-chain state, and no customer-identifying information.The authors release these items under CC BY-NC-SA-4.0.
  • Reproducibility: Per-task results are reproducible without ecoinvent access, but fully reproducing the compositional pipeline requires an ecoinvent licence.The compositional pipeline sums agent outputs using deterministic ecoinvent v3.11 emission factors.
Loading 2608.27716v1…