Source-linked AI summary

Climate-model factor separation with Shapley values and efficient sampling

Abel Jansma

arXiv:2609.11948v1math.OCcs.GTphysics.ao-phphysics.data-an

TL;DR

Nonlinear climate-model responses make it difficult to attribute changes among multiple factors and their interactions. The paper identifies linear-sum and shared-interaction factorisations with the permutation and Harsanyi-dividend forms of the Shapley value, proving their equivalence for any number of factors. It also shows that sampled paths can provide tunable-cost approximations, while important uncertainty sources remain outside the reported error bounds.

  • Problem

    Nonlinear responses make factor attribution depend on the combination and ordering of changed factors, creating a combinatorial problem over 2^n simulations.

  • Method

    The paper uses Möbius inversion for interaction decomposition and identifies linear-sum and shared-interaction attributions with two equivalent Shapley-value representations.

  • Results

    The two factorisations coincide for any number of factors, and reversal-paired sampling reached about 0.5 degree-years median maximum error with 546 cached evaluations.

  • Takeaways & Limitations

    The identification imports established Shapley estimators and variance-reduction schemes into climate-model factor-separation experiments.

  • Takeaways & Limitations

    Reported sampling error bounds cover deterministic models but not internal variability, finite integration, or parameter and structural uncertainty.

Abstract

from arXiv · show

Climate models often feature nonlinear responses to changes in model parameters or boundary conditions. Factor separation asks how much of the resulting change should be assigned to altered factors and their interactions. Lunt et al. (2010) showed that two factor attribution methods, the \emph{linear-sum} and \emph{shared-interaction} methods, coincide for $n\le 4$ factors and conjectured this holds for arbitrary $n$. We show that the linear-sum and shared-interaction formulas are respectively the permutation and Harsanyi-dividend representations of the Shapley value. Their equality therefore follows from a classical theorem in cooperative game theory. To our knowledge, this is the first explicit identification of the Stein--Alpert and Lunt climate-model factor-separation methods with Möbius/Harsanyi coefficients and Shapley values. The result imports many established and efficient Shapley sampling methods into climate-model experimental design, and extends to vector-valued responses.

1 Introduction

Factor separation attributes nonlinear climate-model changes to altered factors and their interactions, but exhaustive attribution becomes combinatorially complex. The paper proves that linear-sum and shared-interaction methods are two equivalent representations of the Shapley value.

  • Motivation: Factor separation applies to climate-model changes involving conditions or parameters such as CO2, ice sheets, orography, vegetation, land cover, and emissions.The method asks how much of a prediction change should be assigned to individual altered factors and their interactions.
  • Motivation: Nonlinear responses make a factor’s assigned effect depend on which other factors changed, requiring 2^n simulations to study all combinations.The resulting model states form an n-dimensional cube, complicating attribution and interpretation.
  • Contribution: The paper distinguishes decomposition of all 2^n interaction terms from attribution of n values to individual factors.This separation clarifies that identifying interactions and assigning factor credit are related but distinct tasks.
  • Contribution: Möbius inversion solves decomposition by rewriting Stein–Alpert interactions, while Shapley values solve attribution.The framework treats interaction identification and factor assignment with corresponding mathematical tools.
  • Contribution: Linear-sum and shared-interaction formulas are respectively the permutation and Harsanyi-dividend representations of the Shapley value.Their equality follows from the classical equivalence of these Shapley representations.

2 Related work and scope of the contribution

Related work connects factor separation to experimental design and applies Shapley values in climate, energy, sensitivity analysis, and machine learning. The paper’s specific identification of the two Lunt factorisations with Shapley representations remains distinct from those contributions.

  • Design-of-experiments connections: Cleveland et al. connect Stein–Alpert factor separation to simple effects, regression coding, reduced designs, and factors with more than two levels.Their work clarifies the statistical structure of factor-separation experiments but does not identify the linear-sum and shared-interaction formulas with Shapley values.
  • Shapley-related applications: Prior work applies Shapley values to historical carbon emissions, climate-mitigation energy indicators, sensitivity analysis, and machine-learning prediction interpretation.These applications establish broader uses of Shapley-based attribution across climate, energy, sensitivity analysis, and machine learning.

3 Stein–Alpert factor separation as Boolean calculus

Stein–Alpert factor separation can be expressed as a set-function decomposition over binary factors. Möbius inversion gives unique interaction coefficients, while Shapley values distribute those interactions into factor attributions.

  • Set-function formulation: The model response is a set function on binary factors, with values in a real vector space and baseline normalization.Scalar climate variables use V = R, while the formulation permits vector-valued responses.
  • Stein–Alpert decomposition: Stein–Alpert decompose the response for each factor subset into irreducible contributions indexed by subsets of factors.The decomposition treats interactions as contributions associated with coalitions of factors.
  • Möbius inversion: Möbius inversion replaces recursive solution of the interaction equations with a closed-form subset-lattice calculation.This set-based formulation rewrites the original Taylor-series equation as a function over factor sets.
  • Möbius inversion: Möbius inversion makes the decomposition unique, invertible, and exactly complete, matching Stein–Alpert’s Equation (16).Completeness follows because the reconstructed decomposition satisfies the response equation exactly.
  • Shapley attribution: In cooperative game theory, the set function is a value function and the irreducible contributions are Harsanyi dividends.Shapley attribution distributes each dividend equally among coalition members, yielding the shared-interaction factorisation.
  • Shapley attribution: The shared-interaction attribution obeys Shapley’s axioms, while endpoint-reversal antisymmetry is a distinct property from player-exchange symmetry.The paper cautions that Lunt et al.’s four properties alone do not uniquely characterize the Shapley value.

4 Linear-sum factorisation as Shapley values

The linear-sum method averages marginal effects over factor orderings. Rewriting that average over subsets shows it is exactly the Shapley value and therefore equals the shared-interaction attribution for any number of factors.

  • Theorem: The resulting expression is the original Shapley value and coincides with the shared-interaction method, proving Lunt et al.’s conjecture.The equivalence holds for arbitrary n rather than only the previously examined small numbers of factors.
  • Theorem: For arbitrary n binary factors, linear-sum and shared-interaction attributions are equivalent.The theorem applies to baseline-normalized responses valued in any real vector space.
  • Permutation representation: The linear-sum attribution averages each factor’s marginal effect across all permutations of the factors.For a permutation π, the relevant predecessor set contains factors appearing before the target factor.
  • Proof: Möbius inversion rewrites the permutation-based expression as a sum over interaction coefficients.This connects the ordering-based construction to the subset-based decomposition.
  • Proof: For a fixed coalition containing factor i, the probability that all other coalition members precede i is 1/|S| under a uniformly chosen permutation.This weighting produces equal sharing of each interaction term among its coalition members.

5 Permutation sampling of Shapley factor attributions

Permutation sampling estimates Shapley factor attributions from sampled orderings rather than reconstructing all interaction dividends. The estimator preserves key attribution properties exactly, while its accuracy and variance depend on sampling design.

  • Each sampled permutation supplies one marginal-contribution observation for every factor, so shared paths use n + 1 model states and can cache repeated coalition responses.Reusing a common ordering across factors also preserves completeness at every finite sample size.
  • Completeness, null-factor behavior, and linearity are exact for sampled estimates, whereas endpoint reversal and factor-exchange symmetry require corresponding sampling-design conditions.Path independence is asymptotic: finite-sample estimates depend on sampled paths, although their expectation is the unique Shapley attribution.
  • M sampled permutations control all n factor-level Monte Carlo errors with high probability using a sample count logarithmic in n, under bounded marginal contributions.This avoids the 2^n configurations required for exact recovery of the complete response table, though sampled paths still require O(n) model states each.
  • Reversal-paired sampling reduces variance for factor i exactly when the paired marginal contributions have negative covariance, but can increase variance when covariance is positive.The paired estimator uses 2K path evaluations, matching the evaluation count of the independent estimator used for comparison.
  • In the 14-factor FaIR experiment, exact attributions require 2^14 = 16384 model evaluations, while sampled estimates track exact values with pointwise 95% Monte Carlo intervals.Figure 1 compares exact full-table attributions with running sampled estimates for CO2 and CH4; positive values increase cumulative degree-years above 2°C, and negative values reduce them.

6 FaIR illustration with a threshold-response metric

The FaIR experiment applies Shapley sampling to a nonlinear cumulative threshold-exceedance response and compares sampled estimates with exact values from the full factorial table. Reversal-paired sampling achieves substantially better cost–accuracy performance while caching repeated coalition evaluations.

  • Experimental design: The experiment uses 14 binary forcing factors under SSP5–8.5, allowing exact Shapley values from 16384 FaIR coalition evaluations for comparison with sampling estimates.The factors include greenhouse-gas species and effective radiative-forcing components.
  • Threshold-response metric: The threshold-response game produced 141.30 degree-years above 2◦C, with 32.45 degree-years of non-additive residual, equal to 23.0% of the full response.Exact Shapley attributions sum to the full response by efficiency.
  • Sampling performance: The two largest exact Shapley values are 126.81 degree-years for CO2 and 17.45 degree-years for CH4, and sampled running estimates quickly converge toward them.Convergence is assessed with pointwise normal Monte Carlo intervals based on observed marginal-contribution variances.
  • Sampling performance: At M = 50, reversal-paired sampling used 546 cached model evaluations, saved 96.7% of full-table runs, and reduced median maximum factor error to 0.52 degree-years.The full benchmark requires 16384 model evaluations; relative errors were about 1.6–2.0% for smaller factors and 1.1% for larger factors.
  • Sampling performance: At M = 25, ordinary sampling used 292 evaluations and saved 98.2% of full-table runs, but its median maximum error was 4.76 degree-years and smaller factors had larger relative errors.Factors below 5 degree-years had about 14% median relative absolute error, compared with about 10% for factors at least 10 degree-years.

7 Conclusion

The paper identifies Stein–Alpert and Lunt factor-separation attributions with Möbius/Harsanyi and Shapley-value representations, proving their equivalence for any number of factors. This connection enables sampled Shapley estimation while highlighting unresolved uncertainty from stochastic and structural model sources.

  • The linear-sum and shared-interaction factorisations are respectively the permutation and dividend representations of the Shapley value, so they coincide for arbitrary n.
  • Shapley sampling estimates factor attributions without reconstructing every interaction dividend, offering tunable computational cost for expensive climate-model simulations.
  • 546 cached evaluations with reversal-paired sampling reduced the median maximum error to about 0.5 degree-years, compared with about 4.8 degree-years at M = 25.
  • The framework connects factor separation to Möbius-inversion and Shapley-value methods, including extensions to vector-valued responses.
  • Sampling error bounds apply to deterministic models, while internal variability, finite integration, and parameter or structural uncertainty remain unaddressed.

A Three-factor sampling example

The three-factor example shows that dividend and permutation routes produce identical attributions, while reversal pairing substantially reduces sampling variance.

  • The dividend and permutation representations both return attributions of (7/3, 10/3, 13/3), summing to the total response change of 10.
  • Each of the six permutations telescopes exactly to the total response change, illustrating completeness of the permutation estimator.
  • Factor 1’s six permutation marginals have population variance 14/9, whereas reversal-pair means have variance 1/18.
  • The negative covariance between each permutation and its reversal explains the variance reduction from reversal-paired sampling.
Loading 2609.11948v1…