Source-linked AI summary
Collapsibility of Performance Metrics in Clinical Predictive AI
João Matos, Ben Van Calster, Richard D. Riley, Paula Dhiman, Gary S. Collins
TL;DR
Population-level predictive-AI assessments can conceal subgroup performance differences, raising a fairness-related question about whether overall metrics aggregate subgroup values validly. The paper analyzes 15 metrics using weighted decompositions or counterexamples and finds that five are non-collapsible, including AUC, whose within- and cross-group terms can produce overall values outside subgroup AUCs.
Problem
Population-level assessments can conceal subgroup performance differences, while non-collapsible metrics do not equal weighted averages of subgroup-specific values.
Method
The paper examines 15 performance metrics by expressing collapsible metrics as weighted combinations of subgroup values and using counterexamples to disprove collapsibility.
Results
Five metrics are non-collapsible, including AUC, whose overall value decomposes into within- and cross-group AUC terms.
Takeaways & Limitations
Reporting collapsibility properties can improve the interpretability and transparency of fairness assessments.
Takeaways & Limitations
The paper presents special cases as theoretical exceptions and recommends generally treating AUC as non-collapsible in the population of interest.
Abstract
from arXiv · showhide
Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI, with a focus on the area under the receiver operating characteristic curve (AUC, also known as c-statistic). Methods: We investigate the collapsibility of 15 performance metrics, either by expressing each metric as a linear combination of its stratum specific values or, where non-collapsible, by providing a counterexample inspired by Simpson's paradox as a formal disproof. Results: Five performance metrics (AUC, calibration intercept, calibration slope, expected calibration error, and Nagelkerke R^2) are shown to be non-collapsible, and ten (O:E ratio, logloss, Brier score, accuracy, F1-score, true positive rate, true negative rate, positive predictive value, negative predictive value, and net benefit) are shown to be collapsible. The AUC is shown to be non-collapsible because it decomposes into within- and cross-group AUC terms when subpopulations coexist, such that its overall value may fall outside the range of subgroup specific AUCs. Conclusions: Non-collapsibility of performance metrics has important consequences for reporting, model appraisal, and fairness evaluation. It can generate spurious differences between subgroup and overall performance, which may mislead fairness evaluations. Explicitly acknowledging and reporting the collapsibility properties of performance metrics improves both the interpretability and transparency of fairness assessments.
1 Introduction
Population-level model performance can conceal subgroup differences, motivating careful interpretation of aggregated and subgroup-specific results. The paper examines whether commonly reported performance metrics support valid aggregation for fairness assessment.
- Population-level evaluations can conceal important differences in model performance across subgroups.
- Collapsibility requires the marginal metric value to equal a weighted average of subgroup-specific values.When this holds, the aggregation artefact described in the paper cannot occur.
- Non-collapsible metrics can make subgroup and overall performance difficult to interpret jointly and potentially distort fairness evaluations.
- The paper examines collapsibility for 15 commonly reported predictive-AI performance metrics, focusing particularly on the AUC.
2 Collapsibility of Performance Metrics
The paper defines collapsibility through metric-specific weighted aggregation of subgroup values and distinguishes collapsible from non-collapsible measures. It reports that ten metrics are collapsible, whereas five—including AUC—are not.
- Definition: A metric is collapsible when its overall value can be written as a linear combination of subgroup-specific values using appropriate metric-specific weights.Weights may depend on subgroup proportions, outcome prevalences, or model-dependent quantities.
- Definition: Non-collapsibility occurs when no metric-specific weights recover the overall value from subgroup values.Such metrics may exhibit an aggregation artefact in which the overall value lies outside the subgroup-value range.
- Definition: Jensen’s gap captures the discrepancy between a non-collapsible metric and the weighted combination of subgroup values.The relationship may be convex or concave; Figure 1 shows zero gap for collapsible metrics and a nonzero gap for non-collapsible metrics.
- Collapsibility results: Strictly collapsible metrics include logloss, Brier score, accuracy, recall, specificity, and net benefit.Their overall values aggregate linearly using subgroup outcome prevalence or subgroup proportion.
- Collapsibility results: Generally collapsible metrics include the O:E ratio, F1-score, PPV, and NPV.These remain collapsible when weighted by model-dependent terms such as true positives.
- Collapsibility results: AUC, calibration intercept, calibration slope, expected calibration error, and Nagelkerke’s R2 are non-collapsible.Their overall values cannot be expressed as weighted averages of subgroup values.
3 AUC mixture decomposition and non-collapsibility
The overall AUC decomposes into within-group and cross-group ranking terms when subpopulations coexist, so it is generally non-collapsible. Special theoretical conditions can restore collapsibility, but the paper recommends treating AUC as non-collapsible in practice.
- Mixture decomposition: The overall AUC combines within-group and cross-group rankings, weighted by subgroup proportions and subgroup-specific outcome prevalences.Within-group terms occupy the decomposition’s diagonal, while cross-group terms occupy its off-diagonal entries.
- Non-collapsibility: Because cross-group terms contribute to the mixture, the overall AUC may fall outside the convex bounds of subgroup AUCs, violating collapsibility.This dependence is the formal reason the AUC is non-collapsible over subgroup membership.
- Non-collapsibility: Collapsibility would require expressing the marginal AUC as a weighted average of subgroup-specific AUCs.The weights depend on the marginal distribution of the stratifying variable.
- Special cases: If cross-group AUCs equal the corresponding within-group AUCs, the overall AUC can be written as a linear combination of subgroup AUCs with weights πg.Equal subgroup outcome prevalences provide an additional special case for simplifying the expression.
- Practical interpretation: In practice, the AUC should generally be treated as non-collapsible in the population of interest.The collapsible scenarios are presented as theoretical exceptions, and differing cross-group AUCs motivate counter-intuitive fairness evaluations.
- Illustrative derivation: The closed-form mixture expression is derived under normality or truncated-normal approximations for illustration and interpretability.The paper notes that normally distributed estimated probabilities may not accurately reflect real-world scenarios.
4 AUC non-collapsibility and fairness implications
The AUC is non-collapsible because pooled performance combines within-group and cross-group comparisons, allowing overall AUC to fall outside subgroup values. Subgroup proportions, prevalence, calibration, and case-mix differences can therefore produce fairness patterns that do not necessarily reflect differences in discrimination.
- The overall AUC entangles within-group and cross-group pairwise comparisons, so it can breach the range defined by subgroup AUCs.This aggregation artefact complicates interpretation and communication of performance across heterogeneous subpopulations.
- Counter-intuitive fairness scenarios: 0.86 versus 0.86 subgroup AUCs can coexist with an overall AUC of 0.78 when cross-group terms are highly asymmetric.The example has xAUCA,B = 0.36 and xAUCB,A = 0.99, producing an overall value below both subgroup values.
- Counter-intuitive fairness scenarios: An overall AUC of 0.76 can match subgroup B while subgroup A reaches 0.92, because cross-group terms and subgroup composition affect the mixture.The cited scenario reports AUCB = 0.76, AUCoverall = 0.76, and AUCA = 0.92.
- Case-mix differences: Differences in subgroup proportions alter the relative influence of AUC components even when model behaviour is identical across subgroups.Subgroup proportions do not change the AUC or xAUC terms themselves, but they change their weights in the overall mixture.
- Case-mix differences: Different subgroup prevalences shift estimated-probability distributions and amplify or attenuate particular strata’s contributions to overall AUC.These prevalence-driven shifts affect the mixture decomposition through both probability means and stratum weights.
- Miscalibration differences: Overall AUC variation may reflect subgroup calibration discrepancies rather than genuine differences in discrimination.The paper therefore emphasizes evaluating discrimination and calibration jointly, alongside subgroup and overall performance.
5 Discussion
Non-collapsibility can obscure the relationship between subgroup and population performance and create disparities driven by population composition. The paper recommends considering collapsibility and reporting population characteristics transparently in fairness evaluations.
- Non-collapsible metrics can create inconsistencies between subgroup and overall performance, undermining confidence in fairness evaluations.Overall values need not faithfully reflect subgroup behaviour or equal a linear combination of subgroup values.
- Differences driven by subgroup proportions or outcome prevalence can make subgroup improvements fail to improve, or even diminish, overall performance.The overall AUC also combines within-group and cross-group comparisons, potentially conflating discrimination with calibration differences.
- Five metrics are non-collapsible—AUC, calibration intercept, calibration slope, ECE, and Nagelkerke R2—while ten listed metrics are collapsible.The collapsible set includes O:E ratio, logloss, Brier score, accuracy, F1-score, TPR, TNR, PPV, NPV, and net benefit.
- AUC non-collapsibility arises from decomposition into within- and cross-group components, so its overall value is not always a weighted average of subgroup AUCs.The paper notes that each subgroup AUC can be smaller than the overall AUC.
- AUC collapsibility holds when non-event estimated-probability distributions are identical across subgroups.The paper attributes non-collapsibility to distributional variation from case-mix or subgroup miscalibration differences.
- Fairness comparisons between subgroup and overall performance are not necessarily directly comparable because differences may arise naturally.Transparent reporting should include population characteristics, outcome distributions, and fairness-related design choices.
A.1 Collapsibility Proofs
The appendix investigates the collapsibility properties of the performance measures summarized in Table 1.
- The proof section examines the performance measures summarized in Table 1.
A.1.1 O:E Ratio
The O:E ratio is collapsible because it can be expressed using group-wise counts with weights proportional to expected outcomes under the model.
- The derivation substitutes group-wise observed and expected counts into the O:E expression.
- The O:E ratio is collapsible with weights proportional to group-expected counts under the model.
A.1.2 Calibration Intercept
The calibration intercept is bounded between subgroup-specific intercepts when groups are pooled, but this bounding relation does not establish collapsibility.
- The pooled outcome mean is the group-size-weighted average of subgroup outcome means.
- The pooled calibration intercept is the unique solution of the pooled estimating equation.Uniqueness follows from strict concavity of the log-likelihood and the resulting unique score-equation root.
- The overall intercept αall lies between the subgroup intercepts αA and αB.The result follows from monotonicity and the intermediate value theorem.
- The bounding result does not imply that the calibration intercept is a convex or linear combination of subgroup intercepts.
A.1.3 Calibration Slope
The calibration slope is non-collapsible because pooled logistic-regression coefficients are generally not weighted averages of subgroup-specific coefficients.
- The calibration slope is estimated as a logistic-regression coefficient fitted to the entire dataset.Because logistic-regression coefficients are generally non-collapsible, the pooled slope is not generally a weighted average of subgroup slopes.
- Pooling subgroups changes the linear-predictor and outcome distributions nonlinearly, preventing the overall calibration slope from being a weighted average of subgroup slopes.
A.1.4 Expected Calibration Error (ECE)
ECE measures average absolute calibration deviation across probability bins, but the absolute value prevents linear aggregation across groups, making ECE non-collapsible.
- ECE measures the average absolute deviation between predicted and observed outcome rates across probability bins.
- The absolute value within each bin prevents ECE from aggregating linearly across groups.
- ECE is therefore expected to be non-collapsible.
A.1.5 Logloss / Cross-Entropy
The paper classifies several predictive-performance metrics by whether pooled performance can be expressed as a weighted combination of subgroup values, using formulas and counterexamples to establish the distinction.
- A.1.5 Logloss / Cross-Entropy: Logloss and Brier score are collapsible because they are empirical averages of per-instance losses.
- Accuracy is collapsible with weights given by group sizes.
- F1-score is collapsible with metric-specific weights proportional to 2TPg + FPg + FNg.
- TPR, TNR, PPV, and NPV are collapsible with weights proportional to their respective group-specific outcome or prediction totals.The weights correspond respectively to actual positives, actual negatives, predicted positives, and predicted negatives.
- Net benefit is collapsible with group-size weights.
- Calibration slope, Nagelkerke R2, and ECE are non-collapsible in toy examples because pooled values are not weighted averages of subgroup values.The pooled calibration slope is 0.979 versus subgroup values 0.934 and 0.963; pooled Nagelkerke R2 is 0.623 versus 0.481 and 0.440; pooled ECE is 0.0297 versus 0.0378 and 0.0379.