Source-linked AI summary
Proxy reliance in large language model decisions is uncalibrated to predictive evidence
Zengqing Wu, Chuan Xiao
TL;DR
Existing audits cannot distinguish discriminatory proxy use from warranted predictive reliance. This study measures causal proxy effects against an evidence-warranted reference in four LLMs and finds that reliance undertracks evidence, while social-label suppression is fragile.
Problem
Existing audits treat decision changes under demographic manipulation as bias, although group-correlated attributes may also carry legitimate predictive signal.
Method
The study compares causal proxy effects in four LLMs with an exactly computed, task-conditional evidence-warranted reliance level.
Results
Reliance severely undertracks predictive evidence across four LLMs; social field names suppress reliance, but in-context examples raise it above zero in every model.
Takeaways & Limitations
A normative reliance reference turns proxy-use audits into structured verdicts and provides a general framework for assessing evidence-calibrated model reliance.
Takeaways & Limitations
The study uses synthetic generating processes, and its evidence sweep covers proxies contributing up to one third of outcome variance.
Abstract
from arXiv · showhide
Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.
1 Introduction · 2 Results
The study benchmarks LLM proxy reliance against an exact, task-conditional evidence-warranted reference. Across four models, reliance is often unwarranted, undertracks predictive evidence, responds fragily to social labels and examples, and is not revealed by accuracy alone.
- 1 Introduction: The audit measures causal proxy effects against the reliance warranted by a prediction-optimal rule, distinguishing over-reliance, warranted reliance and under-reliance.Protected attributes are withheld from prompts; only proxy-mediated changes in patient rankings are measured.
- 2.1 One audit signal, three verdicts: Zero-information proxies produce unwarranted reliance under neutral labels in every model, while informative proxies yield all three audit verdicts.For Sonnet, neutral-label excess effects are +11.1, +19.7 and +14.6 percentage points at k = 6, 12 and 18.
- 2.1 One audit signal, three verdicts: Socially connoted field names lower Sonnet’s reliance below the evidence-warranted reference by −9.6 to −12.1 points across dimensions.With informative proxies, neutral-label reliance is indistinguishable from the reference at k = 6 and k = 18 but exceeds it by +7.1 points at k = 12.
- 2.2 Proxy reliance does not track the evidence: Accuracy gains largely reflect the changing true ranking: a frozen-policy control explains 12.1 of Qwen3.7-max’s 15.2 neutral-label points.Residual evidence-specific gains are small and signed in both directions.
- 2.3 Label-triggered suppression is conditional, graded and fragile: Social-label suppression is conditional and provider-specific, with contrasts of −20.7 points for Sonnet, −15.2 for GPT-5.6 Terra, −9.1 for Qwen3.7 and −6.6 for DeepSeek-V4-Flash-0731.Suppression appears across risk-gap strata and requires an open proxy channel with an ambiguous task.
- 2.4 Dimension and dependence structure act on different channels, and interact: With eighty examples, regret rises from +19.7 to +27.8 to +36.9 points as k increases, while dependence reduces regret by up to 13 points.Under neutral labels, adding attributes changes reliance from +15.7 to +23.2 points at low dependence but +12.1 to −1.5 points at high dependence.
- 2.5 Accuracy and proxy reliance are separable channels: Accuracy and proxy reliance are separable: structural manipulations move reliance by up to 24 points while accuracy remains within estimation noise, and omission, labels and distribution shift bias audits by 13.9, 8.25, 4.6 and 2.6 points.In-context examples nevertheless raise social-label reliance to +15.7 points even with zero proxy-outcome correlation.
3 Discussion
The study reframes proxy audits around deviation from evidence-warranted reliance, finding that model reliance severely undertracks the evidence. Its normative reference, ideal-learner control, and evidence-strength sweep provide a general framework, while synthetic data and limited proxy variance bound the claims.
- Core findings: The audit reference converts one behavioural signal into three verdicts by measuring how reliance deviates from what the evidence supports.The verdict depends on the deployment’s evidence structure relative to the model’s carried level of reliance.
- Core findings: Reliance severely undertracks the evidence in every model tested under both serving channels, undermining the presupposition that reliance responds adequately to evidence.This undertracking makes the warranted-integration question depend on the deployment’s evidence structure and the model’s existing reliance level.
- Framework: The framework combines a normative reliance reference, an ideal-learner control for finite evidence, and a dose-response over evidence strength.The authors present these instruments as applicable wherever researchers ask whether a model uses an input.
- Limitations: The claims are bounded by a synthetic generating process, documented abstraction risks, structural rather than outcome-realism anchoring, and coverage limited to proxies contributing up to one third of outcome variance.The paper identifies semi-synthetic designs as the natural next step and does not cover regimes where proxies dominate.
4 Methods
The study evaluates proxy reliance using simulated patients, fixed paired comparisons, and a Bayes-level reference under identical interventions. It controls task ambiguity, model execution, uncertainty estimation, and external covariance comparisons across clinical datasets.
- Stimuli and structure: Simulated patients contain 6, 12, or 18 standardized numeric attributes, with a fixed two-to-one ratio of legitimate to proxy attributes and controlled latent-factor dependence.Structure concentration rescales the share of correlation carried by the leading eigenvalue from 0 for mutual independence to 1 for a single common factor.
- Experimental design: Evaluation uses 99 fixed patient pairs per dimension and structural condition, presented in both orders across factual and protected-attribute counterfactual arms.Pairs are stratified by true-risk gap and held identical across rule, label, model, and evidence conditions; cross-dimension comparisons use separate between-sample pairs.
- Experimental design: Proxy effects are detectable only with ambiguous rules and near-tie comparisons, because explicit weighted rules or wide risk gaps make decisions deterministic or obvious.Therefore, a null effect under those alternative task designs does not establish an unbiased model.
- Reference model: The Bayes-level reference applies the same intervention to a true-risk ranker, equaling zero under zero information and using conjugate Bayesian linear regression when n = 80 labelled examples are available.A complete-knowledge line is reported only as an upper bound because finite-sample shrinkage of 2 to 4 points could otherwise be misread.
- Model evaluation: Four systems—claude-sonnet-4-5, DeepSeek-V4-Flash-0731, qwen3.7-max, and GPT-5.6 Terra—run under specified temperature, tool-choice, reasoning, and reporting settings.GPT-5.6 Terra is treated as a separate channel rather than pooled with the other systems.
- Statistical analysis: Uncertainty uses 95% pair-level bootstrap intervals with 2,000 to 5,000 resamples, while cross-dimension differences below 8 to 16 points are reported directionally.Pairs are the resampling unit because presentation orders are dependent; contrasts are between-sample and use α = 0.05 with 80% power.
Supplementary Information · Supplementary Note 1: Seed replication
A fresh-seed sensitivity analysis reproduced unwarranted proxy reliance, social-label suppression, and the dose-response verdict across the tested replication cells. The analysis used two models and two new generating-process draws, with evidence-warranted references recomputed for each seed.
- Supplementary Information: The analysis reran core cells for two models under two fresh replication seeds, redrawing the population, evaluation pairs, and calibration sample together.The sensitivity analysis covered DeepSeek-V4-Flash-0731 and Qwen3.7-max on the judgment-grid and dose-response conditions.
- Supplementary Information: 11,880 calls per model were made, and all 40 replication cells completed with zero unparseable responses.This reports the replication workload and response-completeness check.
- Supplementary Note 1: Seed replication: 16 of 16 zero-information cells showed unwarranted proxy reliance, with effects from +7.1 to +26.8 pp versus an evidence-warranted level of exactly zero.Every combination of model, seed, rule condition, and label semantics had an interval excluding zero.
- Supplementary Note 1: Seed replication: 4 of 4 model-by-seed combinations reproduced social-label suppression under no examples, with effects from −6.6 to −11.1 pp.Three of four intervals excluded zero, closely matching the frozen-seed values.
- Supplementary Note 1: Seed replication: 8 of 8 dose-response curves replicated the verdict: slopes ranged from −0.09 to +0.22, none excluded zero, and all intercepts were positive at +10.8 to +27.6 pp.Every slope interval excluded calibrated tracking at γ = 1, while every intercept was clearly positive.
- Supplementary Note 1: Seed replication: The evidence-warranted reference was recomputed for each seed, so each model was compared with the reference matching its own population.This comparison rule applied to the two fresh generating-process draws used for the dose-response fits.
Supplementary Note 2: Internal analysis plan and post-hoc revisions
The analysis plan pre-specified verdict criteria and slope tests before dose-response runs, but two directional hypotheses were revised after seeing data onto pre-written contingency branches. The plan was not externally registered.
- Two directional hypotheses were revised after seeing data onto pre-written contingency branches.The revisions concerned dependence structure’s effect on in-context learning difficulty and the sign of the dimension effect on proxy reliance under high dependence.
- Verdict criteria for the judgment grid and both slope tests—γ versus 0 and γ versus 1—were specified before the dose-response runs.
- The internal analysis plan was not externally registered.
Supplementary Note 3: Real-data anchoring detail
Supplementary anchoring sets the manipulated structure band at [0.114, 0.314] and compares proxy manipulation against real levels 0.068 and 0.429. The two administrative protected attributes fall below the lower proxy level, so the manipulated range covers the upper part of the real range.
- Structure-axis anchors: The manipulated structure band is [0.114, 0.314], with intervals defined by dataset-level bootstrap.These are the structure-axis anchors reported in Supplementary Table S2.
- Proxy-axis anchors: Proxy-axis anchors are compared against manipulated levels 0.068 and 0.429; both administrative protected attributes fall below 0.068.Thus, the manipulated range covers the upper part of the real range.
Supplementary Note 4: Pair-level covariation between error and proxy sensitivity
At the individual-pair level, errors and sensitivity to protected-attribute counterfactuals covary, even after controlling for pair difficulty. This association does not establish a causal direction.
- Pair-level covariation: Incorrectly answered pairs were more sensitive to protected-attribute counterfactuals, with mean r = +0.22 and significance in 8 of 12 cells.The combined association had p < 0.001.
- Pair-level covariation: Controlling for pair difficulty left the association essentially unchanged, from +0.219 to +0.223.The passage reports this as evidence that the covariation persists after accounting for pair difficulty.
- Pair-level covariation: The association cannot identify whether errors cause proxy consultation or proxy sensitivity contributes to errors.The passage states that the observed association is compatible with either causal direction.
Supplementary Note 5: Full audit-error decomposition · Supplementary Note 6: Counterfactual scaling sensitivity
The audit-error decomposition shows that full omission and label mismatch dominate distribution shift, limiting standard correction. Counterfactual scaling changes a neutral-label estimate but does not change the conclusion, and frozen-scaler mode supports future runs.
- Supplementary Note 5: Full audit-error decomposition: Audit-error terms measure bias in the proxy effect caused by auditing conditions differing from deployment, with residuals reported after importance weighting.Reported signs are magnitudes because the underlying gaps are negative, so the audit understates the deployment effect.
- Supplementary Note 5: Full audit-error decomposition: The label-mismatch term compares neutral and social full-panel reference cells, while distribution shift is the covariance between audit-to-deployment weights and pair-level proxy effect.The distribution-shift term excludes zero, but its importance-weighted residual does not separate from zero at this sample size.
- Supplementary Note 5: Full audit-error decomposition: Full omission and label mismatch are 5.3 and 3.1 times the size of distribution shift, respectively.Distribution shift is the only decomposition term targeted by standard correction.
- Supplementary Note 6: Counterfactual scaling sensitivity: Standardizing displayed attributes within the 5,000-person sample makes counterfactual proxy regeneration shift each column’s mean and standard deviation by a factor of order 1/n.This can move a rendered comparator value by one unit in the second decimal.
- Supplementary Note 6: Counterfactual scaling sensitivity: At k = 12, scaling sensitivity affects five of the 99 evaluation pairs.Excluding those pairs moves the zero-information neutral-label proxy effect estimate from +22.2 to +21.3.
- Supplementary Note 6: Counterfactual scaling sensitivity: The scaling adjustment produces no conclusion changes, and released code provides a frozen-scaler counterfactual mode for future runs.That mode reuses the factual sample’s standardization constants in both counterfactual arms.
Supplementary Note 7: Task designs that cannot detect the proxy channel · Supplementary Note 8: Serving-channel sensitivity of the pooled slope · Supplementary Note 9: Order-instability robustness of the proxy-specific effect
The supplementary analyses show that proxy effects require an appropriately ambiguous task design, are insensitive to pooled serving-channel choices, and remain robust to position adjustment and order-consistency checks, with a model-specific decisiveness mechanism.
- Supplementary Note 7: Task designs that cannot detect the proxy channel: Explicit scoring rules and field weights make models execute deterministically, producing near-zero errors, exactly zero proxy effects, and no semantic-manipulation effects.Unstratified evaluation pairs also suppress detectable proxy effects when comparisons are wide enough that legitimate fields settle most decisions.
- Supplementary Note 7: Task designs that cannot detect the proxy channel: Proxy-pathway identification requires withholding the scoring rule, using homogeneous technical field names, and including near-tie evaluation pairs.All confirmatory experiments use a design meeting these requirements.
- Supplementary Note 8: Serving-channel sensitivity of the pooled slope: +0.046 is the pooled slope across three temperature-zero models, while including the non-isomorphic fourth model gives +0.052.Both estimates exclude calibrated tracking and neither excludes zero, so the pooled verdict does not depend on including the separate-channel model.
- Supplementary Note 9: Order-instability robustness of the proxy-specific effect: Every pair is presented in both orders, so additive position preferences average out within each counterfactual arm and cannot bias the paired estimator.A position-adjusted linear probability model reproduces the paired estimator essentially exactly in every cell.
- Supplementary Note 9: Order-instability robustness of the proxy-specific effect: 0.83 pp is the largest difference between the paired estimator and a position-adjusted linear probability model across all model cells.This explicit adjustment confirms that position preference does not materially alter the paired estimates.
- Supplementary Note 9: Order-instability robustness of the proxy-specific effect: 5.9 to 8.2 pp is the median absolute estimate change after restricting to pairs whose two orders agree, with no sign changes above 5 pp.The restriction moves estimates modestly for the three models with little position preference.
- Supplementary Note 9: Order-instability robustness of the proxy-specific effect: DeepSeek-V4-Flash-0731 shows higher position rates in A = 0 than A = 1, 0.75 versus 0.64, while consistent-pairs-only estimates can collapse toward zero.In an example-supplied cell, the estimate changes from +24.2 to +2.7, consistent with part of the proxy effect modulating decisiveness.
Supplementary Note 10: Full cell-level results
The supplementary note reports all sixteen replication cells for DeepSeek-V4-Flash-0731, Qwen3.7-max, and GPT-5.6 Terra, with Sonnet results available in released artifacts. It defines the reported metrics, dependence conditions, and recomputability from raw records.
- DeepSeek-V4-Flash-0731: Supplementary Table S6 reports all sixteen replication cells for DeepSeek-V4-Flash-0731.The cells use LH/HH/LL/HL dependence-structure conditions and report regret, PSE, the ideal-learner line, and excess in percentage points.
- Metrics and replication: Excess is defined as PSE minus the complete-knowledge justified level, while dose-response cells report the ideal-learner line separately.All quantities are recomputable from raw per-call records using the released code.
- Qwen3.7-max: Supplementary Table S7 reports all sixteen replication cells for Qwen3.7-max.The table uses LH/HH/LL/HL dependence-structure conditions and reports regret, PSE, the ideal-learner line, and excess in percentage points.
- GPT-5.6 Terra: Supplementary Table S8 reports all sixteen replication cells for GPT-5.6 Terra.The table uses LH/HH/LL/HL dependence-structure conditions and reports regret, PSE, the ideal-learner line, and excess in percentage points.
- Sonnet: Sonnet’s corresponding condition-level results appear in the released gate summaries and canonical dose-response artifact.The note does not reproduce Sonnet’s cell-level table here.
Supplementary Note 11: Suppression across difficulty strata · Supplementary Note 12: Order-inconsistency rates by model and arm
Social-label suppression is approximately uniform across true-risk-gap strata rather than concentrated among near ties. Order inconsistency is higher in the cfA=0 arm than cfA=1 across models, without affecting paired proxy-effect estimates.
- Supplementary Note 11: Suppression across difficulty strata: Social-label suppression is not concentrated among near ties; evaluation pairs were split into three strata by true risk gap |∆r|.The strata are defined by the true risk gap, |∆r|.
- Supplementary Note 11: Suppression across difficulty strata: All three difficulty-stratum contrasts exclude zero, indicating an approximately uniform reduction in proxy weight.This pattern supports suppression across difficulty levels rather than a near-tie-specific effect.
- Supplementary Note 12: Order-inconsistency rates by model and arm: Within-pair inconsistency is the fraction of evaluation pairs receiving different answers across their two presentation orders.Rates are reported for dose-response cells and all sixteen replication cells.
- Supplementary Note 12: Order-inconsistency rates by model and arm: Order-inconsistency rates are reported across two scopes: dose-response cells and all sixteen replication cells.Dose-response cells are used in Methods to bound decoding-noise attenuation of the slope.
- Supplementary Note 12: Order-inconsistency rates by model and arm: In every model, the cfA=0 arm is more order-inconsistent than the cfA=1 arm.This is a directional asymmetry in within-pair inconsistency rates.
- Supplementary Note 12: Order-inconsistency rates by model and arm: The cfA-arm inconsistency asymmetry does not affect paired estimands because both arms enter every proxy-effect estimate symmetrically.Sonnet’s replication-format cells cover only the dose-response scope, so its all-cell columns are not defined.