Source-linked AI summary
AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs
Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin
TL;DR
Existing evaluations provide limited evidence about how LLM anchoring varies across pathways and anchor relevance. AnchorBench tests five delivery pathways with controlled, relevance-aware prompts, finding that anchoring depends strongly on pathway and remains present even in highly accurate frontier models.
Problem
Existing LLM anchoring evaluations cover limited anchor pathways and rarely distinguish irrelevant from plausible anchors.
Method
AnchorBench evaluates five realistic anchor-delivery pathways using controlled numeric judgment tasks and separates control, irrelevant, and plausible anchors.
Results
Anchoring is strongly pathway-dependent, with plausible anchors generally producing larger shifts and influence weakening as anchor distance increases.
Takeaways & Limitations
High task accuracy does not guarantee robust judgment, so reliable LLM evaluation must test susceptibility to anchoring across delivery pathways.
Takeaways & Limitations
Synthetic deterministic aggregation tasks simplify real judgment, and the suites do not test fully agentic retrieval and tool-use loops.
Abstract
from arXiv · showhide
The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibit similar behavior. However, existing work on anchoring in LLMs typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant from plausible anchors. We introduce AnchorBench, a benchmark for the anchoring effect in LLMs that evaluates multiple anchor pathways under an explicit anchor relevance axis. Across fourteen models, including ten open-weight models and four frontier API models, and a large set of controlled prompts, we find that (1) anchoring is strongly pathway-dependent, (2) plausible anchors usually induce larger shifts than irrelevant ones when introduced through stronger pathways, (3) anchor influence generally weakens as the anchor moves farther from the evidence-supported answer, most clearly on External and RAG, and (4) high task accuracy on the anchor-free control condition (Acc$_{10}$: answers within 10 points of gold) does not guarantee robustness: even frontier API models above 95% control accuracy remain susceptible to plausible anchors.
1 Introduction
AnchorBench is a controlled, multi-pathway benchmark that tests how anchor delivery and relevance affect quantitative judgments in LLMs. Its experiments show pathway-dependent anchoring, stronger effects from plausible anchors in strong pathways, distance-sensitive influence, and a separation between accuracy and robustness.
- Benchmark design: AnchorBench evaluates anchoring through five pathways: External, History, In-Context Learning, Retrieval-Augmented Generation, and Tool.These pathways represent standard ways context reaches a deployed model.
- Benchmark design: The benchmark distinguishes control, irrelevant, and plausible anchors along an explicit anchor relevance axis.Structured numeric evidence and deterministic gold answers allow measurement of output shifts and accuracy.
- Main findings: Across fourteen models, anchoring is strongly pathway-dependent, and plausible anchors usually produce larger shifts than irrelevant anchors in stronger pathways.The model set includes ten open-weight models and four frontier API models.
- Main findings: Anchor influence generally weakens as the anchor moves farther from the evidence-supported answer, most clearly on External and RAG.This distance effect is explicitly reported for the External and Retrieval-Augmented Generation pathways.
- Main findings: High control accuracy does not guarantee robustness: frontier API models above 95% Acc10 remain susceptible to plausible anchors.Acc10 denotes answers within 10 points of gold; susceptibility remains measurable, though at smaller magnitudes.
2 Related work
Prior work establishes anchoring as a human judgment bias and documents analogous effects in LLMs, while showing that existing evaluations often isolate single pathways. AnchorBench extends this literature by varying anchor framing and delivery pathways within a unified setup.
- Human anchoring theory: Human estimates are often drawn toward an initial value, even when the anchor is arbitrary or only weakly informative.Prior accounts propose multiple mechanisms rather than a single explanation, including selective accessibility for externally provided anchors.
- Anchoring effect in LLMs: LLMs shift estimates toward previously mentioned values, while chain-of-thought, reflection, and ignore-anchor instructions provide limited and inconsistent relief.Anchoring also persists in multi-turn price negotiation, although reasoning models are less susceptible.
- Prior benchmarks: SYNANCHORS varies in-prompt anchor magnitude and uses causal tracing, whereas AnchorBench fixes the numeric value and varies framing and delivery pathway.Most prior work studies one pathway at a time, leaving comparisons among External, History, ICL, RAG, and Tool anchoring open.
- Cognitive bias in LLMs: Anchoring is part of a broader literature documenting human cognitive biases in LLMs as systematic model failures and unreliable substitutes for human respondents.These effects have been studied in instruction tuning, decision-making, and survey response.
- Context influence beyond anchoring: Numerical anchoring is narrower than sycophancy and persuasion because it isolates a single numeric value rather than opinions or arguments.Related work also examines shifts toward confidently framed claims and authoritative-sounding sources.
3 Benchmark setup and evaluation design
AnchorBench is a deterministic diagnostic built from numeric aggregation tasks to isolate how anchor pathways and anchor relevance affect LLM judgments. It evaluates five controlled conditions across distinct pathways using task accuracy, anchor susceptibility, and relevance discrimination metrics.
- Task and item construction: Each item asks for a single integer on a shared 0–100 scale, with gold equal to the rounded mean of visible evidence ratings.The benchmark spans six domains and varies difficulty through evidence noise, hidden ratings, and conflicting values.
- Conditions and relevance: Five matched conditions combine a no-anchor control with irrelevant or plausible anchors placed in low or high directions.Matched irrelevant and plausible conditions use the same numeric anchor; only the framing sentence differs, isolating relevance effects.
- Anchor pathways: The five suites hold the judgment task fixed while varying how anchors reach the model through External, History, ICL, RAG, and Tool pathways.History uses the model’s Stage 1 response as the anchor, ICL includes metadata and distribution-matching variants, and RAG embeds an anchor in one document.
- Benchmark scale: 1,800 prompts per model per suite result from 360 items presented in five conditions, with seed-controlled and deterministic data generation.Each suite contains 6 domains × 60 items.
- Evaluation metrics and comparability: The evaluation measures control-only task accuracy, anchor susceptibility, and relevance discrimination, while History and Tool magnitudes are interpreted qualitatively because of format confounds.Accuracy uses MAEc and Acc10; susceptibility uses UAI and TAR, with near-zero anchor–control gaps excluded from UAI.
4 Empirical findings
Anchor susceptibility varies sharply by pathway: plausible anchors usually shift judgments more than irrelevant anchors in stronger pathways, while standard ICL is near-zero. Influence generally attenuates with anchor distance, but high control accuracy does not ensure robustness.
- Pathway dependence: 55/69 model–suite cells (80%) show greater susceptibility to plausible than irrelevant anchors, rising to 48/55 (87%) when ICL is excluded.The pattern is significant on EXTERNAL, RAG, TOOL, and HISTORY, but not ICL (pBH = 0.81).
- Pathway dependence: Standard ICL remains near-zero because anchors appear only in demonstration metadata, whereas ICL-dist raises mean UAIpls from ≈0.03 to 0.16.ICL-dist yields Disc∆ > 0 for 8/10 open-weight models, comparable to RAG.
- Anchor relevance: On EXTERNAL, panel-mean UAI rises from +0.09 for placebo and +0.05 for irrelevant framing to +0.27 for plausible and +0.49 for authority.Credibility modulates anchoring, but the placebo floor exceeds zero and the irrelevant condition.
- Distance sensitivity: Open-weight External UAIpls decreases from 0.32 to 0.26 to 0.18, while RAG decreases from 0.23 to 0.15 to 0.06 as δ increases from 15 to 40.The monotonic attenuation is clearest on External and RAG; ICL stays near zero.
- Accuracy and robustness: All four API models achieve ≥96% EXTERNAL accuracy yet show positive Disc∆ of 0.05–0.16, while accuracy and discrimination correlate only weakly (r = −0.24).Higher accuracy therefore does not prevent anchoring; the association explains little variance.
- Prediction quality: Plausible anchors increase MAE by +3.6 on External and +3.5 on Tool, with error-increasing shifts exceeding error-reducing shifts overall (29% vs. 23%).The gap is driven by plausible anchors (37% vs. 20%) and is largest on EXTERNAL and TOOL (≥22 pp).
5 Conclusion
ANCHORBENCH is a controlled diagnostic that evaluates anchoring in LLMs across five delivery pathways while distinguishing irrelevant from plausible anchors. Across fourteen models, anchoring depends strongly on pathway, with the clearest effects in External and RAG, where plausible anchors produce larger shifts than irrelevant ones and influence generally declines with distance from the supported answer.
- Benchmark design: ANCHORBENCH evaluates anchoring across five delivery pathways and separates irrelevant from plausible anchors to assess whether shifts are justified.The benchmark tests both whether models shift and whether the shift is warranted by anchor relevance.
- Main findings: Across fourteen models, anchoring is strongly pathway-dependent rather than uniform.The conclusion identifies delivery pathway as a major determinant of anchor influence.
- Main findings: In External and RAG, plausible anchors induce larger shifts than irrelevant anchors, while influence generally decreases as anchors move farther from the evidence-supported answer.These pathways show the clearest relevance-sensitive anchoring effects.
Limitations
AnchorBench’s controlled synthetic setting enables clean comparisons but does not capture the complexity of real judgment. Several pathway-specific design differences and prompt-based mitigations limit interpretation and do not eliminate anchoring.
- Synthetic deterministic aggregation tasks make comparisons clean but do not cover the complexity of real judgment.
- History, Tool, and UAI-based exclusions complicate some cross-suite comparisons because of differing interaction structures, cross-family formats, and sensitivity to small effects.
- Prompt-based mitigations weaken anchoring but do not remove it.
LLM usage disclosure
LLMs supported code, LaTeX, draft review, and synthetic-item generation under author-designed processes, while authors retained responsibility for the study’s design, results, analyses, and writing.
- LLM usage disclosure: LLMs assisted with code, LaTeX, draft review, and generating some synthetic items, which the authors verified.The authors’ pipeline governed synthetic-item generation and all such data were author-verified.
- LLM usage disclosure: The authors—not LLMs—produced the design, metrics, analyses, and writing, and no LLM generated or altered results.The fourteen models in Table 3 were experimental subjects rather than authoring tools.
Reproducibility statement
The paper releases its code, benchmark data, and source, with the benchmark mirrored on GitHub and Hugging Face. The release covers five core suites and a partial-evidence variant, totaling 14,400 prompts with metadata sufficient to compute UAI.
- Reproducibility statement: Code, benchmark data, and the paper’s source are publicly released through GitHub and the arXiv version.The benchmark repository is available at https://github.com/Ydrg9989/AnchorBench.
- Reproducibility statement: The Hugging Face dataset covers the five core suites and the partial-evidence variant used behind Table 2.The benchmark is available at https://huggingface.co/datasets/Yiderigun/AnchorBench.
- Reproducibility statement: 14,400 prompts include gold answers and, when fixed in advance, anchor values, making UAI computable from the release alone.The passage notes that HISTORY is an exception, but its description is truncated in the supplied text.
Ethics statement
The paper frames anchoring as an LLM reliability problem with potential consequences for decision-support, while using synthetic scenarios without personal, sensitive, or proprietary data and no human subjects.
- Ethics statement: The authors argue that documenting anchoring vulnerabilities outweighs the risk of revealing settings in which models may be susceptible.The benchmark uses synthetic scenarios, contains no personal, sensitive, or proprietary data, and does not involve human subjects.
A Appendix · A.1 Experimental details
AnchorBench’s appendix specifies a fourteen-model evaluation built from standardized, multi-condition suites and deterministic inference procedures. It also details pathway-specific formatting and the History suite’s self-generated-anchor interaction.
- A.1 Experimental details: Each suite contains 360 items across five conditions, yielding 1,800 prompts per suite per model.This design was applied across the fourteen-model panel.
- A.1 Experimental details: Decoding was standardized across suites and models with temperature = 0 and max tokens = 512.These settings governed inference for both local and API-accessed models.
- A.1 Experimental details: History used a two-stage conversation in which Stage 2 received the model’s own Stage 1 response as context.Predictions were extracted with one deterministic parser shared across all conditions.
- A.1 Experimental details: Tool prompts used structured tool messages when supported and equivalent plaintext tool context otherwise.The formatting was adapted to each model’s capabilities while retaining a shared deterministic parser.
- A.1 Experimental details: Open-weight models ran locally on H100 GPUs, whereas API models were queried through OpenRouter.Table 3 summarizes all fourteen evaluated models.
- A.1 Experimental details: The History revision condition evaluates revision after full evidence is presented.This appendix label identifies the full-evidence revision stage.
- A.1 Experimental details: In the History suite, partial-evidence Stage-1 answers became self-generated anchors before full evidence was revealed in Stage 2.The control presented the full evidence in a single turn.
A.2 Prompt examples … A.17.13 LLM anchoring on the human effect-size scale
Across suites and robustness probes, LLM anchoring is pathway-, relevance-, format-, and evidence-dependent: plausible anchors generally shift judgments more than irrelevant ones, while standard metadata ICL is near zero. The effect persists across models, domains, decoding and scoring variants, but magnitudes are qualified by protocol confounds and are not directly comparable to human effect sizes.
- A.4 Cross-suite patterns: Anchoring concentrates in External, History, and Tool, is moderate in RAG, and is near zero for standard metadata ICL.History and Tool comparisons are qualified by format or protocol confounds, while API models show smaller but non-zero discrimination.
- A.5 Statistical robustness checks: 0.24 mean Disc∆ separates History from ICL across models, whereas Tool discrimination is not significantly different from the other suites.The History–ICL difference has bootstrap 95% CI [0.11, 0.41] and Wilcoxon pBH ≈0.01; Tool pBH ≈0.56.
- A.6 UAI exclusion-threshold sensitivity: ≈80% of computable model–suite cells preserve Disc∆’s sign across ε ∈{1, 3, 5}, with no large External or History cell flipping sign.At ε=3, exclusion rates range from 3.8% for ICL to 11% for History.
- A.7 UAI distribution and extreme values: 0.40 to 0.23 is the capped reduction in History Disc∆, which preserves the direction and suite ranking despite overshoot.The Disc∆ sign remains stable under mean, median, and capped [0, 1] aggregation.
- A.9 Anchored-condition prediction quality: 0.6–3.6 MAE increases show that plausible anchors degrade accuracy on External, RAG, and Tool, while History estimates are distorted by its single-stage control.Matched controls increase mean History MAE from 8.1 to 16.0 and reduce Qwen-1.5B Disc∆ from 0.89 to 0.27.
- A.11 ICL distribution-matching comparison: 0.04 →0.16 is the mean UAIpls increase under numerically relevant ICL demonstrations, with 8/10 models showing positive discrimination.Standard metadata ICL is near zero, indicating that stronger numeric priming—not few-shot exposure itself—drives the effect.
- A.13 Gold-referenced error decomposition: 48.0% versus 19.9% is External’s harmful-shift contrast for plausible versus irrelevant anchors; Tool shows 39.9% versus 17.6%.Across all items, plausible anchors increase error at 37.3% versus 24.1% reducing it, compared with 20.2% harmful shifts for irrelevant anchors.
- A.17.13 LLM anchoring on the human effect-size scale: |dpls| reaches 0.2–0.5 on External, while |dhi-lo| reaches 0.3–0.9; the authors stop short of numerical comparison with human effect sizes.The matched-prompt, within-item UAI and d contrasts are not on a common scale with human between-subject anchoring measures, although the qualitative regularities replicate.