Source-linked AI summary
When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA
Manikandan Ravikiran, Siddharth Vohra
TL;DR
The paper asks whether unverified claims about another source’s answer destabilize multiple-choice responses. It audits paired switches from neutral gold answers to fixed wrong options across models, benchmarks, and languages, finding higher adoption for expert than majority cues while bounding the conclusion to the tested forced-choice setting.
Problem
Clean benchmark accuracy does not test whether models preserve task-supported answers when an outside source claim conflicts with the question.
Method
The audit pairs each neutral response with misleading cues naming one fixed wrong option, and measures valid switches using NC-MCAR across four models and two benchmark families.
Results
41.1% aggregate NC-MCAR for the expert template compared with 12.5% for the majority template across 220,000 outputs.
Takeaways & Limitations
Under the tested forced-choice prompts, bare unverified source claims can destabilize answers that were previously consistent with the task evidence.
Takeaways & Limitations
The findings may not extend beyond four models, two benchmark families, five language settings, one response per prompt, and answer-only forced-choice evaluation.
Abstract
from arXiv · showhide
Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \emph{neutral-conditioned misleading cue adoption rate} (NC-MCAR), which measures switches to that option only on valid cued trials where the same model first selected the gold answer under a neutral prompt. This is a measure of answer instability, not proof that the model knew the answer or that all deference is irrational. We evaluate four instruction-following models on MMLU-Pro and IndicMMLU-Pro in English, Hindi, Bengali, Tamil, and Telugu. Across 220{,}000 outputs, the expert template yields 41.1\% aggregate NC-MCAR, compared with 12.5\% for the majority template. These two conditions use the same wrong option and final instruction. Filler accuracy remains well above expert-wrong accuracy, while correct-cue prompts have high valid-response accuracy. The audit documents answer instability relevant to grounding under the tested forced-choice prompts: a bare, unverified source claim can outweigh an answer that was previously consistent with the task evidence.
1 Introduction
The paper audits whether unverified source claims destabilize multiple-choice answers, focusing on switches from a neutral gold answer to a named wrong option. It introduces a paired measure of this answer instability and finds stronger adoption for expert than majority cues.
- The audit asks whether models preserve answers supported by the item when an attributed source conflicts with the question.
- The controlled design first records a neutral answer, then adds an evaluator-constructed claim that an external source selected a wrong option.The model is not told that the claim was fabricated.
- NC-MCAR counts valid switches from the gold answer under a neutral prompt to the fixed wrong option named by the cue.Neutral correctness identifies the starting point but does not prove knowledge or confidence.
- The audit evaluates four instruction-following models on MMLU-Pro and IndicMMLU-Pro across English, Hindi, Bengali, Tamil, and Telugu.
- 41.1% aggregate NC-MCAR for the expert template exceeds 12.5% for the majority template across 220,000 outputs.The wrong option and final instruction are held constant for this comparison.
2 Related Work
Prior work shows that model judgments can shift with social, group, authority, and externally supplied cues. This paper distinguishes its audit through paired, gold-labeled switches to one fixed wrong option.
- Earlier studies report agreement with user beliefs or social framing and movement toward group responses.
- Other work finds that model judgments change with authority, position, bandwagon cues, or externally supplied information.
- The audit measures a paired change from a gold answer to one fixed wrong option rather than agreement with a user belief or majority.
- The evaluation fixes the question, options, order, gold label, output instruction, and cued wrong option across misleading source conditions.This supports within-protocol comparisons but does not establish transfer to free-form reasoning or unconstrained conversation.
3 Audit Design
The audit uses paired forced-choice perturbations across five language settings, holding one wrong option fixed while varying source cues and comparing each condition with a neutral prompt.
- Every prompt contains a question, fixed options, and an instruction to return one option letter, using MMLU-Pro for English and IndicMMLU-Pro for four Indic languages.
- Each base item is evaluated under a neutral prompt and ten controlled variants, with the same model and item paired before and after a cue.The study is a perturbation audit rather than a leaderboard estimate.
- The variants include filler, four wrong source templates, three stated-reliability claims, and two correct-cue sanity checks.
- One wrong option is reused across all seven misleading conditions, while correct-cue conditions point to the gold option.Fixed distractors prevent distractor plausibility from varying across cue-template comparisons.
- The four source-template comparisons keep the item and final instruction fixed, but wording differences prevent complete isolation of source identity.In particular, the student cue says “guessed,” whereas the other English cues say “chose.”
- The final evaluation contains 5 language groups, 1,000 base examples per group, 11 conditions per example, and 4 models, producing 220,000 outputs.
4 Metrics and Uncertainty
The paper defines NC-MCAR as a neutral-conditioned switch to the cued wrong option and reports it alongside validity, raw adoption, invalid responses, and paired uncertainty estimates.
- Valid-response accuracy, raw misleading cue adoption, invalid response rate, and NC-MCAR are reported as distinct quantities.Neutral accuracy is also reported over all trials.
- Raw misleading cue adoption can overstate cue-induced answer abandonment because it includes responses from models already wrong under the neutral prompt.
- NC-MCAR measures how often a model switches to the cued option after selecting the gold answer under the neutral prompt.It separates observed cue-induced transitions from ordinary benchmark errors.
- Aggregate values pool eligible model-item observations rather than averaging the four model-level percentages.Neutral-correct and valid-response counts differ by model.
- Invalid cued responses are excluded from the NC-MCAR denominator because they are neither valid answers nor valid cue adoptions.
- 95% confidence intervals use bootstrap resampling over base items while preserving paired conditions.They capture item-sampling uncertainty, not run-to-run variation in hosted endpoints.
5 Results
Across models, misleading source cues sometimes induced switches from a neutral-correct answer to the named wrong option, with susceptibility varying by template, stated reliability, and language setting. Controls and design comparisons show that these results are descriptive and that NC-MCAR specifically measures cue-linked answer changes.
- Model-level results: All four models show nonzero NC-MCAR, meaning each sometimes switches from the gold answer to the exact wrong option named by a source cue.NC-MCAR conditions on a neutral-correct response, unlike raw misleading cue adoption.
- Controls: Filler valid-response accuracy is 49.9%, versus 30.0% for expert-wrong cues; correct-cue accuracy reaches 73.8% for majority cues and 84.9% for expert cues.The filler does not establish a target-specific switching floor because that rate is unreported.
- Cue-template results: The expert template has the highest NC-MCAR for every model, while the other template rates do not form a stable ordering.Gemma reaches 66.0% for expert cues, whereas Claude and Qwen each reach 37.0%.
- Stated reliability: 58.6% NC-MCAR is reached by Gemma at 90% stated reliability, compared with 1.3% at 20%, while other models show different response patterns.The experiment manipulates textual reliability claims rather than observed source accuracy, so these results do not establish optimal trust or over-deference.
- Language variation: Per-language NC-MCAR estimates are heterogeneous, with no single pattern shared across all five language settings and models.Pooled non-English comparisons are secondary because the benchmarks differ in item source, subject mix, option structure, translation, and difficulty.
6 Discussion
The audit frames source-cue susceptibility as a grounding concern: models may change a previously gold answer when a bare source claim is added. It therefore motivates evaluations that test unsupported claims while permitting retention, revision, uncertainty, or evidence requests.
- Implications for grounding audits: Grounding evaluations should test how models combine task content with source claims lacking rationale, citation, or verifiable support.The proposed audit concerns answer changes under forced-choice prompts, not unrestricted reasoning.
- Implications for grounding audits: Audits should allow models to retain or revise answers, express uncertainty, and request evidence instead of forcing every response into one option letter.This follows the paper’s stated evaluation implications for unsupported source claims.
7 Conclusion
The audit measures whether source-attributed cues move models from a neutral-correct answer to a fixed wrong option. Across 220,000 outputs, the expert template produced substantially more such switching than the majority template.
- Every evaluated model made the transition from a neutral-correct response to the fixed cued wrong option on some trials.
- 41.1% aggregate NC-MCAR for the expert template versus 12.5% for the majority template.The comparison uses the same fixed wrong option and final instruction.
- Correct cues had high valid-response accuracy, while the audit does not determine when deference is rational.
- The results support testing whether answers remain stable when prompts add bare claims without supporting evidence.
Limitations
The findings are bounded by the forced-choice, answer-only evaluation and by template, denominator, sampling, and cross-language limitations. These constraints primarily affect generalization and interpretation of smaller or language-stratified estimates.
- The findings may not extend beyond answer-only forced-choice prompts to explicit reasoning, repeated sampling, free-form conversation, abstention, open-ended tasks, or other models and benchmarks.
- NC-MCAR lacks a target-specific switching floor and exact effective denominator counts, limiting interpretation of small estimates and weak-source ordering.The missing floor is less likely to alter the qualitative expert-versus-majority separation.
- Neutral accuracy and valid-response accuracy use different trial sets, and invalid cued outputs are excluded from NC-MCAR.This is especially relevant for Claude, whose invalid rate across the four source templates is 11.0%.
- Neutral correctness does not establish stable knowledge because low-margin choices or lucky guesses can enter the NC-MCAR denominator.The study lacks option probabilities, repeated neutral paraphrases, and independent runs for separating stable knowledge from fragile initial choices.
- MMLU-Pro and IndicMMLU-Pro are not parallel, so pooled language summaries remain descriptive rather than evidence that language causes susceptibility.Differences can reflect difficulty, subject mix, option structure, translation, and item origin.
Ethics Statement
The audit uses synthetic prompt perturbations on benchmark questions and does not involve human subjects, private user data, or personally identifying information. Ethical concerns center on multilingual interpretation, misuse for inducing errors, and inference cost.
- The work evaluates model robustness with synthetic prompt perturbations and uses generated responses to benchmark questions.
- The evaluation does not involve human subjects, private user data, or personally identifying information.
- English–Indic differences should not be interpreted as inherent properties of languages or language communities because the evaluations use different benchmark sources.Matched translation is identified as necessary for stronger causal claims about language.
- The perturbation templates are presented as diagnostic probes for robustness evaluation and mitigation, not for deployment-time manipulation.
- The audit reduces inference cost by sampling benchmark subsets, collecting one response per prompt at temperature 0, capping generation length, and evaluating four models.
J Benchmark Mismatch Analysis
English and Indic evaluations use different benchmarks, so their differences are descriptive rather than causal evidence about language. The analysis therefore reports settings separately, uses pooled comparisons secondarily, and applies a coarse cross-model difficulty proxy.
- English uses MMLU-Pro, whereas Indic languages use IndicMMLU-Pro, whose item sources, subjects, difficulty, option structures, and translation processes may differ.
- 48% and 62% are the cross-model neutral accuracies for Indic and English items, respectively, in this sample.
- The study presents per-language values descriptively, reports difficulty-stratified comparisons, and avoids attributing observed gaps to language alone.
- A matched translated benchmark would be needed to isolate language effects.
- The leave-one-model-out difficulty score uses the fraction of other evaluated models answering correctly under the neutral prompt.
- With four models, the score takes values 0, 1/3, 2/3, and 1, with easy items above 0.75 and hard items below 0.25.
- Inter-model Spearman correlations of neutral correctness range from 0.40–0.52, supporting cross-model neutral accuracy as a coarse difficulty proxy.
L Difficulty-Stratified Benchmark Summary
Difficulty-stratified NC-MCAR rates rise as agreement among the other models falls, while pooled and cross-language summaries remain descriptive because the benchmark items are not parallel. Reliability cues also produce model-specific sensitivity patterns, and qualitative cases illustrate transitions without explaining them.
- Difficulty-Stratified Results: NC-MCAR rates increase as agreement among the other three models falls across the reported coarse difficulty strata.
- Difficulty-Stratified Results: The English interval is wide in the hard stratum, and all stratum estimates are descriptive.
- Per-Language Results: Per-language values vary visibly across languages and models, so pooled results are treated as secondary.
- Control Conditions: The filler condition helps assess whether simple added text explains the response pattern, alongside wrong-cue and correct-cue conditions.
- Reliability Cues: Reliability conditions attach numerical claims to wrong options, but the study does not measure whether responses to those claims are normatively calibrated.
- Reliability Cues: Gemma and Claude show stronger sensitivity to stated reliability, especially at 90%, while GPT is comparatively stable and Qwen is similar across 20%, 50%, and 90%.
- Qualitative Examples: Representative cases show models retaining or changing answers after source-attributed cues, but they do not establish why the changes occurred.