Source-linked AI summary

Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy

Yi-Cheng Lai, Hen-Hsen Huang

arXiv:2609.00067v1cs.CLcs.AI

TL;DR

Multimodal contextual sycophancy describes external text overriding conflicting image evidence, but the paper asks whether this depends on when text enters the model’s process. Using a 998-case diagnostic and staged witness–arbiter interventions, it finds large gains from withholding text during visual commitment, alongside model- and generator-dependent exceptions.

  • Problem

    The paper addresses limited evidence about whether external text changes multimodal visual answers because of source preference or because text is introduced before visual commitment.

  • Method

    The authors introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, then move the information boundary around a context-blind visual witness.

  • Results

    Across six models, S2VA improves over Witness-Only by 19.7–44.1 points, while the best information boundary varies by model and context generator.

  • Takeaways & Limitations

    Contextual sycophancy is sensitive to when external text is introduced, with text contaminating some models’ visual reasoning but scaffolding others.

  • Takeaways & Limitations

    The study is a controlled image–text stress test with six models, AI-generated abnormal scenes, single controlled text sentences, and designated visual truth as ground truth by construction.

Abstract

from arXiv · show

External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero. The best information boundary is not uniform: textual context scaffolds some models, and a GPT-4o-regenerated subset changes the relative ordering of joint conditioning, Witness-Only, and S2VA. Contextual sycophancy is therefore sensitive to when text is introduced, as well as to the model and context source.

1 Introduction

The paper defines multimodal contextual sycophancy as external text overriding conflicting visual evidence and asks whether the timing of text exposure changes visual commitment. A 998-case diagnostic and information-boundary interventions reveal substantial, model- and context-dependent effects.

  • Problem: Multimodal contextual sycophancy occurs when external text overrides conflicting visual evidence in an image–text input.The paper distinguishes this failure from ordinary visual hallucination because a plausible competing evidence stream induces it.
  • Research question: The central question is whether text changes the visual answer depending on whether it is admitted before or after the model forms a visual account.The study moves the information boundary around visual commitment rather than only measuring which source ultimately wins.
  • Information boundary: 7.9% under Joint versus 84.2% under S2VA on the main GPT-5.1 condition shows that staging and withholding text can substantially change accuracy.S2VA first forms a visual account and then lets an arbiter reconcile it with external text.
  • Model dependence: Models show contaminant and scaffold response patterns: text can destabilize unusual-image reasoning for some models but structure visual reading for others.The paper treats these as benchmark-specific roles rather than fixed model classes.
  • Diagnostic: 998 cases independently vary visual evidence, commonsense priors, and external text to diagnose context-conditioned image–text behavior.The diagnostic also reports text-following and sycophancy-rate metrics.

2 Related Work

Prior work mainly measures which source models follow during multimodal conflict or mitigates unsupported visual generation. This paper instead varies evidence sources and moves text exposure to test whether timing changes the visual read itself.

  • Prior benchmarks: Existing conflict and sycophancy benchmarks generally ask which source wins when visual, textual, and prior knowledge conflict.They measure source preference under conflict rather than when that preference is formed.
  • Sycophancy: Sycophancy research usually concerns agreement with a user’s stated belief, whereas this study examines pressure from external evidence streams such as captions or retrieved text.The paper also distinguishes its timing question from leading or deceptive query wording.
  • Positioning: The diagnostic independently varies visual truth, commonsense priors, and external text before moving the timing of text exposure.This design extends source-preference benchmarks with controlled evidence and information-boundary manipulations.
  • Mitigation landscape: Existing hallucination mitigations target unsupported generation, object hallucination, or prior-induced errors rather than changes in visual reading after external text is admitted.Examples include correction, contrastive decoding, and causal disentanglement approaches.

3 Context-Conditioned Diagnostic Setup

The diagnostic fixes each image–question pair while varying external text against visual truth and commonsense priors. It uses abnormal scenes and a context-blind witness to make text’s effect on visual integration observable.

  • Evidence variables: The setup represents an image V, visual question Q, external text C, visual truth YV, and commonsense-prior answer YK.The designated visual answer and typical-world answer can agree or conflict.
  • False-text traps: False-text traps pair abnormal images with text supporting the typical-world answer YK rather than the designated visual truth YV.Text-following measures adoption of false text, while sycophancy requires both false-text adoption and visual incorrectness.
  • Dataset: The abnormal split uses 499 WHOOPS! images whose visible scenes conflict with commonsense expectations, making false text strongly adversarial.Each image receives a commonsense-answerable query, visual truth, and three text variants.
  • Controlled comparison: Holding V and Q fixed while varying only C makes answer changes reflect integration of text with visual evidence and priors.Under joint conditioning, text shares the context window with the image and question before conflict resolution.
  • Evidence configurations: The benchmark’s controlled evidence configurations mark equality, non-equality, and absent external text across conditions.Table 1 summarizes these configurations using +, −, and –.

4 Information-Boundary Probe

The information-boundary probe separates visual commitment from later text reconciliation by comparing context exposure during the witness stage with exposure only during arbitration. These conditions isolate the roles of staging, context withholding, and arbitration.

  • Two-stage separation: The witness produces a visual account from the image and question, while the arbiter later reconciles that account with external text.This two-step structure prevents external context from entering the initial visual readout.
  • Witness-Only: Witness-Only withholds external text and uses the context-blind witness output directly as the final answer.This condition tests the isolated visual account without a later arbitration call.
  • Information boundaries: Leaky Witness exposes the witness to external text, whereas S2VA withholds text from the witness and introduces it only through a second arbitration call.The comparison separates context exposure during visual commitment from later context reconciliation.
  • Arbitration rule: S2VA prioritizes the witness when confidence exceeds 0.7 and it contradicts external text, but permits text or general knowledge below 0.4 or when evidence is not visible.Intermediate cases use the supplied evidence under the same visual-preference hierarchy.

5 Experimental Setup

The diagnostic fixes each image–question pair while varying external text and its relation to visual evidence and commonsense priors. It evaluates balanced abnormal and normal cases across multiple controlled inference conditions and model families.

  • Diagnostic dataset: 998 balanced cases vary whether controlled text supports visual evidence, commonsense priors, or neither while holding each image and question constant.The benchmark includes 499 abnormal and 499 normal images.
  • Diagnostic dataset: 499 abnormal WHOOPS! images create false-text traps by pairing visible scenes that conflict with commonsense expectations with text supporting the prior answer.Gemini 3 Flash generates the query, visual truth, and three text variants for each abnormal image.
  • Diagnostic dataset: 499 normal ImageNet images provide congruent controls that test ordinary multimodal performance and distinguish trap-specific failures from general prompt degradation.Their visual truth agrees with prior expectations under the same query and text-condition structure.
  • Text conditions: Each case uses true, false, and irrelevant text, with additional variants changing specificity or removing and shuffling context.Medium and weak text provide dose-response variants generated at decreasing specificity.
  • Robustness control: A 200-case GPT-4o-regenerated held-out subset tests sensitivity to the external-text generator while preserving the original images and visual-truth labels.The subset contains 100 abnormal and 100 normal cases.

6 Main Results

The main results show that arbitration and context isolation can substantially recover visual fidelity under false-text conflict, but their benefits depend on the model and text generator. GPT-5.1 benefits strongly from S2VA, while other models can use context as scaffolding and show different condition orderings across generators.

  • Isolation and arbitration: 7.9% GPT-5.1 abnormal-image accuracy under false-text Joint rises to 49.7% with Witness-Only and 84.2% with full S2VA.S2VA is 34.5 points above Witness-Only, while withholding context adds 20.5 points relative to prompt-matched Leaky Witness.
  • Reasoning variants: CoT improves GPT-5.1 over Joint but remains below S2VA, while degrading Gemini 2.5 and Qwen3-Instruct.Free-form reasoning can surface visual evidence or create room to rationalize contextual text.
  • Model-dependent context: 68.3% is Kimi-K2.5’s best accuracy under Visual Supremacy Only, whereas Qwen3-Instruct is strongest under Joint conditioning.These results show that removing context can hurt when text provides scaffolding.
  • Generator sensitivity: 68.0% is GPT-5.1 Joint accuracy on the GPT-4o-regenerated abnormal subset, versus 61.0% Witness-Only and 85.0% S2VA.The 24-point arbitration gain has a 95% confidence interval of [16, 33], but context-preserving Joint remains best for three other evaluated models.
  • Isolation and arbitration: 19.7–44.1 points is the S2VA improvement over Witness-Only across all six models, with every paired 95% confidence interval excluding zero.The comparison isolates the incremental contribution of the arbiter at a fixed context-blind witness.
  • False-text strength: Weak and medium false text improve GPT-5.1 on abnormal images, whereas original false text reduces accuracy to 7.9%; normal-image accuracy stays approximately 64.0–64.1% under no and weak false text.The variants jointly change specificity, hedging, and assertive framing, so their individual effects cannot be disentangled.

7 Analysis and Discussion

The analysis separates false-text adoption from visual incorrectness and shows that contextual sycophancy varies by model, context type, and information boundary. Text can contaminate visual reasoning for some models but scaffold it for others.

  • Measuring contextual sycophancy: Follow measures adoption of the false textual claim, whereas Syco. counts false-text adoption combined with visual incorrectness.The two rates use separately judged labels; under Joint, GPT-5.1 follows false text in 92.0% of cases, with 89.0% classified as unambiguous sycophancy, versus 8.4% under full S2VA.
  • True-text interference: True text reduces abnormal-image accuracy by 8.4–13.6 percentage points for GPT-5.1, Gemini 2.5, and Qwen3-Thinking despite correctly describing the unusual scene.Claude Sonnet 4.5, Qwen3-Instruct, and Kimi-K2.5 show the opposite pattern, so true text is not inherently helpful.
  • Model-specific context roles: GPT-5.1, Gemini 2.5, and Qwen3-Thinking form a contaminant pattern, while Claude Sonnet 4.5, Qwen3-Instruct, and Kimi-K2.5 form a scaffold pattern.The contaminant group shows true-text interference and large S2VA gains; the scaffold group tends to benefit from context and shows smaller or model-dependent S2VA gains.
  • Qualitative error modes: Abnormal-image failures include prior override, text copying, compromise, abstention, and granularity error.When context contaminates, prior override and compromise dominate; staged S2VA can interrupt this path, while granularity errors remain a witness-stage limitation.
  • Operational implications: Context effects depend on model and source: S2VA outperforms Witness-Only for all six models and Joint for five of six in the Gemini-generated false-text setting, but Joint remains better for three of four models on the GPT-4o subset.The authors therefore recommend model- and source-matched calibration rather than a universal information boundary.

8 Conclusion

Across six models, the diagnostic shows that external text can either contaminate or scaffold visual reasoning, depending on when it is introduced and on the model and context source. S2VA improves over direct witness reports, but the best boundary is not uniform.

  • Conclusion: S2VA improves Witness-Only by 19.7–44.1 points across all six models, with every paired confidence interval excluding zero.The results characterize context use through true-text effects, false-text susceptibility, and isolation benefit.
  • Conclusion: The relative ordering of Joint, Witness-Only, and S2VA changes on the GPT-4o-regenerated subset.This cross-generator result shows that context-source choice affects the observed intervention comparison.

Limitations

The study is a controlled diagnostic with important scope, evaluation, model, and data-generation boundaries. Its findings should not be treated as prevalence estimates or universal rules for resolving multimodal disagreement.

  • Scope: The benchmark is a controlled context-conditioned stress test rather than a prevalence estimate and covers image–text inputs only.It uses AI-generated counter-intuitive scenes, a single controlled external sentence, and an unmatched ImageNet normal split.
  • Baselines: The intervention comparison omits full Self-RAG, Woodpecker, VCD, and search-based systems.The study includes prompting, staged description–answer procedures, Visual Supremacy Only, and a leaky witness variant instead.
  • Evaluation: The rederived Witness-Only cells use the same GPT-4o-mini judge but were not separately re-audited by humans.The broader evaluation includes human validation, but the authors identify multi-annotator validation as a remaining strengthening step.
  • Reproducibility: Several systems are proprietary or preview APIs, so outputs may drift across providers and dates.The authors recommend replication with locally hosted frozen checkpoints and archived inference snapshots.
  • Data generation: The data-generation pipeline may introduce artifacts because generator style and false-text variants jointly change specificity, hedging, and assertive framing.The GPT-4o subset checks generator style but does not replace naturalistic retrieval, and factorial controls remain needed.
  • Interpretation: Contaminant and scaffold labels are benchmark-specific descriptors, not immutable model classes, and the sample contains only six models.The benchmark designates an image-based answer as ground truth by construction, which is appropriate for this diagnostic but not a general rule for source disagreement.

Ethical Considerations

The paper cautions against interpreting constructed image–text conflicts as evidence for universal visual supremacy. Ethical deployment should account for ambiguous or deceptive imagery and recognize the study’s methodological boundaries.

  • Deployment caution: For ambiguous, low-quality, or deceptive imagery, systems may need to signal disagreement, request clarification, abstain, or combine evidence probabilistically.The benchmark deliberately constructs conflicts and should not be treated as a universal prescription to privilege visual answers.
  • Contribution boundary: The paper introduces a context-conditioned diagnostic rather than a general-purpose verification architecture.Its mitigation comparisons separate information boundaries and prompting structures, while several external mitigation families are not implemented directly.
  • Information boundary: S2VA withholds external text from the witness, then exposes the witness report and text to a later arbiter.The witness sees the image and question; arbiter guidance prioritizes the witness when confidence is high and permits contextual fallback when confidence is low or blindness is reported.
  • Sensitivity analysis: Accuracy remains near 79.9% for witness-confidence thresholds τ ∈[0, 0.85] in a retrospective Claude Sonnet 4.5 sweep.This offline counterfactual does not represent an operative fallback frequency because reported S2VA runs always invoke the arbiter.
  • Within-family contrast: Qwen3-Thinking remains sensitive to true text, falling 13.6 percentage points, while Qwen3-Instruct falls into the scaffold pattern.The within-family contrast is suggestive rather than causal because training details are unavailable.
  • Arbitration: Across all six models, S2VA improves judged accuracy over the same context-blind witness account.This separates visual commitment from question-specific answer selection, but does not establish whether gains arise from reconciliation, answer compression, or both.
  • Operational cost: S2VA has about 1.5× the joint baseline’s end-to-end latency because it uses two sequential calls.Exact dollar cost depends on provider pricing and caching.

F Adaptive Routing Analysis

The routing analysis finds that learned routing improves over a baseline but remains below fixed S2VA, so routing is exploratory. Additional controls show that recovery depends on information isolation rather than call budget alone, while robustness and failure analyses reveal model- and generator-dependent effects.

  • Adaptive Routing: 71.3% router accuracy exceeds the 55.8% baseline average but remains below fixed S2VA at 75.3%.A case-level oracle reaches 82.2%, indicating headroom beyond the learned threshold.
  • Prompt Controls: 84.2% S2VA accuracy versus 49.7% Witness-Only shows that isolated visual testimony alone does not explain the full recovery.Two-Call Describe–Answer reaches only 11.8% on abnormal images despite matching S2VA’s call count.
  • Prompt Controls: 73.5% is the strongest single-call baseline, indicating that evidence-separation prompting recovers much of the loss.Its improvement is distinct from the weaker two-call Describe–Answer result on abnormal images.
  • Failure and Generalization Analyses: S2VA failures often reflect granularity mismatches, while cross-condition gains and susceptibility estimates vary across models and remain noisy with six models.The susceptibility proxy correlates 0.45 with S2VA gain and has R2 = 0.16; qualitative anomaly families are not treated as disjoint leaderboard subsets.

M Dataset Construction Details and Prompt Specifications

The benchmark constructs abnormal-image cases by pairing visual truths with commonsense queries and controlled text variants, then evaluates prompts that vary how visual evidence and context are introduced. Its appendix documents text strength, anomaly-family scope, implementation settings, and witness–arbiter scoring procedures.

  • M.1 Annotation Protocol: Each abnormal case pairs a commonsense query with a visual truth, false text describing the plausible but unseen scene, true text, and irrelevant text.False text and parametric priors are aligned against visual evidence, while the benchmark treats visual-truth fields as annotations.
  • M.1 Annotation Protocol: Independent human validation of visual-truth annotations remains an important extension because the audit checks judge decisions against labels, not whether the labels themselves are correct.The anomaly families are non-exclusive and were not exhaustively multi-label annotated, so they are not disjoint subsets.
  • M.2 Dose-Response Text Variants: Weak text hedges the false claim, medium text removes identifying details while retaining it, and original false text is the strongest endpoint.These variants document surface commitment and length rather than constituting a semantic specificity metric.
  • M.5 Representative Dataset Examples: The appendix supplies representative benchmark examples and fixes inference settings at temperature 0.0, 1024-pixel maximum edge, and model-specific output budgets.Experiments ran through provider APIs in April–May 2026, with identifiers listed separately.
  • N.2 Direct Inference Prompts: Direct prompts answer from the image with optional context, while evidence-separation prompts explicitly list visual evidence, text claims, conflict, and a visual-priority answer.Several controls instruct the model to verify the image and reject misleading context.
  • N.2 Direct Inference Prompts: The prompt suite also includes single-call and two-call Describe–Answer, CoVe-style verification, and explicit visual-priority controls.Two-Call Describe–Answer gives the second call the visual description, context, and question.
  • N.3 Witness and Arbiter Prompts: The blind witness reports only question-relevant image content and confidence, whereas the context-aware witness may use external text; Witness-Only scores the blind report directly.S2VA introduces context later through an arbiter that reconciles witness testimony with potentially misleading text.
  • N.4 Judge Prompts: Separate GPT-4o-mini calls judge correctness and text following, with correctness referenced to designated visual truth and allowing compatible broader or narrower answers.The condition-aware judge receives active context and text condition, while the blind variant omits them.
Loading 2609.00067v1…