Source-linked AI summary

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang

arXiv:2608.04244v1cs.CVcs.CL

TL;DR

Existing benchmarks rarely show how MLLMs arbitrate between conflicting visual and textual evidence, despite the importance of choosing which cue to trust. SIGNPOST-Bench addresses this gap with controlled five-condition counterfactual scene-text edits evaluated across geolocation tasks. Conflicting text increases median prediction error 4.8×, and 6.5–20.1% of adversarial predictions fall within 50 km of injected targets.

  • Problem

    Existing benchmarks rarely reveal how MLLMs choose between conflicting visual and textual cues in grounded real-world predictions.

  • Method

    SIGNPOST-Bench uses paired five-condition counterfactual scene-text interventions to measure text-induced localization changes and shifts toward injected geographic targets.

  • Results

    4.8×: conflicting text increases median localization error from 282 km to 1,347 km, while 6.5–20.1% of adversarial predictions fall within 50 km of injected targets.

  • Takeaways & Limitations

    SIGNPOST-Bench provides a reproducible framework for testing when MLLMs use, reject, or follow conflicting scene text.

  • Takeaways & Limitations

    The paper does not determine why conflict robustness varies, offering only possible mechanisms involving pretrained place-name representations and limited disagreement exposure.

Abstract

from arXiv · show

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.

1 Introduction

SIGNPOST-Bench addresses the largely hidden problem of how multimodal models arbitrate between visual context and scene text when the two conflict. It uses controlled, localized counterfactual text edits and continuous geolocation measurements to separate localization degradation from directed movement toward injected geographic targets.

  • Motivation: Visual geolocation provides a shared coordinate space for measuring both localization degradation and shifts toward an injected geographic target.This distinguishes generic performance loss from directed influence by conflicting text.
  • Benchmark design: SIGNPOST-Bench transforms each source image into five counterfactual variants: Original, Blank, Similar, Random, and Adversarial.The interventions are synthetic and localized to scene text while preserving surrounding non-textual content.
  • Research questions: The study asks how conflicting text affects localization, how compatible, unrelated, and conflicting replacements differ, and whether geographic conflict directs predictions toward injected targets.It also tests whether conflict-aware prompting helps two representative models resolve text–vision conflicts.
  • Contributions: 5,111 counterfactual groups and 25,555 image variants from four datasets comprise the benchmark, evaluated across 20 MLLMs from seven providers.The benchmark uses a paired five-condition design for within-scene comparisons.

2 Related Work

Prior work has broadened multimodal evaluation toward specific failure modes, including conflicts caused by scene text. Visual geolocation offers a continuous, geographically interpretable setting for studying these effects, while recent work applies MLLMs to location and urban reasoning.

  • Multimodal evaluation has expanded from aggregate accuracy toward perception, reasoning, integrated capabilities, expert knowledge, hallucination, natural adversarial samples, and image adversarial robustness.
  • Scene text can create cross-modal conflict: typographic class names can override CLIP’s visual recognition, while task competition occurs among scene-text, object, and action recognition.
  • Visual geolocation supports this analysis because outputs are continuous and errors retain geographic meaning.
  • Geolocation systems map visual appearance to locations and evaluate predictions by geodesic distance, using cues such as architecture, scripts, signs, contextual information, and neighborhood context.

3 Benchmark Formulation

SIGNPOST-Bench formulates text–vision conflict resolution as visual geolocation and measures text-induced prediction changes using matched counterfactual quintuplets derived from individual scenes. The framework also classifies native scene text by its geographic identifiability in context.

  • Task Definition: Given an image I, the model predicts coordinates, and localization error is the haversine distance D(f(I), ygt).
  • Task Definition: Each counterfactual group contains ground truth, selected scene-text spans, and five matched images for paired within-scene comparisons.Applying the same model to all five images measures text-induced prediction changes.
  • Counterfactual Quintuplet: The quintuplet consists of Original, Blank, Similar, Random, and Adversarial images.Original is unmodified; Blank removes selected text; Similar uses geographically or linguistically compatible replacements; Random adds unrelated readable text; Adversarial injects geographically conflicting text.
  • Counterfactual Quintuplet: Comparisons against Blank measure the effects of compatible, unrelated, and conflicting text within the same scene.
  • Counterfactual Quintuplet: For geocodable Adversarial images, shared coordinate space reveals whether predictions shift toward the injected target.These diagnostics are formalized in Section 5.
  • Scene-Text Coupling Taxonomy: Native scene text is classified by geographic identifiability in the source-scene context rather than fixed lexical categories.T1 Portable text has little geographic specificity; T2 Cultural text narrows location to a language or cultural region; T3 Geo-Specific text directly identifies a place or distinctive local entity.

4 Benchmark Construction

SIGNPOST-Bench is constructed from four geographically diverse image sources through OCR-guided scene-text replacement, target geocoding, and localized synthesis. The resulting benchmark contains 5,111 counterfactual quintuplets, with construction quality assessed through annotation and image audits.

  • Data sources: Images come from IM2GPS3K, YFCC4K, GoogleSV, and BaiduSV, combining geo-tagged web photographs with international and Chinese street-view imagery.Street-view sources are geographically sampled, and EasyOCR identifies candidate scene-text spans across all four sources.
  • Counterfactual generation: A three-stage pipeline generates Similar, Random, and Adversarial replacements, geocodes adversarial text, and synthesizes localized edits.Gemini-3.1-Flash-Lite selects up to three informative, editable scene-text spans per image and generates one replacement of each type.
  • Counterfactual generation: Adversarial targets are defined by the top-ranked valid Nominatim OpenStreetMap result for place names, landmarks, or addresses in the replacement text.Source-location metadata may be used during construction to ensure conflict with ground truth, but are not shown to evaluated models.
  • Quality assessment: 83.4% agreement (κ = 0.747) was obtained when two annotators independently assigned tiers to 350 stratified source images.A separate audit covered 120 generated Similar, Random, and Adversarial images.
  • Quality assessment: Audit means were 4.00 ± 1.26 for rendered-text naturalness, 1.32 ± 0.78 for artifact severity, and 1.14 ± 0.52 for surrounding-context damage.Naturalness uses a 1–5 scale from highly unnatural to fully natural, while artifact severity and context damage use a 1–5 scale from none to severe.
  • Benchmark composition: 5,111 counterfactual groups yield 25,555 images and 10,084 scene-text spans, distributed across T1, T2, and T3.T1, T2, and T3 contain 347 (6.8%), 3,851 (75.3%), and 913 (17.9%) groups, respectively.

5 Evaluation Protocol

SIGNPOST-Bench evaluates localization quality, text-induced error shifts, and adversarial target-following using paired metrics across controlled image variants. It also aggregates these measures into capability and conflict-robustness scores for 20 MLLMs evaluated equally across four datasets.

  • Primary diagnostic metrics: WLA scores geolocation predictions by geodesic error using exponential decay, with α = 0.005 and values reported as percentages.WLA decays to 0.5 at 138 km and below 0.01 at 1,000 km.
  • Primary diagnostic metrics: TBS measures paired error shifts from Blank, where a positive score indicates edited text worsened localization relative to the text-removed scene.TBS captures change in ground-truth error, not the distance between Blank and edited predictions.
  • Primary diagnostic metrics: TFR measures whether adversarial predictions land within τ = 50 km of injected targets, using only samples whose text resolves to valid coordinates.TFR is macro-averaged equally across datasets.
  • Primary diagnostic metrics: TDR measures paired movement toward the injected target from Blank to Adversarial, with positive values indicating reduced trap distance.It uses the same geocodable subset as TFR and is equally macro-averaged across datasets.
  • Composite scores: Capability Score C averages Original and Blank WLA, while Conflict Robustness Score R combines Random and Adversarial WLA retention with normalized TBS and TFR penalties.Both scores range from 0 to 1 and are reported as percentages; MCRS gives greater exponent weight to R.
  • Evaluation coverage: 20 MLLMs from seven providers were evaluated on all five variants of 5,111 counterfactual groups, producing 511,100 model–image evaluations.Metrics were averaged equally across four datasets, and every model received the same direct-coordinate prompt without chain-of-thought.

6 Results MLLM Performance Degrades under Text–Vision Conflict

Scene-text conflicts substantially degrade MLLM geolocation across models and datasets, while adversarial edits induce measurable shifts toward injected geographic targets. Clean-input capability and conflict robustness diverge, and targeted probing shows that strong localization does not guarantee conflict detection.

  • Overall degradation: 36.6%: Mean WLA falls from 47.11 to 29.89 under Adversarial edits, while median error grows 4.8× from 282 km to 1,347 km.Adversarial edits reduce WLA for every model and dataset, with 79 of 80 individual model–dataset cells declining.
  • Overall degradation: 20.6%, 31.8%, and 50.1%: Adversarial Acc@25, Acc@200, and Acc@750 decline from 34.9%, 50.1%, and 71.6%, respectively.Errors exceeding 2,500 km rise from 11.7% to 31.7%.
  • Semantic intervention effects: 1,577 km: Adversarial replacements increase error relative to Blank on average, compared with a 379 km reduction for Similar replacements and a 959 km increase for Random replacements.Similar has negative mean TBS in every dataset, whereas Random and Adversarial have positive mean TBS throughout.
  • Geographic conflict effects: 6.5%–20.1%: Adversarial Trap-Fit Rate spans models, with TFR and TDR computed on 1,732 groups containing geocodable targets.Adversarial TBS and TFR are strongly correlated (Spearman ρ = 0.836, p < 0.001), and positive mean paired TDR indicates movement toward the injected target for every model.
  • Provider- and family-level differences: 72.70, 72.23, and 69.80: Gemini-3-Flash, Gemini-3.1-Pro, and Gemini-2.5-Pro lead MCRS, while Claude-Haiku-4.5 scores 38.87 and Moonshot-Vision models score 41.46–41.47.No model is immune to conflicting text, and clean-input capability does not determine conflict robustness; named-place representations and limited disagreement exposure are proposed mechanisms.
  • Targeted conflict probing: 14.2%: Gemini-2.5-Flash achieves probing WLA of 38.28 but low Conflict Detection Accuracy, while GPT-4o-mini reaches 32.8% CDA with WLA of 15.10.The targeted analysis compares Gemini-2.5-Flash (R=77.77) and GPT-4o-mini (R=69.02) using structured conflict probing, defense prompting, and cross-task generalization.

7 Conclusion

SIGNPOST-Bench shows that text–vision conflict is a systematic, measurable failure mode across evaluated MLLMs. Its controlled five-condition design provides a reproducible framework for testing whether models use, reject, or follow conflicting scene text.

  • Benchmark: SIGNPOST-Bench is a controlled five-condition benchmark for text–vision conflict resolution across 20 models and four datasets.The benchmark evaluates how MLLMs arbitrate between visual and textual evidence under controlled conflict.
  • Findings: 4.8×: Conflicting scene text increases median prediction error across evaluated models.This identifies text–vision conflict as a systematic, measurable failure mode.
  • Findings: 6.5–20.1%: Adversarial predictions lie less than 50 km from their injected targets, and every model shows positive mean paired Trap Distance Reduction.These results indicate directed shifts toward targets introduced by conflicting scene text.
  • Findings: Compatible, unrelated, and conflicting text produce distinct effects, while Capability and Conflict Robustness reveal behavior not captured by clean-input performance or model scale alone.The benchmark distinguishes general localization capability from robustness to conflicting text.
  • Implications: SIGNPOST-Bench provides a reproducible framework for testing when MLLMs use, reject, or follow conflicting scene text.The framework is designed to measure text–vision conflict resolution directly.

Ethical Statement … Cross-task generalization prompts

The paper documents ethical safeguards for benchmark data and specifies standardized, structured prompts for attack generation, coordinate prediction, probing, defense, and cross-task diagnostics.

  • Ethical Statement: Benchmark metadata contains no personally identifiable information, and imagery is not redistributed beyond low-resolution examples reproduced for illustration.Variants are identified by source IDs and reconstruction instructions, remain subject to third-party source terms, and are synthetic edited images used solely for benchmarking.
  • A Full prompts: The full prompt suite standardizes how models generate attacks, predict coordinates, report evidence, apply defenses, and perform diagnostic tasks.The evaluated models receive task-specific prompts, while attack-generation prompts use construction-time metadata only to create conflicting targets.
  • Attack generation prompt: Attack generation requests structured output for automatic verification of scene-text replacements before image synthesis.Available source metadata may include the ground-truth city, county, province, or country, but this information is never provided to evaluated models.
  • Coordinate prediction prompt: All 20 models use the same standard evaluation prompt for direct coordinate prediction across datasets and image variants.This standardization supports consistent comparison across the benchmark’s conditions.
  • Structured probing prompt: The probing prompt asks models to provide visual evidence, textual evidence, consistency judgments, and final predictions in structured fields.The structured response exposes the evidence and judgments underlying each prediction.
  • Defense prompting: The defense prompt uses a structured format for direct coordinate prediction.This prompt is designated for conflict-aware defense prompting.
  • Cross-task generalization prompts: The generalization experiment evaluates scene-text consistency, country identification, and language detection as three diagnostic tasks.The accompanying prompts cover consistency judgment, country identification under text–vision conflict, and language identification.

B Reproducibility details … C MCRS formula details

The reproducibility details specify controlled sampling, OCR-based intervention construction, auditing, data organization, and deterministic evaluation procedures. The MCRS formulation combines macro-averaged variant performance with retention and behavior-quality terms, emphasizing adversarial conflict.

  • Benchmark construction: Benchmark construction combines dataset-specific sampling, OCR selection, localized text replacements, image synthesis, taxonomy labeling, and post-OCR screening.GoogleSV and BaiduSV sampling use fixed random seed 42; generated records undergo structured-field checks before synthesis.
  • Benchmark construction: The release provides metadata, source identifiers, coordinate labels, replacement specifications, taxonomy labels, and evaluation code, while restricted images require reconstruction from identifiers.Underlying images remain governed by their source terms.
  • Audit protocols: Taxonomy auditing reports 83.4% inter-annotator agreement and unweighted Cohen’s κ = 0.747 across a 350-image, tier-stratified set.The audit contains 100 T1, 150 T2, and 100 T3 automatically assigned images; automatic labels are not treated as ground truth.
  • Audit protocols: Human edit-quality auditing samples 120 completed ratings across datasets and Similar, Random, and Adversarial variants, assessing naturalness, artifacts, visual-context damage, and readability.Naturalness uses a 1–5 scale, while artifact severity and visual-context damage range from none to severe on 1–5 scales.
  • Evaluation and data organization: Predictions are linked to counterfactual quintuplets through shared group identifiers, with standard coordinate evaluation covering all five variants and separate probing, defense, and cross-task modes.The analysis pipeline produces dataset-level outputs from these model-scoped prediction records.
  • Evaluation and data organization: WLA, TBS, TFR, and TDR use equal macro-averaging across IM2GPS3K, YFCC4K, GoogleSV, and BaiduSV before MCRS is computed from model-level quantities.Metric eligibility is determined within each dataset, and per-dataset percentages are reported separately.
  • C MCRS formula details: MCRS defines capability from non-conflicting conditions, computes edited-condition retention against Blank with safeguards, and sets R = 0.22 ρrnd + 0.44 ρadv + 0.17 qTBS + 0.17 qTFR.The denominator floor stabilizes low-Blank ratios, clipping prevents compatible-text gains from offsetting conflict failures, and similar-text retention is omitted because it saturates at 1.0.

D MCRS sensitivity analysis

MCRS rankings are highly stable across score-weight settings and leave-one-component-out ablations. Separating capability and robustness remains important because robustness correlates imperfectly with adversarial performance, while fixed anchors are not saturated.

  • Weight and exponent sensitivity: Kendall τ ≥0.905 and Spearman ρ ≥0.979 across all tested exponent and adversarial-retention weight settings, relative to default score 100C0.40R0.60.Rankings remain stable when varying both the outer Capability/Robustness exponent and the internal adversarial-retention weight.
  • Component ablations: No single component determines the leaderboard when each component is removed and the remaining robustness weights are renormalized.These leave-one-component-out ablations test whether any individual component drives MCRS rankings.
  • Decomposition validity: Spearman ρ is 0.811 between R and Adversarial WLA versus 0.976 between MCRS and Adversarial WLA, supporting separate reporting of C and R.R also correlates negatively with adversarial TBS (ρ = −0.886) and TFR (ρ = −0.939), consistent with higher robustness scores indicating stronger resistance.
  • Anchor checks: The maximum observed model-level TFR is 0.201, below the 0.40 anchor, while maximum mean adversarial TBS is 2,660.5 km, below the 3,000 km anchor.Thus, the fixed anchors are not saturated for either TFR or adversarial TBS.
  • Geocodable evaluation cohort: 1,732 groups define the geocodable cohort for TFR and TDR; TFR uses Adversarial predictions, while paired TDR also uses corresponding Blank predictions.Table 7 reports TFR for every model and dataset, with Macro equal to the unweighted mean of the four dataset-level rates.

E Supplementary figures and analysis … I Additional examples

Supplementary analyses show that scene-text specificity amplifies vulnerability, while sensitivity checks support the robustness of comparative conclusions. Cross-task diagnostics and qualitative examples further show that injected text can redirect both geolocation and higher-level judgments.

  • E Supplementary figures and analysis; Full MCRS leaderboard: MCRS supplementary materials provide a 20-model capability-versus-conflict-robustness breakdown and a full leaderboard of component scores.The full leaderboard omits ρsim because it equals 1.0 for all evaluated models.
  • E Supplementary figures and analysis; Per-dataset probing and defense breakdown: Structured probing and defense metrics are reported for two diagnostic models across four datasets, and the supplementary text cautions against generalizing them to the full suite.Per-dataset results and condition-specific CDA values are provided in the diagnostic tables.
  • F Scene-text coupling stratification: 25.14 WLA points is the largest T3 attack-induced collapse, compared with 11.31 for T1 and 17.23 for T2.T3 also has the highest clean accuracy, with Original WLA 59.65; relative degradation rises from 25.0% for T1 to 42.2% for T3.
  • Cross-task generalization: 35.2% conflict recall and 37.7% text dominance characterize Gemini-2.5-Flash, while GPT-4o-mini marks only 13.9% of adversarial samples as conflicts.GPT-4o-mini’s text-dominance rate reaches 56.8%, compared with Gemini-2.5-Flash’s 37.7%.
  • Ablations and metric sensitivity: 46.31 Original WLA versus 38.14 Blank WLA yields an 8.17-point drop, confirming that native scene text carries geographic signal.Similar recovers 5.94 points relative to Blank, whereas Random and Adversarial are 6.84 and 8.03 points worse than Blank.
  • Ablations and metric sensitivity: Varying α ∈{0.002, 0.005, 0.01} changes absolute WLA but preserves relative model ordering across four models.This supports robustness of comparative conclusions to the WLA decay-constant choice.
  • Additional diagnostic tables: Mean and median TDR are macro-averaged by dataset, while a positive mean can coexist with an attraction rate below 50% when fewer large shifts dominate.Trap-Fit Rate varies gradually with τ on T3, and τ = 50 km is presented as a conservative balance; these values are T3-only, not aggregate all-tier TFR.
  • H Supplementary example figures; I Additional examples; Cross-task generalization breakdown: Representative examples cover source images, taxonomy tiers, per-dataset vulnerability, text-bias scores, and prompt-conditioned arbitration between injected text and visual evidence.In one GoogleSV example, structured probing predicts Sydney from injected text, whereas conflict-aware defense predicts Virginia from visual evidence.

J Full per-dataset results

Tables 20–23 present full per-dataset results for the 20-model suite across four datasets. They report WLA for all five image variants and adversarial TBS.

  • Per-dataset coverage: Tables 20–23 report per-dataset results for the full 20-model suite.The datasets are IM2GPS3K, YFCC4K, GoogleSV, and BaiduSV.
  • Metrics: Each table includes WLA for all five variants together with adversarial TBS.
  • IM2GPS3K: Table 20 reports full per-model results on IM2GPS3K.
  • YFCC4K: Table 21 reports full per-model results on YFCC4K.
  • GoogleSV: Table 22 reports full per-model results on GoogleSV.
  • BaiduSV: Table 23 reports full per-model results on BaiduSV.
Loading 2608.04244v1…