Source-linked AI summary

From Pixels to Concepts: Do Segmentation Models Understand What They Segment?

Shuang Liang, Zeqing Wang, Yuxian Li, Xihui Liu, Han Wang

arXiv:2605.09591v1cs.CV

TL;DR

Promptable segmentation benchmarks do not establish whether models faithfully ground textual concepts when visual attributes are misleading. CAFE addresses this gap with mask-preserving counterfactual edits and paired valid and invalid prompts, revealing that models often retain accurate masks for semantically invalid concepts; agentic verification improves this behavior but remains limited in scene complexity.

  • Problem

    Existing evaluations emphasize mask accuracy or object presence, leaving unclear whether promptable segmentation models distinguish concept-faithful grounding from shortcut-driven responses to misleading attribute cues.

  • Method

    CAFE evaluates edited images whose target regions and masks are preserved while surface appearance, context, or material attributes create paired valid and misleading prompts across three conflict types.

  • Results

    Current open-vocabulary models often fail to reject semantically invalid concepts, while CAFE-SAM3 reduces false positives and concept swaps through MLLM-based reasoning.

  • Takeaways & Limitations

    Strong positive-mask quality does not imply reliable semantic grounding, whereas explicit verification offers a path toward more reliable promptable segmentation.

  • Takeaways & Limitations

    CAFE evaluates a single counterfactually edited target per image and does not test crowded or mixed-instance scenes with multiple counterfactual or related instances.

Abstract

from arXiv · show

Segmentation is a fundamental vision task underlying numerous downstream applications. Recent promptable segmentation models, such as Segment Anything Model 3 (SAM3), extend segmentation from category-agnostic mask prediction to concept-guided localization conditioned on high-level textual prompts. However, existing benchmarks primarily evaluate mask accuracy or object presence, leaving unclear whether these models faithfully ground the queried concept or instead rely on visually salient but semantically misleading cues. We introduce CAFE: \textbf{C}ounterfactual \textbf{A}ttribute \textbf{F}actuality \textbf{E}valuation, a novel benchmark for evaluating concept-faithful segmentation in promptable segmentation models. Our \textbf{CAFE} is built on attribute-level counterfactual manipulation: the target region and ground-truth mask are preserved, while attributes such as surface appearance, context, or material composition are modified to introduce misleading semantic cues. The benchmark contains 2,146 paired test samples, each consisting of a target image, a ground-truth mask, a positive prompt, and a misleading negative prompt. These samples cover three counterfactual categories: Superficial Mimicry (\textbf{SM}), Context Conflict (\textbf{CC}), and Ontological Conflict (\textbf{OC}). We evaluate various model types and sizes on our CAFE. Experiments reveal a systematic gap between localization quality and concept discrimination: models often generate accurate masks even for misleading prompts, suggesting that strong mask prediction does not necessarily imply faithful semantic grounding. Our CAFE provides a controlled benchmark for diagnosing whether promptable segmentation models perform concept-faithful grounding rather than shortcut-driven mask retrieval.

1. Introduction

Promptable segmentation has progressed toward language-conditioned concept localization, but existing evaluations do not establish whether models ground concepts faithfully under misleading visual cues.

  • Motivation: Promptable segmentation extends beyond spatial grounding toward language-conditioned region association and concept segmentation.Earlier systems use visual prompts or modular language-grounding pipelines, while SAM3 directly produces masks from concept prompts.
  • Motivation: Existing benchmarks mainly assess predefined-category mask accuracy or object presence, leaving attribute-level semantic validity insufficiently tested.Counterfactual edits can preserve a visible, localizable target while changing appearance, context, or material identity.
  • CAFE: CAFE evaluates concept-faithful grounding using controlled counterfactual attribute interventions that preserve the target region and annotation mask.Each sample pairs an edited image and ground-truth mask with a semantically valid positive prompt and a visually plausible but invalid negative prompt.
  • CAFE: CAFE contains 2,146 paired test cases spanning superficial mimicry, context conflict, and ontological conflict.The benchmark evaluates end-to-end models, modular grounding-segmentation pipelines, and an agentic verification variant.

2. Related Works

Prior segmentation benchmarks emphasize output accuracy and localization, whereas CAFE targets semantic validity when mask-preserving edits create misleading concept cues.

  • Counterfactual evaluation: Counterfactual evaluation has been applied to fairness, robustness, vision-language understanding, and increasingly segmentation.These approaches test whether predictions rely on causal evidence rather than spurious correlations.
  • Open-vocabulary and promptable segmentation: Classical segmentation uses closed vocabularies, while SAM and SAM2 enable class-agnostic mask prediction from visual prompts.SAM2 extends promptable segmentation to video through a memory-based architecture.
  • Benchmarking segmentation models: Existing benchmarks span semantic, instance, panoptic, visual-promptable, and language-guided segmentation but do not directly test attribute-level semantic validity under mask-preserving edits.Such edits retain a visible annotated target while manipulating appearance or material, exposing misleading-prompt masks.

3. Task Definition

CAFE defines counterfactual attribute factuality by editing target attributes while keeping the region identifiable, then tests whether segmentation follows the valid concept rather than a plausible invalid cue.

  • Task Definition: A counterfactual image changes a target attribute while preserving spatial identifiability, enabling evaluation against a visually plausible but semantically invalid competing concept.The edited attribute may preserve object identity or shift the valid concept toward a material- or substance-defined concept.
  • Counterfactual Categories: Superficial Mimicry changes surface patterns, Context Conflict changes surroundings, and Ontological Conflict changes the target substance or material.Each category creates a distinct source of misleading semantic evidence.
  • Prompt Pair Construction: Each sample is represented by an edited image, target mask, valid positive prompt, misleading negative prompt, and counterfactual category.The positive prompt describes the valid edited-image concept, whereas the negative prompt describes the visually plausible invalid concept.
  • Semantic Validity: Semantic validity assigns v(I, q)=1 to the positive prompt and v(I, q)=0 to the misleading negative prompt.This binary definition formalizes whether a queried concept is supported by visual evidence in the edited image.
  • Evaluation Protocol: The evaluation checks high-confidence target alignment for the positive prompt and rejection of the negative prompt, using IoU and confidence thresholds to classify responses.High-confidence negative predictions are further distinguished by overlap with the target mask as aligned or unaligned false positives.

4. CAFE: Counterfactual Attribute Factuality Evaluation

CAFE evaluates whether promptable segmentation models distinguish concept-faithful grounding from visually misleading attribute cues. It combines target-aware classification, false-positive rates, and concept-swap metrics across counterfactual categories and model paradigms.

  • Dataset Statistics: 2,146 paired counterfactual samples span Superficial Mimicry, Context Conflict, and Ontological Conflict edits.The benchmark draws from COCO-Val2017, SA-Co/Gold, and LVIS-Val, with 1,111 SM, 593 CC, and additional OC samples.
  • Target-aware Classification: CAFE pairs each ground-truth annotation with a positive prompt and a misleading negative prompt for target-aware evaluation.Positive predictions are assessed by confidence and IoU alignment; misleading prompts are classified as true negatives or aligned and unaligned false positives.
  • False-positive Analysis: AFPR and UFPR partition false positives into target-aligned and unaligned responses under misleading prompts.AFPR captures semantically invalid high-confidence responses over the edited target, whereas UFPR captures high-confidence responses elsewhere.
  • Model Evaluation: Table 2 compares end-to-end, multi-model, and agentic systems using cgF1, IL_MCC, and pmF1 across the three counterfactual categories.The table caption states that CAFE-SAM3 (GPT-5.5) substantially improves over direct SAM 3, especially on Ontological Conflict.
  • Concept Swap Analysis: ACSR measures target-level concept swaps, while UCSR captures concept loss paired with hallucinated counterfactual detections elsewhere.ACSR isolates the strictest failure mode: the misleading concept replaces the original on the target itself.

5. Experiments

Experiments expose a systematic gap between positive-mask localization and concept discrimination, especially under ontological conflict. Explicit agentic verification substantially improves rejection of semantically invalid, visually plausible prompts.

  • Non-agentic models retain relatively high pmF1 but show low IL_MCC and cgF1, indicating that positive-case localization is easier than rejecting invalid concepts.Grounded SAM2 maintains stable pmF1 across categories while its IL_MCC remains consistently low.
  • Model evaluation: Ontological Conflict is the hardest category, with most non-agentic models obtaining negative IL_MCC and SAM3 falling from 0.857 on CC to -0.241 on OC.The result indicates that image-level presence prediction does not resolve ontological counterfactuals.
  • Agentic verification: CAFE-SAM3 raises overall cgF1 from 38.5 to 63.3, IL_MCC from 0.590 to 0.843, and pmF1 from 65.4 to 75.1 versus direct SAM3.The largest gains occur on OC, where cgF1 increases from -10.5 to 44.7 and IL_MCC from -0.241 to 0.633.
  • False positives: Agentic verification reduces overall FPR from 21.2% to 13.5%, AFPR from 20.5% to 11.9%, and ACSR from 8.9% to 1.7%.On OC, FPR drops from 66.3% to 29.2%, AFPR from 65.6% to 25.8%, and ACSR from 37.8% to 6.8%.
  • Interpretation: SAM3’s presence head improves robustness for superficial mimicry and context conflict but remains insufficient for ontological conflict.The authors attribute the improvement in semantic rejection primarily to explicit verification rather than positive-case segmentation.

6. Conclusion

The conclusion presents CAFE as a 2,146-sample benchmark for testing concept-faithful promptable segmentation across three attribute-level conflicts. Results show frequent rejection failures for semantically invalid concepts, while agentic reasoning reduces false positives and concept swaps.

  • CAFE comprises 2,146 paired samples with positive and misleading prompts spanning Superficial Mimicry, Context Conflict, and Ontological Conflict.
  • Current open-vocabulary segmentation models often fail to reject semantically invalid concepts under counterfactual cues.
  • SAM3’s image-level presence head improves robustness in some cases but remains insufficient for ontological conflicts.
  • CAFE-SAM3 demonstrates that MLLM-based reasoning can reduce false positives and concept swaps.

7. Limitations

CAFE evaluates one counterfactually edited target per image, so robustness in scenes containing multiple counterfactual or related instances remains untested.

  • CAFE currently evaluates a single counterfactually edited target per image, excluding more complex multi-instance configurations.
  • Counterfactual robustness in crowded scenes or scenes mixing edited and unedited related instances remains untested.

A. Dataset Preparation

The CAFE dataset preparation pipeline transforms source image-annotation pairs, uses Gemini 3 to generate editing instructions, applies nano-banana edits, and performs three stages of human filtering and cross-checking.

  • Source image-annotation pairs come from validation sets of COCO, LVIS, and SA-Co/Gold, with affine transformations applied to fit Gemini’s input resolution.
  • Gemini-3 generates editing instructions from transformed images, annotations, and prompt-engineered inputs containing in-context cases.
  • Nano-banana applies the generated edits across all three counterfactual categories.
  • Human review removes obvious artifacts, assesses edit quality and prompt plausibility, then uses three editors for final cross-checking.

A.2. Prompts and Models for Dataset Generation

CAFE’s generation pipeline specifies structured counterfactual edits, prompt pairs, and strict semantic controls across superficial, contextual, and ontological conflicts.

  • CAFE Auto-Prompt Pipeline – Shared Header & Output Schema: The pipeline requires one specified counterfactual edit for the target instance and returns an exact XML-style schema.The schema includes edit type, instruction, rationale, and prompt fields.
  • CAFE Auto-Prompt Pipeline – Shared Header & Output Schema: Positive prompts identify the valid concept after editing, while negative prompts identify the mimicked, contextually conflicting, or ontologically invalid category.For ontological conflict, the positive prompt names the newly rendered instance category.
  • I. Rules for Superficial Mimicry: Superficial mimicry changes an object’s appearance while preserving its material and structure, with the new appearance covering the visible surface.The conflicting cue must resemble a recognizable category without depicting an image or transforming the object’s volume-level material.
  • A.2.8. More Discussions on the Prompts for Ontological Conflict: Ontological-conflict negatives use restrictive modifiers and expert consensus to avoid ambiguity between visual resemblance and actual category membership.Examples include “real airplane,” “living human,” and “functional blender.”
  • A.2.8. More Discussions on the Prompts for Ontological Conflict: CAFE uses precise negative prompts such as “real airplane” when an edited cloud resembles an airplane without being one.The design distinguishes category ambiguity from semantic invalidity.
  • A.2.8. More Discussions on the Prompts for Ontological Conflict: Ontological-conflict cases undergo strict review, yielding a lower accepted proportion because their semantic acceptance criteria are intentionally stringent.The review aims to minimize semantic controversy while testing hallucination.
  • B. More Examples from CAFE: Additional CAFE samples pair edited images and inherited masks with valid positive prompts and visually plausible invalid negatives across SM, CC, and OC.The examples vary object categories, prompt pairs, and attribute-level conflicts.

C.1. Compute Resources.

The experiments use fixed, evaluation-only inference on CAFE with publicly released segmentation and grounding models, standardized scripts, and calibrated thresholds.

  • C.1. Compute Resources.: All experiments are evaluation-only inference runs on the fixed CAFE benchmark, using identical image-prompt pairs and evaluation scripts.No model training or fine-tuning is performed.
  • C.1. Compute Resources.: The evaluation uses RTX 5090 GPUs with 32GB memory, with compute dominated by benchmark inference and threshold calibration.The agentic diagnostic probe also requires additional tool calls.
  • C.1. Compute Resources.: The evaluated model set includes YOLO-World-Seg-L, SAM3, OpenSeeD, Grounded SAM2, and OWLv2 followed by SAM.The models use official released checkpoints or repositories.
  • C.1. Compute Resources.: CAFE-SAM3 Agent uses SAM3 as a segmentation tool through four tool calls, separately processing positive and negative prompts for each target.The agent can run up to 10 turns per episode with a confidence threshold of 0.5.
  • C.1. Compute Resources.: Baseline detection thresholds are calibrated on LVIS box detection by selecting the value that maximizes LVIS cgF1 before fixed CAFE evaluation.The implementation sweeps thresholds from 0.05 to 0.95 in steps of 0.05.

C.4. Threshold Sensibility of Target-aligned Metrics

Target-aligned metrics are robust to moderate IoU-threshold changes, while the agent’s multi-turn process is designed to inspect concept-relevant visual evidence rather than rely on a single mask proposal.

  • C.4. Threshold Sensibility of Target-aligned Metrics: τ = 0.3 is used throughout because AFPR and ACSR remain stable from τ = 0.3 to 0.7 across all CAFE subsets.For example, OC-AFPR changes by less than 0.025.
  • C.4. Threshold Sensibility of Target-aligned Metrics: At τ = 0.9, some predictions move from TA-FP to UA-FP, but the image-level false-positive rate remains unchanged.This threshold behavior motivates using the lowest threshold that still captures meaningful target alignment.
  • C.4. Threshold Sensibility of Target-aligned Metrics: The sensitivity curves are flat for τ ∈[0.3, 0.7], indicating that wrong predictions generally overlap the source target with high IoU rather than reflecting marginal mask alignment.The figure reports overlap of approximately 0.7 or higher.
  • Environment and Execution Flow: The agent begins with a concept prompt, calls segment_phrase, and compares subsequent tool results with the original image and query.Exactly one tool is called per turn.
  • Visual Reasoning Guidance: Visual reasoning guidance instructs the agent to assess material, substance, physical properties, surface cues, and scene context rather than relying on shape alone.It recommends examine_masks when fine-grained visual details are uncertain.
  • Understanding the User Query: The agent must ground the whole queried target, distinguish specific instances when necessary, and avoid grounding secondary objects.The query determines whether all or only one instance should be selected.
  • SAM3-CAFE Agent System Prompt: segment_phrase accepts concise noun phrases, while examine_masks inspects selected mask regions and select_masks_and_return finalizes only confident masks.report_no_mask is reserved for cases where careful examination finds no matching concept.
  • SAM3-CAFE Agent System Prompt: The agent may report no mask only after re-examining the image and ruling out a clear concept match, not because of minor discrepancies.This rule constrains false absence reports.

D.2. Case Analysis for CAFE-SAM3 Agent

Case analyses show that SAM3 and the agent can follow misleading visual cues, while iterative inspection can also reject an incorrect mask and report that a queried concept is absent.

  • Case Analysis: In the superficial-mimicry example, the original toy bird is rendered with tiger-like stripes, making “toy tiger” a misleading negative prompt.The negative concept should be rejected despite the visible pattern.
  • Case Analysis: In the context-conflict example, the agent receives an ECG Monitor query and initially segments a desktop iMac rather than a medical monitor.The direct prompt returns one incorrect mask.
  • Case Analysis: After trying “patient monitor” and finding zero masks, the agent re-examines the scene and reports that no ECG monitor is present.The reasoning distinguishes the iMac from a device with vital-signs waveforms, leads, or bedside monitoring hardware.
  • Case Analysis: The agent accepts the tiger-like toy mask because it treats stripe pattern as decisive and fails to inspect the bird’s beak and wings.SAM3 alone produces an identical false-positive mask under the same prompt.
  • Case Analysis: The cases illustrate why CAFE evaluates concept-faithful grounding under attribute interventions rather than mask retrieval alone.The benchmark is built from existing datasets and models used for evaluation without redistributing those models.
Loading 2605.09591v1…