Source-linked AI summary
Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
Yuanzhi Xu, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao, Sixue Lin
TL;DR
Large vision-language models can recognize visible objects and attributes yet bind an attribute to the wrong same-class instance, a failure that aggregate evaluations do not identify. InstaBind-Lite formalizes and measures this error across seven models, finding 19.84% average misbinding for open-source systems versus 7.55% for API systems.
Problem
Existing evaluations often collapse wrong-instance binding with other errors, leaving reliable instance–attribute correspondence insufficiently measured.
Method
InstaBind-Lite uses controlled same-class groups, ordered instance annotations, source labels, and four question levels to distinguish misbinding from other failures.
Results
Open-source models average 19.84% MBR versus 7.55% for API models across seven LVLMs.
Takeaways & Limitations
DSCAM is a reproducible, spatially structured failure category, so instance-level reliability cannot be inferred from general accuracy alone.
Abstract
from arXiv · showhide
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metrics may regard both the object and attribute as image-supported; neither reveals the transfer. This study formalizes this blind spot as Dense Same-Class Attribute Misbinding (DSCAM) and presents InstaBind-Lite, a controlled benchmark that makes it directly measurable. Its 524 images contain 529 curated groups of 3-6 same-class entities, 1773 boxed instances, ordered neighbors, distinguishable color-like attributes, and four complementary question levels, yielding 9580 deterministically evaluated questions. Unlike existing protocols, source-instance annotations separate unsupported generation and recognition failure from an attribute copied from another visible entity. Binding-specific metrics further quantify transfer frequency, adjacency, ordinal distance, and intervention effects. Across five open-source and two commercial/API models, the open-source systems average 19.84% Misbinding Rate and the API systems 7.55%; these errors are hidden by aggregate accuracy. Among identifiable transfers, 80.70% and 81.51%, respectively, originate from adjacent instances. Localization and instance-first interventions help selected models but are not universal remedies. InstaBind-Lite therefore turns previously undifferentiated wrong answers into source-identifiable failure categories and tests a reliability dimension that conventional benchmarks cannot determine: whether a model knows not only what is visible, but which instance owns each attribute.
I. INTRODUCTION
The introduction identifies Dense Same-Class Attribute Misbinding (DSCAM) as the hidden failure of assigning a visible attribute to the wrong same-class instance, which existing benchmarks cannot distinguish from other wrong answers. InstaBind-Lite addresses this gap with source-annotated, ordered same-class groups and binding-specific evaluation.
- Evaluation gap: Generic VQA and object-hallucination benchmarks cannot determine whether each entity retains its own property.They collapse these cases into the same incorrect answer or test only whether content is supported somewhere in the image.
- Problem definition: DSCAM is a distinct failure category in which a visible attribute is transferred to the wrong same-class instance.Source identification separates binding failure from unsupported hallucination and attribute-recognition failure.
- Evaluation gap: Existing resources cannot recover DSCAM prevalence, source, or distance because they lack inspectable same-class groups, distinguishable attributes, and directed transfer annotations.Grounding evaluates localization, while compositionality benchmarks do not measure instance-to-instance transfer in dense natural scenes.
- InstaBind-Lite: 524 images, 1773 instances, and 9580 questions define InstaBind-Lite as a diagnostic benchmark for measurable wrong-instance binding.The benchmark uses controlled same-class groups, visible color-like attributes, boxes, ordered relations, and source-instance annotations.
- Evaluation design: Binding-specific metrics and interventions measure transfer source and test whether localization or instance-first prompting reduces misbinding.MBR identifies visible wrong-instance transfers; A-MBR and ordinal Distance-MBR characterize spatial source, while crop/context localization and instance-first prompting target visual competition.
II. RELATED WORK … E. Binding Mechanisms and Multi-Subject Misbinding
Existing benchmarks detect hallucination, recognition, grounding, and compositional failures, but generally cannot identify when a visible attribute is assigned to the wrong same-class instance. InstaBind-Lite targets this gap by testing instance-specific attribution with controlled same-class scenes, ordered attributes, and source-instance labels.
- II. RELATED WORK: Unlike broad evaluation, a focused benchmark can trace a wrong attribute to a competing same-class source instance without claiming general superiority.Its diagnostic value lies in observability of the transfer event.
- A. LVLM Hallucination Evaluation: Presence-based hallucination protocols detect unsupported objects or visual illusions, whereas DSCAM concerns image-supported objects and attributes paired with the wrong instance.CHAIR, POPE, AMBER, HallusionBench, and related methods either target unsupported generation or leave this transfer unspecified.
- A. LVLM Hallucination Evaluation: DSCAM specifically measures false instance correspondence when both the class and attribute are visibly supported, a failure that existing presence-based protocols miss or mark only as unspecified.This distinction separates attribute transfer from unsupported generation and ordinary recognition failure.
- B. Visual Question Answering and General Multimodal Benchmarks: General VQA and multimodal benchmarks broaden answer and capability coverage but do not jointly control same-class density, attribute uniqueness, ordered neighbors, and wrong-source attribution.This leaves adjacent attribute transfers indistinguishable from ordinary errors in standard protocols.
- C. Visual Grounding and Referring Expressions: Grounding benchmarks score region localization, while DSCAM additionally tests whether a correctly located target receives its own attribute rather than its neighbor’s.InstaBind-Lite retains boxes for interventions and adds ordered attributes with source-instance error labels.
- D. Compositionality and Attribute–Relation Binding: Compositionality benchmarks motivate separating component recognition from composition by holding the same-class set and vocabulary fixed while changing the queried instance.Relation-sensitive evaluation further motivates InstaBind-Lite’s instance-level and ordinal testing.
- E. Binding Mechanisms and Multi-Subject Misbinding: Mechanistic studies analyze object–reference binding across model components, while MultiBind examines cross-subject attribute transfer in image generation.The present work asks the complementary behavioral question of whether an LVLM assigns an observed attribute to the correct same-class instance in a natural image.
III. PROBLEM DEFINITION
Dense Same-Class Attribute Misbinding identifies errors where a model attaches a visible attribute or position to the wrong same-class instance. The definition distinguishes these transfers from out-of-set hallucinations and evaluates them across four question levels.
- A same-class group contains entities sharing class c, arranged spatially, with each entity assigned an attribute value for a specified attribute type.
- Four question levels target instance attributes, attribute-to-position retrieval, local neighbor relations, or verification of an instance-attribute proposition.Predictions are correct when they match the canonical answer after normalization.
- A misbinding occurs when an incorrect prediction matches another same-class entity’s attribute or position rather than the specified target.This distinguishes a wrong-instance transfer from a generic incorrect answer.
- Out-of-set hallucination is counted separately when the predicted attribute does not appear in the same-class group, testing whether errors are image-grounded but attached to the wrong instance.The diagnostic focus is the failure mechanism behind an incorrect answer, especially a neighbor’s attribute.
IV. INSTABIND-LITEDATASET … C. Attribute Scope and Design Trade-off
InstaBind-Lite is a high-purity diagnostic benchmark built from curated same-class groups with inspectable ordering, attributes, and instances. Its controlled focus on color-like attributes enables source attribution without an LLM judge while limiting conclusions about other attribute types.
- IV. INSTABIND-LITEDATASET: 524 images, 529 groups, 1773 instances, and 9580 questions comprise the benchmark.Each image contains 3–6 same-class entities with clear order, visible attributes, and low ambiguity.
- A. Data Sources: Manual verification removes incomplete groups, severe occlusion, strong reflection, tiny objects, and uncertain attributes.Images come from COCO, GQA, VAW, manually selected open-license web images, and self-shot images; source metadata and original licenses remain separate from annotations.
- B. Annotation Schema: Each group records its image, class, spatial order, boxes, attributes, and neighbor links.Boxes support quality control and interventions, while questions use natural-language positions and relations.
- B. Annotation Schema: Person questions use upper-clothing color, and transparent or multicolor labels are retained only when unambiguous.This annotation rule preserves interpretable person attributes and excludes ambiguous labels.
- C. Attribute Scope and Design Trade-off: Color and upper-clothing color provide local attributes shared across classes and expressible with a compact vocabulary.This scope supports controlled binding evaluation rather than a general-purpose attribute benchmark.
- C. Attribute Scope and Design Trade-off: L2 additionally requires within-group attribute uniqueness, enabling source attribution without an LLM judge.The benchmark does not assume unchanged misbinding rates for actions, materials, textures, shapes, or states.
D. Question Levels … B. Misbinding and Adjacency
The benchmark uses four question levels to probe position-conditioned reading, reverse binding, relational interference, and instance verification. Its metrics distinguish standard accuracy from misbinding frequency and adjacency-structured attribute transfer.
- D. Question Levels: The benchmark generates four question levels spanning position-to-attribute, attribute-to-position, relation-interference, and instance verification.L1 asks for an attribute by position; L2 asks for a position by attribute; L3 introduces relational interference; L4 verifies an instance-level proposition.
- D. Question Levels: L2 requires the queried attribute to be unique within the same-class group.This constraint supports unambiguous attribute-to-position evaluation.
- D. Question Levels: Together, the levels separate perception errors from instance-level attribute transfer.They test position-conditioned reading, reverse binding, local relational interference, and proposition verification.
- A. Accuracy: Accuracy measures whether the normalized model prediction matches the canonical answer but does not explain how wrong answers relate to image content.It remains the standard task-level score.
- V. METRICS: The metric framework distinguishes overall task correctness from binding-specific transfer and adjacency behavior.This makes attribute-transfer errors measurable beyond aggregate accuracy.
- B. Misbinding and Adjacency: MBR measures transfer frequency, error-conditioned MBR its share among wrong answers, and A-MBR whether transfers follow within-group neighbor structure.The rates are defined over all questions, wrong answers, misbindings, and adjacent misbindings.
C. Out-of-Set Hallucination … B. Main Results: What Aggregate Accuracy Conceals
InstaBind-Lite separates out-of-set hallucination from in-image attribute misbinding and shows that aggregate accuracy conceals substantial wrong-instance transfers. Across models, residual misbindings remain strongly local, while interventions provide model-dependent evidence about same-class visual competition.
- C. Out-of-Set Hallucination: Out-of-set hallucination measures predictions matching neither a same-class instance’s attribute nor position, separating ungrounded attributes from in-image misbinding.A model may therefore show low object hallucination while retaining substantial instance-level misbinding.
- D. Distance-MBR and Confusion Matrix: Distance-MBR uses ordinal distance, with distance 1 denoting adjacent instances in annotated left-to-right order rather than a Euclidean pixel-distance bin.This tests local instance competition robustly to image scale and uneven spacing, but cannot distinguish ordinal adjacency from physical separation as the cause.
- E. Intervention Gap: The binding gap compares intervention and full-image accuracy on the same target-question subset, combining localized visual views with reformulated queries.A positive gap suggests reduced same-class competition helps, whereas a small or negative gap does not establish a universal remedy.
- A. Models and Inference Protocol: Five open-source and two commercial/API LVLMs received identical questions, answer constraints, normalized parsing, and short-answer prompts under deterministic or low-temperature decoding when supported.The protocol is designed to measure visual binding rather than generation style.
- B. Main Results: What Aggregate Accuracy Conceals: 81.52% and 80.13% aggregate accuracy identify Qwen3-VL-Plus and Gemini-3.5-Flash as strongest, yet their MBRs remain 9.31% and 5.79%.Aggregate accuracy cannot explain why remaining answers are wrong, and high general accuracy does not ensure instance–attribute correspondence.
- B. Main Results: What Aggregate Accuracy Conceals: Complementary instance-binding tests are required because these findings do not negate general-benchmark improvements but show that visual-reliability claims need instance-level evaluation.The conclusion follows from aggregate accuracy’s inability to explain residual errors and the prevalence of identifiable wrong-instance transfers.
- B. Main Results: What Aggregate Accuracy Conceals: 13.65% to 34.01% is the open-source MBR range, while 74.54% of LLaVA-1.5’s wrong answers are identifiable wrong-instance transfers.MBR exceeds out-of-set hallucination in six of seven models, so object-presence evaluation underdescribes a substantial error component.
- B. Main Results: What Aggregate Accuracy Conceals: 76.97% to 84.65% is the A-MBR range across all seven models, including 78.38% and 84.64% for the two API systems.Although stronger models reduce misbinding frequency, residual failures preserve a local signature consistent with neighboring same-class representation competition.
C. Statistical Stability · D. Parser Reliability · VII. QUALITATIVE ANALYSIS AND PRACTICAL SIGNIFICANCE
The paper supports stable, parser-validated misbinding measurement and shows that identifiable wrong-instance transfers have practical safety implications despite correct object and attribute recognition. Qualitative interventions recover target attributes in recurring adjacent-transfer cases, but do not establish causality or guarantee safe deployment.
- C. Statistical Stability: Image-cluster bootstrap resampling preserves all questions from each sampled image, avoiding dependence from multiple templates derived from one scene.The evaluation uses 1000 bootstrap resamples, and reported MBR intervals are narrow relative to cross-model differences.
- C. Statistical Stability: 34.01% [32.70, 35.24] is LLaVA-1.5’s full-image MBR, which remains distinctly high under image-cluster bootstrap evaluation.The interval is reported as remaining distinct under the resampling procedure.
- D. Parser Reliability: 100% parser agreement with human judgment is observed across a stratified audit of 200 outputs from the five open-source models.The audit contains 40 outputs per model, all four question levels, and five outcome strata.
- VII. QUALITATIVE ANALYSIS AND PRACTICAL SIGNIFICANCE: Six qualitative cases show adjacent transfers across bags, basins, bottles, cars, cups, and umbrellas, with full-image predictions matching the annotated distance-1 source.Crop and context-crop interventions recover the target value in each described case.
- VII. QUALITATIVE ANALYSIS AND PRACTICAL SIGNIFICANCE: Localized recovery after an identifiable adjacent-source transfer is more consistent with same-class competition than arbitrary color guessing.The passage qualifies this interpretation and does not claim that the intervention reveals an internal causal mechanism.
- VII. QUALITATIVE ANALYSIS AND PRACTICAL SIGNIFICANCE: Aggregate accuracy and object-level hallucination do not reveal visible wrong-instance transfers measured by MBR.This motivates binding-specific evaluation beyond object presence and attribute visibility.
- VII. QUALITATIVE ANALYSIS AND PRACTICAL SIGNIFICANCE: The joint image-and-query intervention does not establish an internal causal mechanism, while complete questions, ordered attribute sequences, and outputs are released as metadata.The release supports reproducibility without resolving intervention causality.
- VII. QUALITATIVE ANALYSIS AND PRACTICAL SIGNIFICANCE: Misbinding can return the wrong belonging, inventory item, robot-picking target, or vehicle, although these are risk mappings rather than measured deployment outcomes.Object-presence checks cannot flag a wrong owner when every class and color is visible.
VIII. INTERVENTION ANALYSIS · A. Protocol
The intervention protocol tests whether isolating a queried target reduces source-attributable misbinding, using oracle and context crops that differ in retained visual context. Results show substantial but model-dependent benefits, with localization helping selected open-source systems but failing as a universal remedy.
- A. Protocol: Interventions evaluate 4261 L1/L3 questions with a single target and attribute, resizing images above 1800 pixels while proportionally scaling target boxes.Crop oracle uses the annotated target box with 15% padding; context crop uses 50% padding.
- A. Protocol: “Oracle” provides the ground-truth target box rather than perfect isolation, so 15% crops may retain overlapping neighbors while 50% crops preserve local context.The protocol therefore compares target isolation against retained nearby entities and scene cues.
- A. Protocol: The protocol tests whether making the target easier to isolate reduces source-attributable misbinding through a joint visual-and-query localization evaluation.This intervention directly probes whether visual competition contributes to binding errors.
- A. Protocol: 20.14 accuracy points and 19.12 MBR points are gained or reduced, respectively, by LLaVA-1.5 under crop oracle; MiniCPM-V-2.6 changes by 9.22 and 10.57 points.LLaVA-OneVision also reduces MBR by 4.95 points, supporting competition-sensitive binding beyond accuracy changes alone.
- A. Protocol: Gemini-3.5-Flash and Qwen3-VL-Plus increase MBR under localized views despite lower full-image MBR, showing that target isolation is not universally safe.For proprietary models, preprocessing prevents separating context removal from crop-induced scale or distribution shift.
- A. Protocol: MBR reduction is reliably positive for LLaVA-1.5, LLaVA-OneVision, and MiniCPM-V, whereas Qwen2.5-VL, InternVL3, and API-model intervals include zero or are negative.Qwen2.5-VL’s context-crop interval is positive, but its crop interval crosses zero.
- A. Protocol: Localization identifies a substantial same-class competition component in selected open-source models, but no single mechanism is shared equally by every system.The image-cluster confidence intervals reinforce the model-dependent nature of intervention effects.
B. Lightweight Inference-Time Mitigation
Lightweight inference-time interventions can reduce same-class attribute misbinding, but their effects are model-dependent and may trade off against accuracy or useful context. Binding-specific evaluation exposes these trade-offs and motivates mitigation that lowers MBR without sacrificing recognition and instruction following.
- Instance-first prompting: Instance-first prompting raises InternVL3 accuracy by 1.54 points and reduces MBR by 4.36 points, while MiniCPM-V gains 2.17 accuracy points and reduces MBR by 3.46 points.The method externalizes a left-to-right entity–attribute map without requiring training or architectural access.
- Model-dependent trade-offs: Qwen2.5-VL slightly reduces MBR while losing accuracy, whereas LLaVA-1.5 loses 22.58 accuracy points and increases MBR under the longer prompt.Binding-specific evaluation distinguishes whether apparent changes repair wrong-instance transfer rather than reporting only aggregate improvement or degradation.
- Localization: Localization sharply reduces source-attributable misbinding for several open-source models but can remove useful context or shift the input distribution for stronger API systems.The intervention effect is reported for L1/L3 MBR.
- Implications: Future mitigation should lower MBR without sacrificing general recognition and instruction following, motivating grounding-aware decoding, contrastive neighbor suppression, and targeted fine-tuning.The mixed intervention outcomes position InstaBind-Lite as a testbed for these approaches.
IX. DISCUSSION … D. Diagnostic Scope and Mitigation
InstaBind-Lite exposes source-specific attribute misbinding that standard accuracy, hallucination, and grounding measures cannot distinguish, revealing unreliable entity-specific responses despite strong aggregate performance. Its neighbor-concentrated errors motivate source-aware evaluation and targeted mitigation, while the controlled color domain and model-dependent interventions limit generalization.
- A. Why a Source-Aware Benchmark Is Necessary: Standard accuracy, object-hallucination, and grounding measures do not determine whether a visible attribute was assigned to its owner.InstaBind-Lite records target and competing source instances, separating unsupported generation from attribute transfer and recognition failure.
- A. Why a Source-Aware Benchmark Is Necessary: API models lead aggregate accuracy yet retain measurable MBR and high A-MBR, while six of seven models transfer attributes more often than hallucinating out-of-set objects.
- A. Why a Source-Aware Benchmark Is Necessary: Broad capability and object-presence scores can mask entity-specific unreliability, so InstaBind-Lite tests instance grounding directly rather than replacing conventional benchmarks.
- B. Adjacent Structure and Mechanistic Implications: Most identifiable transfers originate from ordinal distance 1 across architectures, and bootstrap resampling preserves this neighbor concentration despite substantially different total MBR.The pattern supports local same-class competition rather than uniform answer noise, but does not establish a causal effect of Euclidean separation.
- C. Implications for Multimedia Systems: Wrong-entity identification can cause wrong-item retrieval or action in catalog search, robotic selection, traffic retrieval, and other multimedia systems.Source-aware evaluation is required before treating an LVLM as reliable for entity-specific decisions.
- D. Diagnostic Scope and Mitigation: The controlled color domain enables deterministic source attribution and a clean binding signal, but rates may differ for actions, materials, states, and part attributes.Those semantics introduce temporal and annotation ambiguity, requiring adapted question templates and judges.
- D. Diagnostic Scope and Mitigation: Crop and instance-first prompting are probes rather than complete solutions because their effects depend on the model.This model dependence motivates further binding-aware training research.
X. DATA AVAILABILITY AND ETHICAL CONSIDERATIONS · XI. CONCLUSION · A. Future Work
InstaBind-Lite is designed as a reproducible, ethically bounded benchmark that identifies Dense Same-Class Attribute Misbinding as wrong-instance attribute transfer. The study reports substantial misbinding, proposes controlled mitigation and evaluation extensions, and plans broader leakage-safe development and cross-dataset validation.
- X. DATA AVAILABILITY AND ETHICAL CONSIDERATIONS: InstaBind-Lite is intended for release with annotations, generated questions, evaluation scripts, parser code, and model-output metadata for reproducing tables and figures.Established-dataset images will follow original terms, while manually collected web images will be redistributed only with recoverable license evidence permitting use.
- X. DATA AVAILABILITY AND ETHICAL CONSIDERATIONS: Person images are restricted to visible upper-clothing color evaluation, excluding identity, demographic, and sensitive-attribute questions.Manual review excludes potentially offensive, private, or ambiguous depictions; the dataset is not intended for surveillance, identity inference, or demographic profiling.
- XI. CONCLUSION: Dense Same-Class Attribute Misbinding is formalized as assigning an image-present attribute to the wrong same-class instance.InstaBind-Lite uses controlled groups, ordered instance annotations, and four question levels to identify whether evidence was copied from a specific visible instance.
- XI. CONCLUSION: DSCAM is a reproducible, spatially structured error category supported by confidence intervals, parser audits, real-image cases, and controlled interventions.Mitigation must reduce wrong-instance transfer without discarding useful context or general capability.
- A. Future Work: Future work will use leakage-safe train, validation, and heldout test splits to compare box-conditioned tokens, coordinate embeddings, neighbor-contrastive losses, and ordered supervision.All variants will share a frozen test set and parser; success requires lower MBR and Err-MBR with stable or higher performance.
- A. Future Work: Future expansions will balance underrepresented classes and add auditable material, state, action, part, depth, and temporal-identity attributes.Evaluation will add normalized pixel distance, overlap, depth, and occlusion, broaden parser audits, and test successful mitigations on newly collected and cross-dataset images.