Source-linked AI summary
Conditional Visual Evidence Utility: State-Dependent Rank Reversals in Frozen Vision-Language Encoders
Yunxuan Fang, Xinhe Wang
TL;DR
Static importance scores may fail to capture how visual evidence value changes after other cues are acquired. The paper measures conditional utility in controlled compositional search with frozen vision-language encoders, then tests whether rank reversals follow designed candidate-overlap regimes and matter for decisions. It finds robust state-dependent reversals and positive exploratory matched-first-action utility across evaluator changes.
Problem
The paper asks whether visual evidence has a fixed importance ordering or whether its marginal utility changes with acquired evidence and the semantic query.
Method
The study exposes color, shape, and texture evidence independently in controlled scenes and evaluates conditional utilities and reranking with frozen OpenCLIP and SigLIP encoders.
Results
Robust reversals concentrate in oracle-positive candidate-overlap regimes, persist across evidence constructions and wordings, disappear under query-scene derangement, and show positive cross-condition matched-first-action utility.
Takeaways & Limitations
The findings support evaluating vision-language evidence use conditionally across matched acquisition states rather than relying on a single static ranking.
Takeaways & Limitations
The controlled setup is narrow: both oracle-positive regimes involve texture, whose isolated evidence is decoded less accurately than color or shape, so attribute-role invariance remains unestablished.
Abstract
from arXiv · showhide
Static importance scores compress visual evidence into a single ranking, but the value of remaining evidence can change after one cue has been observed. We study this possibility in controlled compositional visual search, where color, shape, and texture evidence can be independently exposed and their conditional marginal utility measured across acquisition states. In a held-out confirmation on 800 scenes, frozen OpenCLIP and SigLIP exhibit robust state-dependent rank reversals that concentrate in candidate-overlap regimes designed to induce ordering changes. The structure persists across two evidence-accumulation constructions and ten equivalent query wordings, but disappears under query-scene derangement. We also ask whether these reversals matter for decisions. In a post-confirmation exploratory matched-first-action analysis, reranking only after the first acquisition yields positive step-2 utility when decisions are selected under one evidence mode, wording, or backbone and evaluated under another. Together, these results show that evidence importance is state-dependent in this controlled setup and that updating an evidence ordering can retain decision-relevant value across evaluator changes. They motivate evaluating vision-language evidence use conditionally rather than through a single static ranking, while providing a measurable target for future adaptive evidence-selection methods.
1 Introduction
The paper tests whether visual evidence has a fixed importance ordering or whether its marginal utility changes with acquired evidence and the query. Controlled experiments find state-dependent reversals and exploratory evidence that post-acquisition reranking retains decision value across evaluator changes.
- Static importance scores ignore the evidence already acquired, motivating conditional marginal utility as a test of state-dependent visual evidence value.
- The study exposes color, shape, and texture evidence independently in controlled compositional search scenes and measures utilities with frozen vision-language encoders.
- Held-out confirmation finds robust reversals concentrated in oracle-predicted candidate-overlap regimes across backbones and evidence constructions, with stability under wording changes and disappearance after query-scene derangement.
- The analysis treats reversal structure as conditional on both acquired evidence state and semantic target query rather than as an aggregate reversal rate.
- Post-confirmation matched-first-action analysis finds positive decision value when remaining evidence is reranked after the first acquisition and evaluated across evidence modes, wordings, or backbones.
- The contributions include a controlled framework for measuring conditional utility over isolated, nonspatial evidence types and evidence of regime-specific, cross-condition rank reversals.
2 Related work
Related work studies sequential and history-conditioned evidence selection, visual-token routing, and static importance estimation. This paper instead measures how visual evidence utility changes across controlled acquisition states.
- Sequential feature acquisition and active vision choose subsequent observations conditional on the current acquisition state, including spatial, scale, temporal, or information-based choices.
- The paper complements these approaches by directly measuring state-dependent evidence utility with controlled evidence states rather than focusing on policy behavior or token routing.
- Multimodal agents use accumulated observations to guide retrieval, inspection, or refinement through iterative policy-level decision making.
- Visual-token routing methods rank or route spatial tokens, while diagnostic work reports fragility in static importance estimates and cross-modal alignment.
3 Formulation and Setup
The paper formalizes conditional evidence utility and tests rank reversals in controlled 25-object scenes. Frozen-model evaluation compares oracle-positive and oracle-negative overlap regimes across models, accumulation constructions, wording, and controls.
- 3.1 Problem formulation: Each query identifies one target among N = 25 candidates, with acquired evidence S drawn from color, shape, and texture attributes.
- 3.1 Problem formulation: Target utility is U(S, q) = log p(y∗| S, q), and ∆k(S, q) measures the utility increase from adding evidence type k.
- 3.1 Problem formulation: A rank reversal occurs when the utility difference between two available evidence groups changes sign between the empty state and an acquired state.
- 3.2 Controlled compositional setup: The controlled setup renders 5 × 5 scenes with objects sampled from six colors, six shapes, and four textures, using isolated views that vary one named attribute while preserving spatial correspondence.
- 3.2 Controlled compositional setup: Four candidate-overlap regimes define oracle-positive cases where an initially diagnostic attribute becomes redundant and oracle-negative cases without the predicted reversal.
- 3.2 Controlled compositional setup: Regime specificity is summarized by Γrev = R+ − R−, where a positive value indicates more reversals in oracle-positive than oracle-negative regimes.
- 3.3 Evaluation protocol and controls: Frozen confirmation uses 800 held-out scenes with OpenCLIP and SigLIP, while additive-logit and direct-render accumulation share the same initial state and singleton evidence.
4 Results
Held-out confirmation finds regime-specific, robust state-dependent rank reversals across frozen backbones and evidence constructions, with query-sensitive structure. Exploratory matched-first-action results indicate that reranking after the first acquisition retains positive decision value across evaluator changes.
- State-dependent utility reversals are regime-specific: 72.7–81.0 percentage points: Γrev contrasts were positive across all four backbone/mode cells, with every 95% interval excluding zero.Robust reversals concentrated in oracle-positive regimes rather than uniformly across scenes.
- State-dependent utility reversals are regime-specific: Robust reversals localized to the two attribute comparisons predicted by the symbolic construction, while color-versus-shape reversals remained uncommon.The pattern followed the candidate-overlap manipulation designed to change relative evidence value.
- Robustness to evidence construction and query semantics: Reversal labels, conditional utilities, and preferred next actions agreed strongly between additive and jointly rendered accumulation for both backbones.Direct rendering changed some individual decisions but preserved the same regime-level structure.
- Robustness to evidence construction and query semantics: Every wording template preserved a positive Γrev, while query-specific symbolic-oracle action agreement reached 92.06–94.84% versus 61.44% for the query-independent baseline.The semantic-target confirmation covered 20,670 matched scene-state-target observations and yielded a 30.62–33.40 percentage-point improvement.
- Robustness to evidence construction and query semantics: Query-scene derangement destroyed localization and returned oracle alignment and regime contrast to approximately chance-scale behavior.This control preserved recipient images and target indices, isolating the intended query-scene relation as the relevant structure.
- Exploratory value of replanning: Commit-once and replan shared identical costs, first states, and final all-evidence states, differing only in the second acquisition.Replan reranked the remaining evidence after the shared first action, whereas commit-once followed the empty-state order.
- Exploratory value of replanning: All 12 cross-condition scene-bootstrap intervals remained above zero, with larger gains in oracle-positive than oracle-negative regimes.Cross-condition actions were selected under one evidence mode, wording, or backbone and evaluated under another.
5 Discussion and Future Directions
The discussion argues that evidence importance should be evaluated conditionally and that sequential selectors should update priorities after acquisition. It also bounds the conclusion to a narrow controlled intervention setting and identifies role-balanced and natural-image extensions.
- Conditional evaluation: A single evidence ordering cannot express state-specific reorderings, so grounded evaluation should report conditional utility changes across matched acquisition states.Suggested measures include conditional marginal-utility curves, rank-stability statistics, and reversal rates localized to predicted interactions.
- Conditional evaluation: Aggregate reversal counts are insufficient because evaluation should test whether ordering changes occur where the evidence structure predicts an interaction.The controlled positive/negative regimes provide this localization test.
- Adaptive evidence selection: Sequential selectors can be evaluated on whether they update the next choice after the evidence state changes, using matched-first-action comparisons against commit-once policies.The paper proposes learning a low-cost predictor of ∆k(S, q) or the preferred next evidence type.
- Adaptive evidence selection: Cross-condition evaluations preserve scenes and targets, measuring robustness across evaluators within this setup rather than distributional generalization.The replanning signal is presented as a falsifiable target for future routing or active-perception methods.
- Scope and extensions: Both oracle-positive regimes use texture in the reversing comparison, and isolated texture is decoded less accurately than color or shape.Role-balanced constructions are needed to test invariance to permutations of attribute roles.
- Scope and extensions: Color, shape, and texture are controlled input interventions, so the results characterize behavioral evidence use without implying internal processing order, separable concepts, or executable compute units.Extensions include richer natural-image evidence, independently executable evidence groups, and prospective learned-selector tests under real costs.
6 Conclusion
In this controlled setup, visual evidence importance depends on both the acquired evidence state and the semantic target query. Updating the remaining evidence order retains decision value across evaluator changes.
- Robust rank reversals follow candidate-overlap regimes, persist across evidence constructions and query wordings, and disappear when the query-scene relation is broken.
- Preferred next evidence can change with the semantic target query even when the visual scene is fixed.
- Post-confirmation matched-first-action analysis finds decision value from reranking remaining evidence across evidence-mode, wording, and backbone evaluator changes.
- Conditional evidence utility provides a measurable target for evaluating evidence changes after acquisition and for developing future adaptive evidence-selection methods.
A Setup construction and reproducibility
The confirmation uses frozen, reproducible scene and query configurations, while two evidence modes share the same candidate-state lattice and scoring basis.
- The held-out confirmation contains 200 scenes in each of four regimes, with unique target conjunctions and ten equivalent query templates.T1 is primary, while T2–T10 provide repeated wording measurements; metadata, controls, configurations, and hashes were frozen before inference.
- Both evidence modes crop and independently encode 25 spatially corresponding candidate cells, encode the text query once, and score candidates by image-text similarity logits.Additive logits sum separately encoded singleton views, and both modes use the same all-zero empty state.
B Robustness and controls
Robust reversals are concentrated in the pairings predicted by the symbolic construction, while derangement and selectivity controls test whether the pattern depends on the intended query-scene and attribute structure.
- 46.3%, 56.0%, 59.2%, and 59.6% are sample-level primary reversal rates for OpenCLIP-additive, OpenCLIP-direct, SigLIP-additive, and SigLIP-direct.Pair-level rates are 29.8%, 36.0%, 36.7%, and 41.0%, reported as scale summaries rather than the primary diagnostic.
- The oracle predicts no color-shape reversals, but predicts both texture-containing comparisons in the 400 oracle-positive scenes; model reversals match this localization.The table localizes the effect to texture-containing comparisons and does not establish attribute-role invariance.
- Deranged positive-minus-negative reversal contrasts are 0.5, 5.25, −2.25, and 2.0 percentage points across the four backbone-mode combinations.Deranged oracle balanced accuracy also falls substantially relative to normally matched queries.
- Both backbones score 100% on isolated color and shape and 74.86% on isolated texture, while no removed-attribute cell triggers the prespecified strong-leakage flag.The diagnostic is a one-sided screen for detectable strong leakage, not an equivalence test for chance-level decoding.
C Prior frozen semantic-target confirmation
A separate prior frozen confirmation tests whether the preferred next evidence depends on the semantic target while holding scenes fixed. It evaluates matched scene-state units where the symbolic oracle requires different actions.
- The prior confirmation used 400 scenes and 20,670 matched observations and was specified prospectively as criterion P3 in its own frozen protocol.It is separate from the later 800-scene confirmation reported as the main result.
- Evaluation includes only matched scene-state units with at least two semantic target queries requiring different non-tied oracle-optimal actions.Two-attribute states with only one available action are excluded.
- Table 4 compares query-specific agreement with a query-independent modal-oracle baseline, reporting improvements in percentage points with scene-bootstrap intervals.
- 85.6–90.4% model switch rates and 82.0–87.1% agreement with oracle switch direction show target-dependent changes in preferred next evidence.The comparison directly tests semantic-target dependence rather than paraphrase robustness.
D Planned adaptive comparison and its negative interaction
The planned adaptive-versus-global-fixed comparison found positive overall adaptive gains, but its interaction test was negative, so the result does not isolate replanning after acquisition.
- 0.562 [0.504, 0.621], 0.447 [0.394, 0.501], 0.441 [0.400, 0.482], and 0.447 [0.401, 0.491] were the adaptive-minus-best-fixed AUC values across four backbone/mode cells.The best global fixed order by mean target-log-probability AUC was color-texture-shape in all four cells.
- −0.075, −0.072, 0.003, and −0.002 were the planned mechanism-relevant interaction values across the four cells.The interaction tested whether the adaptive advantage was larger in oracle-positive regimes.
- The global order comparison did not isolate the value of updating after acquisition because a global order cannot adapt to information already present at the empty state.
E Full exploratory replanning transfers
The post-confirmation exploratory analysis evaluates whether reranking after an initial acquisition transfers across selection and evaluation conditions. Table 5 reports cross-condition step-2 gains, while same-condition results provide upper-bound context.
- E Full exploratory replanning transfers: The analysis was specified after confirmation results were inspected and was not part of the frozen confirmation protocol.It uses saved transitions, 10,000 bootstrap repeats with seed 20280827, and scenes as clusters.
- E Full exploratory replanning transfers: Table 5 reports exploratory cross-condition step-2 target-log-probability gain for replan minus commit-once.The table distinguishes OpenCLIP and SigLIP and the additive and direct-render evidence-accumulation modes.
- E Full exploratory replanning transfers: Cross-condition evaluations select an order under one condition and evaluate it under another, so positive step-2 utility is not guaranteed by construction.This contrasts with same-condition evaluations, whose opportunity gains are upper bounds because replanning directly observes counterfactual evaluator utilities.
- E Full exploratory replanning transfers: Same-condition opportunity gains were 0.698 [0.650, 0.749], 0.752 [0.704, 0.800], 0.727 [0.682, 0.772], and 0.745 [0.700, 0.790].The corresponding robust on-policy switch rates were 42.5%, 46.6%, 52.9%, and 53.5%.