Source-linked AI summary
Visual Persuasion: What Influences Decisions of Vision-Language Models?
Manuel Cherep, Pranav M R, Pattie Maes, Nikhil Singh
TL;DR
VLMs increasingly make consequential visual choices, yet existing accuracy-focused evaluations provide limited evidence about the visual preferences driving those decisions. The paper treats pairwise choices between systematically edited images as revealed evidence of latent visual utility, using visual prompt optimization, interpretability, and normalization. Its experiments show that naturalistic optimized edits can substantially shift VLM choices and that normalization only partially mitigates these sensitivities.
Problem
Accuracy benchmarks do not characterize the visual sensitivities shaping VLMs’ preference-based decisions, while exhaustive image comparisons are costly and provide no coverage guarantee.
Method
The framework iteratively proposes prompt-mediated, identity-preserving image edits, evaluates them through pairwise VLM choices, and interprets resulting visual changes with automated thematic analysis.
Results
Naturalistic optimized edits substantially shift VLM selection probabilities and reveal recurring visual themes influencing judgments across large-scale choice experiments.
Takeaways & Limitations
Visual prompt optimization provides a methodology for mapping implicit value functions and surfacing visual vulnerabilities in vision-based AI systems before exploitation at scale.
Takeaways & Limitations
The optimization requires substantial computational resources, and the boundary between presentation and substance remains fuzzy because identity-preserving edits can change attributes or minor amenities.
Abstract
from arXiv · showhide
The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (VLMs). These agents make visual decisions at scale, deciding what to click, recommend, or buy. Yet, we know little about the structure of their visual preferences. We introduce a framework for studying this by placing VLMs in controlled image-based choice tasks and systematically perturbing their inputs. Our key idea is to treat the agent's decision function as a latent visual utility that can be inferred through revealed preference: choices between systematically edited images. Starting from common images, such as product photos, we propose methods for visual prompt optimization, adapting text optimization methods to iteratively propose and apply visually plausible modifications using an image generation model (such as in composition, lighting, or background). We then evaluate which edits increase selection probability. Through large-scale experiments on frontier VLMs, we demonstrate that optimized edits significantly shift choice probabilities in head-to-head comparisons. We develop an automatic interpretability pipeline to explain these preferences, identifying consistent visual themes that drive selection. We argue that this approach offers a practical and efficient way to surface visual vulnerabilities, safety concerns that might otherwise be discovered implicitly in the wild, supporting more proactive auditing and governance of image-based AI agents.
1. Introduction
VLMs make preference-based visual decisions at scale, but accuracy evaluations and brute-force image comparisons do not reveal which visual features drive those choices. The paper introduces visual prompt optimization and related analyses to surface, interpret, and mitigate these sensitivities.
- Motivation: VLM agents increasingly make consequential visual decisions under an assumption that their visual values align with people’s interests.Violations of this assumption could compound across automated platforms and shift visual culture toward agent preferences.
- Motivation: Accuracy benchmarks capture task competence but not the contextual sensitivities that shape behavioral outcomes in visual decisions.The paper argues that behavioral systems require behavioral tests.
- Motivation: Naive exhaustive pairwise testing is expensive, slow, and offers no guarantee that naturally varying images probe the visual features that matter.The possible feature space is vast, making coverage difficult.
- Approach: Visual prompt optimization iteratively edits a candidate image with a text-to-image model while preserving its semantic content, revealing visual factors that shift agentic choices.The edits are intended to be naturalistic rather than adversarial and keep the main element recognizable.
- Findings: Visual normalization partially mitigates vulnerabilities but does not eliminate them, raising concerns about VLM robustness in real-world decisions.The paper presents normalization as a promising but incomplete path toward mitigation.
2. Related Work
Prior work largely evaluates model competence or isolated visual capabilities, while emerging behavioral evaluations study inputs that systematically influence decisions. This paper builds on prompt optimization, controllable image generation, iterative preference elicitation, and auto-interpretability to study visual sensitivities in agent-like choices.
- Behavioral evaluation: Functional benchmarks assess task competence, whereas behavioral evaluation examines model behaviors and the inputs driving their successes, failures, and consequences.The related-work discussion positions this paper within the behavioral-evaluation perspective.
- Behavioral evaluation: Behavioral evaluations have exposed systematic influences from psychological nudges and item attributes such as prices and ratings in language-model agents.These sensitivities can distort behavior relative to assumed rationality.
- VLM diagnostics: Prior VLM studies examine shape versus texture, numerosity, feature binding, and visual attribute reliance, while this work focuses on agent-like decisions.The paper inherits the diagnostic lens but applies it to visual properties of inputs in decision tasks.
- Prompt optimization: TextGrad, Feedback Descent, and GEPA optimize textual artifacts for arbitrary objectives, motivating analogous feedback-driven optimization of visual prompts.Recent controllable image models make visually precise prompt-mediated edits feasible.
- Interpretability: Auto-interpretability repurposes language-model interpretive abilities to organize arbitrary artifacts, supporting analysis of optimized and original images.The paper uses this literature to interpret the visual changes that influence decisions.
3. Methods
The method treats a VLM’s pairwise choices as evidence about a latent visual utility and searches over identity-preserving, prompt-mediated image edits. It combines iterative proposal-and-evaluation procedures with interpretability and normalization analyses to identify and mitigate influential visual factors.
- Datasets and tasks: The study evaluates product purchasing, house searching, job candidate screening, and hotel scouting using four datasets of images and task-specific VLM critics.Each dataset contributes 100 images to the optimization and evaluation procedures.
- Visual prompt optimization: Visual prompt optimization searches over editable text prompts that generate naturalistic edits while preserving the underlying object or scene.The framework optimizes prompts rather than image pixels directly.
- Visual prompt optimization: A task instruction and VLM evaluator define pairwise preferences, with optimization seeking edited images whose utility exceeds the baseline.Allowed edits can change background, lighting, framing, color palette, contextual props, and other mutable visual elements.
- Preference-based evaluation: The preference model interprets repeated pairwise choices as an empirical relation and accepts a proposed edit when it wins against the incumbent.Under the logistic utility view, increasing choice probability corresponds to increasing the utility gap.
- Optimization loop: The optimization loop alternates between proposing prompt updates and evaluating edited images, using repeated comparisons, order counterbalancing, and conservative acceptance to reduce false improvements.The evaluator is treated as noisy and potentially order-sensitive.
- Identity constraints: Identity maintenance constrains every generated image to depict the same underlying entity or scene, though post-hoc checks, similarity thresholds, or editing instructions may implement the constraint.Without this constraint, changing the underlying object could confound interpretation of visual preferences.
- Specific methods and analyses: CVPO uses competitive selection between two candidates evaluated by a panel of judges, while the interpretability pipeline hierarchically clusters visual differences into recurring themes.Normalization separately equalizes task-irrelevant properties before VLM judgment to test mitigation.
4. Results
Across four domains, visual edits substantially shifted VLM and human choices, with zero-shot edits producing large gains and optimization often adding further improvements. CVPO generally performed best among optimization methods, while mitigation reduced but did not eliminate sensitivity.
- 1.8M+ API requests, 2.75B+ tokens, and 125k+ images supported the large-scale evaluation.
- Across all four domains, zero-shot edits and subsequent optimization shifted VLM choice probabilities upward relative to original images.Zero-shot edits already produced large gains, while later optimization added gains varying by domain and method.
- 0.2–0.4 was the approximate zero-shot increase in selection probability over original images, with all contrasts statistically significant.In several settings, zero-shot edits more than doubled selection probability in head-to-head comparisons.
- CVPO and VFD often added +0.1–0.3 absolute choice probability beyond zero-shot editing, whereas VTG sometimes added little or nothing.The additional gain depended on optimization method and domain, and was statistically significant across most settings for CVPO and VFD.
- CVPO won most often in final head-to-head comparisons, outperforming VFD by 0.04–0.21 on 7/9 models and VTG by 0.46–0.64.CVPO’s advantage over VFD was modest on average and heterogeneous by task and model; Anthropic models were a partial exception.
- CVPO averaged 17.4 iterations, compared with 24.9 for VFD and 30 for VTG, but generated more images per iteration depending on κ.Iteration-based efficiency therefore does not directly measure total image-generation cost.
- Humans substantially preferred optimized images over originals, while human head-to-head preferences favored CVPO on average but lacked statistically significant post-hoc contrasts.Human CVPO-versus-zero-shot improvement was not statistically significant in the status comparison (p=0.057).
- Different optimization methods converged on broadly similar visual themes within tasks, including botanical integrations, luxury upgrades, warm lighting, and lifestyle product settings.The overlap suggests the discovered edits may reflect stable properties of evaluator VLMs and perhaps human decision-makers.
5. Limitations
The framework has practical scalability, validity, and coverage limits, while its normalization mitigation remains partial and human evidence is directional.
- Substantial computational resources currently limit the scalability of the optimization framework.
- The boundary between presentation and substance can be fuzzy because identity preservation allows changes to attributes and minor amenities.The paper defines identity maintenance as preserving the underlying asset while permitting some malleable visual features.
- Human validation studies (N=154) provide directional evidence but lack statistical power to detect small effect-size differences in head-to-head comparisons.
- The prompt distillation experiment’s use of the same images as optimization limits its external validity.
6. Conclusion
VLMs exhibit strong, structured visual sensitivities, challenging the assumption that human-like performance implies sufficient robustness for delegation.
- VLMs exhibit strong, structured visual sensitivities that make human-derived priors about visual decision-making potentially costly to carry forward uncritically.
Impact Statement
The paper shows that naturalistic presentation edits can substantially shift VLM choices, while also offering tools to interpret and mitigate these sensitivities. These findings support auditing and governance beyond accuracy benchmarks, but indicate that normalization only partially reduces vulnerabilities.
- Naturalistic presentation changes can substantially shift VLM choices even when the underlying object or scene remains fixed.
- The framework provides controlled measurements of VLM sensitivities for auditing, debugging, and evaluation beyond accuracy-oriented benchmarks.
- The same procedures could manipulate marketplace or other high-stakes image-based decisions, creating fairness concerns when presentation changes advantage items without changing substantive qualities.
- Visual normalization partially mitigates vulnerabilities but does not eliminate them, motivating stronger context normalization and checks for irrelevant decision-shifting cues.
- Interpretability surfaces recurring themes such as luxury furniture, warm lighting, botanical additions, professional attire, corporate settings, and lifestyle staging.
- Systematic visual edits span lighting, backgrounds, staging, props, human presence, framing, and product presentation while preserving core visual content.
C.4.3. VTG
VTG-derived visual edits consistently reshape images through contextual staging, lighting, subject changes, and removal of packaging or isolation while preserving the core subject or structure.
- Lifestyle and textured backgrounds move products from isolated white settings into furnished interiors or realistic environments.
- Decorative props such as textiles, greenery, coffee cups, jewelry, and laptops add contextual staging around products.
- Warm, atmospheric lighting uses cinematic tones, sunset illumination, window shadows, and bokeh effects.
- Human interaction edits add models or hands to communicate product scale and usage.
- Packaging and studio isolation are removed to produce more direct, situated product views.
- Zero-shot editing prompts apply these discovered visual themes to hotels, houses, job candidates, and products while keeping the underlying subject or structure unchanged.The prompts target context, lighting, landscaping, staging, clothing, accessories, and participants, with identity or structural preservation constraints.
E.1. Out-of-sample Distillation
Out-of-sample distillation tests whether visual preferences discovered through optimization transfer to held-out images and narrower product categories. The results support category-specific distillation and show that CVPO remains superior to VFD with fewer proposals, although the gap narrows.
- E.1. Out-of-sample Distillation: For 3/4 tasks, distilled images perform significantly better than zero-shot edits on 240 held-out images.The out-of-sample experiment used 20 new base images per task across four tasks.
- E.1. Out-of-sample Distillation: Category-level distillation supports the hypothesis that category heterogeneity drives interpretability performance.
- E.1. Out-of-sample Distillation: Narrowing distillation by product category improves performance rather than acting as a distractor.
- E.1. Out-of-sample Distillation: With k = 1, CVPO still beats VFD in head-to-head comparisons, but the difference becomes smaller.CVPO used fewer image-generation calls than VFD on average, with 17.9 rounds in this setting.
- E.1. Out-of-sample Distillation: Figures 8 and 9 report comparisons against zero-shot editing and CVPO under 50% base-rate conditions, with Figure 9 using one candidate.
G. Verifying the Identity Maintenance Constraint
The identity-maintenance evaluation pairs original and final images for human matching while modeling image-level choices with clustered linear probability models. Most pairs were matched reliably, but errors concentrated in visually ambiguous people and frying-pan cases, and longer human deliberation improved accuracy.
- G. Verifying the Identity Maintenance Constraint: Participants matched 400 original/final image pairs, with each pair receiving two independent matchings.Each participant evaluated ten original/final pairs per task to keep cognitive load manageable.
- G. Verifying the Identity Maintenance Constraint: 11 of 400 pairs, approximately 2.75%, were incorrectly matched by both raters.Ten errors came from the people dataset and one from products; the paper attributes these cases to pre-existing visual similarity or artifacts as unlikely optimization effects.
- G. Verifying the Identity Maintenance Constraint: Each additional minute of task duration increased accuracy odds by approximately 11.2%.The 95% confidence interval was [6.1%, 16.5%].
- G. Verifying the Identity Maintenance Constraint: Model-predicted accuracy increased from 77.8% to 93.1% across the observed 5.22-to-17.90-minute duration range.The predicted average accuracy was 85.5%, and excluding one outlier raised overall accuracy to 89.1%.
- G. Verifying the Identity Maintenance Constraint: The analysis reshapes binary image comparisons into image-level observations and estimates linear probability models with cluster-robust errors.Clustering is by participant and pair ID in human studies, and by pair ID otherwise.
- G. Verifying the Identity Maintenance Constraint: Model specifications include optimization status, strategy, model, image class, task, and mitigation-pass effects with selected interactions.The optimization-status variable distinguishes original, zero-shot, final optimized, distilled, and inconsistent outcomes depending on the experiment.
H.3. Human Experiments
The human experiments recruited 154 participants across three experiments comparing original, edited, optimized, and mitigated images.
- 154 participants took part across three human experiments.The experiments included 64 participants in the first, 50 in the second, and 40 in the third.
- The first experiment collected 1,920 judgment observations from 64 participants making 30 comparisons each.
- The second experiment collected 1,600 observations from 50 participants making 32 head-to-head comparisons each.
I. Effects of Mitigation on Visual Similarity
The mitigation procedure altered visual similarity in three reported ways: aligning images within choice pairs, moving optimized images toward their originals, and transferring properties across paired images.
- Mitigation steps more closely aligned perceptual features of images within a choice pair.Similarity was evaluated with multiple metrics, including CLIP-based measures and SSIM.
- Mitigation steps moved optimized images closer to their original states.
- Mitigation also modified original choice-set images by applying more properties from the other image’s optimized version.
J. Disaggregated Empirical Results
The appendix decomposes the main pairwise experiments across evaluator models, tasks, image classes, and optimization strategies, using the same underlying evaluation protocol.
- Evaluation protocol: The appendix figures use pairwise forced-choice judgments with randomized order, exclusions for inconsistent responses, and linear probability models with estimated marginal means.
- Model and task heterogeneity: The disaggregated analyses examine heterogeneity by evaluator model, task domain, image class, and optimization strategy.Head-to-head comparisons between VTG, VFD, and CVPO are broken out by model and task under identical procedures.
- Model-stratified choice probabilities: Choice probabilities for original, zero-shot edited, and final optimized images are stratified jointly by evaluator model and strategy for each task domain.The plots report estimated marginal means with 95% confidence intervals.
- Image-class heterogeneity: Further decompositions report choice probabilities by image class for hotel image types and product categories, averaged across evaluators.These results also include 95% confidence intervals and remain descriptive decompositions of the main comparisons.
- Human results: Human mitigation results are disaggregated by task, with additional tables detailing human and VLM choice probabilities by strategy, status, and task.