Source-linked AI summary

Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls

Manuel Cherep, Pattie Maes, Nikhil Singh

arXiv:2608.27727v1cs.AI

TL;DR

Models’ perceptual priors remain difficult to measure because fixed stimuli leave high-dimensional input spaces largely unexplored, despite priors influencing behavior and consequential decisions. The paper combines interpretable generative controls with Gibbs sampling and binary model judgments to recover these priors directly. Across four image domains and four VLMs, it recovers canonical biases and novel priors that direct prompting fails to surface.

  • Problem

    Models’ perceptual priors are poorly understood because fixed stimulus sets leave most high-dimensional input spaces unseen, although these priors influence behavior and safety-relevant decisions.

  • Method

    The method steers a controllable, interpretable image generator and runs Gibbs sampling over its axes, using the VLM’s binary preferences as Barker acceptance steps.

  • Results

    Across four image domains and four frontier VLMs, the method recovers full perceptual-prior posteriors, including canonical biases and novel priors invisible to direct prompting.

  • Takeaways & Limitations

    The results support using models’ own pairwise choices, rather than introspective reports alone, to characterize priors relevant to responsible design and governance.

  • Takeaways & Limitations

    Recovered priors are limited by what SliderSpace and FLUX can express, and the method does not attribute preferences to pre-training, alignment, or in-context confounds.

Abstract

from arXiv · show

A model's behavior on a task is jointly determined by the input it receives and the prior it brings in, i.e. the distribution over stimuli it implicitly expects. Interpretability research has traditionally studied models by holding inputs fixed and examining model responses either mechanistically, probing how internal structure represents inputs, or behaviorally, measuring how variation in inputs leads to variation in outputs. Neither reconstructs the prior distribution itself, since internal structure shows what a model can represent, not what it expects, and any fixed stimulus set leaves most of the possible input space unseen. In particular, such an input space in real-world settings, such as images seen by VLMs, is extremely high-dimensional and diverse. These priors thus remain a poorly understood component of models that nonetheless influence real-world behavior. We propose a method to sample from models' perceptual prior distributions directly, by steering a generative model to produce stimuli along controllable axes and running Gibbs sampling over that space with the model under study as the judge. We apply this to a variety of categories and target variables (such as trustworthiness in faces and cheapness in art images) and recover both canonical biases and surprising novel priors invisible to direct prompting, warranting further investigation of their downstream effects.

1 Introduction

The paper introduces a method for directly recovering multimodal models’ perceptual priors in high-dimensional stimulus spaces, where fixed evaluations are inadequate. It combines controllable image generation with Gibbs sampling and applies the framework across four domains.

  • Model priors shape perception, preferences, and behavior, making their alignment with human priors important for safety and alignment.
  • Fixed stimulus evaluations become impractical in high-dimensional spaces because covering controllable dimensions requires combinatorially many stimuli.
  • The proposed method uses Gibbs sampling over interpretable generative dimensions, with the model under study judging candidate stimuli.
  • The framework is modality-agnostic: any controllable, interpretable generator paired with a binary preference judge can use the same scheme.
  • Experiments span faces, affordances, aesthetics, and authenticity using four frontier VLMs, recovering full posteriors that direct prompting fails to surface.

2 Related Work

The paper adapts human perceptual-prior sampling methods to multimodal language models. It combines Gibbs dimension-wise updates with Barker-rule binary acceptance over controllable multimodal generative spaces.

  • Human perceptual-prior methods evolved from adaptive Metropolis-Hastings sampling with binary choices to Gibbs sampling with dimension-wise updates.
  • The paper transfers this sampling family from people to multimodal language models, using categorical responses rather than continuous slider responses.
  • The resulting sampler is Metropolis-Hastings within Gibbs, with the MLLM serving as the acceptance function.
  • Unlike prior model-prior studies over text or abstract numerical quantities, this work probes priors in a controllable multimodal generative space.
  • Interpretable axes are necessary because they connect posterior mass in latent space to meaningful concepts, unlike some entangled latent directions.

3 Methods

The method samples interpretable image-generating axes one coordinate at a time while using binary VLM judgments as Barker acceptance steps. It addresses costly, unstable direct slider selection through Metropolis-Hastings within Gibbs and relies on approximately orthogonal controls.

  • Directly asking a VLM to select slider positions is impractical because image generation is costly and turn-based interaction can cause oscillation and arbitrarily small moves.
  • Metropolis-Hastings within Gibbs replaces direct slider selection with coordinate-wise proposals and binary VLM acceptance decisions.
  • The setup uses a controllable text-to-image generator whose latent space is a product of interpretable, mutually orthogonal slider axes.
  • The implementation uses FLUX.1-schnell with SliderSpace, training n = 64 candidate sliders per domain and selecting interpretable subsets.
  • The VLM receives two generated stimuli under a textual criterion and its binary choice determines whether the proposed state is accepted.
  • Validity requires the Bradley-Terry assumption only along individual slider axes, a weaker condition supported by SliderSpace’s approximate semantic orthogonality.
  • Each Gibbs scan updates all d coordinates and, with independent uniform proposals, costs 2d image generations and d VLM queries before generation reuse reduces the cost.

4 Results

Across four domains and four VLM judges, Gibbs sampling recovered interpretable, target-dependent perceptual priors, including demographic associations and domain-specific consensus. The results also show model-specific extremes, varying cross-model agreement, selective versus diffuse targets, and differences from explicit elicitation baselines.

  • Strongest recovered priors: Eyeglasses was the strongest slider for Intelligent, Hardworking, and Serious faces, while Smiling led for Fun and Trustworthy.Other strongest associations included authenticity edits, Paper/Material for Rough, and Ornate Detail for Beautiful.
  • Strongest and weakest recovered priors: Below-baseline associations included Smiling for Serious and Youthful, Child for Intelligent, Sfumato and Low Key/Dark for Beautiful, and Textured/Dusty for Resonant.Figure 3 surfaces the most negative associations relative to the 0.5 no-preference baseline.
  • Demographic associations: All four judges placed East Asian and Masculine/Receding Hairline below baseline for Attractive, Trustworthy, Intelligent, and Hardworking, but these associations were target-dependent.For example, East Asian was above baseline for Youthful for 3/4 judges, while Feminine often showed the inverse pattern.
  • Top & bottom sliders per model: Gemini 3 Flash produced the strongest per-model extremes, whereas Claude Sonnet 4.6 was comparatively less extreme and often out of step with the others.Gemini assigned 0.87 to Masculine/Receding Hairline for Criminal and 0.16 to East Asian for Attractive; Claude assigned 0.28 to Glasses for Trustworthy.
  • Cross-model rank agreement: Cross-model rank agreement was highest for Authenticity (median ρ ∈[0.42, 0.69]), intermediate for Faces and Affordances ([0.30, 0.56]), and lowest for Aesthetics ([0.07, 0.30]).Sharp visual signatures yielded more consensus than diffuse concepts such as cheap or experimental aesthetics.
  • Selectivity: Selectivity was high for several face targets but low for Authentic (0.05) and Cheap (0.12), indicating diffuse priors for those targets.Face selectivity included Intelligent/Eyeglasses at 0.33 and Serious/Eyeglasses at 0.32.
  • Explicit-prior baselines: Explicit scalar-rating and forced-choice baselines recovered more uniform-like priors, while their disagreement with IMH correlated with inter-slider correlations ignored by explicit methods.IMH samples the joint slider distribution, allowing preferences for feature combinations beyond independent marginal changes.

5 Limitations

The method’s expressiveness is bounded by what SliderSpace can learn and FLUX can express, and recovered preferences are not attributed to particular causes without downstream analysis.

  • Recovered priors span only concepts represented in SliderSpace and expressible by FLUX; absent or poorly represented concepts remain invisible.
  • The method exposes biases hidden by direct prompting but does not determine whether they arise from pre-training, alignment, or in-context confounds.

6 Conclusion

The paper presents Gibbs sampling with model pairwise preferences as a way to recover full perceptual-prior posteriors in models’ native modalities. Across image domains, it provides an auditing tool while warning that surfaced priors could also be exploited.

  • The method uses the model’s pairwise preferences as Gibbs acceptance decisions over inputs in its native perceptual modality, recovering full posteriors.
  • Surfacing implicit priors can support model auditing but may also make those priors legible to actors who could amplify or exploit them.
  • Qualitative examples show recovered priors across faces, affordances, aesthetics, and authenticity, with each column representing a target–model pair.
  • Slider sweeps vary one of 10 selected controls while holding other sliders and the seed fixed, producing coherent monotone semantic changes.

C Comparison to Explicit Priors

The comparison evaluates adaptive joint-prior sampling against explicit one-dimensional slider baselines. IMH can recover dependence among sliders, whereas explicit procedures use fixed sweeps and may disagree in posterior magnitudes or rankings.

  • IMH recovers a joint implicit prior from adaptive pairwise comparisons, while explicit baselines query models on fixed slider sweeps.
  • Explicit baselines: Explicit sweeps vary one slider while fixing the other nine values and generation seed, isolating each control’s visible effect.
  • Explicit baselines: The rating baseline converts 1-to-10 single-image scores into weights, producing one-dimensional posterior summaries for each target–slider pair.
  • Explicit baselines: The Bradley–Terry baseline fits linear latent utility to pairwise choices and derives posterior medians and intervals from slider-value weights.
  • Comparison metrics: Explicit and IMH posteriors are compared by median differences, centered profiles, and within-target slider rank displacement.
  • Dependence and informativeness: IMH estimates joint slider correlations, testing whether stronger interactions coincide with larger disagreement from explicit methods.
  • Dependence and informativeness: KL divergence from a uniform slider prior measures posterior concentration for IMH and explicit methods using their respective sampled representations.

D Auto-Interpretability Pipeline

The auto-interpretability pipeline assigns unique semantic labels to generative sliders using multimodal embedding evidence, endpoint contrasts, chain weighting, global assignment, and language-model refinement.

  • Slider trajectories are matched against 867 visual-descriptor candidates embedded with sampled images in a shared Gemini Embedding 2 space.
  • Images are sampled across five chains per slider, with the focal slider varied while other sliders remain at zero and images sorted by focal value.
  • Label scores use residualized image–text similarities and endpoint contrast, rewarding similarity increases from the low to high slider ends while zeroing negative contrasts.
  • Chains receive weights based on coherent image-embedding movement, so visually salient slider sweeps contribute more to final label scores.
  • A Hungarian assignment selects one label per slider while using each canonical label at most once within a domain.
  • Gemini 3.1 Flash Lite refines assigned and runner-up labels using scores, visual plausibility, and semantic uniqueness across each domain.
  • Table 1 lists the top candidate labels and post-hoc refined labels used to interpret sliders.

E Additional Diagnostics

Additional diagnostics test whether sampled posterior marginals predict later accept/reject decisions within and across chains, while convergence analysis compares proposal strategies and model-domain mixing.

  • Predictive diagnostics: Within-chain checks evaluate whether posterior density fitted on training folds predicts held-out accept/reject decisions from the same chain.High accuracy indicates alignment between the chain’s posterior mass and the model’s later choices.
  • Predictive diagnostics: Cross-chain checks fit density on all but one chain to test whether different random seeds recover compatible posterior structure.Results are reported in Figure 18.
  • Convergence: GPT-5.4 and Claude Sonnet 4.6 produce the most-converged chains across all four domains, with median split-R̂ between 1.02 and 1.20.Qwen3-VL 235B and Gemini 3 Flash are systematically harder to mix, reaching median R̂ = 1.6 on Aesthetics.
  • Convergence: Affordances mixes comparably well for all models, with Qwen3-VL 235B and Gemini 3 Flash near median R̂ = 1.15.Their chains mix more poorly in Faces and Aesthetics.
  • Proposal comparison: The random-walk variant underperforms the uniform proposal, with median R̂ frequently exceeding 1.5.Its scale often collapses to σ ≲ 0.1, while acceptance plateaus at 0.05–0.15 and chains barely move.

G Validation with Color Experiment

The color experiment validates the Gibbs+Barker kernel in a low-dimensional HSL space with unambiguous ground truth, recovering canonical regions for all tested color names.

  • Experimental setup: The validation uses HSL coordinates, where each state directly determines a flat color tile without an intervening generative model.Ten independent chains of 2,000 trials are run with Qwen3-VL 235B.
  • Experimental setup: Eight named-color targets provide a comparison against established human data, including sunset, eggshell, lavender, chocolate, lemon, cloud, strawberry, and grass.The targets span the same color-name set used by Harrison et al. and Zhu et al.
  • Results: Every target’s posterior central tendency falls within its canonical color region.Examples include saturated yellow for lemon, green for grass, light purple for lavender, and near-white high-lightness regions for cloud and eggshell.
  • Results: The kernel recovers the correct color-space region for every target using only binary preferences.Posterior density overlays and rendered samples are shown in Figure 19.

H Compute Resources

SliderSpace training and sampling rely on substantial GPU infrastructure, with approximately nine hours of preparation per domain and parallel FLUX inference.

  • Training: Per-domain data generation, PCA, and slider training take approximately 9 hours on a single NVIDIA H200 GPU.Data generation and PCA take about 3 hours, followed by approximately 6 hours of slider training.
  • Inference: Sampling uses 40 parallel Amazon SageMaker endpoints, each backed by one NVIDIA L40S GPU.Each endpoint runs the compiled FLUX.1-schnell pipeline at approximately 0.7 seconds per image.

I Posterior Summaries

This section reports posterior summaries across the primary image-domain experiments, alongside the fixed prompts and visual formats used to constrain stimulus variation.

  • Posterior summaries: Tables 5–8 report posterior medians with 50% credible intervals across the primary IMH experiments.The tables cover Affordances, Art, Authenticity, and Faces.
  • Prompt design: FLUX prompts are fixed within each domain so slider perturbations remain the dominant source of stimulus variation.The visual formats include person id photos, isolated decorative objects, paintings, and outdoor city photographs.
  • Prompt design: Task prompts ask models to choose which stimulus better matches a target property, adjective, description, or authenticity judgment.The four prompt templates correspond to Faces, Affordances, Aesthetics, and Authenticity.

K Slider Validation

Slider validation shows that the recovered posteriors are constrained by which image regions the trained generative sliders can reach. Most demographic targets were reachable, but some remained poorly accessible, and per-target analyses revealed distinct slider structures across domains.

  • Reachability: Positive reachability does not imply equal accessibility, since demographic configurations dominant in FLUX’s training data may be easier to reach.The validation establishes reachability, not uniform sampling difficulty across reachable regions.
  • Reachability: Eight of ten demographic descriptors achieved perfect agreement, while hispanic person reached 20% and elderly person 0%.The reachability check used separate chains targeting demographic descriptors and judged accepted images for descriptor matches.
  • Reachability: Unreachable concepts are systematically absent from recovered posteriors, regardless of whether the model holds the corresponding association.The authors therefore treat slider reachability as a fundamental constraint on interpreting recovered priors.
  • Per-target slider structure: Face targets often have one dominant slider plus secondary modifiers, whereas other targets show weaker dominance and more cross-cutting cues.The top-3 and bottom-3 slider views expose structure that pooled analyses average over.
  • Per-target slider structure: The two authenticity targets are mostly antipodal, with Manipulated associated with Black, White/Crushed Blacks, and Upscaling Noise, and Authentic with Natural Edit and Sepia.This comparison comes from the per-target top-3 and bottom-3 slider analyses across models.

M Slider Training

The slider-training setup combines controllable image-generation axes with prompt design intended to isolate relevant features while preserving other properties. Training materials span faces, affordances, aesthetics, and authenticity, and slider rankings are summarized across models.

  • Training setup: SliderSpace trains 64 candidate sliders per domain over FLUX.1-schnell using 50,000 CLIP-encoded samples for PCA.Each slider is trained for 3000 iterations with LoRA-based settings and specified inference parameters.
  • Training setup: Diverse prompts isolate target features while keeping other attributes invariant, such as face variation against a fixed camera-facing background.This prompt strategy is used to learn diverse but interpretable sliders.
  • Training domains: Face training prompts vary demographics, age, gender, skin tone, hair, and head coverings under standardized identification-photo conditions.The prompt set includes descriptors such as teenage, elderly, Indigenous, East Asian, dark-skinned, and hijab-wearing people.
  • Training domains: Authenticity prompts use a fixed Victorian alley scene, while Figure 20 and Figure 21 summarize each target’s top-3 and bottom-3 sliders across models.The figures report cross-model means and interquartile ranges; the fixed scene is expanded from a shared placeholder.
  • Training domains: Affordance prompts cover material, shape, size, transparency, texture, and object-function-relevant properties.Examples include metal, wood, glass, plastic, leather, cylindrical, pointed, hinged, padded, and spouted objects.
  • Training domains: Aesthetic prompts span historical styles, composition and lighting conventions, photorealism, geometric design, ink wash, and folk or outsider art.The prompt collection includes Renaissance, Baroque, Romantic, Impressionist, Gothic, Bauhaus, Chinese ink wash, and African tribal art.
Loading 2608.27727v1…