Source-linked AI summary

Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces

Luca Bux, Thiago Rios, Ingo Scholtes, Stefan Menzel

arXiv:2608.25876v1cs.HCcs.CV

TL;DR

Generative design interfaces need to know whether VLM-based affective controls produce stable object rankings across models. The paper audits six VLMs with bipolar Kansei directions on controlled 3D shapes and finds partial convergence above an empirical null but below geometric controls, with category-dependent agreement. It uses this audit to inform which descriptors a design interface might expose, while distinguishing model consistency from human validation.

  • Problem

    It is unclear whether VLMs consistently represent affective qualities of shape, making cross-model stability an important preliminary question for semantic design controls.

  • Method

    The study ranks untextured multi-view ShapeNet objects along adjective-pair directions using six pretrained VLMs, calibrated with geometric positive controls and an irrelevant-adjective null.

  • Results

    Affective axes show higher average agreement than the empirical null but lower agreement than geometric controls, and convergence varies with alignment between category variation and the evaluated semantic direction.

  • Takeaways & Limitations

    Cross-model convergence can provide a first-stage consistency assessment for deciding which affective descriptors to consider exposing as controls in AI-assisted design interfaces.

  • Takeaways & Limitations

    Cross-model convergence does not establish agreement with human affective judgments, and the interface mechanisms were not evaluated with users.

Abstract

from arXiv · show

Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as "more elegant" or "more minimalist," typically encoded by a vision-language model (VLM). A practical question is whether state-of-the-art VLMs represent objects consistently in terms of the same concept. We audit 6 VLMs by ranking untextured 3D objects along Kansei adjective pairs, where Kansei describes affective impressions of product form, with each axis defined as the difference between the text representations of its two poles. Geometric pairs serve as positive controls, and pairs of unrelated adjectives establish an empirical null. Across 10 categories of ShapeNet database, affective axes converge above the null (mean pairwise rank correlation 0.36 vs. 0.14) but below the geometric ceiling (0.44). The agreement between models is partial and highly uneven: on the three axes shared by all categories, mean convergence ranges from 0.21 for bookshelves to 0.51 for jars. Convergence depends primarily on whether a category's representational variation aligns with the semantic direction being evaluated, rather than simply on how much the objects vary in shape overall. Cross-model convergence does not imply agreement with human judgments. Based on our findings, we implement a UI prototype that shows how the audit can inform which Kansei descriptors to expose as controls for a given object class and which to withhold.

1 Introduction

Generative design interfaces increasingly use affective semantic controls, but it remains unclear whether VLMs produce stable rankings for qualities expressed through shape. This work proposes a cross-model audit to assess reproducibility before exposing such controls to users.

  • Motivation: Semantic controls let designers steer generation and refinement with descriptors such as “sleeker,” “elegant,” and “minimalist.”These controls may guide semantic editing, optimization, or iterative design refinement.
  • Research gap: VLM design research has emphasized object recognition, leaving consistency in representing affective qualities of shape less studied.Kansei engineering concerns how product forms relate to people’s affective impressions.
  • Research gap: Large-scale human validation is difficult because affective 3D-shape ratings are limited, slow, and expensive to collect.Photographic evaluations also include color, texture, lighting, and scene context, whereas the audit asks first whether independently developed VLMs agree.
  • Audit rationale: The study evaluates whether independently developed VLMs produce consistent rankings of affective attributes before those descriptors become interface controls.Cross-model convergence is treated as an initial consistency signal rather than evidence of human perceptual correspondence.
  • Contributions: The audit uses geometric positive controls and unrelated adjective pairs as an empirical null to calibrate affective-control consistency.The paper frames this as an offline cross-model consistency audit for AI-assisted design interfaces.
  • Contributions: The work characterizes when affective controls yield reproducible signals and demonstrates a UI prototype exposing convergence information alongside semantic controls.The proposed workflow supports interface design decisions about which descriptors to expose.

2 Related work

The paper connects Kansei engineering, semantic directions, VLM evaluation, interactive interfaces, and model agreement. Its distinctive focus is using cross-model convergence to audit whether affective directions are stable enough to consider as controls for 3D design.

  • Kansei engineering: Kansei engineering relates product forms to affective impressions, but demographic variation can limit the generality of static human-labeled datasets.The paper uses Kansei adjective pairs while testing model consistency rather than collecting new human ratings.
  • Affective VLM evaluation: Prior aesthetic and affective VLM work mainly evaluates photographs or single-model scores, where appearance and scene context influence judgments.The present setting instead targets untextured multi-view 3D form and reproducibility across independently developed VLMs.
  • Semantic directions: Semantic projection represents an attribute as a direction between antonym text embeddings and scores objects by projection in the shared embedding space.This study applies that framework to candidate affective directions for object categories.
  • Model agreement: Model-agreement research motivates convergence as evidence from multiple measurements, including ensembles, representational similarity, and shared latent structures.The paper adapts this rationale to an interaction-oriented audit rather than general model-similarity analysis.
  • Audit criterion: The audit calibrates affective agreement against geometric controls and an empirical null to decide which semantic directions merit consideration as interface controls.This links model convergence directly to transparent, uncertainty-aware intelligent design interfaces.

3 Methodology

The methodology audits affective semantic structure in pretrained VLMs using controlled multi-view ShapeNet objects, bipolar adjective directions, and calibrated convergence measures. Untextured rendering isolates form while geometric and irrelevant probes provide positive and null controls.

  • Directional projection: For each affective concept, subtracting the two pole text embeddings defines a semantic direction, and object projections produce continuous scores along that axis.Cross-model evaluation then tests whether these scores yield stable rankings across encoders.
  • Stimuli: The dataset contains 4,950 objects sampled across chair, table, lamp, sofa, cabinet, bookshelf, bottle, jar, clock, and car categories.Eight categories reach 500 objects; bottle and bookshelf contain 498 and 452 objects, respectively.
  • Stimuli: Each object is rendered from 8 fixed viewpoints at two elevations and four azimuthal arrangements.The views are generated at 512 × 512 resolution on a flat white background.
  • Stimuli: Uniform matte-grey, textureless rendering removes surface-appearance cues so VLM judgments focus on geometric form.This avoids associations such as linking “luxurious” to glossy dark surfaces or “cheap” to colorful plastic-like appearances.
  • Representation extraction: Object representations average the 8 view embeddings and renormalize the result to unit length.Averaging emphasizes shape shared across viewpoints, while normalization makes projections reflect angular alignment rather than embedding magnitude.
  • Probe design: Bipolar probes comprise 7 geometric positive-control pairs, 32 Kansei pairs, and 20 irrelevant pairs forming an empirical null.The Kansei tier includes three pairs shared across all categories and additional category-specific descriptors.

3.5 Directions and Projection

The paper constructs affective semantic directions from antonym text embeddings and evaluates their object rankings through cross-model convergence, calibrated against geometric controls and an empirical null. It also tests whether convergence reflects alignment with category-specific variation.

  • Direction construction: Affective axes are built by subtracting antonym text embeddings and normalizing the resulting direction within each encoder’s native space.The direction points from the negative pole to the positive pole, and absolute scores are therefore compared through rank orderings rather than across-model magnitudes.
  • Convergence metric: Cross-model convergence measures how consistently the six encoders rank objects along a semantic axis using pairwise Spearman rank correlations.The metric follows a convergent-validity logic: agreement across independent assessment methods provides evidence of consistent measurement.
  • Calibration and evaluation: Geometric adjective pairs provide positive controls, while unrelated adjective pairs establish the empirical null used to calibrate affective-axis convergence.The statistical comparison uses one-sided Mann–Whitney tests, common-language effect sizes, and KS statistics for distributional differences.
  • Uncertainty: Because categories and axes share null sets, core pairs, and object renderings, category-dependent claims use bootstrap confidence intervals rather than individual p-values.Objects are resampled with replacement for 1,000 replicates, and differences are described as category-dependent when intervals do not overlap.
  • Variation-subspace alignment: Category-specific convergence is hypothesized to increase when an affective direction lies within the subspace along which that category’s shapes vary.The study estimates this alignment from the top 20 principal components and correlates alignment with convergence, treating the relationship as correlational rather than causal.

3.8 View Reliability

The view-reliability analysis tests whether rankings derived from eight-view object representations remain stable across viewpoint subsets. It corrects split-half reliability and uses those reliabilities to adjust cross-encoder agreement, while also testing category-specific axes on own versus foreign categories and alternative probe constructions.

  • View reliability: Eight views are split into two disjoint four-view halves, and Spearman correlation between their object score vectors measures ranking stability.A Spearman–Brown correction estimates reliability for the full eight-view representation.
  • Reliability correction: Pairwise encoder correlations are disattenuated using the reliabilities of the two encoders, reducing the role of view-sampling noise in convergence estimates.The corrected value is intended to represent agreement beyond measurement unreliability from the sampled views.
  • Own-versus-foreign transfer: Category-specific affective axes are applied to their own category and to foreign categories to test whether convergence is tied to the object class they describe.If the signal were generic to the text embedding rather than object-specific, own and foreign applications would converge equally well.
  • Probe robustness: Robustness checks replace bipolar adjective differences with single-pole embeddings and replace bare adjective pairs with category-specific sentence templates.Convergence is recomputed across all three evaluation tiers for both alternative constructions.
  • Reproducibility: The audit uses publicly released model weights, and the rendering, extraction, analysis code, and precomputed audit values are slated for public release.These materials support reproduction of the view-reliability and convergence analyses.

4 Results

Across 6 encoders, affective axes showed intermediate convergence between geometric controls and the irrelevant null, with substantial variation across categories and descriptors. Convergence was associated with alignment between semantic directions and category-specific shape variation, while remaining distinct from geometric axes.

  • 4.1 Convergence across tiers: 0.364 affective convergence fell between geometric controls at 0.441 and the irrelevant null at 0.135.Affective axes outconverged irrelevant axes with CL = 0.906, while geometric controls beat the null with CL = 0.949.
  • 4.1 Convergence across tiers: Affective orderings approached geometric consistency despite using untextured objects without colour, material, or texture.
  • 4.2 Convergence depends on both the axis and the category: 0.21 to 0.51 was the shared-core convergence range across categories, from bookshelves to jars.The shared-core ranking uses modern–traditional, elegant–messy, and luxurious–cheap; bottles ranked third on the shared core but ninth across their full vocabulary.
  • 4.2 Convergence depends on both the axis and the category: Reliability correction left category differences intact: corrected convergence remained spread across 0.32–0.58 and was uncorrelated with reliability at 𝜌= 0.10.Raw convergence correlated with split-half reliability at 𝜌= 0.47, indicating that measurement quality explained part, but not all, of the category effect.
  • 4.3 Alignment with shape variation: 0.72 was the alignment–convergence correlation for affective axes, compared with 0.78 for geometric and 0.24 for irrelevant axes.After accounting for score spread and measurement noise, alignment still predicted Kansei convergence at 𝜌= 0.47, 𝑝< 0.001.
  • 4.3 Alignment with shape variation: Geometric and affective axes remained distinct: the strongest affective–geometric relationship had mean | cos | = 0.24.The simple–complex and minimalist–ornate pair was the strongest relationship, and the authors report that the observed convergence cannot be explained by the selected geometric pairs alone.

5 Discussion

The audit is presented as a practical consistency check for semantic controls, identifying category- and descriptor-specific agreement and informing which controls to expose, flag, or withhold.

  • High-convergence descriptors can be treated as more reproducible controls, while low-convergence descriptors can be flagged for further evaluation.
  • The audit maps each category and semantic axis to support selective control admission across object classes.The same descriptor may show consistent model agreement for one category but not another, without requiring human labels or model retraining.
  • The interface prototype displays encoder projections and per-object agreement so users can see model convergence during interaction.Higher convergence appears as similarly clustered encoder scores, while low-convergence axes reveal greater disagreement.
  • The proposed admission rule exposes controls whose lower bootstrap bound exceeds the empirical null, marks overlapping intervals provisional, and withholds estimates at or below the null.For example, luxurious–cheap is exposed for jars at ¯𝜌= 0.47 but withheld for bookshelves at 0.10 and cabinets at 0.14.
  • The audit is intended as infrastructure for ranking candidate controls, highlighting near-null descriptors, and communicating uncertainty rather than as an end-user evaluation.
  • Low-agreement controls should be flagged as uncertain rather than automatically removed, with selected descriptors checked against intended users in culturally sensitive or consequential settings.The evaluated VLMs may share training data, linguistic conventions, and representational biases, while Kansei meanings can vary across cultures, groups, and contexts.

6 Limitations

The paper identifies two major limitations: convergence is not human validation, and the interface prototype has not been evaluated with users.

  • Cross-model consistency does not establish agreement with human affective judgments or correspondence between recovered directions and human Kansei perception.The audit is framed as a model-based criterion that can assist, not replace, human judgment.
  • The interface has not been evaluated for effects on designer performance, decision quality, calibration, or trust.Controlled user studies are needed to compare convergence-aware interfaces with conventional semantic-control interfaces.

7 Conclusion

The paper concludes that cross-model convergence can audit affective semantic controls by distinguishing reproducible signals from model-specific behavior. It positions the framework as a first-stage assessment that complements future human validation and interface studies.

  • Affective axes show higher average agreement than an empirical null but lower average agreement than geometric controls across six VLMs.
  • Cross-model agreement depends on whether the evaluated semantic direction captures variation present among the compared objects.
  • The audit requires no human labels, scales to new categories and models, and can be recomputed as VLMs evolve.It is proposed as a first-stage consistency assessment rather than a substitute for human evaluation.
  • Future work should validate convergence against human Kansei judgments, examine richer representations, optimize Kansei vocabulary, and evaluate effects on designer decision-making.

GenAI Usage Disclosure

The authors used LLMs to help construct the Kansei probe vocabulary and to edit language, while retaining responsibility for the experimental design, analysis, interpretation, and scientific claims.

  • LLMs suggested candidate affective adjectives and bipolar antonyms when source literature lacked sufficient vocabulary, after which authors reviewed and curated the suggestions.
  • The authors state that experimental design, data analysis, interpretation of results, and scientific claims are their own responsibility.

A Encoder Agreement Matrix

Table 6 reports encoder agreement using mean pairwise Spearman rank correlation, averaged across all Kansei axes.

  • Mean pairwise Spearman rank correlation measures agreement between every pair of encoders.
  • The agreement statistic is averaged over all Kansei axes.
  • Table 6 therefore summarizes cross-encoder ranking consistency across the affective-axis probe set.

B Vocabulary

The probe vocabulary includes shared geometric and irrelevant tiers, category-specific Kansei axes, and three Kansei axes shared across all categories.

  • Table 7 gives the complete probe vocabulary used in the audit.
  • The paper is titled “Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces.”
  • Geometric and irrelevant probe tiers are shared across all ten categories.
  • Kansei axes are category-specific, with additional axes drawn from Kansei literature and an LLM-assisted expansion.
  • Modern–traditional, elegant–messy, and luxurious–cheap form the shared-core axes present in every category.
Loading 2608.25876v1…