Source-linked AI summary
StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs
Shaghayegh Kolli, Timo Cavelius, Nafiseh Nikeghbal, Samantha Dalal, Jana Diesner
TL;DR
Which specific visual attributes drive social judgments in MLLMs remains unclear because prior comparisons often confound appearance with identity. StylisticBias controls identity while varying one attribute at a time and finds that bias concentrates in a small set of cues, with about 15 attributes explaining nearly 80% of variation.
Problem
Which specific visual attributes drive MLLM social judgments remains unclear because prior studies often confound attribute effects with identity differences.
Method
StylisticBias evaluates attribute-level bias across six MLLMs and 25 scenarios using identity-fixed synthetic faces with one visual attribute varied at a time.
Results
About 15 visual attributes account for nearly 80% of total bias variation, with fashion and other self-presentation cues producing the largest shifts.
Takeaways & Limitations
StylisticBias provides a controlled benchmark for fine-grained auditing and attribution of appearance-driven bias in multimodal systems.
Takeaways & Limitations
Because the benchmark uses controlled synthetic images rather than real photographs, its conclusions characterize model behavior in a controlled visual setting rather than all real-image deployments.
Abstract
from arXiv · showhide
Multimodal large language models (MLLMs) are increasingly deployed in personally and societally consequential settings, yet the visual cues that shape how these models judge people remain poorly understood. Prior work often compares different (groups of) individuals, making it difficult to separate appearance effects from identity differences. We introduce StylisticBias, a controlled benchmark for evaluating attribute-level social bias in MLLMs. We generate 500 photorealistic base faces and create about 50 single-attribute variations per face, producing about 25K images. This design keeps identity fixed and changes one visual attribute at a time. It lets us measure how specific cues shift model judgments. We evaluate six MLLMs across 25 binary social judgment scenarios. We find that age and body type dominate identity-level effects, while fashion style and other visual cues drive the largest attribute-level shifts. We further find that about 15 attributes account for nearly 80\% of the total variation, showing that bias is concentrated in a small set of visual cues. Sensitivity is strongest in judgments that are semantically aligned with appearance, especially socioeconomic and style-related judgments. We release StylisticBias as a benchmark for fine-grained bias evaluation in multimodal models. Code and dataset: https://github.com/timo-cavelius/StylisticBias and https://hf.co/datasets/shaghayegh/stylistic-bias-dataset.
1 Introduction
StylisticBias isolates attribute-level social bias in MLLMs by varying one visual cue while holding facial identity fixed. Its large-scale evaluation shows that bias is concentrated in a small set of cues, especially for appearance-related judgments.
- Findings: VS = 0.075 and 0.069 identify body type and age as the strongest demographic drivers, with obese and elderly perceived identities receiving less favorable warmth and competence attributions.These findings position body type and age as the strongest identity-level effects in the study.
- Findings: Approximately 15 visual attributes account for nearly 80% of total |SBS|, showing that most bias comes from a small number of visual cues.Fashion style produces the largest attribute-level shifts according to the supplied introduction context.
- Benchmark: StylisticBias generates 500 photorealistic base faces and 25K synthetic images through controlled single-attribute edits that keep identity fixed.The benchmark evaluates attribute-level bias while separating appearance effects from identity differences.
- Evaluation: 6 MLLMs are evaluated across 25 binary social judgment scenarios, requiring about 4.72 million judgment calls per model and about 28.3 million in total.The scenarios span personality, interpersonal perception, behavioral, and socioeconomic judgments.
- Findings: Models show a similar overall bias pattern, with effects strongest in judgments semantically aligned with appearance.The introduction specifically highlights appearance-related judgments as especially sensitive.
2 Related Work
Prior research documents social, representational, and reasoning biases across language, multimodal, and generative models. Closest to this work, studies examine attractiveness bias, controlled facial counterfactuals, socially grounded trait inference, MLLM reliability, and appearance-based social judgments.
- Biases in Multimodal and Generative Models: Biases documented in language, multimodal, and generative models can reproduce or amplify societal stereotypes and demographic or representational disparities.The cited literature spans large language models, text-to-image systems, and other multimodal or generative systems.
- Closest Existing Studies: MLLMs show pervasive attractiveness bias, associating beautified faces with more positive traits, with effects interacting with gender, age, and race.Related work also uses face-only counterfactual edits to isolate demographic effects and socially grounded VQA to probe latent trait inferences.
- Cognitive and Reasoning Biases in LLMs: LLMs exhibit anchoring, framing effects, and confirmation bias, while multimodal studies report evaluator inconsistencies and fairness concerns across socially grounded tasks.Related evaluations include image-caption alignment, visual question answering, and multimodal quality assessment.
- Visual Appearance and Social Judgment: Human social judgments organize around warmth and competence, shaping inferences such as trustworthiness and socioeconomic status, with facial features influencing these impressions.This social-psychology foundation motivates studying how visual appearance affects social judgments.
3 StylisticBias
StylisticBias constructs a controlled benchmark by varying one visual attribute at a time while preserving face identity and evaluates resulting prediction shifts in binary social judgments. It comprises 500 synthetic base faces, approximately 25,000 variations, and human-validated image quality.
- Evaluation metric: The benchmark measures each variation’s prediction shift as Δ_i(x_v) = ϕ_i(x_v) − ϕ_i(x_b), comparing favorable-descriptor selection with its unchanged-identity base image.Responses are recoded so 1 always denotes the favorable descriptor, across prompt orderings and random seeds.
- Benchmark generation: 500 synthetic base faces span 90 demographic configurations defined by age, gender, ethnicity, and body type.The configurations are formed by 3 × 2 × 5 × 3 demographic categories, with 274 male and 226 female identities among the sampled faces.
- Benchmark generation: Each base face receives approximately 50 single-attribute variations, yielding approximately 25,000 images while preserving identity and other properties as consistently as possible.Variations cover facial and clothing-related cues, with clothing generated in full-body portraits using a dedicated prompt template.
- Benchmark evaluation: Six MLLMs perform binary forced-choice judgments across 25 scenarios under 3 seeds and 4 prompt orderings.Figure 1 describes evaluation as scenario design and model evaluation using controlled prediction shifts.
- Quality validation: 98% of the 90%-reviewed generated images satisfied demographic plausibility, identity consistency, and intended-attribute criteria.Failed images were regenerated and re-evaluated before downstream evaluation.
4 Evaluation Setup
The evaluation uses 25 binary social-judgment scenarios grounded in established person-perception frameworks. Models are tested with controlled prompt orderings and seeds, filtered attribute variations, and preference-shift metrics that quantify disparity and directional bias.
- Scenarios: 25 binary scenarios ask models to choose between paired social descriptors from visible appearance across four person-perception dimensions.The scenarios draw on the warmth–competence framework, Big Five personality traits, visual stereotype benchmarks, and socioeconomic judgments linked to clothing and presentation.
- Prompting: 12 prompt variants per image–scenario pair combine 4 descriptor orderings with 3 random seeds to mitigate prompt sensitivity.This yields 300 prompts per image, with preference scores computed over valid parsed responses.
- Data filtering: 34 values across 12 attribute categories remain after filtering subtle or semantically inconsistent changes, yielding 15,726 evaluated images.Examples removed include neutral lipstick and certain hairstyles on male faces.
- Metrics: Variation Strength measures between-group disparity, while Signed Bias Shift measures average attribute-induced directional change in preference scores.VS ranges from 0 to 0.5, SBS from −1 to +1, and significance uses corrected nonparametric tests.
5 Results
Visual bias in MLLMs concentrates in a small set of appearance cues and is strongest when judgments are semantically aligned with visible appearance. Demographic effects, cue asymmetries, and model-specific magnitudes are structured across architectures.
- Demographic effects: Body type (VS = 0.069) and age (VS = 0.075) show the largest demographic effects, with significant effects in 76% and 78% of scenarios, respectively.These cues correspond most closely to competence-related judgments and culturally linked social status.
- Attribute-level effects: Fashion (+0.046), facial hair (+0.042), makeup & lips (+0.037), and eyewear (+0.035) produce the largest positive SBS, while hair style and skin irregularities are consistently negative.No significant effects are detected for accessories, and 15 attributes reach the 80% cumulative |SBS| threshold.
- Attribute-level effects: Worn/distressed clothing produces median |SBS| = 0.167 versus 0.121 for formal/business attire, a 1.38× larger effect, while messy hair is 5.5× stronger than slicked-back hair.The fashion comparison uses full-body portraits, unlike the head-and-shoulders framing for other attributes.
- Semantic alignment: Stylish vs. Unstylish (SBS ≈+0.244) and Wealthy vs. Poor (SBS ≈+0.114) show the largest positive shifts, whereas Honest, Loyal, and Trustworthy remain near zero.Sensitivity follows Socioeconomic & Appearance > Behavioral > Personality > Interpersonal, with socioeconomic scenarios reaching |SBS| = 0.109 for Gemma-3.
- Cross-model structure: Pixtral is most reactive (SBS = +0.0273, Cohen’s d = 0.644), Qwen3 is most conservative, and Gemma-3 has |∆| ≥0.25 in 30% of cases.Fashion |SBS| ranges from 0.088 for Gemma-4 to 0.176 for Gemma-3, while category rankings are preserved across all six architectures.
Conclusion
StylisticBias evaluates attribute-level social bias by holding identity fixed while varying one visual attribute at a time. Across six MLLMs and 25 social judgment scenarios, it finds that sensitivity is concentrated in a small set of appearance cues, especially self-presentation cues.
- Benchmark and findings: StylisticBias keeps identity fixed and varies one visual attribute at a time to evaluate attribute-level social bias in MLLMs.This controlled design supports fine-grained visual attribution rather than coarse demographic comparison.
- Benchmark and findings: Across six MLLMs and 25 social judgment scenarios, bias concentrates in a relatively small set of visual cues rather than spreading uniformly across appearance categories.The conclusion highlights self-presentation cues as particularly important.
- Implications: MLLMs are systematically sensitive to how people look, not only to who they are represented as being.The benchmark provides a foundation for fine-grained bias evaluation and future auditing and mitigation of appearance-driven bias.
Limitations
The study evaluates controlled synthetic images rather than real photographs, trading ecological realism for privacy-preserving, fine-grained control over isolated visual attributes.
- Synthetic-image evaluation: The benchmark uses controlled synthetic images instead of real photographs, limiting direct evaluation on naturalistic human imagery.This deliberate choice avoids privacy, consent, and related ethical concerns associated with real human images.
- Synthetic-image evaluation: Synthetic data enables one visual attribute to vary while identity, pose, lighting, and background remain as fixed as possible.This control supports isolating attribute-level effects, which is difficult to achieve reliably at scale with real images.
Ethical Statement · A Model Details
The paper frames StylisticBias as a fairness-auditing benchmark for appearance-driven bias in consequential MLLM applications while acknowledging privacy, stereotyping, and construct-validity concerns. It also discloses limited use of LLM assistants for writing support.
- Ethical Statement: StylisticBias targets visual-attribute effects on social judgments in consequential MLLM applications, including hiring, content moderation, and judicial support.The benchmark is intended to support fairness auditing and bias attribution.
- Ethical Statement: Appearance-driven bias concentrates in a small set of self-presentation cues and is amplified for socioeconomic judgments.These patterns are described as not captured by standard evaluation.
- Ethical Statement: Synthetic face generation reduces privacy risks but may reproduce stereotypical associations from generative training data.The ethical tradeoff combines privacy protection with possible stereotype reproduction.
- Ethical Statement: Some tested categories and category values are social constructs shaped by stereotypical perceptions and normative expectations.The passage notes that these constructs may lack diversified perspectives.
- Ethical Statement: The tested categories and values can themselves be judgmental, creating an ethical concern about how the benchmark frames social perception.This concern follows from the passage’s warning about stereotypical and normative foundations.
- Ethical Statement: LLM-based AI assistants provided limited writing support, such as grammar correction and phrasing improvements.The paper discloses this use explicitly.
B Dataset Generation … B.3 Variation Generation
The dataset uses a two-stage pipeline to generate 500 structured, photorealistic base portraits and identity-preserving single-attribute or fashion-style variations. This controlled design fixes identity while changing one visual feature at a time.
- B Dataset Generation: The generation process comprises base-face synthesis followed by controlled variation generation.The documented process specifies prompt families and feature spaces for both stages.
- B.1 Two-Stage Generation Pipeline: Base portraits are studio head-and-shoulders images generated from structured demographic attributes using a template.The demographic slots include body type, age, gender, and ethnicity.
- B.2 Base-Face Generation: 500 valid base faces constitute the finalized dataset.The set includes 274 male and 226 female faces, with distributions across body type, ethnicity, and age.
- B.2 Base-Face Generation: The base portraits use neutral expressions, white backdrops, controlled lighting, and photorealistic studio prompts.Prompt slots explicitly encode body_type, age, gender, and ethnicity.
- B.3 Variation Generation: Each variation modifies exactly one feature key and one value at a time.Nano Banana produces face-focused outputs for non-fashion variations and full-body outputs for fashion-style variations.
- B.3 Variation Generation: All variation prompts require preserving the reference image’s identity.This requirement supports controlled comparisons in which the edited visual attribute changes while identity remains fixed.
- B.3 Variation Generation: Fashion variations change clothing style and use full-body portraits, whereas other controlled variations remain face-focused.The corresponding prompt templates are shown in Figures 7 and 8.
C Experimental Setup … D Detailed Results
The benchmark preserves a full combinatorial variation space in the dataset but evaluates a curated subset, using controlled forced-choice prompts to make MLLM judgments computationally tractable and aggregation-robust. Two-stage reduction yields 15,726 evaluated images, while each image-scenario pair contributes up to 12 responses summarized as an empirical probability.
- C.1 Face variations.: The full variation grid remains in the dataset, while only a curated subset is forwarded to MLLM judgment because exhaustive evaluation is computationally prohibitive.This separates dataset coverage of the variation space from the tractable judgment evaluation grid.
- C.1 Face variations.: The unreduced pipeline entails about 25,000 images, with each image assessed using 300 prompts per MLLM; six models scale the judgment-call total proportionally.Each variation requires both an image-generation call and forced-choice evaluation calls across scenarios.
- C.1 Face variations.: Two-stage reduction shrinks 55 variation values to 34 values across 12 attribute categories, reducing evaluated images from 25K to 15,726, or almost 40%.A plausibility pass removes incoherent or confounded combinations, followed by curation that drops values with limited additional signal.
- C.1 Face variations.: The plausibility pass excludes male braid and bun hairstyles, female neutral and bold-color lipstick values, and daring/provocative and luxury/high-fashion styles.These exclusions address generation fidelity, implicit baselines, visual redundancy, inconsistent interpretation, moderation refusals, cross-face variability, or limited everyday prevalence.
- C.1 Face variations.: Curation retains visually distinguishable feature combinations, including two nose-piercing values and three hair-style values, while excluded values remain documented in the full variation space.The reduced grid is designed to lower judgment cost without materially shrinking the measured bias signal.
- C.2 Forced-choice judgment protocol.: Each image-scenario pair is judged under M × K = 4 × 3 = 12 prompts spanning four order/label variants and three seeds.The protocol varies option order and letter-to-option mapping to marginalize position and label effects.
- C.2 Forced-choice judgment protocol.: Invalid responses, including refusals, hedged answers, or outputs containing both or neither letters, are excluded before aggregation into option-A probabilities ϕ_i(x).For each pair, the reported bias metrics use empirical probabilities computed from the valid parsed responses, with n_i(x) ≤12.
D.1 Demographic Sensitivity Across Models
Demographic sensitivity varies substantially across models and attributes. Body type and age produce significant shifts most consistently, while Qwen3 is least sensitive overall and Gemma models are comparatively balanced.
- Demographic sensitivity: 78% and 76% average significance rates occur for age and body type, respectively, versus 67% for ethnicity and 51% for gender.These rates measure the percentage of scenarios with statistically significant prediction differences across demographic groups.
- Demographic sensitivity: Qwen3 remains at or below 60% sensitivity across all attributes: age (60%), body type (56%), ethnicity (44%), and gender (44%).This is the lowest overall sensitivity reported across the evaluated models.
- Demographic sensitivity: Gemma-3 spans 60% to 84% and Gemma-4 spans 52% to 72% across the four attributes, with Gemma-4 ethnicity (72%) exceeding gender (52%).The Gemma models show the most balanced sensitivity profiles across demographic attributes.
D.2 Mixed-Effects Model and Partial η2p · D.3 Full Demographic × Variation Prediction Shift Table
The mixed-effects analysis attributes prediction-shift variance primarily to scenario and variation categories, with age and body type showing significant demographic effects. The accompanying shift table defines signed appearance-variation effects across demographic groups using standardized magnitude bands.
- D.2 Mixed-Effects Model and Partial η2p: The linear mixed-effects model uses random intercepts per face identity to account for repeated measurements across variation and scenario categories.The model is fitted jointly for variation and scenario category on N=19,868 observations from 500 faces.
- D.2 Mixed-Effects Model and Partial η2p: Age and body type have significant face-level effects, whereas gender is weaker and ethnicity is not significant.The reported ANOVA results are p<0.001 for age and body type, p<0.01 for gender, and p=0.057 for ethnicity.
- D.2 Mixed-Effects Model and Partial η2p: Variation type and scenario type together explain 59% of prediction-shift variance, while face-level random effects add a further 5%.The marginal R2 is 0.594, and the passage attributes the additional 5% to face-level random effects.
- D.2 Mixed-Effects Model and Partial η2p: Scenario category accounts for more variance in prediction shifts than variation category, supporting a semantic-alignment pattern in model judgments.The reported partial η2p values are 0.248 for scenario category and 0.153 for variation category.
- D.3 Full Demographic × Variation Prediction Shift Table: Table 11 reports mean signed prediction shifts for each appearance variation across demographic groups, averaged over six MLLMs and 25 binary scenarios.Each shift is defined as ∆=φ(xv)−φ(xb), comparing a variation image with its baseline image.
- D.3 Full Demographic × Variation Prediction Shift Table: The table classifies shifts as strong, moderate, or neutral using thresholds from ∆≥+0.10 to ∆≤−0.10, with grey dashes marking inapplicable demographic-variation combinations.Positive shifts indicate movement toward the socially favorable pole, while negative shifts indicate movement toward the unfavorable pole.