Source-linked AI summary

Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta, Shafee Hassan, Macken Murphy

arXiv:2609.02512v1cs.CVcs.HC

TL;DR

The study asks whether commercial MLLMs accurately reflect human judgments of facial attractiveness. It compares four models with ratings from 2,513 human participants and finds strong rank-order correspondence but systematic absolute overrating and narrower score ranges.

  • Problem

    Few studies have tested whether MLLMs judge facial attractiveness like humans, despite increasing reliance on these systems as stand-ins for human judgments.

  • Method

    The preregistered exploratory study compares four off-the-shelf MLLMs with an existing dataset of human attractiveness ratings and examines agreement, predictors, and rater subgroups.

  • Results

    MLLMs reproduce human rank-orderings but rate faces more favorably and within a narrower range, failing to reproduce human ratings in absolute terms.

  • Takeaways & Limitations

    Current commercial MLLMs may approximate relative attractiveness judgments but systematically overrate the beauty of human faces.

  • Takeaways & Limitations

    Agreement comparisons may partly reflect lower measurement error in averaged MLLM scores than in single human ratings.

Abstract

from arXiv · show

Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular amongst users, companies, and aestheticians. This raises the question of whether these AI models can accurately reflect human judgments of attractiveness. In a pre- registered exploratory study, we compared the attractiveness ratings of 2,513 human participants to four widely used commercial AI models: Claude, Gemini, GPT, and Grok. Results showed that MLLMs systematically rate faces more favourably and within a narrower range than humans and, at the time of study, do not reproduce human ratings in absolute terms. However, MLLMs exhibit strong correlations with human attractiveness judgments, accurately tracking the rank-ordering of faces. MLLMs may judge faces by different cues than humans; only face age was a predictor of facial attractiveness in both humans and MLLMs, with inconsistent patterns across models for ethnicity and gender. AI models strongly agree with one another, except for Grok, which also showed the lowest agreement with humans. Our findings suggest that while they may be able to approximate rank-orderings of human attractiveness, current off-the-shelf commercial MLLMs systematically overrate the beauty of human faces.

1. Introduction

Facial attractiveness reflects both substantial shared taste and meaningful variation across people and populations. This complexity makes it important to test whether MLLMs reproduce human judgments rather than assuming AI beauty assessments are objective.

  • MLLMs show strong image-processing capabilities but struggle with fine-grained facial perception and tasks such as age and race estimation.
  • The study compares four MLLMs with human ratings to assess agreement, model correspondence, facial predictors, and subgroup differences.
  • Human attractiveness judgments combine broad consensus with differences between populations, individuals, and contexts.Shared taste has been estimated to account for roughly 60% of rating variance, with individual taste accounting for the remaining 40%.
  • Human research identifies averageness, symmetry, health, youth, and sex-linked facial characteristics as contributors to perceived attractiveness.
  • Variation in human preferences challenges the objectivity and functional accuracy of AI beauty assessment tuned to mimic a subset of people.

1.2 MLLM perception of faces

Prior work suggests some AI-human correspondence in facial judgments, but evidence remains limited and methodologically constrained. The study therefore addresses whether MLLMs agree with humans, with one another, and with particular human subgroups.

  • MLLMs learn visual representations through image-text pairs, so their face judgments are filtered through language-based descriptions.
  • Demographic and cultural skews in web-based training data may shape MLLM judgments of facial attractiveness.
  • AI-human correspondence should not be interpreted as evidence that MLLMs perceive faces like embodied human observers.
  • Previous work also raises concerns about model reliability and non-deterministic outputs, motivating stable measurement procedures.
  • Preliminary studies report correlations between MLLM or facial-analysis ratings and human judgments, but existing methods include synthetic, narrow, or small samples.
  • Open questions concern model agreement, facial-characteristic effects, and whether MLLMs more closely match specific human rater subgroups.

1.3 Current study

The preregistered exploratory study evaluates four widely used frontier MLLMs against human attractiveness judgments. It measures agreement across humans and models, tests face-level predictors, and examines rater-subgroup differences.

  • The study investigates attractiveness-rating agreement between GPT, Gemini, Claude, Grok, and human raters.
  • The analyses address agreement between MLLMs and humans, agreement across MLLMs, face-level predictors, and differential agreement with human subgroups.
  • The study combines ICC and correlational measures to detect both absolute agreement and systematic rating biases.
  • It evaluates four distinct frontier models using 102 non-synthetic human photographs rated by 2513 human participants.
  • Research questions and the analysis plan were preregistered on the Open Science Framework before data collection.

2.1 Stimuli and Human Ratings

Human attractiveness ratings came from a public dataset containing 102 neutral, front-facing facial images rated by 2513 participants. The faces represented both sexes and multiple racial or ethnic groups.

  • The human sample included 2513 ratings of 102 neutral, front-facing facial images from the Face Research Lab London Set.
  • The stimulus set was 48% female and 52% male.
  • The racial or ethnic composition was 67.6% White, 12.7% Black, 9.8% West-Asian, 8.8% East-Asian, and 1.0% mixed East-Asian/White.
  • All images were neutral, front-facing 1350 × 1350 px JPEGs.

2.2 MLLM Ratings

The study collected standardized attractiveness ratings from commercial MLLMs using controlled prompts and repeated API runs. The design also documented exclusions, refusals, and model-specific reliability to estimate stable outputs.

  • Twelve human raters with zero variance across all 102 faces were excluded from some analyses.The study had preregistered no exclusions.
  • Four off-the-shelf commercial MLLMs provided attractiveness ratings through independent, stateless API calls using single face images.The models were GPT-5.3, Claude Sonnet 4, Gemini 2.5 Flash, and Grok 4.202.
  • A minimal system prompt instructed models to return numerical ratings after pilot refusals occurred.The prompt framed the task as providing numerical ratings on standard psychological scales and required responding only with the number.
  • Refusals were rare: 10 of 6,324 API calls, all from Claude and concentrated on three faces.Technical API failures were retried up to 10 times, whereas explicit refusals were coded as missing.
  • Ratings were averaged across repeated runs to estimate each model’s stable output, requiring 6, 6, 9, and 41 runs for Claude, GPT, Gemini, and Grok, respectively.Pilot single-run reliability was ICC(1,1) = .79, .79, .70, and .32 in the same model order.

2.3 Data Analysis

Across 2,513 human ratings and four MLLMs, the models rated faces higher and within a narrower range, yet closely tracked human rank-orderings. MLLMs also largely agreed with one another, showed smaller demographic penalties, and aligned marginally more with female than male raters.

  • Rating distributions: Humans gave the lowest average rating (M = 3.02), while AI-pooled ratings averaged M = 4.71 and individual MLLM means ranged from M = 4.41 to 5.21.Grok had the highest mean rating (M = 5.21, SD = 0.87), whereas GPT had the lowest among the models (M = 4.41, SD = 0.66).
  • Agreement with humans: Pairwise rank agreement with individual humans was ρ = .43–.46 for three models and ρ = .33 for Grok, compared with human-human agreement of ρ = .35.These correlations indicate comparable or stronger tracking of relative facial-attractiveness ordering than human raters showed with one another.
  • Agreement with humans: Absolute agreement with humans was poor, with ICC(2,1) values from .08 to .19 versus the human-human benchmark of .27.The models’ higher means, approximately M = 4.4–5.2 versus M = 3.02 for humans, depressed absolute-agreement coefficients despite preserved ordering.
  • Agreement with humans: Against the mean human rating, models showed ρ = .74–.76 except Grok at ρ = .58, while individual human raters averaged ρ = .59.The models therefore matched the human consensus more closely than a typical individual human rater, except Grok, which was approximately at the individual-rater level.
  • Agreement among MLLMs: MLLMs showed strong to very strong inter-model rank agreement (ρ = .69–.86), while pairings involving Grok had lower absolute agreement (ICC = .33–.49).Against the leave-one-out consensus of the other models, all models remained strongly aligned (ρ = .75–.87), with Claude highest at ρ = .87 and Grok lowest at ρ = .75.
  • Demographic predictors: Only face age significantly predicted attractiveness across all raters, while MLLMs gave more equal ratings across face gender and ethnicity than humans.Humans showed a 0.65-point male-versus-female gap, compared with approximately 0.1–0.2 points in MLLMs; MLLMs also showed smaller demographic penalties overall.
  • Rater subgroups: Male raters tracked MLLM ratings less closely than female raters across all five model conditions, although subgroup differences were generally marginal.Agreement also rose slightly with rater age on ICC, with raters aged 31 and older showing the highest values for every model.

4.1 Human-MLLM agreement (RQ1)

MLLMs align poorly with individual humans in absolute attractiveness ratings because they rate faces more highly and use a narrower scale, despite stronger associations with average human ratings.

  • Absolute agreement: All models showed poorer absolute agreement with individual human raters than humans showed with one another.MLLMs gave higher ratings overall and did not use the full rating scale: none assigned the lowest rating, and only Grok assigned the highest.
  • Interpretation: Possible explanations include sycophantic behavior, safety moderation, and differences between model training distributions and humans’ reference distributions.The moderation account remains speculative because the evaluated models’ pipelines were not probed.
  • Rank-order association: All models except Grok showed strong associations with average human ratings under a leave-one-out analysis.Averaging human ratings reduces measurement error relative to noisier single-rater outputs.
  • Rating level: The models’ higher ratings were consistent with previous findings from AI-based facial-attractiveness websites.The supplied discussion identifies this as a recurring difference between AI and human ratings.

4.2 MLLM-MLLM agreement (RQ2)

MLLMs generally agree more strongly with one another than with humans in facial-attractiveness ratings, while Grok shows the weakest agreement and association across comparisons.

  • Agreement strength: MLLM-MLLM absolute agreement ranged from .33 to .81, compared with .08 to .18 for human-MLLM agreement.Thus, models agreed far more with one another in absolute ratings than with human raters.
  • Novelty: The study describes itself as the first evaluation of frontier MLLMs against one another on facial attractiveness.This places the cross-model comparison in a novel research context.
  • Model differences: Grok showed the lowest agreement and association with both other MLLMs and human raters across various tests.The discussion suggests inconsistent image processing may contribute to Grok’s results.

4.3 Face-level predictors (RQ3)

All MLLMs reproduced face-age effects found in human ratings, but gender and ethnicity patterns were inconsistent across models and did not fully reflect human predictors.

  • Age: All MLLMs were sensitive to face-age effects, mirroring the human data and prior literature.Age was the only facial predictor identified as shared by humans and all evaluated models in the supplied results.
  • Gender and ethnicity: GPT and Claude displayed the documented human penalty to male faces, whereas only Gemini showed face-ethnicity effects.These demographic patterns were not consistent across the models.
  • Cross-model divergence: Previously identified human predictors of attractiveness were not entirely reflected in MLLM ratings, and some models showed novel patterns.The passage specifically points to Gemini as an example of a model with novel patterns.
  • Interpretation and scope: A speculative RLHF account proposes that socially sensitive gender and ethnicity judgments may be suppressed while age effects remain.Interpretation is constrained by 102 stimuli, with 67.6% depicting White faces and few images representing other ethnicities.

4.4 Differential agreement with human raters (RQ4)

MLLM-human agreement differed only marginally across rater subgroups, with small regression effects favoring female raters and, except for Grok, raters attracted to either gender.

  • Overall subgroup comparison: Correlation differences across human-rater subgroups were minimal, providing no strong evidence of subgroup-specific agreement.The authors therefore distinguish small statistical effects from strong subgroup preference matching.
  • Observed effects: All MLLMs tracked female raters more closely, and all except Grok aligned more strongly with raters attracted to either gender than with those attracted to men.These were small effects revealed by regression models rather than strong correlation differences.
  • Alternative explanation: The apparent subgroup effects may reflect similar face rankings rather than models sharing particular raters’ preferences.Because MLLMs differentiated male and female faces less than humans, raters minimizing that distinction could resemble the models structurally.
  • Limitations: Interpretation is limited because the rater sample was predominantly young, female, and attracted to men and lacked ethnicity or cultural-background information.These design features left little room to compare demographic preferences and omitted a potentially relevant dimension of attractiveness judgments.

4.5 Limitations and future directions

The study’s comparisons are constrained by unequal measurement reliability, model-version specificity, rating-wording uncertainty, and limited demographic coverage of raters and faces.

  • Measurement reliability: Averaged MLLM ratings achieved reliability of at least .95, whereas human ratings were single judgments per face.This asymmetry means model scores carry less measurement error than individual human ratings.
  • Measurement reliability: MLLM–human and inter-model agreement may partly reflect reduced model measurement error rather than stronger underlying signal.
  • Model scope: The findings concern GPT-5.3, Claude Sonnet 4, Gemini 2.5 Flash, and Grok 4.20 evaluated at one point in time.Frequent, opaque updates may limit generalization to other or future versions.
  • Rating wording: The human survey’s exact wording was unavailable, so MLLMs received wording from another survey using the same 1–7 scale.
  • Human sample: The human benchmark was a single online sample that was predominantly young, female, and attracted to men, with cultural and ethnic background unrecorded.
  • Face stimuli: The face set was dominated by White faces, leaving relatively few observations for non-White groups and limiting precision of ethnicity effects.Reported effects may reflect the particular sampled faces rather than those ethnic groups broadly.
  • Future directions: Future studies should record human raters’ ethnic backgrounds to test whether rater–face ethnicity matching influences attractiveness judgments.

4.6 Conclusion

Across four frontier MLLMs, the models tracked the rank ordering of human facial-attractiveness judgments and agreed closely with one another, but they did not reproduce human absolute ratings.

  • MLLMs reproduced the rank ordering of human facial-attractiveness judgments about as well as, or better than, an individual human rater.
  • Models rated faces more favorably and within a narrower range than humans, so their absolute ratings did not match human ratings.
  • Only face age predicted attractiveness in both humans and MLLMs; models did not replicate humans’ penalty against male faces.
  • MLLMs may serve as coarse proxies for relative attractiveness rankings, but not as substitutes for human ratings on an absolute scale.

View pre-registration & data 13

The supplied passages are bibliography entries referencing prior work on image-text datasets, multimodal transformers, research-quality evaluation, and facial attractiveness.

  • One passage lists a Royal Society B article spanning pages 1913–1917.
  • One passage cites LAION-5B, an open large-scale dataset for training next-generation image-text models.
  • One passage cites work on brain encoding models based on multimodal transformers transferring across language and vision.
  • One passage lists references on ChatGPT research-quality evaluation and facial attractiveness.
Loading 2609.02512v1…