Source-linked AI summary

Beauty is in the ELBO of the Beholder: A Variational Account of Processing Fluency in Face Perception

Francisco M. López, Jochen Triesch

arXiv:2608.24219v1cs.CVq-bio.NC

TL;DR

The paper asks whether facial attractiveness reflects processing fluency and tests this without attractiveness supervision by training VAEs on four face datasets and evaluating 597 CFD faces. Human attractiveness aligned with ELBO geometry across datasets, while independently learned latent spaces shared an attractiveness direction and attractive faces were more prototypical. These findings support a variational interpretation of processing fluency in face aesthetics, within the study’s constrained models and viewing conditions.

  • Problem

    The paper investigates whether facial attractiveness, traditionally associated with statistical regularities, reflects the ease with which faces are represented.

  • Method

    The authors trained unsupervised convolutional VAEs on four face datasets and evaluated their rate-distortion and latent representations on 597 attractiveness-rated CFD faces.

  • Results

    Across datasets, human attractiveness aligned with ELBO direction in rate-distortion space, and independently learned latent spaces contained strongly conserved attractiveness directions.

  • Takeaways & Limitations

    The findings support interpreting facial attractiveness through processing fluency in unsupervised generative representations, alongside greater prototypicality of attractive faces.

  • Takeaways & Limitations

    The simple VAEs are not mechanistic models of human vision, and frontal neutral-expression CFD images limit generalization to naturalistic viewing, pose variability, and dynamic expressions.

Abstract

from arXiv · show

Facial attractiveness has been linked to statistical regularities such as symmetry and averageness, suggesting that beauty may depend on the ease with which a face is perceived. We empirically test this hypothesis by training variational autoencoders on four face datasets without attractiveness supervision and evaluating their representations on the 597 faces from the Chicago Face Database. Across models, human attractiveness ratings closely aligns with the direction defined by the VAE evidence lower bound (ELBO) in rate-distortion space. Independently learned latent spaces contain an attractiveness direction that transfers strongly across random initializations and training data. We also find that attractive faces are more prototypical in both shape and latent space. Our results connect classic accounts of aesthetics with learned generative models and provide empirical support for a variational interpretation of the processing fluency theory of aesthetic pleasure.

1. Introduction

Classic accounts associate facial attractiveness with symmetry, averageness, sexual dimorphism, and prototypicality, motivating the idea that beauty may reflect processing ease. This paper models processing fluency in generative face representations using the VAE ELBO without attractiveness supervision.

  • Facial attractiveness has been linked to symmetry, averageness, sexual dimorphism, and population-level prototypicality.
  • Processing fluency theory proposes that stimuli represented or retrieved more easily are typically evaluated more positively.
  • Prior computational accounts connect aesthetic value to efficient coding, predictability, uncertainty reduction, or probabilistic scalar models.
  • The paper addresses a mismatch between neuroaesthetic theories and detailed generative face models by treating the VAE ELBO as an explicit measure of processing fluency.
  • Across four training datasets and demographic groups, human attractiveness gradients match ELBO geometry, latent attractiveness directions align across models, and attractive faces are more prototypical.

2. A variational interpretation of processing fluency

Processing fluency can be represented variationally as the ease of explaining a face while balancing reconstruction fidelity against latent regularization. The ELBO combines these costs into a scalar measure predicting that higher fluency accompanies greater aesthetic pleasure.

  • Processing fluency theory links easier encoding, retrieval, or interpretation to more positive evaluations.
  • Probabilistic generative models frame perception as inference over latent causes that explain sensory observations.
  • VAEs learn an approximate posterior and generative model under a prior, minimizing negative ELBO because exact marginal likelihood is generally intractable.
  • The ELBO balances reconstruction distortion against the KL-based rate measuring posterior deviation from the prior; β controls this rate-distortion trade-off, with β = 1 as standard ELBO.
  • High processing fluency corresponds to high-fidelity explanation at low representational cost, yielding low distortion and low rate.

3. Methods

The study trained unsupervised convolutional VAEs on four diverse face datasets and evaluated their representations on 597 attractiveness-rated CFD faces. Images were standardized through a common landmark-based preprocessing pipeline, with analyses conditioned on demographic groups.

  • Convolutional VAEs were trained without attractiveness supervision on FairFace, FFHQ, CelebA, and UTKFace.
  • The 597 neutral-expression CFD faces with attractiveness ratings were analyzed within eight ethnicity-by-gender demographic groups.
  • Training and evaluation images underwent shared landmark-based translation, rotation, and uniform scaling to 224 × 224 while preserving relative facial morphology.
  • Each VAE compressed 3×224×224 RGB inputs through convolutional layers into a 128-dimensional latent representation with a standard Gaussian prior.
  • Three independently initialized baseline models were trained for each dataset, producing 12 baseline representations for cross-model analyses.
  • Figure 2 evaluates attractiveness–ELBO alignment using regression, rate-distortion landscapes, and directional comparisons across training datasets.

4. Results

Across independently trained VAEs, attractiveness consistently aligns with variational geometry: higher-rated faces occupy directions toward lower distortion and lower rate, while latent attractiveness directions transfer across models. Attractive faces are also more prototypical in landmark and latent spaces, although the latent direction captures additional predictable structure.

  • 4.1. Attractiveness aligns with VAE’s ELBO in rate-distortion space: Attractiveness correlated positively with ELBO across all four training datasets and all 96 within-group model combinations.Dataset-level correlations were r = 0.249 to 0.312, and every combination was positive with p < 0.001.
  • 4.1. Attractiveness aligns with VAE’s ELBO in rate-distortion space: Both lower distortion and lower rate were associated with higher attractiveness, indicating a joint preference direction in rate-distortion space.Distortion correlations ranged from r = −0.257 to −0.300, while rate correlations ranged from r = −0.117 to −0.256.
  • 4.1. Attractiveness aligns with VAE’s ELBO in rate-distortion space: 0.929–0.988 per-dataset cosine similarity showed close alignment between empirical attractiveness gradients and ELBO directions across trained models.Across 12 models, cosine similarity ranged from 0.905 to 0.997, corresponding to angular deviations of 4.4° to 25.2°.
  • 4.1. Attractiveness aligns with VAE’s ELBO in rate-distortion space: 0.896–0.998 mean cosine similarity across FairFace configurations showed that attractiveness–ELBO alignment persisted across latent sizes, bottlenecks, decoder scales, and random seeds.Mean attractiveness–ELBO correlations remained positive at 0.199 to 0.254, with p < 0.001.
  • 4.2. Attractiveness geometry is conserved across different latent spaces: Attractiveness directions were conserved across independently learned latent spaces, beyond the shared structure expected from aligned face manifolds.The comparison used 12 aligned latent spaces; permuted labels produced a lower mean pairwise direction cosine of 0.632 ± 0.038.
  • 4.3. Attractive faces have prototypical shapes and latent codes: Attractive faces were more prototypical in both landmark and latent spaces, but latent prototypicality remained distinct from landmark-space prototypicality.Shape prototypicality correlated with attractiveness at r = 0.268; latent correlations ranged from r = 0.281 ± .012 to 0.315 ± .002, while cross-space correlations ranged from r = 0.455 to 0.544.
  • 4.3. Attractive faces have prototypical shapes and latent codes: The latent attractiveness direction predicted held-out attractiveness better than either prototype measure, while adding both prototype measures yielded only moderately higher correlations.The latent direction alone reached r = 0.526 to 0.541; the full model reached r = 0.546 to 0.570.

5. Discussion

Across four datasets, attractiveness aligned with higher processing fluency in VAE representations, including lower reconstruction distortion and regularization cost. Latent-space transfer and prototypicality analyses extended this result, while the authors emphasize correlational and scope limitations.

  • Attractiveness increased along a higher-processing-fluency direction characterized by lower reconstruction distortion and lower regularization cost.
  • After latent-space alignment, attractiveness directions transferred across models and training datasets with little loss of predictive information.
  • Attractive faces were more prototypical in landmark and latent space, but prototypicality explained only part of the latent attractiveness geometry.
  • Both ELBO components correlated with attractiveness, but their rate-distortion trade-off aligned best with attractiveness ratings.Rate measures representation complexity, whereas distortion measures representation fidelity.
  • The correspondence does not imply that human observers explicitly optimize an ELBO; the model provides a computational measure of fluency in a learned generative representation.
  • The study’s conclusions are limited by simple VAEs, neutral frontal CFD images, coarse demographic categories, population-average ratings, and correlational evidence.

Appendix A. Extended methods

The study trained fully unsupervised face representations on four datasets selected for differing demographic distributions and image statistics. Dataset construction varied substantially in scale, coverage, and balance, while annotations were ignored during training.

  • Face representations were trained fully unsupervised on four datasets with different demographic distributions and image statistics.All available labels and metadata were ignored during training.
  • CelebA provided the largest training set, with 202,599 images and substantial variation in pose, appearance, and background, but implicit demographic biases.163,456 images were retained after preprocessing.
  • FairFace was constructed to reduce racial imbalance and included seven race categories, two gender categories, and nine age ranges.It contained 108,501 images, with 52,801 retained after preprocessing.
  • FFHQ contained 70,000 high-quality images designed for generative face modeling, with broad variation in age, ethnicity, and background.57,741 images were retained, and resizing to 224 × 224 pixels reduced its high-resolution advantage.
  • UTKFace contained more than 20,000 internet face images spanning ages 0–116 and five ethnicity categories, with annotations ignored during training.18,492 images remained after preprocessing.

A.1.2. Evaluation data

Evaluation used all 597 neutral-expression CFD faces, whose morphology and attractiveness ratings were available across eight ethnicity-by-gender groups. Images and scalar measures were standardized before analysis using landmark-based preprocessing and within-group normalization.

  • The CFD evaluation set contained 597 individuals across eight ethnicity×gender groups, and all retained images were neutral-expression faces.The dataset also provided objective facial morphology measurements and subjective attractiveness ratings.
  • CFD images and attractiveness ratings were excluded from VAE training and hyperparameter selection, while demographic differences in ratings motivated standardization.Female faces yielded significantly higher attractiveness scores.
  • The example pipeline progresses from the original FFHQ image through landmark detection and geometric alignment to a masked 224 × 224 input.
  • Each scalar quantity was transformed using the mean and standard deviation of faces in its corresponding demographic group.
  • The preprocessing pipeline detected 468 facial landmarks, rejected images with missing required geometry or non-frontal faces, and applied translation, rotation, and uniform scaling.

A.3. VAE architecture

The study uses convolutional VAEs to encode faces into probabilistic latent representations and evaluates attractiveness through rate-distortion geometry, latent directions, prototypicality, and reconstruction fidelity.

  • VAE architecture: The convolutional VAE processes 224 × 224 RGB faces through convolutional layers into a latent representation with baseline dimensionality dz = 128.The architecture mirrors the encoder in the decoder and uses a standard Gaussian prior.
  • VAE architecture: Baseline experiments trained three independently initialized models on each of four datasets, producing 12 representations for analysis.Additional FairFace experiments varied latent dimensionality, β, or decoder scale σ.
  • Rate-distortion analysis: For each CFD face, the encoder provides a posterior distribution, from which rate and Monte Carlo-estimated distortion are computed.Distortion uses 16 independent posterior samples.
  • Analyses: Attractiveness directions are estimated from standardized within-group rate and distortion coordinates and compared with the ELBO direction using cosine similarity.The same representations support latent-space alignment, prototype analyses, reconstruction-shape comparisons, and held-out predictive tests.
  • Analyses: The latent-space analysis aligns centered representations with generalized orthogonal Procrustes analysis before comparing attractiveness directions across models.Predictive analyses use training identities to estimate prototypes, normalization parameters, and directions before evaluating held-out identities.

B.1. Robustness across hyperparameters

The attractiveness–ELBO relationship remains positive across broad variations in latent dimensionality, regularization, and decoder scale, although alignment varies with hyperparameters.

  • Hyperparameter robustness: Pearson correlations between human ratings and ELBO ranged from r = 0.190 to r = 0.262, all p < 0.001, across individual models.This indicates significant correspondence under every tested configuration.
  • Hyperparameter robustness: The lowest average correlation occurred at β = 4, with r = 0.199 ± 0.015, possibly reflecting over-regularization and degraded reconstructions.The attractiveness–processing-fluency correspondence was retained despite this reduction.
  • Hyperparameter robustness: The highest direction alignment was 0.998 ± 0.003 at σ = 1/2, where attractiveness and ELBO vectors were virtually collinear.Figure 5 reports Pearson correlations in panel A and rate-distortion direction cosine similarity in panel B.

B.2. Robustness across demographic groups

Attractiveness–ELBO correlations were positive across all demographic-group comparisons, but their magnitudes differed substantially among ethnicity×gender groups.

  • Demographic robustness: All 96 within-group ELBO–attractiveness comparisons were positive, with mean correlation r = 0.286 ± .The comparisons covered 12 baseline models and eight ethnicity×gender groups.
  • Demographic robustness: Figure 6 organizes demographic-group correlation results by training-dataset seed, enabling comparison of robustness across groups and independent model initializations.The caption specifies that each column represents a seed within a training dataset.
  • Demographic robustness: The strongest within-group correlations occurred for Latino-Male at r = 0.431 and White-Female at r = 0.408.These values were higher than those reported for the lowest-correlation groups.
  • Demographic robustness: The weakest within-group correlations occurred for Latino-Female at r = 0.171 and White-Male at r = 0.178.The authors speculate that training and evaluation data imbalances may contribute to these differences.
  • Demographic robustness: Landmark preservation was moderately associated with attractiveness in all 12 models, indicating that reconstruction fidelity covaried with attractiveness.The analysis used Procrustes discrepancy between original and reconstructed facial landmarks.

B.4. Objective and subjective facial properties

Human and VAE assessments agree substantially on objective facial properties, but diverge more for subjective social and affective judgments. The combined properties predict human attractiveness more accurately than ELBO.

  • Objective properties: Objective facial properties showed substantial agreement between human attractiveness judgments and VAE ELBO scores.The correlation between their correlation profiles was r = 0.760.
  • Objective properties: Upper-face ratio correlated positively with human attractiveness at r = 0.297 and ELBO at r = 0.306.
  • Subjective properties: Trustworthiness correlated with attractiveness at r = 0.559 but with ELBO at only r = 0.209.
  • Subjective properties: Subjective properties predicted human attractiveness more accurately than ELBO, with out-of-fold correlations of r = 0.704 and r = 0.262, respectively.
  • Combined properties: Combining objective and subjective properties increased prediction accuracy to r = 0.765 for human ratings and r = 0.523 for ELBO.
  • Interpretation: The results support a variational account in which VAEs capture facial statistical regularities, while unsupervised learning leaves social and affective judgments unavailable.
Loading 2608.24219v1…