Source-linked AI summary

Holistic Evaluation of Text-To-Image Models

Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Benita Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei, Jiajun Wu, Stefano Ermon, Percy Liang

arXiv:2311.04287v1cs.CVcs.LG

TL;DR

Existing evaluations provide limited evidence about the capabilities and risks of increasingly prevalent text-to-image models. HEIM addresses this gap with a 12-aspect benchmark spanning 62 scenarios and 26 models, combining human and automated metrics. The results show that no single model excels across all aspects, and different models exhibit different strengths.

  • Problem

    The central gap is a comprehensive quantitative understanding of text-to-image models’ capabilities and risks despite their rapid proliferation and adoption.

  • Method

    HEIM evaluates 12 aspects across 62 scenarios and 26 models using standardized comparisons that combine crowdsourced human and automated metrics.

  • Results

    No single model excels in all aspects; different models show different strengths across general, creative, societal, and multilingual evaluation dimensions.

  • Takeaways & Limitations

    The benchmark provides holistic insights for comparing models and identifies reasoning, multilinguality, originality, toxicity, and bias as areas needing further attention.

  • Takeaways & Limitations

    Crowdsourced judgments for aesthetics and originality may vary more than judgments for alignment, photorealism, and subject clarity, so strong conclusions based solely on these metrics are avoided.

Abstract

from arXiv · show

The stunning qualitative improvement of recent text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (HEIM). Whereas previous evaluations focus mostly on text-image alignment and image quality, we identify 12 aspects, including text-image alignment, image quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency. We curate 62 scenarios encompassing these aspects and evaluate 26 state-of-the-art text-to-image models on this benchmark. Our results reveal that no single model excels in all aspects, with different models demonstrating different strengths. We release the generated images and human evaluation results for full transparency at https://crfm.stanford.edu/heim/v1.1.0 and the code at https://github.com/stanford-crfm/helm, which is integrated with the HELM codebase.

1 Introduction

HEIM addresses the limited and potentially incomplete evaluation of text-to-image models with a comprehensive benchmark spanning capabilities, risks, human judgment, and standardized model comparisons. Its findings show that models have different strengths rather than one model excelling across all aspects.

  • Recent text-to-image models have proliferated, gained broad applications, and attracted substantial adoption, while their full capabilities and risks remain insufficiently understood.The cited passage identifies both technical and safety or ethical gaps in current understanding.
  • HEIM evaluates 12 aspects across 62 scenarios using 25 metrics, including both crowdsourced human and automated evaluations.The benchmark covers alignment, quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency.
  • HEIM uniformly evaluates 26 recent accessible text-to-image models across all aspects to enable standardized comparisons.The evaluation targets models accessible as of July 2023.
  • No single model excels in all aspects; different models show different strengths, while human and automated metrics generally correlate weakly, especially for photorealism and aesthetics.The findings also identify poor performance in reasoning and multilinguality and continuing imperfections in originality, toxicity, and bias.
  • The authors release the evaluation pipeline, code, generated images, and human evaluation results for transparency and reproducibility.The framework is extensible to additional aspects, scenarios, models, adaptations, and metrics.

2 Core framework

HEIM decomposes text-to-image evaluation into standardized components that connect evaluative dimensions and use cases with model execution procedures and quality metrics. The framework emphasizes consistent evaluation conditions and combines zero-shot model runs with human and automated assessment.

  • HEIM decomposes each evaluation run into an aspect, scenario, adaptation, and metric.These components represent the evaluative dimension, specific use case, model-running procedure, and quality measurement.
  • Each aspect is evaluated through a scenario-metric pair, allowing multiple characteristics of generated images to be assessed.Examples of aspects include image quality, originality, and bias.
  • Scenarios represent use cases through textual inputs and optional reference images, covering domains such as common-object descriptions and logo design.A scenario may contain multiple instances and can include sub-scenarios representing category variations.
  • The benchmark focuses on zero-shot prompting while also exploring prompt engineering methods that refine inputs before generation.Other adaptation strategies include few-shot prompting and finetuning, but zero-shot prompting is the primary focus.
  • Metrics may be human or automated, enabling HEIM to combine subjective judgments with objective measurements.Examples include human alignment ratings on a 1–5 scale and CLIPScore.

3 Aspects

HEIM broadens text-to-image evaluation beyond alignment and quality to include creative, cognitive, societal, linguistic, robustness, and efficiency dimensions. These aspects are operationalized through targeted scenarios and metrics, including human judgments and performance changes under input variations.

  • HEIM evaluates 12 aspects, including alignment, quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency.The aspects are presented as dimensions crucial for deploying text-to-image models.
  • Alignment and quality use established automated metrics alongside human-rated alignment and photorealism measures.Examples include CLIPScore, FID, Inception Score, human-rated alignment, and human-rated photorealism.
  • Aesthetics and originality extend evaluation toward visual-art applications, with originality motivated partly by copyright infringement concerns.Art-generation scenarios include oil painting, vector graphics, landing-page, and logo-design prompts.
  • Knowledge and reasoning assess entity-specific understanding and visual composition, respectively, using CLIPScore and human-rated alignment.Historical Figures scenarios target knowledge, while PaintSkills targets reasoning about visual composition.
  • Toxicity, bias, fairness, multilinguality, and robustness address inappropriate content, demographic associations, and performance across social, linguistic, and input variations.Modified captions introduce gender or dialect changes, translations, and typos or misspellings; performance changes are measured against unmodified captions.
  • Efficiency is evaluated through inference time because it affects model usability and can be assessed with any scenario.Efficiency is treated as a general aspect rather than one tied to a specific scenario.

4 Scenarios

HEIM evaluates 12 aspects through 62 diverse, practical scenarios, including existing datasets and newly created prompts for previously underexplored areas.

  • Scenario Coverage: 62 scenarios cover the 12 evaluation aspects using textual inputs, sub-scenarios, existing datasets, and newly created prompts.Examples include MS-COCO for alignment, quality, and efficiency, and I2P for toxicity.
  • Scenario Design: Scenarios represent specific use cases through sets of textual inputs and, optionally, reference output images.
  • New Scenarios: New scenarios target originality, aesthetics, bias, and fairness, which previously lacked dedicated datasets.Originality scenarios test artistic creativity through landing pages, logos, and magazine covers.

5 Metrics

HEIM combines human and automated metrics to evaluate 12 aspects with broader, more realistic coverage, including newly introduced measures for previously neglected dimensions.

  • Metric Coverage: HEIM uses human and automated metrics to assess the quality of image generations across 12 aspects.Human metrics capture subjective judgment, while automated metrics provide objective assessments.
  • Human Evaluation: At least 5 participants evaluate each human-rated image, with at least 100 image samples used for each aspect.Human evaluation follows a crowdsourcing methodology with concrete English definitions for questions and rating choices.
  • Human Evaluation: Human metrics cover alignment, photorealism, aesthetics, and originality, supporting evaluation of several additional aspects.These metrics are also used for assessing quality, knowledge, reasoning, fairness, robustness, and multilinguality.
  • New Metrics: New metrics address fairness, robustness, multilinguality, and efficiency, which received limited attention in prior evaluations.

6 Models

HEIM evaluates 26 recent text-to-image models spanning different architectures, parameter scales, organizations, and accessibility conditions.

  • Model Coverage: 26 recent text-to-image models include diffusion, autoregressive, and GAN models ranging from 0.4B to 13B parameters.The evaluation covers models from different organizations and includes both open and closed models.

7 Experiments and results

HEIM evaluates 26 text-to-image models across 12 aspects using 62 scenarios and 25 metrics, revealing distinct strengths, persistent weaknesses, and trade-offs across models. Results span alignment, realism, aesthetics, originality, reasoning, knowledge, societal risks, multilinguality, efficiency, and evaluation methodology.

  • Evaluation setup: 26 models were evaluated across 12 aspects, using 62 scenarios and 25 metrics.The benchmark reports a model win rate as the probability of outperforming a uniformly random opponent in head-to-head comparisons for a metric.
  • Core visual capabilities: DALL-E 2 leads human-rated alignment, while Dreamlike Photoreal 2.0 achieves high photorealism and Promptist achieves the highest human-evaluated aesthetics win rate.Openjourney performs strongly on aesthetics but tends to score lower on photorealism; Promptist plus Stable Diffusion v1-4 improves human-rated aesthetics with comparable alignment.
  • Originality and societal aspects: GigaGAN virtually never generates watermarks, while Openjourney and Dreamlike Diffusion 1.0 achieve the highest human-rated originality win rates at 86% and 82%.CogView2 exhibits the highest frequency of watermark generation; both high-originality models are Stable Diffusion variants fine-tuned on high-quality art images.
  • Reasoning and knowledge: DALL-E 2 reaches only 47.2% object-detection accuracy on PaintSkills, and all models perform poorly on reasoning involving counts and spatial relations.DALL-E 2 also scores below 4 for Relational Understanding and reasoning sub-scenarios of DrawBench, while DeepFloyd-IF XL remains below 4 across all reasoning scenarios.
  • Evaluation methodology: Human and automated metrics correlate weakly, with coefficients of 0.42 for alignment, 0.59 for image quality, and 0.39 for aesthetics.The weak correlations are particularly pronounced for aesthetics, supporting the use of human ratings in evaluating image-generation models.

8 Related work

Existing image-generation benchmarks largely assess quality and text-image alignment, while human evaluations cover limited aspects. HEIM extends evaluation toward broader capabilities and ethical and societal dimensions.

  • Holistic benchmarking: Holistic and meta-benchmarking approaches in NLP have supported comprehensive assessments across multiple scenarios or tasks.The related-work discussion positions this broader evaluation strategy as a precedent for image generation.
  • Benchmarks for image generation: Existing benchmarks primarily evaluate image quality and text-image alignment using automated metrics.Common metrics include FID, Inception Score, and CLIPScore.
  • Human evaluation: Crowdsourced human evaluations have expanded perception-based assessment but remained focused mainly on alignment and quality.HEIM extends this approach to aesthetics, originality, reasoning, and fairness.
  • Ethical and societal evaluation: Ethical and societal evaluations had covered only a select few models, leaving most models unevaluated on these dimensions.HEIM standardizes evaluation across models and includes ethical and societal aspects.
  • Art and design: HEIM incorporates aesthetic evaluation and design principles, including composition, color harmony, balance, clarity, legibility, hierarchy, and consistency.These criteria are used to assess whether generated images are visually pleasing and thoughtfully composed.

9 Conclusion

The paper introduces HEIM, a benchmark for evaluating text-to-image models across 12 aspects, and applies it to 26 recent models. Results show differing model strengths rather than one universally superior model, with released artifacts supporting transparency and reproducibility.

  • Conclusion: HEIM evaluates 26 recent text-to-image models across 12 aspects, including alignment, quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency.The benchmark is designed to assess multiple dimensions of text-to-image generation.
  • Conclusion: Different models excel in different aspects, opening research avenues for developing models that perform well across multiple dimensions.The conclusion does not identify a single model as best overall.
  • Conclusion: The authors release the evaluation pipeline, generated images, and human evaluation results to enhance transparency and reproducibility.The benchmark is presented as a resource for future model development and evaluation.

10 Limitations

The authors identify limitations in HEIM’s aspect coverage, metrics, and crowdsourced human evaluations. These constraints bound how comprehensively and conclusively the benchmark can characterize text-to-image models.

  • Scope: The 12 identified aspects may not be exhaustive, leaving potentially important dimensions of text-to-image generation unconsidered.The authors describe this as an ongoing research area.
  • Metrics: Current metrics may omit relevant demographic factors, and wall-clock time is only a surrogate for actual energy consumption.The bias evaluation focuses on binary gender and skin tone representations.
  • Human evaluation: Crowdsourced workers can evaluate alignment, photorealism, and subject clarity effectively, but aesthetics and originality show greater response variance.These latter judgments may differ between the general public and professional artists or legal experts.
  • Human evaluation: The authors refrain from drawing strong conclusions based solely on crowdsourced aesthetics and originality metrics.Those metrics rely on subjective judgments and may not represent expert opinions.
  • Human evaluation: General-public judgments remain relevant because generated images are intended to be visually pleasing and original to a wide audience.This explains why the authors retain value in considering crowdsourced opinions despite their variance.

A Datasheet

The HEIM datasheet describes a benchmark of prompts and generated images spanning 62 scenarios, with releases covering 26 models and multiple evaluation resources. It also documents sensitive-content, privacy, and human-evaluation considerations.

  • A.1 Motivation: HEIM was created to evaluate text-to-image models across 12 aspects relevant to real-world deployment, beyond the previously typical focus on alignment and quality.The listed aspects include aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency.
  • Dataset composition: The benchmark provides prompts or captions covering 62 scenarios and releases images generated by 26 text-to-image models.The benchmark contains 500K prompts in total.
  • Dataset composition: Each dataset instance consists of an input prompt and generated images; the MS-COCO scenario additionally includes a reference image for every prompt.Other scenarios do not have reference images.
  • Safety and privacy: Released images may contain sexually explicit, racist, abusive, or otherwise disturbing content, and the website displays warnings for scenarios likely to produce inappropriate images.The authors call for discretion and responsible research use.
  • Safety and privacy: People may be identifiable in generated images through face recognition or through associated prompts.The datasheet therefore records a direct privacy-related risk associated with the released data.
  • Scenario construction: Scenarios include artist-sourced prompts from DailyDall.e and 50 magazine-cover prompts designed around headline text.These additions broaden coverage toward artistic image generation and design-oriented tasks.

C.2 Automated metrics

HEIM combines automated metrics spanning alignment, image quality, aesthetics, safety, bias, fairness, robustness, multilinguality, efficiency, and model capabilities. The evaluation uses both standard image-quality measures and task-specific procedures across multiple model families.

  • CLIPScore measures alignment between generated images and corresponding natural-language descriptions using a pretrained CLIP model.
  • FID compares generated images with reference images using InceptionNet features, based on 30,000 MS-COCO prompts and 512×512 images.
  • IS, LAION Aesthetics, and the fractal coefficient assess image quality or visual complexity with established classifiers, predictors, and image-statistics procedures.
  • Safety and bias metrics detect watermarks, inappropriate content, nudity, blackouts, API rejections, gender representation, and skin-tone distributions.
  • Fairness, robustness, and multilinguality measure performance changes under social-group substitutions, semantic-preserving perturbations, and translations into Spanish, Chinese, and Hindi.
  • Efficiency is evaluated with raw and denoised inference runtime, while reasoning and knowledge include object recognition and counting accuracy.
  • The evaluated models include Stable Diffusion variants, DALL-E 2, Lexica Search, Dreamlike models, and other text-to-image systems with differing architectures and training histories.

E Human evaluation procedure

Human evaluation used Amazon Mechanical Turk to collect ratings of generated images through multiple-choice questions. The procedure required screened adult annotators, five annotators per sample, and a standardized survey interface.

  • Annotators had to be over 18, accept potentially offensive content, and hold AMT Master status.
  • Five different annotators were required for each sample.
  • Each multiple-choice question paid $0.02 at a stated hourly wage of $16, with total annotation spending of $13,433.55.
  • The Stanford Research Compliance Office determined that the protocol did not meet the regulatory definition of human subjects research.
  • MTurk workers rated generated images with multiple-choice questions about alignment, photorealism, aesthetics, and originality.
Loading 2311.04287v1…