Source-linked AI summary

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

Jingjing Chang, Yixiao Fang, Peng Xing, Shuhan Wu, Wei Cheng, Rui Wang, Xianfang Zeng, Gang Yu, Hai-Bao Chen

arXiv:2506.07977v3cs.CV

TL;DR

Existing T2I benchmarks do not comprehensively evaluate capabilities such as reasoning, text rendering, style, and diversity. OneIG-Bench addresses this gap with a multidimensional benchmark and scenario-specific metrics, enabling granular comparisons and targeted evaluation, while its reasoning and aesthetic-related metrics retain acknowledged limitations.

  • Problem

    Existing T2I evaluation systems provide limited coverage of reasoning, text rendering, style, diversity, and other multidimensional capabilities despite rapid model advances.

  • Method

    OneIG-Bench organizes over 1000 prompts into six dimensions and applies customized quantitative metrics, including prompt-alignment, text-rendering, reasoning, stylization, and diversity measures.

  • Results

    OneIG-Bench provides a framework that identifies performance bottlenecks and strengths of current T2I models through systematic multidimensional comparison.

  • Takeaways & Limitations

    The benchmark supports nuanced cross-model analysis and allows users to evaluate only the prompt subset associated with a selected dimension.

  • Takeaways & Limitations

    Knowledge-and-reasoning evaluation may admit more rational approaches, while aesthetic and body-quality assessment models show bias, limited discriminative power, or limited generalizability.

Abstract

from arXiv · show

Text-to-image (T2I) models have garnered significant attention for generating high-quality images aligned with text prompts. However, rapid T2I model advancements reveal limitations in early benchmarks, lacking comprehensive evaluations, for example, the evaluation on reasoning, text rendering and style. Notably, recent state-of-the-art models, with their rich knowledge modeling capabilities, show promising results on the image generation problems requiring strong reasoning ability, yet existing evaluation systems have not adequately addressed this frontier. To systematically address these gaps, we introduce OneIG-Bench, a meticulously designed comprehensive benchmark framework for fine-grained evaluation of T2I models across multiple dimensions, including prompt-image alignment, text rendering precision, reasoning-generated content, stylization, and diversity. By structuring the evaluation, this benchmark enables in-depth analysis of model performance, helping researchers and practitioners pinpoint strengths and bottlenecks in the full pipeline of image generation. Specifically, OneIG-Bench enables flexible evaluation by allowing users to focus on a particular evaluation subset. Instead of generating images for the entire set of prompts, users can generate images only for the prompts associated with the selected dimension and complete the corresponding evaluation accordingly. Our codebase and dataset are now publicly available to facilitate reproducible evaluation studies and cross-model comparisons within the T2I research community.

1 Introduction

OneIG-Bench addresses the limited multidimensional coverage of existing T2I evaluations by organizing over 1000 prompts into six assessment categories. It provides standardized quantitative evaluation and supports dimension-specific model assessment.

  • Existing T2I benchmarks often assess single dimensions and inadequately cover reasoning, text rendering, and style.
  • OneIG-Bench contains over 1000 prompts organized into six categories spanning general objects, portraits, anime and stylization, text rendering, knowledge and reasoning, and multilingualism.
  • The benchmark provides standardized quantitative metrics for objective capability ranking and direct cross-model comparison.
  • Users can evaluate a selected dimension by generating images only for prompts associated with that evaluation subset.
  • The benchmark evaluates both open-source and proprietary image-generation methods.

2 Related Works

T2I evaluation has evolved from traditional image-quality metrics and core-prompt restoration toward broader assessment needs. Existing approaches still provide limited coverage of higher-order semantics and the rapidly changing capabilities of newer models.

  • GAN-based T2I methods often suffered from mode collapse and struggled to generate even simple subject matter accurately.
  • Traditional metrics such as FID, SSIM, and PSNR do not comprehensively capture T2I capabilities or higher-order semantic understanding.
  • Many existing benchmarks focus on restoring core prompt elements, making comprehensive measurement of evolving model performance difficult.

3 Benchmark

OneIG-Bench broadens T2I evaluation across six dimensions, diverse prompt formats, and scenario-specific automated metrics. Its construction pipeline combines real-world sourcing, balanced clustering, deduplication, prompt rewriting, complexity control, and manual review.

  • 3.1 Benchmark Overview: Existing evaluations centered on compositional content do not adequately cover broader language understanding and other quantitative image-assessment dimensions.
  • 3.1 Benchmark Overview: OneIG-Bench evaluates scene coverage, prompt diversity, evaluation content, multilingualism, automation, and leaderboard availability.
  • 3.1 Benchmark Overview: The benchmark covers style generation, text rendering, reasoning-based drawing, and long, short, tag-based, phrase-based, and natural-language prompts.
  • 3.1 Benchmark Overview: OneIG-Bench contains 2,440 prompts across English and Chinese subsets, including an additional Multilingualism category in the Chinese subset.
  • 3.2 Benchmark Construction: Prompt construction begins with real-world-oriented data curation, then balances scenes through clustering and removes redundant prompts using cosine similarity to cluster-center embeddings.
  • 3.2 Benchmark Construction: LLM rewriting and three word-length bands structure prompts for analysis across concise, mid-complexity, and elaborate text scenarios.
  • 3.2 Benchmark Construction: Manual review removes sensitive or semantically conflicting prompts to support reliable and fair evaluation.
  • 3.2 Benchmark Construction: The construction pipeline enables granular assessment of model performance across multiple dimensions.

4 Evaluation

OneIG-Bench evaluates text-to-image models across semantic alignment, text rendering, diversity, stylization, and knowledge-driven reasoning using specialized metrics and broad model comparisons. Results show distinct strengths across models, with closed-source systems generally strongest in overall and reasoning performance, while specific models excel in stylization, diversity, or text-rendering submetrics.

  • Metrics: Semantic alignment uses question dependency graphs and vision-language model answers to score prompt-image matching across objects, portraits, and stylization.GPT-4o generates questions covering overall information, spatial relationships, and object attributes; Qwen2.5-VL-7B answers them, and prompt scores average validated question results.
  • Metrics: Diversity is measured by averaging pairwise cosine similarities among images generated from the same prompt, then aggregating across prompts.DreamSim is used to compute the image cosine similarity underlying the diversity calculation.
  • Results: Imagen4, GPT-4o, and Imagen3 lead alignment, while natural-language prompts generally achieve higher alignment accuracy than tag-based or phrase-based prompts.Longer prompts tend to reduce alignment scores, whereas models incorporating T5 or other large language models appear more robust to long prompts.
  • Results: Seedream 3.0 performs best across nearly all text-rendering completion and word-accuracy subdimensions, while Recraft V3 excels in edit distance, especially for long prompts.Recraft V3’s layout-first strategy is associated with fewer severe rendering errors, but its completion ratio and word accuracy are less impressive.
  • Results: GPT-4o substantially outperforms other models in knowledge retention and reasoning, while closed-source models generally exceed open-source models in this dimension.Reasoning performance is grouped into five tiers, with GPT-4o and Imagen4 in the first tier and other models distributed across lower tiers.

5 Conclusion

OneIG-Bench provides an omni-dimensional framework for nuanced text-to-image evaluation by categorizing generation themes and tailoring metrics to different scenarios. The authors also identify unresolved limitations in reasoning evaluation and assessment-model robustness.

  • OneIG-Bench categorizes generation themes and applies scenario-specific metrics for comprehensive text-to-image evaluation.The framework covers general scenarios, human figures, conventional objects, text rendering, and anime/style scenarios.
  • The benchmark supports comparative analysis of model strengths and limitations and helps identify technical bottlenecks.
  • Knowledge and reasoning evaluation remains limited because existing image-generation models often lack robust reasoning capabilities.The authors report that their metric rankings align closely with human evaluations, while acknowledging that better evaluation approaches may exist.
  • Aesthetic models may exhibit unexpected biases, while body-quality assessment models may lack discriminative power and generalizability.

A.1 Word Count Distribution Statistics

OneIG-Bench prompts are organized into short, medium, and long length categories, generally following an approximately 1:2:1 distribution. Portrait prompts deviate slightly because they explicitly require non-anime human figures, increasing their average length.

  • The overall prompt-length distribution follows an approximate Short:Middle:Long ratio of 1:2:1.Short, Medium, and Long denote fewer than 30, 30–60, and more than 60 words, respectively.
  • Short prompts contain fewer than 30 words, Medium prompts contain 30–60 words, and Long prompts exceed 60 words.
  • Portrait prompts depart slightly from the 1:2:1 distribution because their explicit requirements increase average prompt length.The requirements exclude stylized figures such as anime characters.
  • The prompt word-count distribution ranges from 0 to 200 words.
  • Figure 6 presents prompt word-count distributions across the Short, Middle, and Long categories.

A.2 Implementation

Experiments use method-specific default CFG and step settings, with selected adjustments for consistency and image quality. Configuration tables document parameter settings for unified multimodal and open-source methods and release dates for closed-source methods.

  • CFG and inference-step parameters follow each method’s default settings, except where consistency or image quality motivates an adjustment.Stable Diffusion 3.5 Large uses 50 rather than 40 default steps, while Flux.1-dev matches the official API setting.
  • Table 8 records method size, guidance scale, image resolution, and inference-step configurations.
  • Closed-source methods are aligned with experimental results using their corresponding release or update dates.These dates are listed in Table 9.

A.3 The Details on Prompts Rewriting

The prompt-rewriting procedure reshapes an initial prompt set to a target length distribution while preserving each prompt’s intent and improving its specificity. It sorts prompts, samples target lengths, pairs them, and calls GPT-4o to produce rewritten prompts.

  • Algorithm 1 sorts initial prompts by word count, samples target lengths, pairs prompts with those lengths, and rewrites them through GPT-4o.The sampled lengths use a Beta distribution with parameters (2.37, 2.86).
  • The target-length sampling distribution approximately follows a 0–0.3:0.3–0.6:0.6–1 ratio of 1:2:1.
  • Prompts longer than the target are shortened by removing non-essential details, while shorter prompts are expanded with specific and vivid details.
  • The rewriting instructions require coherent, natural, fluent, logically structured prompts that preserve the initial tone and intent.
  • The rewriting assistant outputs only the final rewritten prompt without additional words.

A.4 The Details and Analysis on Stylization

OneIG-Bench organizes Anime & Stylization into Traditional, Media, and Anime subcategories, covering classical art movements, material-based media, and stylized anime aesthetics.

  • Style categories: Table 10 lists the styles corresponding to the benchmark’s specific style categories.
  • Style categories: Anime & Stylization is divided into Traditional, Media, and Anime subcategories.Traditional covers classical and historical art movements; Media covers artistic media and material techniques; Anime covers stylized visual aesthetics.

A.5 Visualization Results

Visualization results compare state-of-the-art models across alignment, text rendering, reasoning, diversity, and stylization. The comparisons reveal distinct strengths and failure modes across these capabilities.

  • Overall comparison: Closed-source methods outperform the other method categories overall in the normalized polar visualization.The visualization is presented in Table 12, with representative examples selected largely by overall performance.
  • Alignment: Imagen4 and GPT-4o demonstrate strong semantic alignment, while many methods confuse subject-level attributes in multi-person prompts.Some methods also overlook fine-grained details while completing the primary generation task.
  • Text rendering: Seedream 3.0 demonstrates high text-rendering accuracy and aesthetic quality, whereas competing methods show case, word-level, typography, or visual-quality limitations.GPT-4o can miss case sensitivity; Recraft V3 can make word-level and layout errors; Imagen4 has good accuracy but relatively poor overall visual quality.
  • Reasoning: Only GPT-4o demonstrates both logical coherence and textual accuracy in the reasoning visualization.Other methods provide varying degrees of readable, limited, redundant, or incorrect content.
  • Diversity: Diversity-score rankings align well with visual inspection, although some observed diversity may partly reflect insufficient alignment.
  • Stylization: GPT-4o performs well across most styles but struggles with some anime styles, while unified multimodal methods show particularly strong anime-style performance.Stable Diffusion 1.5 captures traditional style features despite lower visual quality, and Imagen4 performs notably well in media and anime styles.

GPT-4o

The GPT-4o visualizations show its evaluation across alignment, text rendering, reasoning, diversity, and multiple stylization categories. The figures pair model outputs with prompts and metric scores for these capabilities.

  • Alignment: Figure 7 evaluates alignment using Imagen4, GPT-4o, Imagen3, HiDream-I1-Full, and Kolors 2.0.The figure includes tag/phrase and short prompts, with variation influenced by specified prompt elements.
  • Overall comparison: Table 12 presents the best normalized polar visualization of state-of-the-art methods on OneIG-Bench.
  • Metrics: The visualization materials report ED, CR, and WAC values alongside generated examples.One displayed set includes ED:0, CR:1, and WAC:1.00 for several examples, with another showing ED:24, CR:0, and WAC:0.25.
  • Prompt examples: The text-rendering examples include prompts requiring centered captions, technical diagrams, and detailed stylized scenes.Displayed examples include a rainy-street caption, Bernoulli’s-principle diagram, minimalism, clay style, cyberpunk, and 3d rendering.
  • Text rendering: Figure 8 evaluates text rendering across short, middle, and long prompts for five state-of-the-art methods.Seedream 3.0, GPT-4o, Imagen4, Recraft V3, and HiDream-I1-Full are compared.
  • Reasoning: Figure 9 evaluates reasoning through coral-reef formation, series and parallel circuits, and butanone and butane-2,3-diol diagrams.GPT-4o is compared with Imagen4, Recraft V3, HiDream-I1-Full, and Imagen3.
  • Diversity: Figure 10 evaluates diversity for Stable Diffusion 1.5, Janus-Pro, Kolors 2.0, Stable Diffusion XL, and Seedream 3.0.
  • Stylization: Figures 11–13 evaluate traditional, media, and anime styles across different model sets and style prompts.The visualizations include pointillism, minimalism, pencil sketch, stone sculpture, cyberpunk, and 3d rendering.
Loading 2506.07977v3…