Source-linked AI summary
UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
Yibin Wang, Zhimin Li, Yuhang Zang, Jiazi Bu, Yujie Zhou, Yi Xin, Junjun He, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang
TL;DR
UniGenBench++ addresses the limited prompt diversity, multilingual support, and coarse sub-dimension coverage of existing T2I benchmarks. It introduces a 600-prompt bilingual, length-varied benchmark with MLLM-assisted evaluation and an offline model, revealing strong style and world-knowledge performance but persistent reasoning weaknesses.
Problem
Existing T2I benchmarks lack diverse real-world scenarios and multilingual coverage while evaluating only limited semantic dimensions and sub-dimensions coarsely.
Method
UniGenBench++ uses 600 prompts across 5 themes, 20 subthemes, 10 primary dimensions, and 27 sub-dimensions in English and Chinese short and long forms, with Gemini-2.5-Pro-assisted evaluation and an offline model.
Results
Benchmarking shows strong performance on style and world knowledge but consistent difficulty with causal, contrastive, and other complex relational reasoning, especially for open-source models in grammar and action.
Takeaways & Limitations
The benchmark provides fine-grained, bilingual, length-varied diagnosis of semantic consistency across open- and closed-source T2I models.
Abstract
from arXiv · showhide
Recent progress in text-to-image (T2I) generation underscores the importance of reliable benchmarks in evaluating how accurately generated images reflect the semantics of their textual prompt. However, (1) existing benchmarks lack the diversity of prompt scenarios and multilingual support, both essential for real-world applicability; (2) they offer only coarse evaluations across primary dimensions, covering a narrow range of sub-dimensions, and fall short in fine-grained sub-dimension assessment. To address these limitations, we introduce UniGenBench++, a unified semantic assessment benchmark for T2I generation. Specifically, it comprises 600 prompts organized hierarchically to ensure both coverage and efficiency: (1) spans across diverse real-world scenarios, i.e., 5 main prompt themes and 20 subthemes; (2) comprehensively probes T2I models' semantic consistency over 10 primary and 27 sub evaluation criteria, with each prompt assessing multiple testpoints. To rigorously assess model robustness to variations in language and prompt length, we provide both English and Chinese versions of each prompt in short and long forms. Leveraging the general world knowledge and fine-grained image understanding capabilities of a closed-source Multi-modal Large Language Model (MLLM), i.e., Gemini-2.5-Pro, an effective pipeline is developed for reliable benchmark construction and streamlined model assessment. Moreover, to further facilitate community use, we train a robust evaluation model that enables offline assessment of T2I model outputs. Through comprehensive benchmarking of both open- and closed-sourced T2I models, we systematically reveal their strengths and weaknesses across various aspects.
I. Introduction
UniGenBench++ addresses limited prompt diversity, multilingual coverage, and coarse semantic evaluation in existing T2I benchmarks. It introduces a hierarchical bilingual benchmark and streamlined evaluation resources for fine-grained assessment.
- I. Introduction: Existing benchmarks provide coarse coverage of limited dimensions and omit multilingual, diverse real-world prompt scenarios.They also lack systematic fine-grained assessment of sub-dimensions such as grammar, action, relation-similarity, and inclusion.
- I. Introduction: 600 prompts span 5 real-world themes and 20 subthemes, with English and Chinese versions in short and long forms.Each prompt targets multiple test points to balance coverage and efficiency.
- I. Introduction: UniGenBench++ evaluates semantic consistency across 10 primary and 27 sub-dimensions, with each prompt assessing multiple explicit test points.The hierarchical design targets diagnostic detail without requiring a large prompt count.
- I. Introduction: Both open- and closed-source models perform strongly on style and world knowledge but struggle with causal, contrastive, and other complex relational reasoning.Open-source models additionally show larger fluctuations, particularly in grammar and action.
- I. Introduction: A streamlined point-wise pipeline supports consistent, fine-grained, interpretable judgments, while a dedicated offline model enables community assessment of T2I outputs.The paper also reports extensive bilingual and prompt-length-varied benchmarking across closed- and open-source models.
II. Related Work
Prior T2I benchmarks examine selected domains or prompt lengths but remain limited in semantic coverage, prompt diversity, and multilingual support. UniGenBench++ is positioned as a unified benchmark for fine-grained evaluation across these gaps.
- II. Related Work: Existing benchmarks often use short, repetitive prompts or focus on narrow domains, limiting semantic richness and expressiveness.Other work introduces dense prompts or matched short and long variants, but does not resolve the broader coverage gaps.
- II. Related Work: Prior benchmarks still provide coarse evaluation across limited dimensions and insufficient sub-dimension coverage.They also lack diverse prompt scenarios and multilingual support for real-world application settings.
- II. Related Work: UniGenBench++ covers 5 themes, 20 subthemes, 10 primary dimensions, and 27 sub-dimensions using only 600 prompts.Each prompt targets 1–10 explicit test points, balancing diagnostic coverage and efficiency.
- II. Related Work: The benchmark provides English and Chinese prompts in short and long forms and uses Gemini-2.5-Pro to streamline model evaluation.A trained offline evaluation model further supports community use.
B. Prompt Themes and Subject Categories
The benchmark organizes prompts across diverse themes and subject categories, then evaluates generated images through ten semantic dimensions and fine-grained test points. Short and long prompts differ in their distribution of evaluated attributes.
- B. Prompt Themes and Subject Categories: The themes include Creative Divergence, Art, Illustration, Film & Story, and Design, spanning imaginative, artistic, narrative, and commercial generation.Design includes advertising, e-commerce, spatial layouts, game and UI prototyping, posters, logos, fashion, and other resources.
- B. Prompt Themes and Subject Categories: Subject categories cover animals, objects, anthropomorphic characters, scenes, and atypical entities, allowing evaluation on common and unusual subjects.The categories are intended to reveal model strengths and weaknesses across entity types.
- C. Evaluation Dimensions: The benchmark decomposes major semantic dimensions into explicit sub-dimension test points because coarse metrics can mask fine-grained weaknesses.It organizes evaluation into 10 major categories, including style, world knowledge, attributes, compound concepts, action, layout, relationships, reasoning, grammar, and text generation.
- C. Evaluation Dimensions: Attribute evaluation covers quantity, expression, material, color, shape, and size, assessing object and scene characteristics.These criteria address counts, affect, surface properties, visual appearance, geometry, and relative dimensions.
- C. Evaluation Dimensions: Compound and action dimensions test concept integration, feature matching, physical and non-physical interaction, hand and full-body actions, state, and animal behavior.Together they assess both compositional integration and dynamic content.
- C. Evaluation Dimensions: Entity layout and relationship dimensions assess spatial arrangement, composition, similarity, comparison, and inclusion between objects.They cover both two- and three-dimensional layouts and semantic connections such as containment.
- C. Evaluation Dimensions: Logical reasoning, grammar, and text generation evaluate causality, contrast, language expressions, pronoun reference, consistency, negation, and prompt-aligned text.Long prompts tend to contain more attribute-related test points than short prompts.
D. Bilingual and Length-variant Prompt Construction
The benchmark constructs bilingual short prompts from sampled themes, subjects, and fine-grained testpoints, then expands them into aligned long prompts. This process preserves core semantics while updating testpoints to match added content.
- Bilingual Short Prompt Generation: Short prompts sample a theme, subject category, and 1–5 fine-grained testpoints, producing English and Chinese prompts with explicit testpoint descriptions.The generated outputs are bilingual prompt pairs plus descriptions explaining how each selected testpoint is instantiated.
- Expanded to Long Prompt: Long-prompt rewriting preserves themes, core subjects, and key attributes while allowing additional attribute, scene, and background details.
- Pipeline Overview: The construction and evaluation workflow is organized into benchmark construction, offline evaluation-model training, and evaluation cases.
- Expanded to Long Prompt: Testpoint alignment removes targets unsupported by expanded prompts and adds newly emerged targets, with at most five additions.The resulting long prompt remains paired with a dynamically updated, semantically coherent target set.
E. T2I Model Evaluation
The evaluation pipeline feeds generated images, prompts, and fine-grained testpoint descriptions to an MLLM that returns binary judgments and rationales. Scores aggregate hierarchically to support both detailed diagnosis and concise reporting.
- Point-wise Evaluation: Each evaluation instance pairs a generated image and prompt with testpoint descriptions, which the MLLM uses to produce judgments and explanatory outputs.
- Point-wise Evaluation: Rationales expose why testpoints are satisfied or violated, enabling failure-mode analysis beyond scalar correctness.
- Hierarchical Aggregation: Sub-dimension scores equal satisfied instances divided by total occurrences, while primary-dimension scores average their constituent sub-dimensions.
- Hierarchical Aggregation: Hierarchical aggregation combines fine-grained capability trends with holistic reporting and supports quantitative comparison alongside qualitative interpretability.
F. Offline Evaluation Model Training
The paper trains a local offline evaluator to reproduce proprietary MLLM assessment without external API calls. It learns both testpoint judgments and explanatory reasoning from MLLM-generated supervision.
- Offline Evaluator: The offline evaluator distills proprietary MLLM scoring behavior into a compact model executable locally without external API calls.
- Supervision Construction: Training targets consist of MLLM-generated judgments and rationales assembled for each image–prompt pair and testpoint description.
- Training Objective: A language-modeling objective trains the evaluator to learn binary judgments and explanatory reasoning from tokenized target sequences.
- Offline Evaluation: At evaluation time, the offline model follows the original proprietary-model workflow and produces decisions with explanatory rationales.
A. Implementation Details
The benchmark evaluates a broad set of open- and closed-source text-to-image models, spanning multiple contemporary model families.
- Evaluated Models: The evaluation includes numerous closed-source and open-source T2I models, covering diffusion, autoregressive, and other contemporary model families.
1) Benchmarking Models:
The benchmark uses approximately 375K evaluation samples from Gemini-2.5-Pro to train and evaluate an offline assessment model.
- Approximately 375K evaluation samples were collected from Gemini-2.5-Pro for offline evaluator development.
- The evaluation model uses UnifiedReward-2.0qwen-72b as its base model.
- 300K samples are used for training and 75K are reserved for evaluation.
2) Offline Evaluation Model:
Across prompt languages and lengths, the benchmark reveals distinct strengths and weaknesses among closed- and open-source T2I models. Closed-source systems remain strongest overall, while leading open-source models are competitive on several visual-understanding dimensions but lag in reasoning and consistency.
- GPT-4o-1.5 achieves the best overall performance among closed-source models across nearly all dimensions.Nano Banana Pro is also competitive and balanced, particularly in grammar and logical reasoning.
- FLUX.2-dev is the strongest open-source model under English short prompts, with FLUX.2-Klein variants and Qwen-Image forming the next tier.The FLUX.2-Klein family is particularly strong in relation modeling and grammar consistency.
- Closed-source models remain ahead at the frontier, especially on grammar consistency, logical reasoning, and robust compositional generation.The strongest open-source model can match or surpass several mid-tier closed-source models.
- GPT-4o-1.5 reaches the best overall performance for English long prompts, while Nano Banana Pro and GPT-4o remain top-tier.These models show particularly strong logical reasoning and grammar consistency.
- GPT-4o-1.5 achieves the strongest overall performance for Chinese short prompts, followed by Nano Banana Pro and GPT-4o.Nano Banana Pro and Seedream-4.0 excel in Chinese text rendering.
- For Chinese short prompts, closed-source models dominate overall evaluation while leading open-source models remain competitive on many non-text visual dimensions.A gap remains in Chinese grammar consistency, complex compositional reasoning, and stable Chinese text rendering.
- For Chinese long prompts, GPT-4o-1.5 leads closed-source models, while Z-Image leads open-source models.Hunyuan-Image-2.1 and Qwen-Image also show strong attribute, layout, and text performance.
4) Chinese Long Prompt (Tab. V):
The Chinese long-prompt evaluation reports detailed 27-dimension results, compares offline evaluator reliability, and summarizes UniGenBench++ extensions and overall conclusions. It identifies strong closed-source performance alongside persistent challenges in multilingual text rendering and complex reasoning.
- 4) Chinese Long Prompt (Tab. V):: The Chinese long-prompt results are reported across 27 dimensions in Tables VII, VIII, IX, and X.
- C. Offline Evaluation Model: Qwen2.5-VL-72b performs reasonably on simple dimensions but becomes unreliable on grammar-consistency and action-contact.
- C. Offline Evaluation Model: The dedicated evaluation model significantly outperforms Qwen2.5-VL-72b across short and long English and Chinese prompt evaluations.English and Chinese qualitative evaluation cases are also provided.
- D. Compared with UniGenBench: Compared with UniGenBench, UniGenBench++ adds bilingual and length-variant prompts to improve evaluation of sensitivity and robustness to language and prompt length.
- V. Conclusion: The benchmark reveals strengths and weaknesses of open- and closed-source T2I models across semantic consistency dimensions.