Source-linked AI summary

Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

Niantong Li, Guangzheng Hu, Weixu Qiao, Ying Ba, Qichen Hong, Shijun Shen, Jinlin Wang, Fan Zhou, Jianye Kang, Xin Shang, Ziyi He, Wei Wang, Dalin Li, Jiahao Li, Jie Zhang, Kaiyuan Gao, Kun Yan, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiao Xu, Xiaoyue Chen, Yuxiang Chen, Yan Shu, Yanran Zhang, Yilei Chen, Yixian Xu, Zekai Zhang, Zhendong Wang, Zihao Liu, Zikai Zhou, Hongzhu Shi, Yi Wang, Bing Zhao, Hu Wei, Lin Qu, Chenfei Wu

arXiv:2605.28091v2cs.CV

TL;DR

Existing T2I benchmarks do not adequately capture the real-world fidelity and creative expression required in professional workflows. Qwen-Image-Bench introduces a creator-centric hierarchical benchmark with expert-designed prompts and fine-grained Q-Judger evaluation, and it distinguishes leading models while exposing shared weaknesses in world-grounded capabilities. Its planned evolution into a living benchmark defines an explicit scope boundary for keeping pace with changing models and workflows.

  • Problem

    Existing T2I benchmarks remain centered on basic alignment and quality criteria, limiting their coverage of nuanced capabilities needed for professional creative practice.

  • Method

    Qwen-Image-Bench combines a five-pillar, three-level taxonomy with 1,000 structured prompts and Q-Judger, which scores images across 56 expert-defined facets.

  • Results

    Qwen-Image-Bench separates 18 leading T2I models into five performance tiers and reveals shared ceiling facets where even the strongest model scores below 44.

  • Takeaways & Limitations

    The benchmark provides fine-grained, rubric-grounded diagnostics that expose capability gaps beyond what coarser evaluation can detect.

  • Takeaways & Limitations

    The benchmark is currently static, motivating future prompt refreshes, taxonomy extensions to video and interactive editing, and continuously updated evaluation.

Abstract

from arXiv · show

Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. To address the gap, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Qwen-Image-Bench enriches conventional evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1000 stratified prompts with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model Q-Judger based on Qwen3.6-27B, supervised by 80 professional annotators from global art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.

1 Introduction

Qwen-Image-Bench addresses the limits of coarse T2I evaluation by combining creator-centered dimensions, hierarchical capability analysis, expert prompts, and fine-grained judging. It evaluates leading models across these dimensions and exposes shared weaknesses in several physically grounded facets.

  • Benchmark design: Qwen-Image-Bench jointly evaluates Quality, Aesthetics, Text-Image Alignment, Real-world Fidelity, and Creative Generation through 1,000 structured prompts.The benchmark uses a three-level taxonomy designed for broad, balanced coverage and evaluation efficiency.
  • Benchmark design: Its top-down taxonomy mirrors professional workflows and decomposes five pillars into 23 sub-capabilities and 56 third-level evaluation facets.The hierarchy follows ideation, styling, and iterative refinement, from creator goals to reviewer-inspected details.
  • Evaluation pipeline: Q-Judger scores every prompt-image pair across all 56 facets using a unified model trained on expert annotations, producing attributable capability diagnostics.The model is built on Qwen3.6-27B and supervised by professional experts under blind labeling and strict review protocols.
  • Empirical findings: GPT Image 2 achieves the highest overall score among 18 evaluated T2I models, with ZH: 64.7 and EN: 65.2.The evaluation covers both Chinese and English prompts.
  • Empirical findings: Four facets—Physical Logic, Anatomical Fidelity, Animals, and Contact Interaction—remain systemic ceilings, with even the best models scoring below 44.These weaknesses persist across current T2I models despite differences in overall performance.

2 Related Work

Existing T2I benchmarks increasingly use structured judges and richer semantic tests, but they remain limited in their treatment of creativity and authentic artistic practice. Qwen-Image-Bench responds with creator-centered, fine-grained evaluation of real-world fidelity and creative generation.

  • T2I benchmarks and evaluation methods: Existing benchmarks and evaluators focus largely on object properties, prompt correctness, semantic coverage, or instruction following rather than the full demands of creator workflows.These approaches include detector-based metrics, VQA/VLM judges, and complexity-oriented benchmark designs.
  • Creativity and aesthetics: Creativity and aesthetics remain difficult to assess because they are subjective, multifaceted, and poorly covered by limited-scale, culturally narrow aesthetic datasets.Such datasets hinder generalization across styles, domains, and creative intents.
  • Creativity and aesthetics: Most benchmarks do not adequately disentangle technical quality from artistic creativity, leaving creative intent insufficiently represented.Aggregate human ratings often combine distinct properties instead of diagnosing them separately.
  • Remaining evaluation gap: Single or homogeneous judges remain vulnerable to systematic bias and benchmark drift while failing to assess knowledge-consistent realism and authentic creative expression together.These are identified as twin pillars of professional T2I applications.
  • Proposed direction: Qwen-Image-Bench addresses this gap with an expert-co-designed taxonomy and unified judge providing rubric-grounded scoring at third-level facet granularity.The framework explicitly evaluates Real-world Fidelity and Creative Generation from a creator-centric perspective.

3 Qwen-Image-Bench

Qwen-Image-Bench is a creator-centric evaluation framework combining a hierarchical capability taxonomy, expert-authored prompts, and a human-supervised judge to produce fine-grained diagnostics for T2I models.

  • The benchmark combines a hierarchical taxonomy, an expert-in-the-loop prompt factory, and a unified judge model for interpretable T2I evaluation.These components respectively define capabilities, instantiate them in realistic prompts, and score generated images at fine granularity.
  • 3.1 A Hierarchical Capability Taxonomy: The taxonomy mirrors professional artistic reasoning across five pillars, decomposing evaluation into 23 sub-capabilities and 56 third-level facets.Each third-level facet belongs to exactly one second-level sub-capability and one first-level pillar, enabling deterministic upward aggregation and diagnosis.
  • 3.2 Expert-in-the-Loop Prompt Construction: The prompt factory uses facet-targeted sampling, bilingual drafting, expert review, and length-variant expansion to create realistic, rubric-consistent evaluation cases.Professional artists can discard or rewrite weak prompts, while length variants test robustness to linguistic complexity without changing facet alignment.
  • 3.2 Expert-in-the-Loop Prompt Construction: The final set contains 1,000 Chinese-English prompts, evenly divided into short and long variants, with each prompt testing facets across 3 to 5 pillars.This design supports controlled comparisons across language, prompt length, and capability composition.
  • 3.3 Unified Judge Model with Rubric-Grounded Fine-Grained Scoring: Q-Judger scores each prompt-image pair independently on all applicable third-level facets using rubric-grounded 0, 1, 2, or N/A labels.Its scoring protocol is designed to prevent dominant signals from obscuring subtler capabilities and to make weaknesses traceable to specific facets.
  • 3.4 Evaluation Pipeline and Multi-granularity Scoring: Non-N/A facet scores are normalized before aggregation, preserving traceability from overall scores to pillar, sub-capability, and facet-level diagnostics.The resulting aggregates can be decomposed into the specific capabilities where a model succeeds or falls short.

4 Experiment

Qwen-Image-Bench evaluates 18 T2I models through a hierarchical, fine-grained scoring pipeline and reveals clear overall and capability-profile differences across models.

  • Evaluation Setup: The benchmark scores every prompt-image pair across 56 third-level facets and aggregates them bottom-up into pillar and overall scores.Images are generated for all 1,000 prompts for each model before Judge Model evaluation.
  • Overall Ranking: GPT Image 2 leads overall by nearly 5 points over Nano Banana 2.0, while GLM Image trails by 16.5 points and the models separate into five tiers.Qwen Image 2.0 Pro ranks fifth overall.
  • Fine-grained Capability Profiles: GPT Image 2 forms the outermost radar contour, with five peak facets above 84 and Graphic Design and Game Design exceeding 90.The radar chart represents model capability profiles across all 56 facets, with larger polygons indicating stronger overall capability.
  • Fine-grained Capability Profiles: Text Accuracy, Cross-lingual Generation, and Information Visualization separate tiers, with only a few models sustaining scores above 60 while lower-tier models often fall below 35.These facets create a visible inward collapse beside the radar chart’s apex.
  • Fine-grained Capability Profiles: Anatomical Fidelity, Physical Logic, Objects, Animals, and Contact Interaction remain shared bottlenecks, with the best model below 44 on each.These weaknesses form the radar chart’s top-center indentation.
  • Fine-grained Capability Profiles: Qwen Image 2.0 Pro retains competitive visual-style coverage but collapses on several creative-precision facets, producing an asymmetric radar shape.The same structural patterns hold under English prompts.

4.5 Discriminative Power of Creator-Centric Dimensions

Qwen-Image-Bench’s creator-centric dimensions provide stronger model separation than conventional criteria, especially for creative and knowledge-grounded capabilities.

  • Variance Analysis: Of the 15 highest-variance third-level facets, 12 belong to Creative Generation or Real-world Fidelity, led by Text Accuracy, Information Visualization, and Cross-lingual Generation.These facets test creative imagination, logical reasoning, and execution precision.
  • Sub-capability Rankings: At the second level, Text Rendering has the highest variance, followed by Style Control, Logical Resolution, and World Knowledge; four of the top six belong to the two application-driven pillars.Figure 6 highlights Text Rendering, World Knowledge, and Visual Storytelling as creator-relevant high-variance sub-capabilities.
  • Variance Analysis: Creative Generation variance exceeds Quality variance by over 11× and Aesthetics variance by over 4×, showing where models diverge most sharply.Quality exhibits low variance, indicating that basic image quality has become a table-stakes capability.
  • Sub-capability Rankings: On World Knowledge, GPT Image 2 leads the second tier by 8 points, while models below rank 8 score under half the leader’s score.The dimension measures faithful reproduction of real-world objects, animals, and structured visual information.
  • Design Applications: Game Design shows a threshold effect, with GPT Image 2 at 91.8, ranks 2–4 near 80, and rank 7 onward below 66.Graphic Design similarly has a leader-to-median gap exceeding 30 points.
  • Design Applications: Design Applications vary substantially: Product Design has a 32-point spread, Fashion Styling only 13 points, while Art Design’s ranks 2–4 differ by barely one point.These patterns indicate uneven maturity across fine-grained creative capabilities.

4.7 Per-Pillar Model Rankings

Per-pillar rankings show that application-driven dimensions, especially Creative Generation and Real-world Fidelity, separate models more sharply than conventional pillars. The capability landscape further reveals mature visual-pattern skills alongside persistent world-knowledge and reasoning bottlenecks.

  • Per-pillar rankings: Creative Generation has the largest leader-to-bottom spread at 30.6 points, making it the most discriminative pillar; GPT Image 2 leads and Qwen Image 2.0 Pro ranks sixth.The pillar also produces the largest ranking shifts across models.
  • Per-pillar rankings: Real-world Fidelity places GPT Image 2 first, with a roughly 3-point lead over the next cluster and a 14-point gap to the lowest-performing models.The result highlights faithful reconstruction and knowledge-grounded content as major differentiators.
  • Capability landscape: Most facets fall in the 40–60 developing range, while visual aesthetics, material rendering, and scene-level fidelity include several mature facets above 60.The benchmark therefore distinguishes relatively solved surface-pattern capabilities from active improvement frontiers.
  • Capability landscape: Five systemic ceilings—Physical Logic, Anatomical Fidelity, Animals, Objects, and Contact Interaction—remain below 35 in mean, with even the best model below 44.These facets span four pillars but share dependence on implicit structural, physical, or biological world knowledge.
  • Capability landscape: Creative Generation shows the greatest within-pillar capability range, with more than 32 points separating its highest and lowest cross-model means.High-mean style facets rely on learned aesthetic priors, whereas low-mean facets require precise execution or reasoning.

4.9 Cross-Tier Gap Analysis and Upgrade Pathways

Cross-tier differences are driven primarily by Creative Generation and other application-driven capabilities rather than converged basic quality. The analysis identifies concrete upgrade priorities, with language understanding preceding visual execution in Qwen Image 2.0 Pro’s pathway.

  • Cross-tier gaps: Creative Generation produces the largest inter-tier gaps: +8.68 from T1 to T2 and +4.29 from T2 to T3, about 1.8× conventional-pillar gaps.Under English prompts, the T1–T2 gap reaches +9.98, approximately 2.0× the conventional-pillar average.
  • Cross-tier gaps: Quality and Aesthetics are largely converged, with T2–T3 gaps of only 1.35 and 2.09, respectively, making basic image quality a table-stakes capability.The contrast with Creative Generation identifies application-driven skills as the stronger tier separators.
  • Upgrade priorities: The ten largest T2–T3 facet gaps cluster around design execution, reasoning and knowledge, and precise style and rendering control.Information Visualization leads at +10.4, followed by Logical Resolution at +9.1 among the named examples.
  • Upgrade pathway: Qwen Image 2.0 Pro exceeds the T2 mean on language-intensive facets such as Text Accuracy (+11.4), Storyboard Creation (+10.2), Comic Creation (+5.3), and Font (+3.1).Its remaining deficits concentrate on visual-execution facets including Anatomical Fidelity, Game Design, Feature Matching, and Objects.
  • Upgrade pathway: The tier analysis points to a two-phase upgrade strategy: improve language understanding first, then address visual execution.The broader development implication is to target design precision, knowledge-driven reasoning, and logical-causal expression rather than already-converged basics.

5 Conclusion

Qwen-Image-Bench provides fine-grained, creator-centered evaluation that separates leading T2I models and aligns closely with expert judgment. Its results expose a perception-to-cognition frontier and motivate future expansion into a living multimodal benchmark.

  • Conclusion: Qwen-Image-Bench uses 5 pillars, 23 sub-capabilities, and 56 facets, achieving Spearman ρ = 0.92 with human rankings and separating 18 models into five tiers.The benchmark and Judge Model remain consistent across Chinese and English prompts.
  • Conclusion: Five facets spanning four pillars remain below 44 even for the strongest model, revealing a perception-to-cognition frontier rooted in missing implicit world knowledge.The affected capabilities are Physical Logic, Anatomical Fidelity, Animals, Objects, and Contact Interaction.
  • Future work: The planned living benchmark will refresh prompts, extend coverage to video generation and interactive editing, and provide automated real-time evaluation with a continuously updated leaderboard.These directions are intended to preserve discrimination as models and creator workflows evolve.

A.1 Human Rating Results

Human expert ratings provide 1–10 holistic scores for 18 T2I models across 1,000 prompts per pillar, with models sorted by overall score.

  • Human rating results: Table 4 reports mean professional-annotator ratings for 18 T2I models on a 1–10 scale across 1,000 prompts per pillar.Models are sorted by overall score, and the best score in each pillar column is bolded.

A.2 Score Heatmaps

The heatmaps reveal strong specialization across sub-capabilities and fine-grained facets, while several physical and biological capabilities remain systemic ceilings. GPT Image 2 leads consistently across nearly all evaluated dimensions.

  • Model rankings shift considerably across sub-capabilities, with leaders on Resolution dropping to mid-tier positions on Text Rendering or World Knowledge.
  • Text Rendering, World Knowledge, and Design Applications show the largest visual contrasts among models at the L2 level.
  • GPT Image 2’s near-uniformly dark heatmap column confirms that its overall lead reflects consistent dominance across virtually all 56 facets.
  • Creative facets around ranks 5–6 show an abrupt transition from moderate scores to near-white, indicating little middle ground between capability and incapability.
  • Physical Logic, Anatomical Fidelity, and Animals remain uniformly pale across models, marking systemic capability ceilings rather than model-specific weaknesses.
  • Material Properties under Alignment forms a consistently dark row, with a leader score of 84.1, indicating reliable adherence to material-attribute instructions.

A.3 English Prompt Evaluation Results

English-prompt evaluation preserves the benchmark’s main Chinese-prompt findings while revealing limited within-tier rank changes and stronger performance for some models on English-sensitive capabilities. Creative Generation remains the most discriminative pillar, whereas Quality is comparatively converged.

  • GPT Image 2 leads across all five pillars under English prompts, and the five-tier structure is preserved across languages.
  • GPT Image 1.5 ranks second under English prompts with 60.4, overtaking Nano Banana 2.0 at 59.6, while tier membership remains unchanged.
  • Qwen Image 2.0 Pro improves under English prompts in Quality and Aesthetics and rises in Graphic Design, Game Design, and Fashion Styling.
  • English radar, variance, and heatmap analyses preserve the Chinese-prompt capability landscape, including dominant Creative Generation variance and the same five systemic ceilings.
  • Creative Generation produces the largest inter-tier gaps, including +9.98 at T1–T2 and +4.19 at T2–T3, while Quality’s T2–T3 gap is only +0.91.

A.4 Judge Model Prompt Templates

The Judge Model prompt templates operationalize Qwen-Image-Bench as structured, pillar-specific evaluation. Each prompt supplies the generation context, scoring rules, and checklist needed to produce facet-level judgments.

  • A.4 Judge Model Prompt Templates: A.4 Judge Model Prompt Templates define an evaluator role and provide the text prompt, generated image, evaluation dimension, scoring rules, checklist, and JSON output format.
  • A.4.1 Prompt Template: The scoring rubric maps each criterion to Fail, Pass, Excel, or N/A, corresponding to scores 0, 1, 2, or non-applicability.
  • A.4.2 Evaluation Checklists: A.4.2 Evaluation Checklists pass pillar-specific L2 and L3 criteria through the format_checklist field for each evaluation dimension.
  • A.4.2 Evaluation Checklists: The Quality and Aesthetics checklists cover lighting, anatomy, emotional expression, style, detail, resolution, composition, and color harmony.
  • A.4.2 Evaluation Checklists: The Alignment checklist evaluates attributes such as quantity, expression, material, color, shape, and size, plus contact, spatial, and bodily interactions.
  • A.4.2 Evaluation Checklists: The Real-world Fidelity checklist covers layout, composition relations, containment, real and virtual scenes, fairness, physical logic, safety, animals, objects, information visualization, time, and culture.
  • A.4.2 Evaluation Checklists: The Creative Generation checklist evaluates imagination, feature matching, material texture, logical resolution, text accuracy, text layout, noise, edge clarity, naturalness, and resolution.
Loading 2605.28091v2…