Source-linked AI summary
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
Samin Mahdizadeh Sani, Max Ku, Nima Jamali, Matina Mahdizadeh Sani, Paria Khoshtab, Wei-Chieh Sun, Parnian Fazel, Zhi Rui Tam, Thomas Chong, Edisy Kin Wai Chan, Donald Wai Tong Tsang, Chiao-Wei Hsu, Ting Wai Lam, Ho Yin Sam Ng, Chiafeng Chu, Chak-Wing Mak, Keming Wu, Hiu Tung Wong, Yik Chun Ho, Chi Ruan, Zhuofeng Li, I-Sheng Fang, Shih-Ying Yeh, Ho Kei Cheng, Ping Nie, Wenhu Chen
TL;DR
Existing image-generation benchmarks are fragmented across tasks and domains and often provide opaque scores without explaining failure modes. ImagenWorld addresses this gap with a broad benchmark and structured human evaluation, finding that editing and text-heavy tasks remain difficult, while closed-source systems lead overall and targeted data curation improves text-heavy performance.
Problem
Existing benchmarks are fragmented across isolated tasks or narrow domains and often provide opaque scores without explaining model failure modes.
Method
ImagenWorld evaluates 3.6K condition sets across six tasks and six domains using 20K fine-grained human annotations for object-level and segment-level errors, complemented by VLM-based metrics.
Results
Models perform better on generation than editing, struggle with symbolic and text-heavy domains, and show a persistent closed-source advantage, while Qwen-Image excels on textual graphics through targeted data curation.
Takeaways & Limitations
ImagenWorld provides both broad cross-model benchmarking and localized, interpretable diagnosis of image-generation failures beyond scalar scores.
Takeaways & Limitations
Image generation and editing systems create risks of fraud, deception, deepfakes, political disinformation, and privacy violations when misused.
Abstract
from arXiv · showhide
Advances in diffusion, autoregressive, and hybrid models have enabled high-quality image synthesis for tasks such as text-to-image, editing, and reference-guided composition. Yet, existing benchmarks remain limited, either focus on isolated tasks, cover only narrow domains, or provide opaque scores without explaining failure modes. We introduce \textbf{ImagenWorld}, a benchmark of 3.6K condition sets spanning six core tasks (generation and editing, with single or multiple references) and six topical domains (artworks, photorealistic images, information graphics, textual graphics, computer graphics, and screenshots). The benchmark is supported by 20K fine-grained human annotations and an explainable evaluation schema that tags localized object-level and segment-level errors, complementing automated VLM-based metrics. Our large-scale evaluation of 14 models yields several insights: (1) models typically struggle more in editing tasks than in generation tasks, especially in local edits. (2) models excel in artistic and photorealistic settings but struggle with symbolic and text-heavy domains such as screenshots and information graphics. (3) closed-source systems lead overall, while targeted data curation (e.g., Qwen-Image) narrows the gap in text-heavy cases. (4) modern VLM-based metrics achieve Kendall accuracies up to 0.79, approximating human ranking, but fall short of fine-grained, explainable error attribution. ImagenWorld provides both a rigorous benchmark and a diagnostic tool to advance robust image generation.
1 INTRODUCTION
ImagenWorld addresses fragmented and opaque evaluation with a unified benchmark spanning diverse tasks and domains, combining broad model testing with explainable human error analysis.
- Existing benchmarks are fragmented across isolated tasks and have not kept pace with increasingly capable multi-task image-generation systems.
- ImagenWorld evaluates 3.6K condition sets across six task types and six topical domains using structured human annotations and VLM-based metrics.
- 14 models are evaluated under one protocol, including unified generation-and-editing systems and task-specific baselines.
- Editing exposes two distinct failure modes—regenerating a new image or returning the input unchanged—indicating limited localized control.
- The explainable schema identifies object-level and segment-level failures with textual descriptions and localized masks, extending evaluation beyond scalar scores.
2 RELATED WORKS
Related work spans advances in conditional image synthesis and increasingly semantic evaluation, while existing benchmarks and metrics provide incomplete coverage or interpretability.
- Diffusion remains dominant in conditional image synthesis, while autoregressive and hybrid architectures increasingly support editing, structural control, and personalization.
- Traditional metrics assess fidelity or alignment, whereas VLM-based metrics better capture semantic relevance but may introduce bias and depend on proprietary models.
- Existing benchmark comparisons motivate broader evaluation across task coverage and properties.
3 THE IMAGENWORLD BENCHMARK
ImagenWorld unifies instruction-driven generation and editing across six task configurations, six topical domains, and examples that expose both successful and failed outputs.
- All tasks use natural-language instructions with optional source or reference images, reflecting practical interactions with generative systems.
- Illustrative examples show successful and failure cases for each task, while diagnostic examples distinguish missing or distorted objects from flawed image regions.
- The benchmark distinguishes generation without a source image from editing that modifies an existing source while following an instruction.
- Six task types cover text-guided generation, single- and multiple-reference generation, and corresponding editing settings.
- The dataset spans artworks, photorealistic images, information graphics, textual graphics, computer graphics, and screenshots.
4 EVALUATION SETUP
ImagenWorld combines multidimensional quality scoring, independent human and automated assessment, explainable object and segment issue labels, and broad architectural coverage.
- Four criteria—Prompt Relevance, Aesthetic Quality, Content Coherence, and Artifacts—are rated on a 5-point Likert scale and rescaled to [0,1].
- Prompt Relevance measures instruction fidelity, while Aesthetic Quality assesses visual appeal and design.
- Content Coherence evaluates logical and semantic consistency, including relationships among labels, charts, and depicted content.
- Artifacts capture technical flaws such as distorted text, warped edges, extra limbs, unnatural eyes, and repeated patterns.
- Each image receives ratings from three independent annotators, alongside Gemini-2.5-Flash scores and auxiliary CLIPScore and LPIPS metrics.
- Annotators mark expected objects that are missing, incorrectly rendered, or distorted, and select flawed SoM image segments responsible for score deductions.
- The evaluation includes diffusion, autoregressive, and hybrid models, spanning unified systems and experts specialized for task subsets.
5 RESULTS AND ANALYSIS
Across tasks and topics, human evaluations show that editing and text-heavy or symbolic domains remain the main weaknesses, while VLM judgments broadly track human rankings but miss some fine-grained artifacts.
- Task Level: Models score roughly 0.1 lower on editing tasks than on corresponding generation tasks, making localized modification a major bottleneck.Editing tasks include TIE, SRIE, and MRIE; generation counterparts include TIG, SRIG, and MRIG.
- Task Level: GPT-Image-1 leads overall, outperforming Gemini 2.0 Flash by about 0.1–0.2 points on average, although some open-source models exceed Gemini on selected tasks.No open-source unified model catches up with closed-source systems overall.
- Topic Level: Artworks and Photorealistic Images average near 0.78, whereas Screenshots and Information Graphics average closer to 0.55.Textual Graphics and Computer Graphics are intermediate, averaging near 0.68.
- Evaluation Criteria: Prompt Relevance varies most across tasks, peaking at 0.72 in TIG and falling to 0.46 in editing, while artifacts are especially problematic in text-heavy domains.Aesthetic Quality and Content Coherence are strongest for Artworks and Photorealistic Images and lower for Screenshots and Information Graphics.
- Evaluation Criteria: Human inspection exposes skipped instruction steps, corrupted text, and numerical inconsistencies that scalar metrics do not fully capture.These recurring errors motivate explainable analysis beyond aggregate scores.
- Evaluation Criteria: VLM–human Kendall accuracy ranges from 0.74 to 0.79, but VLMs under-penalize unreadable text, boundary glitches, and misaligned layouts.VLM agreement is strongest for prompt relevance and weaker for artifact judgments.
6 DISCUSSION
Discussion links editing failures to model behavior and shows that targeted data curation can improve text rendering, while the benchmark’s annotations support future diagnostic and preference-based training.
- Targeted Data Curation: Qwen-Image’s synthetic text-rich data, progressive layout curriculum, and balanced domain coverage improve text rendering in text-heavy tasks.It can outperform even closed-source systems in textual graphics for text-guided generation.
- Editing Behaviors: Autoregressive–diffusion hybrids generate entirely new images in 17% of editing cases, suggesting language-driven pathways can override source conditioning.The authors contrast this behavior with diffusion-only editors optimized for localized modification.
- Editing Behaviors: Diffusion-only editors show lower source-disregard rates—0.6% for IC-Edit and 3.4% for InstructPix2Pix—but may offer narrower task coverage.Flux.1 Kontext reaches 14.4%, consistent with its broader mixture of local, global, and reference-based tasks.
- Future Work: ImagenWorld’s human scores and object-level tags can support preference optimization, ranking-based fine-tuning, and diagnostic self-correction.The dataset combines human scores, localized error tags, and full model outputs.
7 CONCLUSION
ImagenWorld combines broad task and domain coverage with explainable human evaluation to benchmark and diagnose image-generation failures. Across 14 models, it reveals consistent performance gaps between generation and editing, symbolic and non-symbolic domains, and closed- and open-source systems.
- ImagenWorld unifies six generation and editing tasks across six topical domains, supported by 20K fine-grained human annotations.
- The explainable evaluation schema labels object-level and segment-level issues, enabling localized diagnosis beyond scalar scores.
- Across 14 models, generation outperforms editing, while symbolic and text-heavy domains remain difficult.
- Closed-source models show a persistent performance gap over open-source models, and structured annotations reveal errors that modern VLM-based metrics miss.
ETHICS STATEMENT
The paper frames image generation and editing as raising safety and ethical concerns because realistic forgeries are increasingly accessible and can support fraud, deception, and privacy violations. ImagenWorld therefore releases sanitized data and annotations while excluding sensitive or personally identifiable content.
- AI-powered image systems lower the technical barrier to realistic forgeries, enabling misuse such as fraudulent refunds, fake receipts, and fabricated identification documents.
- Potential harms include deepfakes, non-consensual imagery, political disinformation, privacy violations, and reduced trust in digital media.
- ImagenWorld addresses these concerns by releasing only sanitized data and annotations and excluding sensitive or personally identifiable content.
REPRODUCIBILITY STATEMENT
The reproducibility statement describes standardized model evaluation and public release plans. Experiments use recommended configurations, a shared seed for open-source models, and publicly accessible code, data, and sanitized annotations.
- Experiments used eight NVIDIA A6000 GPUs and the ImagenHub inference library, with latest models integrated when necessary.
- For fairness, implementations followed default configurations from the respective papers or official releases, and open-source models used seed 42.
- Code, the dataset, and sanitized human annotations will be released on publicly accessible platforms such as GitHub and HuggingFace.
A.1 USE OF LLM IN WRITING
The paper defines human evaluation criteria, annotation procedures, issue labels, benchmark examples, and statistical comparisons for assessing image-generation systems. These materials cover prompt relevance, visual quality, coherence, artifacts, localized errors, model comparisons, and task- and topic-level performance.
- Human evaluation criteria: Human evaluation rates prompt relevance, aesthetic quality, content coherence, and artifacts on five-point scales with explicit behavioral descriptions.The criteria address instruction alignment, visual appeal and readability, logical consistency, and generation flaws.
- Annotation procedure: Annotators flag object-level, segmentation, and other issues, explaining any non-perfect score when predefined labels do not apply.The interface supports both ratings and issue selection, including localized object- and segment-level problems.
- Benchmark diagnostics: The benchmark includes instruction-following, numerical, labeling, plot, diagram, text, and editing cases, alongside topic-wise and model-wise performance summaries.
- Statistical comparisons: Closed-source models outperform open-source models overall by mean difference 0.6821, with t = 38.38 and p = 6.6 × 10−204.
- Statistical comparisons: Generation scores exceed editing scores: generation mean = 3.9152 versus editing mean = 3.5244, with t = 13.71 and p = 6.6 × 10−42.
- Statistical comparisons: Non-symbolic topics outperform symbolic topics: mean = 4.1067 versus 3.1905, with t = 34.26 and p = 7.3 × 10−227.