Source-linked AI summary

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Rongyao Fang, Aldrich Yu, Chengqi Duan, Linjiang Huang, Shuai Bai, Yuxuan Cai, Kun Wang, Si Liu, Xihui Liu, Hongsheng Li

arXiv:2509.09680v1cs.CVcs.CL

TL;DR

Open-source text-to-image research lacks large-scale reasoning-focused data and comprehensive evaluation, limiting its ability to handle complex prompts relative to leading closed-source systems. The paper introduces FLUX-Reason-6M with multidimensional GCoT supervision and PRISM-Bench with seven human-aligned tracks, finding that closed-source models lead overall while complex text rendering and long instruction following remain difficult.

  • Problem

    Open-source text-to-image models lack large-scale reasoning-focused datasets and comprehensive benchmarks aligned with human judgment, especially for complex prompts.

  • Method

    The paper builds FLUX-Reason-6M, a 6-million-image dataset with 20 million bilingual descriptions across six characteristics and GCoT supervision, and PRISM-Bench, a seven-track VLM-based benchmark.

  • Results

    Closed-source models lead PRISM-Bench overall, with GPT-Image-1 scoring 86.3, while models struggle particularly with text rendering and long, complex instructions.

  • Takeaways & Limitations

    The released dataset, benchmark, and evaluation suite provide shared resources for training and diagnosing reasoning-oriented text-to-image systems.

Abstract

from arXiv · show

The advancement of open-source text-to-image (T2I) models has been hindered by the absence of large-scale, reasoning-focused datasets and comprehensive evaluation benchmarks, resulting in a performance gap compared to leading closed-source systems. To address this challenge, We introduce FLUX-Reason-6M and PRISM-Bench (Precise and Robust Image Synthesis Measurement Benchmark). FLUX-Reason-6M is a massive dataset consisting of 6 million high-quality FLUX-generated images and 20 million bilingual (English and Chinese) descriptions specifically designed to teach complex reasoning. The image are organized according to six key characteristics: Imagination, Entity, Text rendering, Style, Affection, and Composition, and design explicit Generation Chain-of-Thought (GCoT) to provide detailed breakdowns of image generation steps. The whole data curation takes 15,000 A100 GPU days, providing the community with a resource previously unattainable outside of large industrial labs. PRISM-Bench offers a novel evaluation standard with seven distinct tracks, including a formidable Long Text challenge using GCoT. Through carefully designed prompts, it utilizes advanced vision-language models for nuanced human-aligned assessment of prompt-image alignment and image aesthetics. Our extensive evaluation of 19 leading models on PRISM-Bench reveals critical performance gaps and highlights specific areas requiring improvement. Our dataset, benchmark, and evaluation code are released to catalyze the next wave of reasoning-oriented T2I generation. Project page: https://flux-reason-6m.github.io/ .

1 Introduction

The paper addresses the gap between strong closed-source and limited open-source text-to-image systems by introducing a reasoning-focused dataset and a human-aligned benchmark. FLUX-Reason-6M combines six-dimensional supervision with generation chain-of-thought, while PRISM-Bench evaluates seven tracks and exposes remaining model weaknesses.

  • 1 Introduction: The authors introduce FLUX-Reason-6M and PRISM-Bench to provide reasoning-oriented training data and comprehensive, human-aligned evaluation for text-to-image generation.FLUX-Reason-6M is a 6-million-scale dataset, while PRISM-Bench contains seven independent tracks.
  • 1 Introduction: PRISM-Bench includes six characteristic tracks plus a Long Text track using GCoT captions, with advanced vision-language models assessing alignment and aesthetics.Each track contains carefully selected prompts designed around its specific evaluation challenge.
  • 1 Introduction: FLUX-Reason-6M organizes images across Imagination, Entity, Text rendering, Style, Affection, and Composition, using overlapping labels to combine different reasoning types.The multi-label design intentionally represents the multifaceted nature of complex scene synthesis.
  • 1 Introduction: Generation chain-of-thought captions explain how and why image elements, layouts, artistic choices, and semantic relationships are constructed, providing intermediate supervision beyond standard captions.The GCoT synthesis uses image content together with category-specific captions to produce detailed reasoning templates.

3 PRISM-Bench

PRISM-Bench is a seven-track benchmark designed to evaluate prompt-image alignment and image aesthetics across distinct T2I capabilities. Its prompts combine representative dataset samples with targeted constructions, while VLMs use track-specific alignment criteria and uniform aesthetic scoring.

  • Prompt Design and Construction: PRISM-Bench provides seven tracks, each containing 100 prompts, to evaluate six FLUX-Reason-6M characteristics plus Long Text.The tracks cover Imagination, Style, Entity, Text rendering, Affection, Composition, and Long text.
  • Prompt Design and Construction: Each track splits its prompts between representative FLUX-Reason-6M samples and curated challenges targeting specific characteristics.The 100 prompts comprise 50 systematically sampled prompts and 50 carefully constructed prompts.
  • Prompt Design and Construction: Curated prompts combine category-specific attributes such as entities, styles, emotions, text properties, and spatial relationships.Long Text prompts expand complete GCoT captions into dense, multi-sentence instructions.
  • Evaluation Protocol: Alignment is assessed with track-specific VLM instructions that focus on each task’s defining requirements rather than generic prompt correspondence.Criteria include text legibility, entity fidelity, stylistic fidelity, emotion, spatial arrangement, and detail density in Long Text prompts.
  • Evaluation Protocol: Aesthetic quality uses one unified VLM instruction set across tracks and scores images from 1 to 10 with a one-sentence rationale.The unified standard evaluates lighting, color harmony, detail, and overall visual appeal.

4 Experiments

PRISM-Bench evaluation shows that closed-source models lead overall and across most tracks, while Long Text and English text rendering remain difficult. Chinese evaluation preserves GPT-Image-1’s lead but reveals stronger Chinese text-rendering performance for several models.

  • Results on PRISM-Bench: GPT-Image-1 achieves the highest overall PRISM-Bench score at 86.3, followed by Gemini2.5-Flash-Image at 85.3.The two closed-source models outperform others across nearly all evaluation tracks.
  • Results on PRISM-Bench: GPT-Image-1 leads Entity at 88.2, Style at 93.1, and Composition at 92.8, while Gemini2.5-Flash-Image leads Imagination at 88.6 and Affection at 92.1.Qwen-Image is nearly tied with Gemini2.5-Flash-Image on Composition, and FLUX.1-dev has the highest Affection aesthetic score.
  • Results on PRISM-Bench: Text rendering receives the lowest overall scores across tracks, with Bagel and JanusPro performing particularly poorly.The benchmark identifies text rendering as a persistent challenge for T2I models.
  • Results on PRISM-Bench: Long Text remains a major weakness: Gemini2.5-Flash-Image leads with 81.1, while all models score notably lower than on other tracks.The result indicates substantial room for improvement in following complex, multi-layered instructions.
  • Results on PRISM-Bench-ZH: On PRISM-Bench-ZH, GPT-Image-1 leads with a total score of 87.5, while SEEDream 3.0 and Qwen-Image remain competitive across tracks.SEEDream 3.0 and GPT-Image-1 share the highest average score, and Chinese text rendering is stronger than the general English weakness.

Alignment score: 6

The example’s alignment assessment separates textual correctness from visual integration and physical plausibility. Although the text is legible, its placement fails to conform to the watch face and compromises the object’s realism.

  • Alignment score: 6: The text is perfectly spelled and fully legible on the watch face.
  • Alignment score: 6: Its integration is unrealistic and flat because it does not follow the watch dial’s curvature.
  • Alignment score: 6: Covering the dial with text makes the wristwatch physically implausible by removing numerals and obstructing hand readability.

Aesthetic score: 2

The aesthetic assessment likewise finds the lettering legible but visually pasted onto the watch face. It fails to integrate with the dial’s curvature and lighting.

  • Aesthetic score: 2: The watch-face text is perfectly spelled and legible.
  • Aesthetic score: 2: The lettering appears unnaturally flat and pasted onto the watch face.
  • Aesthetic score: 2: The text does not follow the dial’s curvature or lighting, weakening its visual integration.

Alignment score: 6

The watch is generally well-rendered, but illegible branding and awkward sentimental-text placement undermine realism and practicality.

  • Illegible branding and awkward sentimental-text placement undermine the watch’s realism and practical design.

Aesthetic score: 5

The watch face contains garbled, nonsensical characters that deviate from the requested heartfelt message and leave parts illegible.

  • Garbled characters substantially deviate from the requested message, making parts of the watch-face text illegible.

Alignment score: 2

Although the watch has realistic materials and lighting, distorted characters and misaligned text reduce its visual quality.

  • Distorted, overlapping, or nonsensical characters and misaligned dial text detract from an otherwise realistic watch rendering.

Aesthetic score: 4

The watch-face text is garbled, contains substantial errors relative to the prompt, and is partly illegible.

  • Garbled text with major spelling and wording errors creates a severe failure of text accuracy.

Alignment score: 1

PRISM-Bench includes a Text rendering track, illustrated as part of the benchmark’s evaluation framework.

  • PRISM-Bench includes a dedicated Text rendering track.
  • The Text rendering track is presented within the benchmark’s broader evaluation of text-to-image models.
  • The benchmark is positioned to support the development of more capable text-to-image models.

5 Conclusion

The paper introduces FLUX-Reason-6M and PRISM-Bench to address gaps in reasoning-oriented text-to-image generation and evaluation. Evaluation across 19 models finds persistent difficulty with text rendering and long instruction following, while the released resources support further research.

  • Across 19 models, text rendering and long instruction following remain difficult tasks, including for leading closed-source systems.
  • FLUX-Reason-6M provides 6 million images and 20 million reasoning-focused prompts with generation chain-of-thought across six characteristics.
  • PRISM-Bench provides seven-track, advanced-VLM-based evaluation for fine-grained, human-aligned assessment.
  • The dataset, benchmark, and evaluation code are publicly released as tools for training and evaluating future text-to-image models.
Loading 2509.09680v1…