Source-linked AI summary

Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation

Yutong Liu, Nan Huang, Xu Cao, James M. Rehg

arXiv:2609.02864v1cs.CV

TL;DR

Existing generative models show strong perceptual fluency but lack systematic evaluation of reasoning from visual inputs to logically constrained images. RIG-BENCH introduces a 2,000-sample benchmark and finds a significant reasoning–generation gap, motivating diagnostic evaluation for logically consistent generative models.

  • Problem

    High-level visual reasoning-to-generation is underexplored because existing benchmarks separate text-based reasoning from aesthetic image generation.

  • Method

    RIG-BENCH evaluates visual rule induction through 2,000 curated samples spanning four task families, using rubric-based assessment weighted toward reasoning correctness.

  • Results

    Advanced generation models exhibit a significant reasoning–generation gap, often producing locally plausible but globally illogical outputs.

  • Takeaways & Limitations

    RIG-BENCH provides a diagnostic framework for developing world models and unified generative models that are logically consistent as well as visually capable.

  • Takeaways & Limitations

    The benchmark prioritizes 2,000 hand-curated items over larger procedurally generated datasets, and automated judges may sometimes favor stylistic coherence alongside logical adherence.

Abstract

from arXiv · show

Recent advancements in unified generative models (UGMs) and world simulators have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level event alignment, leaving the capacity for high-level visual reasoning underexplored. True visual generative intelligence demands "Reasoning-to-Generation", an ability to infer latent rules from visual inputs and manifest solutions through precise, logically constrained visual outcomes. We introduce RIG-BENCH, a novel comprehensive benchmark that systematically evaluates Reasoning-driven Image Generation (RIG) across four cognitively demanding domains: Concept-based, Transformation-based, Pattern & Structure, and Scenario-based. Featuring 2000 curated samples, RIG-BENCH serves as a rigorous stress test for RIG. Our extensive evaluations of state-of-the-art UGMs and image/video generation models reveal a significant reasoning-generation gap, wherein models frequently produce locally plausible but globally illogical outputs. RIG-BENCH provides a vital diagnostic framework to guide the development of next-generation, logically grounded UGMs and world simulators.

1 Introduction

RIG-BENCH addresses the underexplored ability to infer latent visual rules and generate logically constrained images. It evaluates this capability across four task families and reveals a reasoning–generation gap in current models.

  • Recent UGMs achieve strong visual perception and high-fidelity synthesis, but high-level reasoning over visual inputs remains underexplored.
  • RIG-BENCH targets closed-loop Image-to-Reasoning-to-Image evaluation, unlike benchmarks that separately assess text reasoning or aesthetic image generation.
  • The benchmark spans Concept-based, Transformation-based, Pattern & Structure, and Scenario-based reasoning task families.
  • Extensive experiments find that advanced generation models often produce locally plausible but globally illogical results, exposing a significant reasoning–generation gap.
  • RIG-BENCH contains 2,000 curated items across four task families and 11 fine-grained subtasks, with targets induced from visual context alone.
  • The benchmark provides a diagnostic framework for identifying current models’ failure modes and informing future thinking generative agents.

2 Related Works

Prior work advances visual generation, visual reasoning, and unified multimodal models, but commonly separates language-based reasoning from image generation. RIG-BENCH combines visual-context inference and answer-image synthesis in one protocol.

  • Visual Generation: Most visual-generation benchmarks evaluate prompt-conditioned rendering alongside perceptual quality, prompt faithfulness, compositionality, or temporal consistency.
  • Datasets for Visual Reasoning and Cognition: Visual reasoning benchmarks study abstract rule induction, matrix reasoning, visual analogies, concept learning, and scientific or multidisciplinary reasoning.
  • Unified Multimodal Generation Models: Unified multimodal models make it practical to connect visual understanding with visual synthesis, motivating evaluation beyond image quality toward structural inference.
  • Unified Multimodal Generation Models: RIG-BENCH advances prior work through a unified answer-image generation protocol that requires models to parse visual context and synthesize visual answers.

3 Benchmark Design

RIG-BENCH formalizes visual reasoning as a unified task that maps visual context and optional demonstrations to a single logically correct output image. Its 2,000 samples span four families and eleven subtasks, requiring perception, rule induction, and grounded synthesis under a common evaluation interface.

  • 3.1 Task Formalization: RIG supplies visual context, optional demonstration pairs, and a neutral instruction, while withholding the underlying logic and step-by-step solution.The model must synthesize a single output image that satisfies both perceptual fidelity and logical consistency.
  • 3.2 Data Engineering and Curation: The benchmark pipeline reformulates established visual reasoning items into image-output tasks, generates answer images directly, and scores them against ground truth with rubric-based and perceptual measures.An additional human study calibrates the evaluation.
  • 3.1 Task Formalization: RIG-BENCH requires visual perception, latent rule induction, and grounded generation simultaneously, unlike benchmarks that provide text or categorical answers.The synthesized image serves as the primary evidence of reasoning and exposes the reasoning–generation gap.
  • 3.2 Data Engineering and Curation: 2,000 samples are distributed across four task families and eleven fine-grained subtasks, with varied context cardinality but a strictly single-image output.Families include concept, transformation, pattern and structure, and scenario-based reasoning.
  • 3.2 Data Engineering and Curation: The dataset covers concept induction, geometric and attribute transformations, matrix and visual-spatial reasoning, and scenario-based scientific, temporal, analogy, and style inference.Tasks also vary between open-form outputs with multiple valid realizations and closed-form outputs determined by explicit relations.
  • 3.3 Comparison with Existing Benchmarks: RIG-BENCH fills a benchmark gap by unifying visual context, rule induction, and pixel-level synthesis across hand-curated cognitively demanding tasks.This design extends beyond standard text-to-image, text-based reasoning, and procedurally synthesized video-reasoning evaluations.

4 Experiments and Results

RIG-BENCH evaluates multimodal generation models with perceptual metrics, rubric-based reasoning scores, human validation, and matched diagnostics. Results show sharp model and subtask stratification, weak alignment between perceptual similarity and reasoning correctness, and distinct reasoning–rendering integration gaps.

  • Evaluation setup: Nine multimodal generation models are evaluated across all eleven RIG-BENCH subtasks under a unified answer-image protocol.Video models are included by treating their final frames as predicted visual states.
  • Evaluation setup: The evaluation reports DINO, CLIP-I, LPIPS, and FID for visual proximity, alongside a rubric composite weighted toward reasoning correctness.The rubric combines visual quality, structural alignment, and reasoning correctness, with weights 0.15, 0.20, and 0.65 respectively.
  • Main results: 64.6/100 is Gemini 3 Pro Image Preview’s composite score; all open-source image generators score below 32, and no model consistently exceeds 70.Humans solve the tasks reliably, while no evaluated model approaches saturation.
  • Main results: FLUX.2 scores 31.3 versus 43.8 for GPT Image 1 and 64.6 for Gemini 3 Pro, while open-source systems remain below 25 on all three Transformation subtasks.Concept-based and style-preference tasks are substantially easier than uniquely rule-determined transformations.
  • Diagnostic findings: DINO and CLIP similarity do not establish reasoning correctness, because video frames can be perceptually close to ground truth despite incorrect underlying reasoning.A perceptual-only leaderboard could therefore misidentify video generators as strong reasoners.
  • Diagnostic findings: Matched T/D/H/O diagnostics distinguish text reasoning, direct generation, self-conditioned generation, and oracle-conditioned generation.T evaluates text correctness separately, while D, H, and O are judged against the ground-truth image.
  • Diagnostic findings: Explicit inferred answers improve BAGEL by 6.1, Qwen3.5→Emu3.5-Image by 13.4, and Gemini by 7.4, but change Qwen3.5→Qwen-Image-Edit by −0.9.The benefit of making reasoning explicit depends strongly on the image generator.
  • Diagnostic findings: Oracle conditioning improves BAGEL by 12.4 points and the Emu3.5 pipeline by 19.3, while Qwen-Image-Edit changes by only 1.8 points with a confidence interval crossing zero.The contrast isolates renderer limitations when reasoner and textual plans are held constant.

5 Conclusion

RIG-BENCH reframes visual generation as rule induction and evaluates whether models can infer latent structure from visual context and express it as answer images. Its task families cover concept learning, transformations, pattern and structure, and scenario-based reasoning through varied image-output protocols.

  • Conclusion: RIG-BENCH contains 2,000 samples spanning four task families and eleven fine-grained subtasks.The benchmark covers concept, transformation, pattern, and scenario domains.
  • Conclusion: The benchmark shifts evaluation from instruction following toward rule induction and requires answer images to be induced from visual context rather than selected from language options.Multiple-choice options are withheld so models must synthesize the answer.
  • Conclusion: The benchmark is intended to expose the gap between high-level logical inference and low-level visual synthesis.Its diagnostic role is to guide development of world models that are both visually capable and logically consistent.
  • Concept-based: Concept-based tasks ask models to identify a shared abstract interaction or relational concept and depict it in a visually distinct new scenario.The generated image must remain clearly within the same conceptual category as the examples.
  • Transformation-based: Transformation tasks use demonstration pairs to infer geometric, attribute, or pixel-grid rules and generate a standalone transformed query image.Multi-pair grid items require applying an abstract rule to a test input while matching the example output format.
  • Pattern & Structure: Matrix reasoning requires inferring a missing cell from row and column patterns and outputting only that cell as a standalone image.The output must match the surrounding diagram style without reproducing the full grid.
  • Pattern & Structure: Visual-spatial tasks preserve the input image and add the required annotation, such as a path, circle, line, or numeric label.Maze solving specifically draws the correct entrance-to-exit path in red without crossing walls.

B Human Validation of Automatic Evaluation

Two human studies serve distinct validation purposes: calibration tests automatic scoring reliability, while a pilot checks whether benchmark items are well-posed and solvable. The supplied passages identify the study roles and their summarized table outputs.

  • Study purposes: The calibration study validates automatic scoring on model-generated outputs, whereas the pilot study checks benchmark solvability and item quality.The studies are designed for different validation purposes.
  • Study reporting: Table 8 summarizes the two human studies, and Table 9 reports inter-annotator agreement in the human calibration study.The passages identify the tables’ reporting roles but do not provide their numerical contents.
  • Study reporting: Table 10 reports subtask-level agreement between Gemini judge scores and independent human ratings.This table concerns agreement at the subtask level.

C.1 Text Thinking and Self-Reflection

The study tests explicit text thinking and image self-reflection on a matched 200-item subset. Both interventions show heterogeneous, model-dependent effects rather than uniform gains.

  • Interventions: Explicit text thinking and one or two rounds of image self-reflection are evaluated on the same matched 200-item subset.Outputs are blindly scored with the paper’s image rubric on a 0–100 scale.
  • Explicit text thinking: Explicit text thinking improves BAGEL, Gemini, and the Emu3.5 pipeline but does not significantly improve Qwen-Image-Edit.The intervention effects depend on the model.
  • Intervention reporting: Table 11 reports intervention deltas against the matched direct-generation baseline from the corresponding evaluation pass.The table pairs each delta with its direct-generation comparison.
  • Self-reflection: Two rounds of self-reflection improve Gemini but degrade BAGEL, showing that visual feedback is not uniformly reliable across current generators.Self-reflection is therefore also model-dependent.

C.2 Per-Subtask Cascade Decomposition

The cascade decomposition evaluates Gemini 3.1 Flash Image Preview by separating perception, textual reasoning, and image generation into scored stages.

  • Gemini 3.1 Flash Image Preview is decomposed into perception, reasoning, and generation stages on a stratified n=100 subset.Perception describes inputs without solving, reasoning commits to a textual answer, and generation is the original image output.
  • Each stage is scored on a 0–5 scale, with passes defined at τ=3.
  • Table 12 reports pass rates P(P), P(R | P), and P(G | R), identifying the lowest value in each row as the bottleneck.

D Limitations

RIG-BENCH identifies future expansion needs around dataset scale, automated evaluation, and modality scope while retaining its current curated benchmark design.

  • Dataset Scale vs. Quality: RIG-BENCH prioritizes 2,000 hand-curated items for logical complexity and annotation accuracy over larger procedurally generated alternatives.The authors describe expanding volume while maintaining human-level precision as an ongoing objective.
  • Automated Evaluation Nuances: Automated LLM-as-a-judge metrics may sometimes favor stylistic coherence alongside strict logical adherence.The framework uses Gemini-3.1-Flash and is reported to align strongly with human experts, but qualitative analysis is encouraged.
  • Scope of Modality: The current benchmark focuses on static image synthesis rather than generative dynamics.Future versions are intended to extend evaluation to temporal and spatial consistency in video and 3D reasoning.

E Broader Impact

RIG-BENCH provides a diagnostic framework for assessing whether multimodal models produce logically grounded visual answers, while stronger downstream systems could introduce misuse and overtrust risks.

  • RIG-BENCH studies whether multimodal generative models produce visually correct answers through reasoning rather than surface-level plausibility.The framework may support future work on more reliable systems for education, scientific visualization, and diagrammatic reasoning.
  • Stronger reasoning-driven image generation systems could be misused for misleading visual evidence, persuasive disinformation, or plausible but incorrect diagrams.The paper recommends human oversight, transparent disclosure, and task-specific validation in sensitive settings.

F Dataset Provenance

RIG-BENCH contains 2,000 examples across 11 subtasks, with provenance records describing source datasets, splits, source-pool sizes, and final inclusion counts; supplementary figures show pilot and qualitative examples.

  • Dataset Provenance: The 11 subtasks contain 2,000 examples in total.
  • Dataset Provenance: Table 13 records each subtask’s source datasets, original split, considered source-pool size, and final benchmark count.
  • Human Pilot: Figure 6 shows the formative human pilot interface, where participants infer target images and draw or retrieve matching images.
  • Qualitative Examples: Figures 7, 8, and 9 provide additional qualitative examples of representative successes and failure modes.
Loading 2609.02864v1…