Source-linked AI summary

OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang

arXiv:2509.24900v2cs.CVcs.AI

TL;DR

Existing image-generation and editing datasets often lack systematic coverage of challenging real-world capabilities, limiting the training data available for unified multimodal models. OpenGPT-4o-Image introduces a hierarchical taxonomy and automated GPT-4o-based pipeline that produces 80,000 instruction-image pairs across 11 domains and 51 subtasks. Fine-tuning leading models yields gains across benchmarks, including 18% on ImgEdit-Bench and 13% on GenEval.

  • Problem

    Existing datasets cover basic capabilities but often lack systematic structure and challenging scenarios such as scientific imagery and complex multi-operation editing.

  • Method

    OpenGPT-4o-Image uses a hierarchical taxonomy and automated GPT-4o-based construction pipeline to generate 80,000 instruction-image pairs across 11 domains and 51 subtasks.

  • Results

    Fine-tuning leading models on OpenGPT-4o-Image improves performance across multiple benchmarks, including 18% on ImgEdit-Bench for UniWorld-V1 and 13% on GenEval for Harmon.

  • Takeaways & Limitations

    Systematic task decomposition and automated data construction provide a broad resource for advancing image generation and editing capabilities.

  • Takeaways & Limitations

    Complex prompts with many fine-grained details remain difficult to curate reliably because aesthetic quality is a poor proxy for semantic correctness and instruction following.

Abstract

from arXiv · show

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic structure and challenging scenarios required for real-world applications. To address this bottleneck, we introduce OpenGPT-4o-Image, a large-scale dataset constructed using a novel methodology that combines hierarchical task taxonomy with automated data generation. Our taxonomy not only includes fundamental capabilities such as text rendering and style control but also introduces highly practical yet challenging categories like scientific imagery for chemistry illustrations and complex instruction editing requiring simultaneous execution of multiple operations. Through an automated pipeline leveraging structured resource pools and GPT-4o, we generate 80k high-quality instruction-image pairs with controlled diversity, covering 11 major domains and 51 subtasks. Extensive experiments show that fine-tuning leading models on our dataset achieves significant performance gains across multiple benchmarks, with improvements of up to 18\% on editing tasks (UniWorld-V1 on ImgEdit-Bench) and 13% on generation tasks (Harmon on GenEval). Our work demonstrates that systematic data construction is key to advancing multimodal AI capabilities.

1 Introduction

OpenGPT-4o-Image addresses gaps in existing image-generation and editing datasets through a hierarchical taxonomy and automated data construction. It provides broad coverage of challenging capabilities and improves fine-tuned model performance across multiple benchmarks.

  • Dataset Taxonomy: OpenGPT-4o-Image expands coverage beyond style control and text rendering to scientific imagery, complex instruction editing, and spatial reasoning with causal inference.These additions target technical illustrations, simultaneous editing operations, and relational understanding beyond basic object recognition.
  • Dataset Taxonomy: The editing taxonomy covers Subject Manipulation, Text Editing, Complex Editing, Multi-Turn Editing, Other Challenging Editings, and Global Editing.These groups span local, textual, compositional, iterative, difficult, and holistic image modifications.
  • Experimental Results: 18% relative improvement on ImgEdit-Bench for UniWorld-V1 and 13% gain on GenEval for Harmon demonstrate benefits from fine-tuning on OpenGPT-4o-Image.The evaluation also covers four models and four standardized benchmarks, with consistent improvements across tested configurations.
  • Dataset Taxonomy: The dataset organizes image generation and editing into 51 fine-grained sub-capabilities across 11 major domains.Generation includes five core modules, while editing contains six categories with 21 subtasks.
  • Implications: The authors position systematic data construction as a resource for broader real-world applications and future multimodal-AI research.The stated goal is to provide a broader application-oriented resource and inspire further systematic dataset construction.

2 Related Work

Prior datasets provide large-scale image resources and editing data but often lack comprehensive instruction coverage and robust interpretation of complex edits. OpenGPT-4o-Image responds with structured task definition, hierarchical categorization, and GPT-4o-based generation of diverse, high-quality edited images.

  • Unified Multimodal Models: Unified multimodal models motivate datasets that support both semantic understanding and controllable image generation.The related-work discussion describes synergy between multimodal understanding, controllable synthesis, and reasoning with generated images.
  • Image Generation Datasets: Existing image datasets emphasize aesthetic curation, dense descriptions, or large-scale real-world and public-domain image collections, but their coverage reflects the capabilities available when created.Representative resources include LAION-Aesthetics-UMAP, DenseFusion-1M, Megalith, and Public Domain 12M.
  • Image Editing Datasets: Prior instruction-driven editing datasets often rely on open-source models, while limited human quality control can leave complex instructions inadequately interpreted.The passage specifically identifies MagicBrush and SEED-Data-Edit as incorporating varying degrees of human quality control.
  • Data Construction: The data-construction pipeline defines capabilities hierarchically, assigns difficulty grades, and generates prompts from structured resource pools and syntactic templates.Its two phases are task definition and scoping, followed by structured prompt generation for final image production.
  • Image Editing Datasets: OpenGPT-4o-Image combines high-quality source-image selection, comprehensive editing categories, and GPT-4o generation of diverse instructions and edited images.Its construction is presented as a structured alternative to datasets relying mainly on open-source generation or limited human quality control.

3 Method

OpenGPT-4o-Image organizes generation and editing capabilities into a hierarchical taxonomy and constructs training data through scalable, quality-controlled pipelines. It covers diverse practical and challenging tasks, including scientific imagery, complex instruction editing, and multi-turn interaction.

  • The generation taxonomy includes style control, complex instruction following, text rendering, spatial reasoning, and scientific imagery.
  • Scientific Imagery: Scientific Imagery contains 10k samples covering disciplines such as mathematics, physics, engineering, astronomy, and Earth science.
  • The dataset defines six editing categories and 21 subtasks spanning practical instruction-based image editing scenarios.
  • Subject Manipulation: Subject Manipulation covers local object changes while preserving background consistency through Add, Remove, Replace, Alter, and Object Extraction operations.
  • Complex Instruction Editing: Complex Instruction Editing contains 4k samples combining two to four operations drawn from Text Editing and Subject Manipulation.
  • Other Challenging Editing: Other challenging editing tasks include reference image editing, motion modification, material transformation, and object movement with spatial-coherence requirements.
  • The construction pipeline transforms capability targets into trainable data through task definition and scoping followed by automated data generation.

4 Experimental Analysis

The experiments evaluate OpenGPT-4o-Image across multiple models, benchmarks, and dataset scales, finding consistent gains for image editing and generation after fine-tuning.

  • Experimental Setup: The evaluation covers diffusion-based and autoregressive models using GenEval, DPG-Bench, GEdit-Bench, and ImgEdit-Bench.The benchmarks assess generation compositionality and semantic alignment alongside editing quality.
  • Data Scaling: Performance increases consistently as the sampled dataset size grows from 20K to 40K in the scaling experiments.The study computes an overall average across two benchmarks.
  • Performance Improvement on Image Editing Model: 18.4% and 12.0% improvements are achieved by UniWorld-V1 on ImgEdit-Bench and GEdit-Bench, respectively, after fine-tuning.MagicBrush, OmniGen, and OmniGen2 also improve across the two editing benchmarks.
  • Performance Improvement on Image Generation Model: 13.2% and 5.3% improvements are achieved by Harmon on GenEval and DPG-Bench, while OmniGen2 gains 2.5% and 1.9%.These gains are obtained using 40k samples.
  • Performance Improvement on Unified Dataset: 3.2%, 1.7%, 1.2%, and 1.1% improvements over ShareGPT-4o are observed on ImgEdit-Bench, GEdit-Bench, GenEval, and DPG-Bench, respectively.The comparison uses the same UniWorld-V1 training configuration.
  • Qualitative Evaluation: Qualitative examples show fine-tuned UniWorld-V1 executing object replacement and action modification instructions that the original model fails to perform.Fine-tuned Harmon also improves in-image text rendering, multi-entity composition, temporal and spatial reasoning, and relative size comparison.

5 Conclusion

The paper presents OpenGPT-4o-Image as a systematically constructed dataset for image generation and editing, combining hierarchical task decomposition with automated data construction. Its experiments report effectiveness across models and benchmarks, while noting limitations from GPT-4o-based generation.

  • Dataset and Contributions: OpenGPT-4o-Image contains 80,000 instruction-image pairs across 11 domains and 51 subtasks, including scientific imagery, complex instruction following, and multi-turn editing.Its taxonomy separates generation into five modules and editing into six categories with 21 subtasks.
  • Dataset and Contributions: The dataset uses a hierarchical taxonomy and a GPT-4o-based automated pipeline to provide scalable generation with controlled diversity and difficulty.
  • Experimental Validation: Experiments across multiple model architectures and benchmarks demonstrate the dataset’s effectiveness for multimodal image generation and editing.

A.1 Distribution of Image Editing Data

The image editing corpus is distributed across six categories, with larger allocations for practical subject manipulation and global editing tasks and a moderate allocation for complex instruction editing.

  • Editing Data Distribution: Subject Manipulation and Global Editing receive comparatively larger allocations because they represent practical and canonical editing tasks.
  • Generation Data Distribution: The generation corpus is organized into Style Control, Scientific Imagery, Spatial Reasoning, Complex Instruction Following, and In-Image Text Rendering.Style Control and Scientific Imagery receive the largest allocations, at 13k and 10k respectively.

B Extended Experiments

Table 4 compares fine-tuned models on DPG-Bench, with results reported for models trained on OpenGPT-4o-Image and marked baselines without fine-tuning.

  • DPG-Bench Comparison: Table 4 compares fine-tuning results for different models on the DPG-Bench benchmark.The table caption identifies DPG-Bench as the evaluation scope.

B.1 Supplementary Quantitative Experiments

Supplementary experiments show that increasing training data improves UniWorld-V1, while fine-tuning achieves strong results across editing benchmarks and outperforms ShareGPT-4o-Image in the reported unified-dataset comparison.

  • Increasing training-data volume consistently improves UniWorld-V1 performance on GEdit-Bench and ImgEdit-Bench.Table 6 reports results from incrementally larger samples and uses Two Avg to summarize both benchmarks.
  • UniWorld-V1 attains state-of-the-art results among closed-source systems on representative GEdit-Bench editing tasks.The reported tasks include color change and tone transfer.
  • Fine-tuned Harmon achieves the best reported performance on DPG-Bench.
  • The fine-tuning gains extend across additional models and benchmarks, suggesting stronger generalization rather than benefits confined to one configuration.

B.2 Supplementary Qualitative Experiments

Qualitative experiments illustrate improvements after fine-tuning, including generation gains and more faithful multi-operation image editing.

  • Fine-tuning improves image fidelity across editing types such as Add Subject and Change Background.The study also reports gains for hybrid instructions.
  • A qualitative editing example shows simultaneous laptop removal and light-blue-sofa addition after fine-tuning.

B.3 Supplementary Quantitative Experiments on Unified dataset

Comparisons of unified and task-specific fine-tuning reveal a benchmark-dependent trade-off: neither strategy dominates across all editing and generation evaluations.

  • Separate fine-tuning surpasses unified training on ImgEdit-Bench, whereas unified training outperforms separate fine-tuning on GEdit-Bench.
  • Separate fine-tuning matches unified training on GenEval-Bench but surpasses it on DPG-Bench.

C.1 Data Curation and Quality Control Strategy

The dataset uses proactive curation, hierarchical categorization, and calibrated difficulty to produce semantically accurate, visually high-quality training data for diverse generation and editing tasks.

  • Motivation: Complex prompts expose instruction-following weaknesses in both GPT-4o-generated data and existing open-source models, making simple post-hoc filtering ineffective.
  • Curation strategy: The curation strategy defines hierarchical modules and granular subclasses before generation to ensure thematic coherence and targeted data collection.
  • Curation strategy: Difficulty calibration targets instructions that challenge current open-source models while remaining solvable by GPT-4o.
  • Quality objective: The resulting approach is intended to combine high visual quality with semantic accuracy and a meaningful, well-defined challenge.
Loading 2509.24900v2…