Source-linked AI summary

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz, Hanhe Lin

arXiv:2608.25933v1cs.CVcs.AI

TL;DR

State-of-the-art text-to-image models show pronounced defects for complex compositional prompts involving multiple entities, attributes, spatial relations, and interactions, motivating study of how humans identify these defects. The paper conducts a subjective study across 651 compositional prompts and three T2I models, producing the CO-AID dataset; proof-of-concept experiments show that models trained on it can predict defects and optimize image generation.

  • Problem

    State-of-the-art T2I models degrade on complex compositional views, while composition in AI-generated images lacks an explicit definition and human defect identification requires systematic study.

  • Method

    The authors generate 651 compositional prompts from reference images across people, hand, object, and scene categories, use three T2I models, and have 29 participants identify defect types and locations.

  • Results

    Training a deep model on CO-AID enables prediction of defects in AI-generated images and optimization of AI image generation.

  • Takeaways & Limitations

    CO-AID provides a benchmark for systematically identifying and categorizing compositional defects in AI-generated images.

Abstract

from arXiv · show

*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple entities and multiple attributes. In this paper, we investigate how humans identify such defects. Specifically, we manually select 651 reference images from the four categories of people, hand, object, and scene that exhibit complex compositional characteristics, from which prompts emphasizing compositional factors are derived by manually editing ChatGPT-generated prompts. We then feed the prompts into three selected T2I models to generate AI images and conduct a comprehensive subjective study to identify their defects. For each image, 29 participants provide multi-label assessments specifying defect types and locations. The study yields the compositional AI-generated image defect (CO-AID) dataset, including reference images, prompts, AI-generated images, and information on defect locations and types. Experimental results show that training a deep model on CO-AID can both predict defects in AI-generated images and optimize AI image generation, demonstrating its usability and effectiveness. The database and supplementary materials are available at: https://github.com/Future-IQA/CO-AID .

I. INTRODUCTION

T2I models perform well on simpler prompts but degrade on complex compositional scenes, motivating a human-centered study and the CO-AID dataset for identifying defects.

  • T2I performance degrades significantly when prompts involve complex views and compositional factors.
  • Highly compositional images involve multiple entities, multiple attributes, multidimensional spatial relations, and complex interactions.
  • The study generates 651 compositional prompts from reference images in the people, hand, object, and scene categories and evaluates outputs from Midjourney, Imagen, and FLUX.
  • Twenty-nine participants identify compositional defect types and locations in generated images through a subjective study.
  • The paper proposes a framework that systematically studies compositional defects in AI-generated images.
  • The CO-AID dataset supports defect prediction and AI image-generation optimization when used to train a deep model.

II. AI-GENERATED COMPOSITIONAL IMAGE COLLECTION

The collection workflow selects compositional reference images, converts them into refined prompts, generates AI images with three T2I models, and assesses the outputs in a human study.

  • 651 compositional reference images were manually selected across people, hand, object, and scene categories.
  • ChatGPT-generated descriptions were manually edited into compositional prompts before image generation.
  • The prompts were fed into Midjourney, Imagen, and FLUX to generate AI images for subjective study.

A. Reference Image Selection

Reference images were selected from four categories likely to exhibit compositional characteristics, with hands treated separately because of their structural and interaction complexity.

  • The reference set covers people, hand, object, and scene, four categories likely to exhibit compositional characteristics.
  • Hands are treated as an independent category because they typically combine high structural complexity with rich interaction semantics.

B. Compositional Prompt Generation

Prompts were created by refining ChatGPT descriptions so they explicitly encode entities, attributes, spatial relations, and interactions.

  • Manual refinement required prompts to describe multiple entities and their attributes, including appearance, posture, and state.
  • Prompts explicitly specify two- or three-dimensional spatial relations, including relative positions and occlusion.
  • Prompts explicitly describe interactions among entities or between entities and their environments, such as physical contact and hand-object interaction.

C. Compositional Image Generation

The study generated compositional AI images with selected state-of-the-art T2I models after evaluating candidate models on sampled compositional prompts. The resulting images supported a human defect-identification study.

  • Study procedure: The selected T2I models generated the compositional images used for the human perceptual defect-identification study.The study examined defects in images produced from the constructed compositional prompts.
  • Model selection: A preliminary comparison evaluated six T2I models using 80 randomly sampled compositional prompts.The considered models included Stable Diffusion, Imagen, DALL·E, Midjourney, Firefly, and FLUX.
  • Image generation: The 651 compositional prompts were randomly divided into three groups, each assigned to one selected T2I model.The generated images were allocated to training, pilot, and main-study subsets.
  • Study procedure: The study overview comprised participant introduction, training, and an experiment for visually identifying global- and local-level defects.A pilot study refined instructions, optional defect reasons, and the graphical interface before the main study.

A. Study Procedure

Participants were introduced to the study, trained with example images, and then inspected generated images for global and local compositional defects. The protocol defined defect reasons and localized defects through category-specific annotations.

  • Study procedure: Participants completed introduction, training, and experiment stages before providing comprehensive feedback on compositional defects.Training used 11 images and displayed researcher-suggested defects as references without requiring conformity.
  • Defect identification: Participants marked no noticeable defect or reported defects at global and/or local levels after carefully inspecting each image.Global reporting was used when defects were difficult to locate or too numerous for practical individual annotation.
  • Global defects: Global defect reasons covered visual style, blur or missing details, abnormal text, commonsense violations, abnormal entities, anomalous interactions, and spatial anomalies.The composition-related categories addressed entity attributes, inter-entity interactions, and spatial relationships.
  • Local defects: Local defects were grouped into face, hair/fur, hand, body, and object categories that accommodate instances across four content categories.Hand-specific reasons included abnormal structure, finger count, posture, and nail structure.

A. Experimental Setup

The experiment used pilot and main studies with participant reliability checks, fatigue mitigation, and outlier removal. The main analysis retained 23 participants after excluding unreliable responses.

  • Participants and images: 40 images and 15 participants were used in the pilot, while 600 images and 29 participants were used in the main study.The main-study images were divided into four batches, with breaks allowed after each batch.
  • Participants and images: Twelve images were repeated to assess participant reliability, producing 153 images in each batch.The repeated images supported verification of participant responses during the main study.
  • Reliability control: Participant reliability was evaluated using engagement, self-consistency, and conformity criteria.These criteria considered annotation time, agreement on repeated images, and agreement with the majority.
  • Proof of concept: The proof-of-concept results visualized input images, human ground-truth defect maps, model detections, and defect-guided repairs.The figure compared GPT-Image-1 and a fine-tuned TranSalNet model for defect localization.
  • Reliability control: 6 participants were identified as outliers, leaving 23 participants for subsequent analysis.Participants failing the defined reliability thresholds were removed from the analysis.

C. Result Analysis

Human feedback grouped the 600 generated images into no-defect, global-defect, and local-defect categories. Defect-free generation was especially difficult for people and hands, while model outputs showed different defect profiles.

  • Overall categorization: 600 images were grouped into No defect, Global defect, and Local defect categories based on feedback from 23 participants.Each category reflected the majority participant judgment at the corresponding defect level.
  • Category statistics: 18 defect-free images depicted people and hands, compared with 136 depicting objects and scenes.The reported counts indicate a much smaller defect-free set for people and hands.
  • Model comparison: Imagen generated nearly 70 defect-free images, outperforming Midjourney and FLUX on that count.Participants nevertheless labeled half of Imagen’s generated images as Global defect because of unrealistic and CG-style appearance.
  • Model comparison: 72 Midjourney images were labeled as Local defect despite its greater tendency to produce realistic images.The reported local defects were associated with noticeable detail issues.
  • Model comparison: FLUX performed poorly for people, hand, and scene images, except for object images.This category-specific pattern contrasts with the stronger result reported for its object outputs.

D. Proof-of-concept Experiment

The proof-of-concept experiment uses human-annotated defect locations to train a saliency model for defect prediction and evaluates whether predicted maps can guide image repair. Results indicate that CO-AID supports both more human-like defect localization and noticeable quality improvement after repair.

  • Experiment setup: Human-annotated local defects are converted into ground-truth defect maps for model training.The experiment uses AI-generated images and corresponding defect maps to fine-tune TranSalNet, originally trained for attention saliency prediction.
  • Defect prediction: The fine-tuned TranSalNet produces more human-like defect localization patterns than GPT-Image-1 in sampled comparisons.Figure 5 compares results across four sampled images, including hand defects.
  • Image repair: Repaired images guided by predicted defect maps correct defects and show noticeable quality improvement.This demonstrates that the collected data can support image-generation optimization in addition to defect prediction.
  • Implications: The study presents CO-AID as a usable and effective benchmark for identifying defects in compositional AI-generated images.The authors describe the experiment as a proof of concept and identify generalized defect prediction as future work.
Loading 2608.25933v1…