Source-linked AI summary

AGIQA-3K: An Open Database for AI-Generated Image Quality Assessment

Chunyi Li, Zicheng Zhang, Haoning Wu, Wei Sun, Xiongkuo Min, Xiaohong Liu, Guangtao Zhai, Weisi Lin

arXiv:2306.04717v2cs.CVcs.AIeess.IV

TL;DR

AGI quality varies substantially across models, prompts, and parameters, motivating quality models aligned with human subjective ratings. The paper builds AGIQA-3K with diverse generated images and fine-grained perceptual and alignment scores, then benchmarks assessment models and proposes StairReward. The database supports comparisons of assessment consistency, while StairReward performs especially well for alignment on long prompts.

  • Problem

    Wide variation in AGI quality creates a need for quality models consistent with human subjective ratings.

  • Method

    The paper constructs AGIQA-3K from diverse models, prompts, and parameters, collects perceptual and alignment scores, and benchmarks assessment metrics.

  • Results

    StairReward far outperforms other methods on text-to-image alignment for long prompts and leads the alignment index of AGIQA-3K.

  • Takeaways & Limitations

    Fine-grained AGIQA-3K scores support evaluating and comparing AGI quality across perceptual and text-to-image alignment dimensions.

Abstract

from arXiv · show

With the rapid advancements of the text-to-image generative model, AI-generated images (AGIs) have been widely applied to entertainment, education, social media, etc. However, considering the large quality variance among different AGIs, there is an urgent need for quality models that are consistent with human subjective ratings. To address this issue, we extensively consider various popular AGI models, generated AGI through different prompts and model parameters, and collected subjective scores at the perceptual quality and text-to-image alignment, thus building the most comprehensive AGI subjective quality database AGIQA-3K so far. Furthermore, we conduct a benchmark experiment on this database to evaluate the consistency between the current Image Quality Assessment (IQA) model and human perception, while proposing StairReward that significantly improves the assessment performance of subjective text-to-image alignment. We believe that the fine-grained subjective scores in AGIQA-3K will inspire subsequent AGI quality models to fit human subjective perception mechanisms at both perception and alignment levels and to optimize the generation result of future AGI models. The database is released on https://github.com/lcysyzxdxc/AGIQA-3k-Database.

I. INTRODUCTION

AGI quality varies across models, parameters, and prompts, creating a need for standardized, fine-grained subjective evaluation. AGIQA-3K addresses this need with a multi-model database, perceptual and alignment scores, and benchmark assessment.

  • Motivation: AGI quality varies widely across models and can also differ substantially within the same model.Training data, epoch iterations, and prompt design affect generated results.
  • Challenges: Large-scale subjective assessment must select models, parameters, and prompts that cover diverse AGIs under limited data.The diversity of T2I models and generated images makes fine-grained scoring across all combinations difficult.
  • Contributions: AGIQA-3K contains 2,982 AGIs generated from 6 different models spanning GAN, autoregressive, and diffusion-based models.The database carefully designs and adjusts input prompts and internal model parameters.
  • Contributions: A standardized laboratory experiment collected Mean Opinion Scores for perceptual quality and text-to-image alignment.These dimensions support comparisons of different AGI models across quality aspects.
  • Contributions: A benchmark evaluates perceptual-quality and alignment metrics, while StairReward is proposed to improve text-to-image alignment assessment.The benchmark compares current assessment models with subjective evaluation results.

II. RELATED WORK

Existing AGI quality databases and metrics provide useful scale or subjective scoring but have limited model coverage, scoring granularity, or multidimensional characterization. These limitations motivate a fine-grained database covering both perception and alignment.

  • Quality assessment: Traditional perceptual metrics such as IS, FID, and KID generally evaluate groups of images rather than individual AGI quality.IQA methods are used for individual images, but AGI-specific quality factors limit their reliability.
  • Quality assessment: CLIP-based alignment metrics connect text and images, but training difficulty often restricts users to pretrained parameters and small-scale databases.The diverse morphology of generated content makes alignment assessment challenging.
  • Existing databases: DiffusionDB offers more than 1.8 million text-image pairs but provides no subjective scoring and uses only Stable Diffusion.Its scale nevertheless supports later subjective database construction.
  • Existing databases: AGIQA-1K provides fine-grained MOS but uses only 180 simple prompts, limiting representation of diverse AGIs.Its prompts are combinations of real-world image labels.
  • Existing databases: Pick-A-Pic and HPS expand images and prompts but use diffusion-only data and combine perception and alignment into one overall score.A single score cannot characterize AGI quality across multiple dimensions.
  • Existing databases: ImageReward includes more model types and subjective testing, but its discrete scores and one rating per image produce coarse-grained quality characterization.Its model coverage also omits GAN-based models and includes only four strong diffusion models.
  • Research gap: A fine-grained database for both perception and alignment should cover more AGI models, performances, parameters, and refined scoring granularity.This scope directly addresses the limitations identified in prior databases.

III. DATABASE CONSTRUCTION

AGIQA-3K constructs a diverse database by sampling six representative generative models, varying generation settings, and designing prompts around subject, detail, and style. Distribution analysis compares AGIs with natural-scene images across five quality-related attributes.

  • AGI Model Collection: Six representative models cover GAN, autoregressive, and diffusion-based generation, including AttnGAN, DALLE2, GLIDE, Midjourney, Stable Diffusion, and Stable Diffusion XL.Four selected models are diffusion-based, supplemented by one GAN-based and one autoregressive-based model.
  • Distribution Analysis: AGIs and natural-scene images have similar distributions for four attributes, while blur differs because insufficient iterations frequently produce blurred images.AGIQA-3K has a sharper distortion distribution than AGIQA-1K because it adds insufficient-iteration data.
  • AGI Model Collection: Midjourney and Stable Diffusion use reduced iteration settings to simulate distortions caused by insufficient generation iterations.These adjustments broaden the database’s represented quality range.
  • Distribution Analysis: AGIQA-3K compares normalized distributions of lighting, contrast, color, blur, and spatial information between natural and generated images.Color denotes colorfulness, while spatial information denotes content diversity.

B. AGI Prompt Collection

AGIQA-3K addresses fine-grained scoring constraints by combining real and human-designed prompts, then collecting standardized perception and alignment ratings. Its prompt design covers subject, detail, and style while its laboratory protocol produces normalized MOS values.

  • B. AGI Prompt Collection: Fine-grained scoring limits AGIQA-3K to relatively few prompts, making broad coverage of real user inputs a central collection challenge.The database cannot conduct generation and scoring on more than ten thousand prompts as in prior coarse-grained databases.
  • B. AGI Prompt Collection: The prompt collection uses a ‘real’ + ‘human designed’ mechanism that combines real AGI prompts with manual composition.
  • B. AGI Prompt Collection: Prompts are divided into subject, detail, and style items following the Stable Diffusion official prompt-book structure.Subject is present in all prompts and is treated as the most important item.
  • C. Subjective Experiment: The experiment scores AGIs separately at perception and alignment levels, with alignment measuring compatibility between images and prompts.The prompt includes subject, detail, and style, with subject generally more critical than the other items.
  • C. Subjective Experiment: The interface presents images in random order and records two slider ratings from 0 to 5 with a minimum interval of 0.1.The standardized laboratory setup used an iMac monitor with resolution up to 4096 × 2304.
  • C. Subjective Experiment: Twenty-one graduate students rated 2,982 images across 14 sessions, producing 125,244 quality ratings before MOS post-processing.
  • C. Subjective Experiment: Raw viewer scores are converted to Z-scores using session statistics, then rescaled and averaged to compute each image’s MOS.

D. Subjective Data Analysis

The analysis examines how AGI quality varies with models, prompt length, style, and internal parameters. It finds trade-offs between perception and alignment, strong style dependence, and sensitivity to CFG and iteration count.

  • AGI model choice strongly affects generation quality even when the input prompt is unchanged.
  • Midjourney maintains satisfying perception across prompt lengths but loses alignment as it ignores more prompt items.
  • Stable Diffusion XL keeps alignment relatively stable across prompt lengths while showing a downward trend in perception quality.
  • Longer prompts reduce subjective quality in mainstream AGI models, particularly at the alignment level.
  • Specialized styles such as ‘Abstract’ and ‘Sci-fi’ receive poorer results, while Midjourney shows comparatively stronger versatility across styles.
  • Subjective quality follows ‘Baroque’ > ‘Anime’ and ‘Realistic’ > ‘Abstract’ and ‘Sci-fi’ for both perception and alignment.The supplied figure passage summarizes the ordering as ‘Baroque’>‘Anime&Realistic’>‘Abstract&Sci-fi’.
  • CFG and iteration count materially affect AGI quality; increasing CFG favors alignment over perception, while insufficient iterations leave intermediate results.
  • Halving Midjourney’s iterations causes a significant quality drop, with perception nearly reaching GLIDE’s level.

A. Framework

StairReward models text-to-image alignment at the morpheme level rather than treating the entire prompt as one unit. It splits prompts into differently weighted morphemes to better match human saliency.

  • StairReward decomposes alignment assessment into morphemes instead of using the entire prompt as the assessment unit.
  • Earlier prompt morphemes receive greater importance in subjective alignment scores than later morphemes.
  • The method uses a prompt segmentation function based on prepositions and punctuation to produce morphemes with different importance.

C. Image Cutting

StairReward locates morpheme-related image regions through center-based crops of increasing size, aligns each crop with its morpheme, and combines weighted scores with a whole-image score. The selected stair images improve agreement between objective and subjective alignment scores.

  • The method assumes image centers contain more information than edges and samples centered boxes to create stair-images.
  • For three morphemes, the first-morpheme clip score’s growth slows after the box length reaches 0.5.
  • For three morphemes, SRoCC turning points occur at box lengths 0.5, 0.75, and 1.
  • ImageReward computes alignment scores between each morpheme and its corresponding sub-image.
  • Later morpheme scores receive half the weight of preceding scores, reflecting their lower impact on overall alignment.
  • The final score combines one-to-one morpheme–stair alignment scores with the alignment score between the full image and full prompt.

V. EXPERIMENT RESULTS

The benchmark evaluates perceptual and text-to-image alignment metrics against subjective MOS using correlation measures and fitted score mapping. It compares broad families of no-reference perception models and established alignment models.

  • Evaluation protocol: SRoCC and KRoCC measure prediction monotonicity, while PLCC measures prediction accuracy against subjective MOS.A five-parameter logistic function maps predicted scores to MOS-compatible fitted scores.
  • Perception metrics: No-reference perceptual metrics are selected because T2I AGI tasks lack reference images.The benchmark therefore focuses on metrics that assess generated images without reference inputs.
  • Perception metrics: Handcrafted models extract image-quality features using prior knowledge.The evaluated models are CEIQ, DSIQA, NIQE, and Sisblim.
  • Perception metrics: FID, ICS, and KID are loss-function metrics, with FID and KID measuring distance from the MS-COCO database.These metrics are commonly used during AGI model iteration.
  • Perception metrics: SVR-based models combine handcrafted features with Support Vector Regression to represent perceptual quality.The benchmark includes BMPRI, GMLF, and HIGRADE.
  • Compared models: Deep-learning perception metrics learn quality-aware information from labeled data, while alignment comparisons include CLIP, ImageReward, HPS, PickScore, and StairScore.The DL perception models are DBCNN, CLIPIQA, CNNIQA, and HyperNet.

B. Experiment Results and Discussion

Experiments show that perceptual models align better with human scores overall than loss-function metrics, but performance varies across model quality, prompt length, and style. Alignment models remain weaker in challenging subsets, although StairReward leads on long prompts and overall alignment.

  • Perception results: About 0.8 overall SRoCC is achieved by DL perception models, but the best model reaches only about 0.5 on individual AGI-model subsets.Perception performance differs substantially across low- and high-quality content.
  • Prompt-length analysis: Perceptual quality prediction generally decreases as prompts become longer.The reported explanation combines increased generation difficulty with weaker prediction for low-quality content.
  • Style analysis: Perceptual models predict quality well for unpopular styles but less satisfactorily for popular styles.The style groupings reflect similarities between Abstract and Sci-fi, and between Anime and Realistic styles.
  • Alignment results: Alignment models need substantial improvement for good-model images, long prompts, and popular styles.These conditions are identified as difficult alignment-assessment settings.
  • Alignment results: StairReward far outperforms other methods on long-prompt alignment and leads the overall AGIQA-3K alignment index.The reported advantage is attributed to reasonable prompt disassembly.
  • Limitations: Current perception and alignment models distinguish excellent from poor AGIs better than AGIs with similar subjective quality.The paper identifies fine-grained discrimination among similarly rated images as an urgent problem.
  • Implications: DL-based IQA models show excellent agreement with human subjective scores, motivating their consideration in future T2I generation.The paper contrasts these models with traditional loss functions.
  • Limitations: Existing alignment models lag behind perception assessment, and more accurate quality models are still needed.StairReward improves alignment assessment only to some extent.

C. Ablation Study

The ablation study tests prompt segmentation and image cutting within StairReward. Removing either module degrades performance, supporting contributions from both factors.

  • Ablation findings: Removing either prompt segmentation or image cutting degrades StairReward performance.The ablation factors are identified as word-level prompt segmentation and image-level cutting.
  • Ablation findings: The ablation results confirm that both modules contribute to the performance reported in the alignment experiment.The comparison is linked to the results in Table III.
  • Conclusion: The conclusion reports that current perception and alignment models cannot handle the AGIQA task well, especially existing alignment models.StairReward is proposed to objectively evaluate alignment quality in response to this limitation.
Loading 2306.04717v2…