Source-linked AI summary

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen

arXiv:2608.14546v1cs.CV

TL;DR

Existing image-editing benchmarks are too narrow to assess multi-image editing, real-world deployment, and demanding reasoning. CPI-Bench addresses this gap with three complementary subsets and task-specific VLM evaluation, and its results amplify model performance differences while aligning most closely with human Arena rankings.

  • Problem

    Existing benchmarks focus mainly on simple single-image tasks and omit reliable evaluation of multi-image, practical deployment, and reasoning-based editing.

  • Method

    CPI-Bench combines general, practical, and intelligent benchmark subsets with task-specific VLM scoring across editing-quality dimensions.

  • Results

    CPI-Bench amplifies performance discrepancies among models and achieves the highest alignment with human evaluator rankings on the Arena Image Edit Leaderboard.

  • Takeaways & Limitations

    CPI-Bench provides a broader assessment of general editing capability, practical deployment efficacy, and reasoning-based editing, supporting model optimization.

  • Takeaways & Limitations

    Proprietary model parameter counts are not publicly disclosed, and CPI-Overall is an unweighted mean across the three subset scores.

Abstract

from arXiv · show

With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.

1 Introduction

Existing image-editing benchmarks inadequately cover complex, multi-image, real-world, and reasoning-based editing. CPI-Bench addresses these gaps with three complementary subsets and a VLM-based evaluation framework designed to differentiate model performance.

  • Motivation: Existing benchmarks focus mainly on simple single-image editing and omit important high-difficulty and multi-image capabilities.Reported omissions include viewpoint change, single-subject-driven editing, subject re-orientation, and cross-image consistency.
  • Benchmark design: CPI-Bench combines general editing, practical deployment, and intelligent reasoning evaluation in three complementary subsets.CPI-General-Benchmark covers diverse editing tasks, CPI-Practical-Benchmark targets real-world consumer scenarios, and CPI-Intelligent-Benchmark evaluates demanding reasoning-based editing.
  • Benchmark design: CPI-General-Benchmark covers 30 editing tasks and includes 20 single-image and 10 multi-image tasks across 2,039 evaluation samples.The multi-image component directly addresses a gap in existing evaluations.
  • Benchmark design: CPI-Practical-Benchmark evaluates 51 application types with 558 samples across portrait enhancement, advertising, interior design, and content creation.Examples include ID photo generation, product rendering, virtual furniture placement, and multipanel story generation.
  • Evaluation: A VLM-based framework scores outputs using task-specific prompts across instruction adherence, visual naturalness, and physical and detail consistency.The framework is used to evaluate mainstream open-source and closed-source image-editing models.
  • Evaluation: CPI-Bench amplifies performance discrepancies and more sharply captures shortcomings in multi-image, real-world, and demanding reasoning scenarios.The contribution summary presents this broader evaluation as guidance for future model optimization.

2 Related work

Prior image-editing benchmarks provide incomplete coverage of multi-image, real-world, and reasoning-based editing. CPI-Bench combines broader capability coverage with practical scenarios and task-specific VLM evaluation.

  • Existing benchmarks: Existing general benchmarks emphasize narrow, abstract single-image tasks and completely overlook multi-image editing.Their taxonomies often neglect real-world deployment nuances.
  • Existing benchmarks: Existing benchmarks also lack reasoning-based evaluation, while reasoning-specific benchmarks typically focus on narrower cognitive dimensions or domains.Examples include temporal, causal, spatial, logical, knowledge-grounded, and game-world interaction evaluations.
  • CPI-Bench: CPI-Bench is designed to evaluate general editing capability, practical deployment, and reasoning-based editing intelligence together.This provides a holistic framework spanning capability, real-world use, and complex reasoning.
  • Evaluation methods: Traditional metrics such as CLIP Score, PSNR, and SSIM often correlate weakly with human perceptual judgments.Recent benchmarks therefore increasingly use Vision-Language Models as evaluators.
  • Evaluation methods: CPI-Bench extends VLM evaluation with task-specific scoring prompts tailored to the nuances of each editing category.The approach is intended to support finer-grained and more precise assessment.

3 CPI-Bench

CPI-Bench is constructed through a multi-stage process combining taxonomy design, curated imagery, human- and model-assisted instruction creation, validation, and privacy preservation. Its subsets cover general, practical, and reasoning-focused editing, with VLM-based metrics for evaluation.

  • Construction pipeline: CPI-Bench uses a rigorous four-stage construction pipeline covering taxonomy definition, image curation, instruction quality assurance, and privacy preservation.The stages are applied to ensure comprehensiveness, diversity, and ethical compliance.
  • Taxonomy definition: The taxonomy combines top-down theoretical planning with bottom-up analysis of authentic editing requests and expert classification.This process supplements standard editing capabilities with emerging and long-tail user needs.
  • Image collection: The image repository contains over 100,000 images from public datasets and legally licensed commercial channels, organized into task-specific image pools.The collection spans portraits, animals, plants, landscapes, and stylized artistic works.
  • Instruction quality: Instruction generation combines VLM-based automated drafting with human filtering, correction, augmentation, and expert validation.Reviewers check logical consistency, executability, diversity, and quality, while manually authoring difficult instructions when needed.
  • Subset coverage: CPI-Intelligent-Benchmark contains 1,181 reasoning instances spanning 8 expert domains and 67 sub-disciplines.The benchmark comparison also reports 67 intelligent-benchmark sub-tasks for demanding reasoning-based editing.
  • Evaluation metrics: The evaluation framework uses specialized VLM prompts and 1-to-5 scores across instruction adherence, visual naturalness, and physical and detail consistency, with task-specific adaptations.Text editing and style transfer receive customized dimensions suited to their requirements.

4 Evaluation

CPI-Bench produces larger performance differences than established benchmarks by testing complex editing, real-world deployment, and reasoning-intensive capabilities. Its rankings also align most closely with human-preference rankings on the Arena Image Edit Leaderboard.

  • Benchmark differentiation: CPI-Bench amplifies performance discrepancies because it evaluates multi-image editing, real-world applications, and highly demanding reasoning tasks beyond saturated single-image benchmarks.Prior benchmarks show significantly lower variance among models because they focus mainly on simple single-image editing.
  • Single-image vs. multi-image tasks: Most models score around 4.0 on fundamental single-image tasks, while only top-tier closed-source models consistently exceed 4.0 on multi-image editing.Open-source models generally plateau around 3.0 on the multi-image subset.
  • Real-world application performance: Open-source models perform substantially worse than closed-source models on CPI-Practical-Benchmark real-world deployment scenarios.The result identifies real-world business deployment as a critical improvement area for current open-source solutions.
  • Intelligence performance: CPI-Intelligent-Bench exposes weaker performance from open-source models lacking prompt engineering on tasks requiring domain knowledge, intent inference, and visual layout planning.These tasks rely heavily on model reasoning intelligence.
  • Comparison with Arena rankings: CPI-Bench rankings most closely match Arena rankings, with every model except FLUX.2-klein-9B occupying the same position or differing by one rank.It also achieves the highest Spearman correlation coefficients and lowest Mean Absolute Error among evaluated benchmarks.

5 Conclusion

CPI-Bench evaluates image editing models across editing capability, practical deployability, and complex reasoning proficiency. By incorporating multi-image, real-world, and reasoning-based challenges, it improves assessment comprehensiveness and aligns closely with human preferences on Arena.

  • Framework scope: CPI-Bench evaluates image editing models across editing capability, practical deployability, and complex reasoning proficiency.The framework integrates multi-image editing, authentic real-world scenarios, and reasoning-based challenges.
  • Evaluation contribution: CPI-Bench bridges evaluation gaps left by benchmarks confined to simple single-image tasks and helps identify model deficiencies.Its design addresses complex interactions and breaks through the performance saturation bottleneck.
  • Human-preference alignment: CPI-Bench rankings exhibit the highest alignment with human preferences among the evaluated benchmark rankings.The comparison uses the Arena Image Edit Leaderboard as the reference for human perceptual judgments.
  • Practical implication: CPI-Bench offers a roadmap for future model optimization grounded in human consensus.The paper presents it as a tool for precisely identifying model deficiencies.
Loading 2608.14546v1…