Source-linked AI summary

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

Xuehai Bai, Yang Shi, Yi-Fan Zhang, Xuanyu Zhu, Yuran Wang, Yifan Dai, Xinyu Liu, Yiyan Ji, Xiaoling Gu, Yuanxing Zhang

arXiv:2605.13062v1cs.CV

TL;DR

Existing benchmarks provide limited task coverage and coarse or unstable evaluation that can diverge from human judgment, while reward-model benchmarks may not reflect practical RL scenarios. This paper introduces unified benchmarks with progressively challenging tasks, fine-grained rubric-based evaluation, and realistic preference pairs, revealing substantial performance differences across model types and capabilities.

  • Problem

    Existing image-editing benchmarks have limited task coverage and evaluation reliability, making fine-grained quality and human-aligned assessment difficult for complex editing scenarios.

  • Method

    The paper introduces Edit-Compass and EditReward-Compass, combining progressively challenging annotated editing tasks, rubric-guided multidimensional evaluation, and realistic preference pairs for reward-model assessment.

  • Results

    Closed-source models substantially outperform open-source models on image editing, with the best proprietary model scoring 3.99 versus 2.69 for the strongest open-source model.

  • Takeaways & Limitations

    The benchmarks provide more reliable and interpretable assessment of complex image editing and reward-model capabilities across tasks and evaluation dimensions.

Abstract

from arXiv · show

Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for strong frontier models, due to limited task difficulty and coarse-grained evaluation protocols. In parallel, reward models have become increasingly important for RL-based image editing optimization, yet existing reward model benchmarks still rely on unrealistic evaluation settings that deviate from practical RL scenarios. These limitations hinder reliable assessment of both image editing models and reward models. To address these challenges, we introduce Edit-Compass and EditReward-Compass, a unified evaluation suite for image editing and reward modeling. Edit-Compass contains 2,388 carefully annotated instances spanning six progressively challenging task categories, covering capabilities such as world knowledge reasoning, visual reasoning, and multi-image editing. Beyond broad task coverage, Edit-Compass adopts a fine-grained multidimensional evaluation framework based on structured reasoning and carefully designed scoring rubrics. In parallel, EditReward-Compass contains 2,251 preference pairs that simulate realistic reward modeling scenarios during RL optimization.

1 Introduction

The paper introduces Edit-Compass and EditReward-Compass to address gaps in evaluating advanced image editing and reward models. The unified suite expands task coverage and annotation while targeting realistic reinforcement-learning evaluation settings.

  • Motivation: Existing benchmarks increasingly diverge from human judgment as image editing models gain multimodal understanding, complex reasoning, and multi-image editing capabilities.The paper frames accurate evaluation as increasingly challenging for frontier models.
  • Motivation: Existing reward model benchmarks suffer distribution mismatch between evaluation samples and edited images encountered during RL training.This mismatch limits faithful assessment of reward models supporting FlowGRPO-based image editing optimization.
  • Proposed benchmark: Edit-Compass contains 2, 388 carefully annotated instances spanning six progressively challenging task categories and capabilities including general editing, world perception, dynamic manipulation, visual reasoning, and multi-image understanding.It is presented as part of a unified evaluation suite for image editing and reward models.
  • Validation: Evaluations cover 29 image editing models and 21 reward models, including proprietary frontier systems and leading open-source image editing models.Named proprietary models include Nano-Banana Pro, Wan2.7-Image, and Seedream 4.5; named open-source models include Qwen-Image-Edit and Joy-Image-Edit.

2 Related Work

Related work is limited by narrow image-editing task coverage, unreliable evaluation metrics, and reward-model benchmarks that often use unrealistic preference-pair settings. These gaps motivate broader, more reliable assessment of editing quality and reward models.

  • Image Editing Benchmarks: Existing image-editing benchmarks mainly cover narrow tasks and rely on automated metrics such as CLIP-I and DINO-I.These limitations are identified as insufficient task coverage and evaluation reliability.
  • Image Editing Benchmarks: CLIP-I and DINO-I often miss fine-grained editing quality in world-knowledge, visual-consistency, and complex-instruction-following tasks.The passage specifically highlights these capabilities as difficult for such metrics to capture.
  • Reward Model Benchmarks: Existing reward-model benchmarks typically construct preference pairs from limited editing tasks or outputs generated by different models.The passage states that these evaluation settings often deviate from practical reward-model scenarios.

3 Edit-Compass

Edit-Compass organizes image editing evaluation across 36 diverse tasks spanning general, dynamic, world-knowledge, algorithmic, multi-image, and complex editing settings. It combines these tasks with an MLLM-as-judge framework evaluating instruction awareness, visual consistency, and visual quality.

  • Task Coverage: 36 diverse image editing tasks span single-image and multi-image settings, covering general editing and algorithmic visual reasoning.Figure 1 indicates the number of examples for each representative task type.
  • Task Coverage: World Knowledge Reasoning tasks test temporal, causal, game, mathematical, and chemical reasoning for complex image edits.These tasks require models to leverage real-world and domain-specific knowledge to infer intended changes.
  • Task Coverage: Algorithmic Visual Reasoning tasks require interpreting visual structures, performing multi-step reasoning, and rendering solutions through image editing.Examples include optimal path, convex hull, maximum submatrix sum, and knapsack selection identification.
  • Task Coverage: Multi-Image-Aware Editing introduces reference-driven edits based on fine-grained attributes such as object properties, actions, orientations, and colors.The category also includes Multi-Image Composition and Virtual Try-On.
  • Evaluation Framework: Evaluation uses three dimensions—Instruction Awareness, Visual Consistency, and Visual Quality—through an MLLM-as-judge pipeline producing scalar scores and fine-grained rationales.The dimensions assess instruction adherence and world knowledge, preservation of unrelated content and identity, and visual plausibility, coherence, and artifact severity.

4 EditReward-Compass

EditReward-Compass is a benchmark for systematically evaluating image-editing reward models using 2,251 preference pairs and the same rubric-based judging framework as Edit-Compass. Its construction simulates reward modeling during RL optimization and applies multi-dimensional human annotation to ensure preference-pair quality.

  • Benchmark design: 2,251 preference pairs form EditReward-Compass, each combining an editing instruction with two candidate edited images.The benchmark is designed for systematic reward-model evaluation.
  • Construction: The benchmark reuses Edit-Compass instructions and simulates RL optimization with a FlowGRPO-inspired sampling strategy, stochastic differential equations, and outputs from six image editing models.Denoising steps are controlled during candidate sampling.
  • Human annotation: A two-stage human annotation pipeline selects preference pairs across instruction adherence, visual consistency, and visual quality.The benchmark particularly emphasizes instruction adherence and visual consistency because of image-editing evaluation complexity.
  • Human annotation: Eight image-editing experts participate in the annotation process.The supplied passage identifies the experts as part of the two-stage quality-assurance pipeline.

5 Experiments

Experiments benchmark 29 image editing models and evaluate Edit-Compass and EditReward-Compass across model performance, human alignment, visual reasoning, language transfer, prompting, and thinking-enabled inference. Results identify Qwen-Image-Edit as the strongest open-source editor, persistent reasoning and perception challenges, and consistent gains from tailored prompts and thinking-enabled evaluation.

  • Experimental Setup: 29 models—25 open-source and 4 proprietary—are benchmarked across diverse recent image editing paradigms.The open-source set spans diffusion-based, unified multimodal, and other architectural families.
  • Image Editing Model Results: Qwen-Image-Edit achieves the best overall open-source performance under both English and Chinese instructions on Edit-Compass.The paper attributes this result to integrating a 20B diffusion transformer with a 7B Qwen-VL model.
  • Human-Aligned Evaluation Protocol: Edit-Compass agrees more strongly with human preferences than existing benchmarks in benchmark-level evaluation.Experts ranked OmniGen2-generated edits from ImgEdit-Bench, GEdit-Bench, RISE-Bench, and Edit-Compass.
  • Visual and Algorithmic Reasoning: Open-source models struggle most on complex perception and algorithmic visual reasoning tasks, particularly Object Swap, Complex Paint, and derived edits.Closed-source models show some potential but remain limited overall, while visual consistency and quality are harder reward-model dimensions than instruction awareness.
  • Impact of System Prompts: 12.93% is the largest gain from EditReward-Compass system prompts over EditScore prompts, achieved on Qwen3-VL-8B.The prompts improve performance across all evaluation dimensions on corresponding single-image subsets.
  • Effect of Thinking-Enabled Inference: 9.83 points and 10.56 points are the largest thinking-enabled gains for Qwen3.5-9B and Qwen3.6-35B-A3B, respectively.Thinking-enabled inference consistently improves reward-model evaluation performance across the reported model groups.

6 Conclusion, Discussion, and Limitations

The paper introduces Edit-Compass and EditReward-Compass as a unified benchmark suite for evaluating frontier image editing systems and reward models. Edit-Compass combines broad, progressively challenging task coverage with fine-grained multidimensional evaluation using structured reasoning and scoring rubrics.

  • Benchmark Suite: Edit-Compass and EditReward-Compass form a unified benchmark suite for evaluating frontier image editing systems and reward models.The contribution is framed as a joint evaluation resource for both editing systems and reward models.
  • Edit-Compass: 2,388 carefully annotated instances span 36 progressively challenging task categories in Edit-Compass.The categories cover general editing, world knowledge reasoning, visual reasoning, dynamic manipulation, and multi-image editing.
  • Edit-Compass: Edit-Compass covers general editing, world knowledge reasoning, visual reasoning, dynamic manipulation, and multi-image editing.These capabilities are represented across the benchmark’s progressively challenging task categories.
  • Evaluation Framework: The benchmark uses a fine-grained multidimensional evaluation framework based on structured reasoning and scoring rubrics.This framework is designed to support detailed evaluation beyond task coverage alone.

A Edit-Compass Data Construction

Edit-Compass is constructed through three main components, as illustrated in Figure 2.

  • A Edit-Compass Data Construction: Edit-Compass construction consists of three main components.The paper presents this construction overview in Figure 2.

A.1 General and Complex tasks.

The General and Complex task categories use permissively licensed, high-quality images reviewed by five experts and generate diverse editing instructions through a Gemini 3 Pro-based platform.

  • Data collection and review: Images are collected from Unsplash, Pexels, Pixabay, and Freepik under permissive licenses and reviewed for safety, quality, and editing suitability.Only real, high-quality images are considered for inclusion.
  • Data collection and review: An image is retained only when all five human reviewers approve it.The reviewers assess the image from multiple perspectives, including safety, image quality, and suitability for editing.
  • Instruction generation: A Gemini 3 Pro-based instruction generation platform is established to produce diverse editing instructions.

A.2 Dynamic Manipulation, World Knowledge Reasoning, and Multi-Image Tasks

For Dynamic Manipulation, World Knowledge Reasoning, and Multi-Image tasks, experts define tasks and editing instructions, refine source-image descriptions with Gemini 3 Pro, and validate image–instruction pairs through human assessment.

  • Task Construction: Experts in image editing conduct in-depth discussions to define each task and construct coarse-grained source-image descriptions with corresponding editing instructions.The process begins with expert task definition and source-image characterization.
  • Source-Image Generation: Gemini 3 Pro refines the coarse-grained source-image descriptions, which are then used to generate source images.The refined descriptions guide source-image generation.
  • Quality Assessment: Multiple human experts assess whether each generated image–instruction pair is feasible and valid.Human review follows image generation.

A.3 Algorithmic Visual Reasoning tasks.

Algorithmic Visual Reasoning tasks are constructed through expert-defined editing problems, programmatic image rendering with ground-truth annotations, and category-specific instruction templates. This process ensures each instance has a well-defined visual structure and unambiguous evaluation target.

  • Algorithmic Visual Reasoning task construction: Human experts define the editing problems, Python programs render source images and annotations, and experts create category-specific instruction templates for each case.The templates are applied to corresponding rendered cases to structure the task instances.

A.3.1 Longest Word Discovery … A.3.10 Global Word Path Recovery

Appendix A.3 defines ten procedurally generated visual reasoning and editing tasks spanning constrained word search, combinatorial optimization, geometric identification, path planning, and grid-based recovery. The tasks use algorithmic construction and verification to ensure valid, difficult, and well-defined targets, including global variants that remove local guidance.

  • A.3.1 Longest Word Discovery: Longest Word Discovery embeds curated, morphologically or semantically complex English words in grids and verifies the longest reachable word under downward/rightward traversal.A Python reconstruction pipeline provides precise control over textual content, layout, and visual attributes while reducing character-level inaccuracies.
  • A.3.2 Global Longest Word Discovery: Global Longest Word Discovery searches from any grid position and retains samples only when the uniquely verified global longest word exactly matches the embedded target.This extends the fixed-start setting while preserving dictionary-based verification and well-defined ground truth.
  • A.3.3 Knapsack Selection: Knapsack Selection asks models to select visually presented objects maximizing total value within a budget, with procedurally sampled instances solved by dynamic programming and backtracking.Capacities range from 10 to 20, item counts from 6 to 8, weights from 2 to 8, and values from 10 to 50.
  • A.3.4 Optimal Path Identification: Optimal Path Identification requires overlaying a minimum-cost 4-neighbor path between designated endpoints on terrain grids with road, grass, water, and impassable wall costs.Procedural grids have side lengths uniformly drawn from 6 to 10, with traversal costs 1, 3, 8, and +∞, respectively.
  • A.3.5 Convex Hull Identification / A.3.6 Maximum Submatrix Sum Identification: Convex Hull Identification recovers the polygon enclosing procedurally sampled planar points, while Maximum Submatrix Sum Identification locates the fixed-size rectangular region with the greatest grid-value sum.Convex-hull instances sample n ∼U{8, 15} points on a 10 × 10 canvas; submatrix instances use n ∼U{6, 10}, values in [−9, 9], and kernel dimensions from 2 to min(4, n −1).
  • A.3.7 Maximum Bonus Identification: Maximum Bonus Identification finds the highest-reward monotone path between designated cells in integer grids, computing ground truth with dynamic programming and recovering the path by backtracking.Instances use n ∼U{5, 12}, entries in [−5, 9], and endpoints separated by Manhattan distance at least 3.
  • A.3.8 Numberlink Path Identification: Numberlink Path Identification constructs one non-intersecting path for each colored endpoint pair on procedurally generated grids using sequential self-avoiding random walks.Grid sizes are n ∼U{6, 9}, and the number of path pairs is m ∼U{3, 5}.
  • A.3.9 Word Path Recovery / A.3.10 Global Word Path Recovery: Word Path Recovery locates a specified word through downward/rightward moves, whereas Global Word Path Recovery removes the initial-letter cue and requires inferring the complete path.Target-word lengths vary to control difficulty, with longer words producing larger search spaces.

B EditReward-Compass Data Construction … Complex Tasks

EditReward-Compass constructs preference data through stratified sampling across diverse image-editing models, while Edit-Compass organizes image editing into six categories spanning general, dynamic, reasoning, multi-image, and complex tasks. These categories cover increasingly demanding edits involving subjects, interactions, domain knowledge, multiple inputs, and composite multimodal instructions.

  • B EditReward-Compass Data Construction: EditReward-Compass uses stratified sampling and candidate images from diverse editing models because open-source models produce weak outputs on world-knowledge and algorithmic visual reasoning tasks.This strategy supports forming reliable and informative preference pairs when a single model is insufficient.
  • C Detailed Design of Edit-Compass Categories: Edit-Compass defines six distinct categories of image editing tasks, with representative examples and detailed task definitions provided throughout the design.The categories structure the benchmark’s coverage of editing capabilities.
  • General Tasks: General tasks include subject addition and removal under spatial, semantic, and attribute constraints, including copy-based addition and attribute-guided removal.Subject addition covers inserting objects at specified locations or with specified attributes, while removal includes cases with multiple matching objects.
  • Dynamic Manipulation Tasks: Dynamic manipulation tasks require moving or swapping subjects while preserving relevant properties, introducing interactions, and changing a specified subject’s emotion.The task definitions emphasize spatial changes, identity and appearance preservation, and interactions among multiple subjects.
  • World Knowledge Reasoning Tasks: World knowledge reasoning tasks cover temporal, causal, mathematical, chemical, game, and algorithmic visual reasoning applied to image editing.Temporal reasoning includes predicting future changes or inferring past appearance, while causal reasoning considers changes under conditions or external factors.
  • Multi-Image Tasks: Multi-image tasks require transferring subject attributes across reference and source images, composing coherent scenes from multiple inputs, and performing virtual try-on.Transferred attributes can include action, color, function, and other relevant properties.
  • Complex Tasks: Complex tasks combine multiple edits under composite instructions and interpret multimodal signals such as English text, arrows, circles, and cross marks.Complex Instruction integrates tasks from General, Dynamic Manipulation, and World Knowledge Reasoning categories, while Complex Paint(en) follows embedded visual annotations.

D Image Editing Model Evaluation … Evaluation Dimensions

The paper evaluates 29 image editing models with five rubric-based metrics and category-specific aggregation, while using task-specific prompts to assess instruction adherence, reference fidelity, consistency, and visual quality. Its evaluation protocols decompose complex edits, compare source and edited images, enforce spatial and attribute requirements, and score results on a 1–5 scale.

  • D Image Editing Model Evaluation: 29 mainstream image editing models are evaluated, spanning open- and closed-source systems plus Chinese and English variants.Closed-source models are accessed through official API services because their weights are unavailable.
  • D Image Editing Model Evaluation: Five metrics define Edit-Compass evaluation: Instruction Following, World Knowledge Awareness, Unedited Region Consistency, Identity Consistency, and Visual Quality.Instruction Awareness combines Instruction Following and World Knowledge Awareness, while Visual Consistency combines URC and IC.
  • D Image Editing Model Evaluation: Gemini-3.1-Pro rates every metric on a 1–5 scale using dimension-specific prompts, with URC, IC, and VQ applied according to task requirements.URC excludes style transfer, IC targets identity-preserving tasks such as Object Movement and Object Swap, and VQ covers all tasks.
  • Step 1: Instruction Decoupling: Complex instructions are decomposed into atomic tasks, then each target’s source and edited states are described before checking attributes and spatial compliance.The protocol checks color, quantity, state, action, and material while requiring interaction objects to change synchronously with the edited subject.
  • Step 2: Strict Visual Comparison: A score of 5 requires every atomic task to satisfy target, attribute, and spatial requirements perfectly, without omissions or errors.For Replace operations, the new object must occupy the exact same spatial coordinates as the original.
  • Step 2: Comparative Verification (Source vs. Edited): Complex Paint tasks extract every visual marker’s region, arrow, and instruction, then verify each edit strictly against the source image and annotated instruction.Evaluation ignores visual quality and unintended changes outside the marked edit regions.
  • Step 2: Reference Fidelity & Content Consistency Verification: Multi-image evaluation focuses on transferring requested objects or attributes from reference images, ignoring visual quality, unintended changes, and identity consistency.Reference fidelity checks semantic equivalence plus details such as text, logos, textures, patterns, actions, and orientation.
  • Step 1: Analyze Edit Instruction Requirements: Other-task evaluation requires complete execution of compound instructions and synchronized object interactions while prohibiting auxiliary objects for pose, expression, or attribute edits.Instructions are represented as subject, edit type, attribute requirements, and spatial or location requirements.

E Compute Resources

The experiments used a four-machine distributed setup with eight NVIDIA H800 GPUs and 1000 GiB of system memory per machine, requiring no additional compute beyond reported experiments aside from preliminary runs.

  • E Compute Resources: Four machines each provided 8 NVIDIA H800 GPUs and 1000 GiB of system memory, with no additional compute required beyond reported experiments excluding preliminary runs.
Loading 2605.13062v1…