Source-linked AI summary

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model

Shibo Hong, Boxian Ai, Jun Kuang, Wei Wang, FengJiao Chen, Zhongyuan Peng, Chenhao Huang, Yixin Cao

arXiv:2602.23622v2cs.CVcs.AI

TL;DR

Small-object editing remains underexplored in IIEM benchmarks despite its importance for precise local corrections. DLEBench addresses this gap with a dedicated benchmark and human-aligned dual-mode evaluation, finding that current models struggle with editing fidelity and consistency.

  • Problem

    Existing IIEM benchmarks largely evaluate salient objects, leaving small-object editing underexplored despite its importance for precise local corrections.

  • Method

    DLEBench constructs 1,889 small-object editing samples and evaluates them with refined rubrics and Tool-driven and Oracle-guided modes.

  • Results

    Across 10 IIEMs, current models struggle to maintain editing fidelity and consistency on small-scale objects.

  • Takeaways & Limitations

    DLEBench establishes a specialized testbed for measuring small-scale object editing and more reliably assessing it against human judgments.

  • Takeaways & Limitations

    The benchmark has an unbalanced instruction-type distribution and lacks moving, scaling, rotating, and multi-instruction editing scenarios.

Abstract

from arXiv · show

Significant progress has been made in the field of Instruction-based Image Editing Models (IIEMs). However, while these models demonstrate plausible adherence to instructions and strong reasoning ability on current benchmarks, their ability to edit small objects remains underexplored, despite its importance for precise local editing and refining details in both real and generated images. In this paper, we introduce DeepLookEditBench (DLEBench), the first benchmark dedicated to assessing the abilities of IIEMs in editing small-scale objects. Specifically, we construct a challenging testbed comprising 1889 samples across seven instruction types. In these samples, target objects occupy only 1%-10% of the image area, covering complex scenarios such as partial occlusion and multi-object editing. To ensure robust evaluation on this benchmark, we propose an evaluation protocol with refined score rubrics to minimize subjectivity and ambiguity in two criteria: Instruction Following and Visual Consistency. This protocol also introduces a dual-mode evaluation framework (Tool-driven and Oracle-guided Modes) addressing the misalignment between LMM-as-a-Judge and human judgements on DLEBench. Empirical results on 10 IIEMs reveal significant performance gaps in small-scale object editing, highlighting the need for specialized benchmarks to advance this ability.

1. Introduction

DLEBench addresses the underexplored challenge of editing small objects in instruction-based image editing by providing a dedicated benchmark and evaluation protocol. It combines semi-automated data construction with dual-mode assessment to improve the reliability and practicality of evaluation.

  • Data Construction: A semi-automated three-stage transformation pipeline converts visual reasoning samples into image-editing samples through counterfactual synthesis and crop-and-edit ground-truth generation.Counterfactual synthesis produces metadata such as instructions, instruction types, target objects, and captions; crop-and-edit addresses reliable reference-image generation.
  • Evaluation Protocol: The evaluation protocol uses Instruction Following and Visual Consistency criteria with refined rubrics, plus Tool-driven and Oracle-guided modes.Tool-driven Mode enables iterative assessment with external tools, while Oracle-guided Mode uses human-annotated bounding boxes to eliminate localization errors.
  • Empirical Evaluation: Experiments evaluate 10 representative IIEMs and report that the proposed methods align better with human judgments than LMM-as-a-Judge.The protocol is designed to provide reliable automated assessment for small-scale editing.
  • Benchmark: DLEBench is introduced as the first benchmark dedicated to evaluating small-scale object editing abilities in IIEMs.The benchmark establishes a foundation for assessing this capability.

2. Related Work

Related-work benchmarks have evolved from basic single-object editing toward more complex, reasoning-intensive and grounded scenarios. Evaluation has likewise shifted from similarity metrics toward LMM-based assessment across instruction adherence, consistency, and quality.

  • Editing Benchmarks: Benchmarks progressed from single-turn, single-object editing in I2EBench to more complex multi-object and grounded editing in PIE-Bench++ and GIE-Bench.This progression reflects an increasing focus on complex and realistic editing settings.
  • Evaluation Protocols: Similarity-based metrics such as CLIP scores often correlate poorly with human judgments in complex editing scenarios.Early evaluation protocols relied primarily on these similarity measures.
  • Evaluation Protocols: Recent benchmarks adopt LMM-as-a-judge evaluation across Instruction Following, Visual Consistency, and Visual Quality.The paradigm uses strong Large Multimodal Models to assess editing results across three core dimensions.

3. DLEBench

DLEBench is a benchmark for small-scale object editing that transforms visually dense reasoning samples into verified image-editing instances. It contains 1,889 samples spanning seven instruction types, primarily single-object edits alongside multi-object cases with up to eight targets.

  • Benchmark Construction: DLEBench selects visually dense scenes containing small, sometimes partially occluded objects to test precise perception and editing.The source images come from visual reasoning benchmarks whose questions target single or multiple small-scale objects in cluttered environments.
  • Benchmark Construction: A three-stage pipeline converts reasoning samples into structured instances containing source and reference images, captions, target objects, editing types, and instructions.The stages are metadata construction, reference image generation, and human verification; model outputs are assessed against the ground-truth edited reference image.
  • Reference Image Generation: Adaptive bbox expansion combines manually annotated target boxes with size-dependent context to balance local focus and seamless reintegration.The expansion ratio uses λmax = 6.0, λmin = 0.3, Smin = 32, and Smax = 256, with linear interpolation between size thresholds.
  • Benchmark Statistics: 1,889 samples span seven editing instruction types, including 1,771 single-object edits and 118 multi-object edits involving up to 8 targets.The target-object area distribution approximately follows a log-normal distribution.

4. Evaluation Protocol

The evaluation protocol scores Instruction Following and Visual Consistency with hierarchical rubrics that isolate failure severity, while dual Tool-driven and Oracle-guided modes address LMM perceptual limitations and localization errors.

  • Instruction Following: Instruction Following measures target localization, correct editing, and preservation of irrelevant intrinsic attributes.The criterion replaces subjective, vague holistic scoring with distinct failure modes.
  • Instruction Following: The Instruction Following hierarchy ranks outcomes from Score 1 (Localization Failure) through Score 4 (Flawless Execution), enabling reproducible bottleneck diagnosis.Intermediate levels are Score 2 (Wrong Action) and Score 3 (Over Modification).
  • Visual Consistency: Visual Consistency evaluates whether non-target regions remain intact using a severity-based four-point scale.The levels range from Score 1 (Scene Collapse) to Score 4 (Perfect Consistency), with Multiple Anomalies and Single Anomaly between them.
  • Dual-mode evaluation: Tool-driven Mode lets LMMs invoke Grounding, Zoom-In, Difference, and Enhancer tools to compensate for weak visual perception.Zoom-In supports iterative spatial search because GroundingDINO often struggles with small-scale objects.
  • Dual-mode evaluation: Oracle-guided Mode uses human-annotated bounding boxes to decouple evaluation from localization, cropping edited regions for IF and masking targets for VC.The mode provides pre-processed inputs and retains the Tool set for VC evaluation.

5. Experiments

Experiments evaluate 10 IIEMs on DLEBench using Oracle-guided scoring of Instruction Following (IF) and Visual Consistency (VC), with results revealing substantial weaknesses in small-object editing. Performance generally degrades as object scale decreases, while Oracle-guided evaluation aligns most closely with human judgments.

  • Experimental Setup: The evaluation covers autoregressive, LMM-based diffusion, hybrid, and proprietary IIEMs, with all IF and VC scores normalized to 100 points.Models are assessed using the Oracle-guided Mode across both criteria.
  • Overall Performance: 65.55: Gemini-3-Pro achieves the highest average score, while open-source Bagel-Think reaches 61.00 and surpasses proprietary GPT-Image-1.These results challenge the assumption that closed-source models consistently outperform open-source models.
  • Overall Performance: Change Count is consistently the most challenging instruction type because it requires isolating and enumerating multiple small target objects.Its scores degrade across all evaluated models relative to instructions involving a single small-scale instance.
  • Analysis by Evaluation Criteria: 48.97: even Gemini-3-Pro’s average IF score remains low, whereas VC is generally stronger but varies sharply between GPT-Image-1 and Bagel-Think at 35.17 versus 86.43.The VC disparity is attributed to GPT-Image-1 modifying non-target regions when it fails to localize targets.
  • Impact of Object Scale on Performance: Most models show scale-dependent performance, with weaker small-object editing ability as target object size decreases.Correlation analysis confirms positive scale-performance relationships for most competitive models, while three models show notably weak correlations.
  • Validation of the Evaluation Framework: Oracle-guided Mode achieves the strongest agreement with human judgments, recording the highest ρ and r and the lowest MAE among evaluation methods.Tool-driven mode follows, while Gemini-3-Pro- and GPT-4.1-based LMM-as-a-Judge baselines show lower correlations and higher MAE.

6. Conclusion, Limitation and Future Work

DLEBench evaluates small-scale object editing with 1,889 samples targeting objects occupying 1%–10% of image area, revealing that current IIEMs struggle with fidelity and consistency. Its main limitations are imbalanced instruction types and missing complex, coupled editing scenarios, motivating automated data expansion.

  • Conclusion: DLEBench contains 1,889 samples evaluating IIEMs on objects occupying only 1%–10% of image area.The benchmark targets small-scale object editing under constrained spatial conditions.
  • Conclusion: Current models struggle to maintain editing fidelity and consistency under small-object editing constraints.This finding comes from evaluating 10 IIEMs on DLEBench.
  • Conclusion: The proposed dual-mode evaluation framework with refined rubrics aligns better with human judgments than LMM-as-a-Judge.The framework is intended to enable more reliable assessment.
  • Limitation: The benchmark has an unbalanced distribution of instruction types, leaving some editing operations underrepresented.This imbalance stems from converting samples from visual reasoning datasets.
  • Future Work: DLEBench lacks complex manipulations such as moving, scaling, and rotating objects, along with multi-instruction coupling scenarios.These scenarios are common in real-world image editing workflows and are planned for future expansion.

A. Annotation Document · A.1. Definitions of Criteria · A.2. Score Rubrics

The annotation framework evaluates edits along Instruction Following and Visual Consistency. It defines criterion-specific checks and four-level rubrics for judging localization, action accuracy, preservation, and anomalies.

  • A. Annotation Document: The annotation task presents a Source Image, Editing Instruction, Edited Image, and Reference Image for evaluating edit quality.The interface is shown in Figure 11.
  • A.1. Definitions of Criteria: Instruction Following assesses target localization, requested action execution, and preservation of unspecified target-object attributes.It is the primary metric for small-scale editing ability, including difficult cases with small, occluded, or blurred targets.
  • A.1. Definitions of Criteria: Target Object Location checks whether the modification occurs at the correct target location.
  • A.1. Definitions of Criteria: Action Alignment checks whether the model performs the requested modification, such as replacing, removing, or altering the object.
  • A.1. Definitions of Criteria: Visual Consistency measures preservation of instruction-irrelevant environments and visual elements across input and output images.It considers global preservation, including the distinction between grounded editing and scene regeneration.
  • A.1. Definitions of Criteria: Global Scene Stability checks preservation of the scene’s fundamental visual context, including artistic and scene style.
  • A.1. Definitions of Criteria: Local Anomaly Detection checks that non-target objects remain unchanged, without deletions, distortions, or added objects.
  • A.2. Score Rubrics: Instruction Following uses four levels from Score 4 (Flawless Execution) to Score 1 (Localization Failure), while Visual Consistency uses Score 4 (Perfect Consistency) to Score 1 (Scene Collapse).Intermediate levels are Over Modification and Wrong Action for Instruction Following, and Single Anomaly and Multiple Anomalies for Visual Consistency.

A.3. Evaluation Process

The evaluation protocol uses staged checks for Instruction Following and Visual Consistency, assigning specific failure labels based on localization, action correctness, over-modification, scene stability, and anomalies.

  • Instruction Following: Instruction Following first checks whether the requested edit is localized to the intended object, component, or attribute; otherwise it is labeled Localization Failure.Changes outside the target, severe blur preventing verification, or modification of the wrong sub-component also trigger this label.
  • Instruction Following: After localization succeeds, evaluators verify that the edit performs the instructed action and exact object-count change, labeling contradictions or incorrect quantities Wrong Action.The text instruction is primary, while a reference image may provide visual context.
  • Instruction Following: The final Instruction Following check tests whether untargeted shape, texture, and structural details remain intact, distinguishing Over Modification from Flawless Execution.Structural distortion, unrequested style changes, or blur that prevents confirming original features are Over Modification; a sharp, faithful object is Flawless Execution.
  • Visual Consistency: Visual Consistency first checks whether the non-target background preserves the original global environment and artistic style before scanning for local anomalies.A completely different place, time of day, or style fails the global stability check.
  • Visual Consistency: Evaluators count distinct non-target errors and label the result Multiple Anomalies for two or more, Single Anomaly for one, or Perfect Consistency for zero.Errors include missing, distorted, or changed existing elements and newly introduced objects or details.
  • Special Cases: When multiple target instances appear, each is evaluated independently and the image receives the worst result among them.Blurry targets are Localization Failure when localization cannot be verified, or Over Modification when instruction adherence is clear but feature preservation cannot be judged.

B. More Quantitative Results

Table 4 reports ten IIEMs across two evaluation dimensions and seven instruction types. Scores use oracle-guided mode and are normalized to a 100-point scale for direct cross-model comparison.

  • Table 4 reports the performance of ten IIEMs.
  • The evaluation covers two dimensions and seven instruction types.
  • Oracle-guided mode is used to ensure assessment accuracy.
  • All scores are normalized to a 100-point scale for direct cross-model comparison.

C. More Qualitative Comparison Results · D. Distribution of Editing Instruction Types

The paper supplements DLEBench with qualitative comparisons and score rubrics, then characterizes its 1,889 samples across seven editing instruction types. Attribute-level modifications are dominated by color and OCR, while object-level modifications most frequently involve removal.

  • C. More Qualitative Comparison Results: DLEBench includes qualitative comparisons between Figure 13 and Figure 19 to illustrate model behavior on the benchmark.
  • C. More Qualitative Comparison Results: The Instruction Following evaluation uses explicit score rubrics to structure assessment.
  • D. Distribution of Editing Instruction Types: The benchmark’s instruction statistics distinguish attribute-level modifications from object-level modifications.
  • D. Distribution of Editing Instruction Types: Removal is the most frequent operation among object-level modifications.
  • D. Distribution of Editing Instruction Types: 1,889 high-quality image editing samples span seven distinct editing instruction types in DLEBench.
  • D. Distribution of Editing Instruction Types: Color and OCR are the predominant attribute-level modification categories.These categories inherit characteristics of source visual reasoning benchmarks, where such visual properties are frequently queried.

E. Calculation of Target Area Ratios

The paper estimates target-area ratios for mainstream editing benchmarks with a two-stage pipeline because these datasets lack ground-truth target bounding boxes. GPT-4.1 first extracts target object names from editing instructions, supporting comparison of target-scale distributions.

  • Target Area Ratio Calculation: The pipeline compares target-scale distributions with ImageEdit, UniREditBench, RISE, and KRIS-Bench despite missing ground-truth target bounding boxes.It addresses this limitation through a two-stage process.
  • Target Area Ratio Calculation: GPT-4.1 extracts the specific target object name from each editing instruction using prompts detailed in Table 6.Table 6 provides the prompt used for target-object extraction.
  • Target Area Ratio Calculation: The benchmark includes examples of color, OCR, material, object replacement, count, removal, and shape edits for target-object analysis.Examples include changing a white handbag to black, altering bus text, replacing a spider with a dog, and removing a telephone.

F. The Qualitative Comparison Example of Evaluation Method · G. The Prompt Template

The qualitative examples show that Oracle-guided and Tool-driven evaluation better localizes target edits and detects non-target anomalies than LMM-as-a-Judge baselines. The appendix also documents prompt templates for counterfactual synthesis and all evaluation modes and criteria.

  • F. The Qualitative Comparison Example of Evaluation Method: Oracle-guided Mode uses human-annotated bounding boxes to crop target objects, focusing Instruction Following evaluation on the relevant region.Tool-driven Mode instead uses external tools to localize differences and generate visual comparisons of target and surrounding regions.
  • F. The Qualitative Comparison Example of Evaluation Method: The Oracle-guided example verifies that the blue dustpan changed to red and evaluates whether the requested action and visual details were preserved.Its reasoning compares the original and edited images before judging modification occurrence and action alignment.
  • F. The Qualitative Comparison Example of Evaluation Method: Tool-driven analysis correctly localized the blue dustpan while leaving the adjacent blue broom untouched, avoiding edit bleed into similar objects.The example reports that the requested color change succeeded but object-detail preservation still failed.
  • G. The Prompt Template: Table 15 provides prompts for GPT-4.1-based counterfactual synthesis that transforms visual reasoning tasks into instruction-based image editing tasks.The prompt template is presented as part of the benchmark’s synthesis strategy.
  • G. The Prompt Template: Tables 27 through 31 provide LMM-as-a-Judge prompts for evaluating Instruction Following and Visual Consistency.These templates complete the prompt documentation for the dual-mode evaluation framework’s comparison baseline.
Loading 2602.23622v2…