Source-linked AI summary
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
Yongliang Wu, Zonghui Li, Xinting Hu, Xinyu Ye, Xianfang Zeng, Gang Yu, Wenbo Zhu, Bernt Schiele, Ming-Hsuan Yang, Xu Yang
TL;DR
Instruction-based image editors can produce visually plausible outputs while their knowledge-based reasoning remains under-explored. KRIS-Bench addresses this gap with a cognitively grounded benchmark and knowledge-aware evaluation protocol, finding persistent reasoning limitations across current models.
Problem
Knowledge-based reasoning in otherwise visually plausible instruction-based image edits remains under-explored.
Method
KRIS-Bench organizes 22 editing tasks across three knowledge types and seven reasoning dimensions, evaluating them with VLM-based metrics and knowledge hints.
Results
Experiments on 10 state-of-the-art models reveal persistent limitations in knowledge-grounded reasoning across knowledge types, reasoning dimensions, and editing tasks.
Takeaways & Limitations
KRIS-Bench provides a knowledge-centric framework for fine-grained, human-calibrated evaluation of intelligent image editing systems.
Takeaways & Limitations
The benchmark may contain biases from its relatively small scale, knowledge-type imbalance, and culturally specific task assumptions.
Abstract
from arXiv · showhide
Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, we introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on 10 state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems.
1 Introduction
KRIS-Bench addresses the under-explored knowledge reasoning behind visually plausible image edits by organizing evaluation around a structured knowledge taxonomy. It introduces a benchmark, knowledge-aware evaluation protocol, and experiments exposing persistent reasoning limitations.
- Motivation: Visually coherent and semantically aligned edits can still lack reasoning grounded in real-world knowledge.The sodium-in-water example illustrates a plausible-looking edit that fails to reflect chemistry knowledge.
- Benchmark: KRIS-Bench organizes image-editing evaluation around three foundational knowledge types and seven reasoning dimensions spanning 22 tasks.The benchmark contains 1,267 high-quality annotated instances.
- Evaluation: The benchmark introduces a VLM-based protocol with Knowledge Plausibility, manually curated knowledge hints, and human-study validation.Knowledge hints are reported to significantly enhance VLM plausibility assessment.
- Findings: Experiments across 10 state-of-the-art models reveal substantial limitations in knowledge-grounded reasoning for image editing.The reported limitations span knowledge types, reasoning dimensions, and editing tasks.
- Contributions: The paper positions its contribution as a cognitively grounded taxonomy, a 22-task benchmark with 1,267 annotated instances, and a knowledge-aware evaluation protocol.The contributions explicitly define Factual, Conceptual, and Procedural knowledge as the evaluation foundation.
2 Related Work
Prior instruction-based image-editing work has advanced through diffusion-based control and increasingly diverse benchmarks. Existing evaluations cover editing actions, free-form instructions, editing types, and metrics, motivating broader reasoning-focused assessment.
- Instruction-based Image Editing Methods: Diffusion-based methods improve instruction-based editing through test-time controllability, attention control, CLIP guidance, and latent inversion.These approaches target controllability, localized editing, or fidelity preservation.
- Benchmarks for Instruction-based Image Editing: Existing benchmarks evaluate canonical editing subtasks, diverse free-form instructions, multiple editing types, and varied metrics.Examples include inpainting, attribute manipulation, layout adjustment, and free-form user instructions.
3 KRIS-Bench
KRIS-Bench frames image editing as knowledge-based reasoning, using an educational taxonomy to organize tasks and support evaluation across increasing reasoning demands. Its task design spans factual, conceptual, and procedural knowledge across seven dimensions and 22 tasks.
- Benchmark Overview: KRIS-Bench is a comprehensive reasoning benchmark with 1,276 samples across 22 tasks and knowledge hints for knowledge-based cases.The passage describes the benchmark as having the largest size in its comparison with prior reasoning-based benchmarks.
- Taxonomy of Knowledge Types: The taxonomy organizes required knowledge into Factual, Conceptual, and Procedural Knowledge, focusing on what models must represent and apply.The framework is inspired by the revised Bloom taxonomy and differs from approaches centered primarily on editing actions.
- Factual Knowledge: Factual Knowledge covers directly observable visual attributes, spatial relations, and temporal cues without abstract inference.It serves as a prerequisite for more complex reasoning.
- Conceptual Knowledge: Conceptual Knowledge connects perception to generalizable physical, biological, or social principles and supports plausible real-world outcomes.The banana-ripening example requires understanding a natural process rather than factual recognition alone.
- Procedural Knowledge: Procedural Knowledge concerns multi-step reasoning, task decomposition, and rule-based execution within image-editing contexts.It supports coordinated multi-element edits and complex logical reasoning.
- Knowledge-Based Task Formulation: The task formulation derives seven reasoning dimensions from the three knowledge types and spans 22 tasks organized by knowledge requirements.The tasks are designed as knowledge-specific evaluations rather than isolated editing actions.
- Taxonomy Scope: The benchmark excludes Metacognitive Knowledge because current large models do not demonstrate self-monitoring and learning regulation within one-turn image editing.This is an explicit scope assumption in the taxonomy.
Factual Knowledge
KRIS-Bench evaluates knowledge-grounded editing using representative tasks and four complementary metrics. The reported analyses distinguish visual quality from instruction fulfillment and real-world plausibility across models and reasoning dimensions.
- Task Examples: Figure 2 presents representative examples from 22 tasks grounded in factual, conceptual, or procedural knowledge.The examples cover diverse reasoning dimensions.
- Procedural Knowledge: The benchmark’s procedural tasks include logical reasoning and instruction decomposition involving symbolic structures, numerical relationships, and sequential instructions.Examples include solving puzzles, applying logical rules, designing posters, and combining objects from different images.
- Data Collection: The data combines internet-collected images with smaller portions generated by models or drawn from existing datasets, followed by annotator creation and expert review.Three human annotators curated the data and three Ph.D. experts reviewed the annotations.
- Evaluation Metrics: The four evaluation dimensions are Visual Consistency, Visual Quality, Instruction Following, and Knowledge Plausibility.Knowledge Plausibility is added to assess consistency with real-world knowledge and domain-specific principles.
- Evaluation Metrics: Visual Consistency measures preservation of unrelated image content, while Visual Quality measures realism, natural appearance, structural coherence, and artifacts.Together they assess localized editing and perceptual output quality.
- Evaluation Metrics: Instruction Following measures literal and complete execution of the user instruction independently of perceptual quality or real-world plausibility.A wooden block’s presence in a tank is sufficient for instruction following even if its physical behavior is implausible.
- Evaluation Metrics: Knowledge Plausibility evaluates consistency with real-world and domain-specific principles, requires instruction fulfillment, and is available for Natural Sciences, Social Sciences, and Logical Reasoning.Each metric is rated from 1 to 5 using GPT-4o as the evaluation model.
4 Experiments and Analysis
KRIS-Bench evaluates 10 image editing models across knowledge types, reasoning dimensions, tasks, and metrics, revealing substantial differences between model families and persistent knowledge-reasoning gaps.
- Evaluation Models & Settings: 10 models are evaluated on KRIS-Bench, including three closed-source and seven open-source systems.Open-source models other than OmniGen and Emu2 receive the lowest score on tasks requiring multiple input images because they support only single-image inputs.
- Overall Performance: Closed-source models substantially outperform open-source models overall, while BAGEL-Think approaches the performance of Gemini 2.0 and Doubao among open-source systems.GPT-4o leads nearly all knowledge types and evaluation dimensions, except visual consistency for temporal prediction, where it slightly trails Gemini 2.0.
- Analysis by Knowledge Types: Procedural knowledge is generally the weakest area, indicating difficulties with multi-step reasoning and task decomposition.Models do not consistently perform worse on conceptual than factual knowledge; several strong models perform slightly worse on factual tasks.
- Analysis by Reasoning Dimensions: Models perform relatively well on attribute perception and commonsense or cultural tasks but struggle with spatial, scientific, and symbolic reasoning.Examples include failures to interpret chemical reactions, recognize acid-sensitive color changes, and solve abstract pattern tasks, although GPT-4o occasionally succeeds at logical reasoning.
- Analysis by Editing Tasks and Metrics: Knowledge Plausibility scores are consistently lower than Instruction Following, exposing difficulty applying external knowledge accurately during editing.BAGEL-Think surpasses nearly all other open-source models on Knowledge Plausibility and outperforms Gemini 2.0 and Doubao in Biology and Chemistry tasks.
- Evaluation Validation: Knowledge-enhanced prompts produce stronger human–VLM score correlations and lower mean absolute errors, especially for Knowledge Plausibility.The reliability study involved 12 human experts rating three closed-source models.
5 Conclusion
The paper presents KRIS-Bench as a cognitively grounded, knowledge-centric benchmark for image-editing reasoning and reports persistent gaps across diverse knowledge types. It also acknowledges limitations in benchmark scale, distribution, and cultural assumptions.
- Conclusion: KRIS-Bench evaluates image-editing reasoning through Factual, Conceptual, and Procedural knowledge using fine-grained tasks and human-calibrated evaluation.The benchmark is positioned as a principled alternative to task-based or content-driven evaluation.
- Conclusion: Results reveal persistent gaps in models’ ability to reason across diverse knowledge types.
- Limitations: KRIS-Bench may contain biases from its relatively small scale, over-representation of particular knowledge types, and culturally specific task assumptions.
A Detailed Tasks Explanation
The benchmark organizes 22 tasks across Factual, Conceptual, and Procedural knowledge, covering seven reasoning dimensions from direct perception to multi-step instruction execution.
- Overview: The task suite contains 22 representative tasks spanning seven capability dimensions.
- Factual Knowledge: Factual tasks test direct visual and temporal understanding through Attribute Perception, Spatial Perception, and Temporal Prediction.Examples include count change, color change, size adjustment, part completion, position movement, viewpoint change, reverse prediction, and intermediate prediction.
- Conceptual Knowledge: Conceptual tasks apply external real-world knowledge through Social Science and Natural Science dimensions.They include commonsense, cultural, biological, chemical, geographical, and related domain-specific transformations.
- Procedural Knowledge: Procedural tasks require planning, rule-following, and coordinated multi-operation outputs through Logical Reasoning and Instruction Decomposition.Tasks include abstract and rule-based reasoning, multi-instruction execution, and multi-element composition.
B Data Distribution
KRIS-Bench contains 1,267 instances across 22 task types, distributed across knowledge types, reasoning dimensions, and individual editing tasks.
- Dataset Overview: 1,267 instances span 22 task types, each combining knowledge requirements with reasoning dimensions.Figure 6 presents the distribution by knowledge type, reasoning dimension, and individual editing task.
- Knowledge Type Breakdown: Conceptual Knowledge has the largest share with 518 instances (40.9%), followed by Factual Knowledge with 449 (35.4%) and Procedural Knowledge with 300 (23.7%).
- Reasoning Dimension Breakdown: Natural Science dominates the reasoning dimensions with 393 instances (31.0%), followed by Attribute Perception with 275 (21.7%).Logical Reasoning and Instruction Decomposition each contribute 150 instances (11.8%).
- Editing Task Breakdown: Nine of the 22 editing tasks have the highest count of 75 instances (5.9%), while Biology has 68 (5.4%) and Color Change and Size Adjustment each have 50 (3.9%).
C Data Source
KRIS-Bench combines licensed, generated, and existing-dataset images, with specialized 3D assets enabling ground-truth evaluation for Viewpoint Change. Its user study used trained undergraduate-level annotators whose results were reviewed in pairs.
- Data Sources: Most benchmark images came from Creative Commons internet sources, supplemented by generated images and existing datasets.Viewpoint Change additionally uses ABO and Sketchfab 3D assets for accurate ground-truth views.
- Data Sources: The Viewpoint Change task uses ABO and Sketchfab 3D assets to support evaluation against ground-truth views.
- Human Annotation: 12 human annotators with at least undergraduate-level education conducted the user study after training and trial annotation.Their results were reviewed and discussed in pairs to improve alignment with evaluation criteria.
- Benchmark Evaluation: Figure 7 presents KRIS-Bench performance across editing tasks and four metrics, separating closed-source and open-source models.
E Open-source VLM Evaluation
The open-source Qwen2.5-VL-72B serves as a transparent proxy judge for model predictions, with scoring trends closely matching GPT-4o. Experiments report substantial open-source compute while noting that closed-source compute is not user-controllable.
- Evaluation Setup: Qwen2.5-VL-72B is used as an open-source proxy judge to score evaluated models’ predictions.Its cross-task scoring trends align closely with GPT-4o results.
- Evaluation Setup: Qwen2.5-VL-72B’s assessments summarize performance across knowledge dimensions and evaluation metrics in Table 3.
- Compute: Open-source experiments used dual Xeon Platinum 8468 CPUs, 960 GB RAM, and 8×NVIDIA H100 80GB GPUs.Each model required approximately 2 hours to complete all 1,267 editing tasks.
- Compute: Closed-source models were accessed through official APIs or web platforms, where compute details were not user-controllable.No additional large-scale pretraining or auxiliary runs were performed beyond the reported experiments.
G More Visualization Results
Qualitative results show uneven reasoning abilities: models struggle with counting, anomaly correction, conceptual knowledge, and logical reasoning, while some perform better on color changes, instruction decomposition, or temporal prediction. The evaluation protocol jointly assesses Instruction Following and Knowledge Plausibility where separate scoring can be inconsistent.
- Qualitative Results: Models often fail Count Change and Anomaly Correction tasks, while Color Change performance is generally strong across models.Many models also cannot infer missing Part Completion components unless explicitly instructed.
- Spatial Perception: GPT-4o performs consistently well on Viewpoint Change and Position Movement but is relatively weak on Size Adjustment.
- Temporal Prediction: GPT-4o and Gemini 2.0 produce logically coherent Temporal Prediction outputs, unlike Doubao, OminiGen, and Emu2.
- Conceptual Knowledge: Open-source models rarely succeed on Conceptual Knowledge tasks across Humanities, Practical Knowledge, Biology, Chemistry, Geography, Medicine, Mathematics, and Physics.The passage suggests domain-specific content may fall outside their training distributions.
- Logical Reasoning: Closed-source models show some Instruction Decomposition capability but remain limited in the Logical Reasoning dimension.Figures 24–27 visualize these results across Multi-element Composition, Multi-instruction Execution, Abstract Reasoning, and Rule-based Reasoning.
- Evaluation Prompts: Instruction Following and Knowledge Plausibility are jointly evaluated because separate assessment can introduce inconsistencies and inaccurate model evaluations.Temporal Prediction and other specialized tasks use customized evaluation prompts, while some tasks incorporate ground-truth images or knowledge hints.