Source-linked AI summary
GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
Mingxin Liu, Ziqian Fan, Zhaokai Wang, Leyao Gu, Zirun Zhu, Yiguo He, Yuchen Yang, Changyao Tian, Xiangyu Zhao, Ning Liao, Shaofeng Zhang, Qibing Ren, Zhihang Zhong, Xuanhe Zhou, Junchi Yan, Xue Yang
TL;DR
Existing image-editing benchmarks provide limited evidence about whether unified multimodal models can reason over structured, discipline-specific knowledge while editing images. GRADE addresses this gap with a 520-sample, 10-domain benchmark and a three-dimensional evaluation protocol, revealing large model differences and persistent limitations under implicit discipline-informed reasoning.
Problem
Existing image-editing benchmarks are predominantly grounded in natural images and do not adequately assess coordinated knowledge, structured reasoning, and editing under domain constraints.
Method
GRADE introduces the first discipline-informed image editing benchmark with 520 curated samples across 10 academic domains and evaluates Discipline Reasoning, Visual Consistency, and Logical Readability.
Results
46.2% vs. 16.0% separates Nano Banana Pro and GPT-Image-1.5 under discipline-informed editing, while implicit discipline-informed reasoning remains a major bottleneck.
Takeaways & Limitations
GRADE provides actionable directions for advancing unified multimodal models toward discipline-informed image editing and reasoning.
Takeaways & Limitations
Error analysis identifies reasoning-process failures in multi-step procedures and image-recognition failures in parsing dense structured visuals.
Abstract
from arXiv · showhide
Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability under structured, domain-specific constraints. In this work, we introduce GRADE, the first benchmark to assess discipline-informed knowledge and reasoning in image editing. GRADE comprises 520 carefully curated samples across 10 academic domains, spanning from natural science to social science. To support rigorous evaluation, we propose a multi-dimensional evaluation protocol that jointly assesses Discipline Reasoning, Visual Consistency, and Logical Readability. Extensive experiments on 20 state-of-the-art open-source and closed-source models reveal substantial limitations in current models under implicit, knowledge-intensive editing settings, leading to large performance gaps. Beyond quantitative scores, we conduct rigorous analyses and ablations to expose model shortcomings and identify the constraints within disciplinary editing. Together, GRADE pinpoints key directions for the future development of unified multimodal models, advancing the research on discipline-informed image editing and reasoning. Our benchmark and evaluation code are publicly released.
1 Introduction
GRADE introduces a discipline-informed image editing benchmark to test whether unified multimodal models can coordinate academic knowledge, structured reasoning, and precise visual modification. Across 20 state-of-the-art models, it exposes substantial performance gaps and persistent difficulty with implicit discipline-informed reasoning.
- Motivation: Existing image-editing benchmarks mainly use natural images and shallow prompt-based reasoning, leaving structured academic knowledge insufficiently assessed.The paper frames discipline-informed editing as a more demanding setting because models must reason under domain constraints while preserving visual structures.
- Benchmark: GRADE is the first benchmark explicitly designed to evaluate discipline-informed reasoning in image editing across 520 curated samples and 10 academic domains.Its scenarios include correcting geometric diagrams, modifying chemical structures, and refining data visualizations.
- Evaluation: The evaluation jointly measures Discipline Reasoning, Visual Consistency, and Logical Readability, extending assessment beyond aesthetic quality and realism.These dimensions target knowledge reasoning, task-dependent visual constraints, and clear, logically structured academic representations.
- Results: 46.2% vs. 16.0% separates Nano Banana Pro and GPT-Image-1.5 under discipline-informed editing despite comparable performance on other benchmarks.The result is presented as evidence that GRADE discriminates models on implicit academic knowledge and structured reasoning.
- Results: Implicit discipline-informed reasoning remains a major bottleneck, while the benchmark’s analyses provide actionable directions for future unified multimodal models.The authors examine representative failures and compare implicit with explicit instruction formulations.
2 Related Work
Prior work has improved unified image generation and editing, but existing benchmarks largely emphasize visual quality, semantic alignment, or explicitly stated operations. GRADE addresses the resulting gap by evaluating reasoning-intensive editing grounded in structured, domain-specific knowledge.
- Image Generation Models: Diffusion-based and unified generative frameworks have advanced language-aligned synthesis by combining language understanding, visual perception, and image generation.The related work describes unified systems as increasingly built on multimodal large language models.
- Image Editing Models: Recent image editing models accept an input image and textual instruction for targeted modification, but reasoning-intensive domain-specific scenarios remain largely unexplored.The paper distinguishes visual plausibility from the ability to handle structured disciplinary knowledge.
- Image Editing Benchmarks: Most existing editing benchmarks assess visual quality or semantic alignment, while many rely on explicit instructions that specify the required operations directly.ImgEdit is cited as targeting traditional editing tasks where reasoning is not central.
- Disciplinary Knowledge Benchmarks: Prior disciplinary benchmarks study multimodal understanding across broad subjects or high-difficulty academic visual comprehension, rather than discipline-informed image editing.MMMU and HLE provide related examples of disciplinary knowledge evaluation in understanding tasks.
3 GRADE Benchmark
GRADE constructs a curated, hierarchically organized dataset spanning 10 disciplines and evaluates edits through reasoning, consistency, and readability criteria. Its protocol combines structured expert-informed judging with task-specific visual checks and joint score aggregation.
- Data Construction: Each GRADE sample is an image-editing triplet containing an input image, textual instruction, and corresponding ground-truth image.Most samples are sourced and manually edited by academically trained annotators, then cross-validated by additional experts.
- Taxonomy: The dataset covers mathematics, physics, chemistry, biology, history, geography, sports, music, computer science, and economics with hierarchical sub-disciplines.The hierarchy targets distinct domain-specific knowledge and reasoning patterns, such as plane geometry and graph statistics.
- Evaluation Overview: The evaluation protocol measures Discipline Reasoning, Visual Consistency, and Logical Readability as complementary dimensions of editing quality.Together they assess reasoning correctness, low-level visual quality, and clear academic representation.
- Discipline Reasoning: Discipline Reasoning uses weighted, question-guided evaluation aligned with required disciplinary knowledge and cross-validated by human experts.GPT-5 generates weighted binary questions, while Gemini-3-Flash assesses edited results against them.
- Visual Consistency: Visual Consistency distinguishes localized edits, style-preserving global edits, and cases requiring independence from the original image.The corresponding prompts evaluate whether edits preserve unrelated elements, representation style, or domain-specific standards.
- Logical Readability: Logical Readability evaluates whether discipline-specific content is clear, logically consistent, structurally sound, and correctly annotated.Each edited result receives a readability score of 0/1/2.
- Score Aggregation: A sample counts as correct only when it achieves the maximum score in all three evaluation dimensions.This joint aggregation makes overall accuracy require simultaneous satisfaction of every criterion.
4 Experiments
GRADE evaluates 20 state-of-the-art models across discipline-informed editing dimensions, disciplines, human alignment, instruction explicitness, qualitative cases, and error types. Results show substantial gaps between models, persistent difficulty with implicit reasoning, and distinct failure modes.
- Experimental Setup: 20 state-of-the-art models, including 10 closed-source and 10 open-source systems, are evaluated across unified multimodal and specialized image editing models.The benchmark uses Gemini-3-Flash as the automated judge and includes human-alignment and ablation analyses.
- Main Results: 46.2% overall accuracy makes Nano Banana Pro the strongest model, but every evaluated model remains below 50%.Nano Banana Pro surpasses Seedream 5.0 at 24.7%, while comparable existing-benchmark performance differs substantially on GRADE.
- Main Results: 53.1% in Physics and 55.6% in Biology show strong differentiation, whereas Nano Banana Pro reaches only 29.6% in History and Geography remains challenging.Open-source models largely fail across disciplines, with accuracy frequently at or near zero.
- Human Alignment: MAE values around 10% across all three dimensions indicate consistent alignment between automated evaluation and averaged ratings from five human experts.The alignment study uses 68 samples across disciplines and reports normalized MAE and STD.
- Ablation Study: 89.7% versus 67.9% in Discipline Reasoning shows Nano Banana 2 improves when implicit instructions become explicit, while Qwen-Edit-2511 rises from 18.9% to 44.7%.Overall Qwen-Edit-2511 accuracy increases from 1.5% to 8.8%, indicating greater reliance on explicit guidance.
- Qualitative Analysis and Error Analysis: Qualitative cases reveal failures in spatial reasoning, missing-text placement, perception, knowledge grounding, multi-step reasoning, and constraint-preserving synthesis.The error analysis categorizes representative bottlenecks across perception, knowledge grounding, procedural execution, and synthesis.
5 Conclusion
GRADE evaluates discipline-informed image editing across ten academic disciplines using 520 samples and a multi-dimensional protocol. Experiments reveal substantial performance gaps under implicit instructions and persistent limitations in applying structured academic knowledge.
- GRADE evaluates editing performance grounded in discipline-informed knowledge across ten academic disciplines using 520 carefully constructed samples.
- Its multi-dimensional evaluation protocol demonstrates strong alignment with human judgments.
- Extensive experiments reveal substantial performance gaps, particularly when models must apply discipline-informed reasoning under implicit instructions.
- The results expose persistent limitations in handling structured academic knowledge beyond visual realism.
- GRADE points toward future unified multimodal models integrating discipline-informed knowledge, reasoning, and editing.
A.1 Data Sources
GRADE combines educational resources, open-source datasets, and programmatic or interface-based generation to obtain broad disciplinary coverage and structurally precise image pairs. Public or openly licensed materials are used, with preprocessing intended to improve clarity without changing semantic content.
- GRADE draws benchmark images from open educational resources, open-source datasets, and programmatic or interface-based generation.
- Open educational resources include textbook illustrations, teaching slides, instructional video frames, and concept-oriented image websites across academic subjects.
- Open-source datasets supplement disciplines with standardized visual resources, including Geometry3k for mathematics and When-in-Rome for music-related samples.
- Programmatic generation uses tools such as GeoGebra and MathCanvas to create precise structural modifications and randomized visual configurations.
- Images with longer sides no greater than 512 pixels undergo super-resolution, followed by manual checks for semantic distortion.
A.2 Taxonomy Distribution
GRADE organizes its disciplines through a two-level taxonomy, whose distribution is presented in Figure 6.
- GRADE uses a detailed two-level taxonomy of disciplines and sub-disciplines.
- Figure 6 presents the distribution of disciplines and sub-disciplines in GRADE.
A.3.1 Comparison of Relaxed Score
The relaxed score aggregates Discipline Reasoning, Visual Consistency, and Logical Readability after normalization. It retains a similar performance gap between closed-source and open-source models.
- The relaxed score is a weighted average of the three evaluation dimensions after each is normalized to [0, 100].
- The weights are 0.6 for Discipline Reasoning, 0.3 for Visual Consistency, and 0.1 for Logical Readability.
- The relaxed score still shows a similar performance gap between closed-source and open-source models.
A.3.2 Ablation on GT Input in Discipline Reasoning Evaluation
Including GT inputs improves the reliability of Discipline Reasoning evaluation across agreement, error, and variance measures.
- 0.8505 vs. 0.7642 Pearson correlation indicates higher agreement with human evaluation when GT inputs are provided.
- 0.0975 vs. 0.1311 MAE shows lower evaluation error with GT inputs than without them.
- 0.2340 vs. 0.2841 variance demonstrates reduced score variability when GT inputs are included.
A.3.3 Ablation on Instruction Explicitness
The supplementary ablation extends instruction-explicitness results across more models, while the evaluation prompts define structured inputs, scoring targets, consistency criteria, and strict answer formats.
- A.3.3 Ablation on Instruction Explicitness: 72.0% to 76.4% relaxed score shows Seedream 5.0 improves when instructions become more explicit.
- A.3.3 Ablation on Instruction Explicitness: The observed trend remains consistent across additional models: making instructions more explicit generally improves performance to varying degrees.
- A.3.4 Human Alignment of Relaxed Score: 0.8505 Pearson correlation is reported between relaxed scores and mean human ratings, supporting the selected judge model.
- A.3.4 Human Alignment of Relaxed Score: Gemini-3-Flash reaches 0.8505 correlation, exceeding GPT-5 at 0.7192 and Qwen3-VL-235B-Instruct at 0.6105.
- A.4 Prompts: The prompt pipeline uses an original image, edited image, GT image, and scoring-point questions as evaluation inputs.
- A.4 Prompts: Scoring requires strict Yes-or-No answers based on visible evidence, while text and mathematical relations allow meaning-equivalent or visually apparent matches.
- A.4 Prompts: The primary target is the edited image, while the original and GT images may verify the intended edit and correct outcome.
- A.4 Prompts: Visual consistency evaluation checks unchanged academic content and visual style, including axes, labels, values, colors, fonts, and structure.