Source-linked AI summary
OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang
TL;DR
Existing benchmarks provide limited support for comprehensive multimodal physics evaluation across educational stages and output generation. OmniPhys addresses this gap with a unified Chinese benchmark and evaluates both reasoning and diagram-related capabilities. Experiments show persistent challenges, with leading models below 70% strict mastery, while the Chinese-only scope limits global representativeness and complicates separation of language alignment from physics reasoning.
Problem
Existing physics benchmarks rarely combine cross-stage knowledge, complex multimodal comprehension, and multimodal output generation.
Method
OmniPhys is a unified Chinese benchmark spanning secondary education to university, with joint reasoning evaluation and a Physics Diagram Editing task.
Results
Leading models remain below 70% strict mastery, and experiments reveal persistent challenges in conceptual mastery and multimodal physics reasoning.
Takeaways & Limitations
OmniPhys provides a benchmark for evaluating multimodal physics understanding and generation beyond surface-level answer matching.
Takeaways & Limitations
The data are predominantly from Chinese educational materials, so weaker performance may reflect limited physics reasoning, insufficient Chinese multimodal alignment, or both.
Abstract
from arXiv · showhide
Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a comprehensive benchmark. To fill this gap, we introduce OmniPhys, a large-scale benchmark for multimodal physics understanding and reasoning, covering middle school through university-level problems from Chinese Educational Corpora. OmniPhys consists of 15,246 questions and 19,850 images, accompanied by detailed annotations that support fine-grained analysis of reasoning processes and knowledge usage. Beyond conventional evaluation, OmniPhys is a benchmark that systematically evaluates multimodal outputs in the physics domain, including models' ability to generate structured physics diagrams, which constitute a fundamental component of authentic physics problem solving. Extensive evaluations reveal critical gaps in the capabilities of current MLLMs, especially in complex reasoning and visual generation. To address this, we release OmniPhys to serve as a foundational resource for advancing multimodal intelligence in physics and scientific domains. Codes and data are available at https://github.com/ECNU-RAIL/OmniPhys-EMNLP2026.
1 Introduction
OmniPhys addresses the lack of a comprehensive benchmark for multimodal physics reasoning across educational stages, inputs, and outputs. It combines broad physics coverage with diagram synthesis and editing evaluation, while experiments expose persistent reasoning challenges.
- Physics requires integrating textual descriptions, visual diagrams, and symbolic logic for accurate reasoning.
- Existing physics datasets rarely combine middle-school-to-university knowledge fusion, complex multimodal comprehension, and multimodal output generation.
- OmniPhys covers physics mastery from secondary education to university using textual, visual, and symbolic inputs across five disciplines.
- The benchmark introduces a subset assessing models’ ability to synthesize and edit physics diagrams.
- Extensive evaluations find persistent challenges in conceptual mastery and substantial room for improvement across proprietary and open-source MLLMs.
2 Related Works
Related work shows that multimodal models and physics datasets have advanced, but rigorous multimodal physics reasoning remains insufficiently evaluated. Existing benchmarks commonly emphasize text or limited image inputs and incomplete educational coverage.
- Frontier MLLMs support complex interleaved inputs and outputs but remain vulnerable to hallucination, logical deduction failures, and numerical errors.
- These vulnerabilities motivate challenging benchmarks that probe the upper bounds of deep multimodal reasoning.
- Existing physical reasoning datasets predominantly use text-only tasks or general K-12 benchmarks with limited physics subsets.
- Recent image-input datasets still have incomplete educational coverage and lack multimodal output evaluation.
3 The OmniPhys Benchmark
OmniPhys is a multimodal physics benchmark spanning junior high to university curricula and combining diagrams, formulas, and text for visually grounded reasoning. Its construction uses curated educational sources and multi-stage quality, difficulty, leakage, and human-validation procedures.
- OmniPhys spans junior-high-to-university curricula and challenges MLLMs to synergize diagrams, formulas, and text for visually grounded reasoning.
- The benchmark uses authoritative Chinese pedagogical resources and authentic examinations, releasing structured annotations rather than verbatim source PDFs.
- Its difficulty pyramid scales from foundational K-12 concepts to advanced university challenges.
- Data Preprocessing and Quality Filtering: The preprocessing pipeline checks schema completeness, removes short or incomplete samples, and deduplicates problems using a 0.95 cosine-similarity threshold.
- Data Preprocessing and Quality Filtering: 12.1% of redundant instances were eliminated through semantic deduplication to produce the final valid dataset.
- Data Preprocessing and Quality Filtering: Visual-dependency filtering distinguishes text-solvable, text-descriptive, and image-essential problems according to whether diagrams add necessary information.
- Data Preprocessing and Quality Filtering: Adversarial difficulty screening discards problems correctly solved by both Qwen2.5-VL-3B and MiniCPM-V-2.6.
- Data Preprocessing and Quality Filtering: Data-leakage prevention combines source selection, web-search filtering, and manual verification, while expert cross-validation assesses pipeline consistency.
4 Experiments
OmniPhys evaluates proprietary and open-source MLLMs on multimodal physics reasoning across educational stages, using separate result and process measures. Experiments show persistent gaps in conceptual mastery, advanced-stage performance, and physically correct diagram generation.
- Experimental Setup: The evaluation covers diverse proprietary and open-source MLLMs, including multiple Qwen2.5-VL and Qwen3-VL model scales.Proprietary models are accessed through official APIs, while most open-source models are evaluated locally under identical inference settings.
- Dual-Track Reasoning Evaluation: Dual-Track Reasoning Evaluation separates objective tasks with deterministic outputs from open-ended tasks requiring step-by-step derivations.The framework is designed to assess both final-answer correctness and reasoning quality across different task types.
- Dual-Track Reasoning Evaluation: Result Score measures objective-task answer accuracy after normalization, while Process Score measures recall of correctly recovered key reasoning steps.For multi-select and multi-slot tasks, partial credit is awarded only when predicted answers contain no incorrect options.
- Main Results Analysis: Accuracy peaks at Junior High and decreases monotonically toward University across both closed-source and open-source models.The consistent decline is presented as evidence that OmniPhys captures increasing cognitive complexity across educational stages.
- Main Results Analysis: Strict mastery rates Pobj remain below 70% even for leading models, indicating that perfect alignment of answers and reasoning is uncommon.Pobj requires S1 = S2 = 1.0, while Popen requires flawless derivation with S3 = 1.0.
- Main Results Analysis: Multimodal generation exposes failures that text-only metrics miss: models may produce visually faithful diagrams while violating physical constraints.Reported errors include incorrect vector directionality, broken topology, incorrect convex-lens focusing, and redundant line segments.
5 Ablation Study
The adversarial Test-Mini subset probes difficult multimodal physics reasoning under three input settings. Visual inputs usually help, but some models treat visual information as distractor noise, revealing cross-modal alignment challenges.
- 5 Ablation Study: Test-Mini uses 10% of the data, selected by five-model consensus from samples weighted toward empirical failure rates and reasoning complexity.The selection weighting is 75% failure rate and 25% reasoning complexity.
- 5 Ablation Study: The ablation compares Text+Img, Text+Caption, and Text-only settings to quantify models’ dependence on visual information.Text+Img retains the original multimodal input, Text+Caption replaces diagrams with descriptions, and Text-only removes visual content.
- 5 Ablation Study: Test-Mini is substantially harder than the main benchmark and serves as a rigorous probe of deep reasoning capabilities.The passage explicitly characterizes the subset as significantly more difficult.
- 5 Ablation Study: On average, Text+Img achieves the highest performance, with MLLMs consistently outperforming LLMs across the evaluated settings.This pattern supports the role of visual constraints in the tested physics problems.
- 5 Ablation Study: Some models perform better with Text-only inputs, suggesting that visual inputs or captions can be misinterpreted as distractor noise by certain architectures.The authors identify persistent cross-modal alignment challenges requiring further investigation.
6 Conclusion
OmniPhys is presented as a unified benchmark spanning junior-high to university physics and covering both understanding and generation. Its evaluations expose reasoning and physical-fidelity gaps, with leading models below 70% strict mastery.
- 6 Conclusion: OmniPhys evaluates physics reasoning across the full educational spectrum from junior high to university, covering both understanding and generation.The benchmark is positioned as a unified multimodal evaluation suite.
- 6 Conclusion: The DTRE protocol jointly scores final answers and reasoning fidelity, exposing reasoning shortcuts that conventional answer-matching metrics overlook.This moves evaluation beyond surface-level final-answer correctness.
- 6 Conclusion: OmniPhys introduces Physics Diagram Editing and reveals that frontier image-generation models systematically violate physical laws.These violations are described as a blind spot of text-only benchmarks.
- 6 Conclusion: Leading models remain below 70% strict mastery, while ablations confirm the indispensable role of visual grounding.Experiments include both proprietary and open-source MLLMs.
- 6 Conclusion: The paper positions OmniPhys as both an evaluation suite and a diagnostic instrument for advancing physically grounded multimodal intelligence.This conclusion follows the benchmark’s combined reasoning and generation assessments.
7 Limitations
The study’s main limitations concern its predominantly Chinese educational data and the scalability of multimodal-output evaluation. These constraints complicate separating language alignment from physics reasoning and require continued human oversight.
- 7 Limitations: The data are predominantly drawn from Chinese educational materials, limiting linguistic and curricular representativeness.Future iterations are intended to incorporate international curricula such as A-Level and IPhO.
- 7 Limitations: English instructions paired with Chinese problems can confound weaker performance between limited physics reasoning and insufficient Chinese multimodal alignment.This issue arises under the paper’s cross-lingual prompting protocol.
- 7 Limitations: Methods for evaluating multimodal physical outputs remain underexplored, especially because MLLM judges may overestimate generated-content quality.Such judges can overlook subtle physical inconsistencies.
- 7 Limitations: The structured rubric mitigates judge weaknesses, but a fully scalable evaluation solution is still lacking.The authors therefore continue to target a robust framework for multimodal physics outputs.
A Prompt Usage
The appendix specifies prompt and annotation procedures for visual-dependency labeling, model inference, reasoning evaluation, diagram grading, and human validation. These procedures standardize both textual reasoning assessment and multimodal-output judgments.
- A Prompt Usage: Prompts are translated into English for readability, but experiments use Chinese prompts to align with OmniPhys and preserve cross-modal consistency.This applies to inference and evaluation.
- A.1 Visual Dependency Annotation: Visual-dependency annotation assigns Level 1, 2, or 3 according to whether images are decorative, textually described, or essential for solving.Level 3 applies when critical quantities or relationships appear only in the image.
- A.2 Model Inference Prompt: The model-inference prompt requires step-by-step reasoning with each step limited to 30 words, followed by a final answer in a fixed format.The prompt also instructs models to avoid redundant explanations and extra content.
- A.3 Automated Evaluation Prompts: Automated reasoning evaluation decomposes standard solutions into m key steps, counts correctly included steps n, and computes process score as n/m.The evaluator receives the question, standard answer, explanation, and student response.
- A.3 Automated Evaluation Prompts: Objective tasks use dual-track evaluation: process validity is scored separately from final-result correctness using task-specific rules.Single-choice and true-false matches receive 1.0, while multi-select and fill-in tasks use correct-slot fractions.
- A.4 Automated Evaluation Prompts for Multimodal Generation: A vision-language judge scores generated physics diagrams for physical fidelity using the problem, optional context image, and student answer image.Diagram grading distinguishes fully correct, partially correct, and incorrect outputs based on physical elements and laws.
- B Annotation Interface Details: Human annotation uses Label Studio to compare problem statements, input images, references, and generated images, with final scores averaged across three physics-trained graduate students.Scores reflect correctness, completeness, and visual faithfulness.
C.1 Hardness-Based Selection Methodology
OmniPhys constructs Test-Mini by ranking questions with a hardness score combining empirical model failure and reasoning complexity, then selecting the top 10%. This produces a substantially more difficult subset whose challenge is consistent across model architectures.
- Hardness scoring: Test-Mini ranks candidate questions using empirical model failure and reasoning complexity, combining them as H(xi) = w1 · F(xi) + w2 · C(xi).The full dataset contains N = 12,885 questions without multimodal outputs.
- Hardness scoring: Empirical Failure Rate F(xi) aggregates normalized scores from five baseline models, while Reasoning Complexity C(xi) uses the percentile rank of ground-truth explanation length.
- Hardness scoring: Weights w1 = 0.75 and w2 = 0.25 prioritize empirical difficulty while retaining logical depth in the hardness score.
- Subset construction: The top 10% of questions ranked by H(xi) form Test-Mini, using five state-of-the-art closed-source models as the baseline set M.
- Difficulty validation: All-Model Failure Rate rises from 1.6% to 15.5% in Test-Mini, alongside a substantial increase in average reasoning-chain length.
- Difficulty validation: Baseline-model performance drops by 56.9% to 72.5% across all models, indicating that Test-Mini difficulty is not specific to one architecture.
- Data authenticity: Original screenshots from Chinese examinations and authoritative textbooks illustrate the benchmark’s middle-school Mechanics, high-school Electromagnetism, and university Optics problems.
E Human-Machine Alignment and Quality Control in Data Filtering Process
A blind review by two physics experts assessed 300 instances to examine automated visual-dependency filtering and difficulty screening. The reported alignment indicates that the adversarial filtering strategy targets non-trivial problems while preserving pedagogical integrity.
- Study design: Two physics experts manually reviewed a random sample of 300 instances in a blind correlation study of visual-dependency filtering and difficulty screening.
- Quality-control result: High alignment scores show that the adversarial filtering strategy effectively targets non-trivial problems while maintaining pedagogical integrity across physics domains.
F Human-Machine Alignment on Evaluation
OmniPhys validates its LLM-as-a-Judge evaluation through comparison with independent physics-expert scoring of generated responses. The judges rank performance consistently with experts, while humans score slightly more stringently and strict-correctness agreement remains high.
- Study design: Three physics experts independently scored 300 model-generated responses across five physics domains and three task types using the same rubric as the LLM judges.
- Study design: Table 9 compares expert consensus with the DeepSeek-V3 and GPT-4 LLM-as-a-Judge framework using correlation, agreement, and mean-score-difference metrics.
- Alignment results: Pearson correlation r exceeds 0.85 for S1 (Accuracy) and S2 (Objective Process), showing strong ranking consistency between LLM judges and human experts.
- Alignment results: Human experts score 1.2% to 4.2% lower on average than the LLM consensus because they are more sensitive to subtle conceptual inaccuracies in complex reasoning chains.
- Alignment results: Agreement on strict correctness rates Pobj and Popen reaches 91.2%, supporting the reliability of the perfect-performance threshold for detecting model mastery.
- Implication: Despite minor absolute-value bias, strong cross-dimensional correlation supports scaling the automated evaluation protocol to OmniPhys’s large-scale assessment.