Source-linked AI summary
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao
TL;DR
Scientific video generation needs evaluation of mechanistic and causal fidelity beyond visual plausibility, but existing benchmarks lack multidisciplinary, expert-grounded coverage. Sci-VBench addresses this with 1,253 expert-authored tasks and rubric-based evaluation, finding that perceptual quality varies little while scientific reasoning performance reveals a substantial proprietary–open-source gap.
Problem
Existing video-generation benchmarks provide limited multidisciplinary, expert-grounded evaluation of scientific mechanisms, causal relations, and temporal dependencies beyond perceptual plausibility.
Method
Sci-VBench combines 1,253 expert-authored tasks across 60 subjects and four disciplines with reference guides, scoring rubrics, and automatic and human evaluation protocols.
Results
Across 16 models, perceptual-quality scores are nearly flat, while Prompt Grounding and Scientific and Causal Correctness vary substantially with a pronounced proprietary–open-source gap.
Takeaways & Limitations
Visual realism alone does not reliably indicate scientific validity, highlighting the need to evaluate mechanistic reasoning and temporally coherent causal dynamics.
Abstract
from arXiv · showhide
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.
1 Introduction
Sci-VBench addresses the lack of knowledge- and reasoning-intensive evaluation for scientific video generation with an expert-annotated benchmark and scalable rubric-based protocol. Evaluations of 16 models reveal a proprietary–open-source gap in scientific and causal dimensions despite nearly flat perceptual-quality scores.
- 1 Introduction: The benchmark targets knowledge- and reasoning-intensive, visually verifiable video generation, extending evaluation beyond perceptual plausibility toward faithful events, constraints, and temporal dependencies.Scientific settings are consequential because perceptual plausibility can mask fundamental errors.
- 1 Introduction: Existing benchmarks emphasize perceptual quality, prompt alignment, compositionality, or generic physical commonsense, leaving knowledge- and reasoning-intensive scientific video generation under-evaluated.Knowledge- and reasoning-intensive evaluation has primarily been studied for video understanding rather than generation.
- 1 Introduction: Sci-VBench contains 1,253 expert-authored and independently reviewed tasks spanning 60 subjects across Natural Science, Healthcare, Humanities & Social Sciences, and Engineering.Each task pairs a minimal generation prompt with an expert-authored evaluation specification.
- 1 Introduction: The rubric-based protocol uses expert ratings as reference labels and combines VBench Vision Tools with rubric-conditioned MLLM judging across perceptual, grounding, scientific-causal, and spatiotemporal dimensions.A controlled human study shows that evaluation specifications substantially improve non-expert agreement with experts.
- 1 Introduction: Evaluations of 16 frontier models show a pronounced proprietary–open-source gap concentrated in Prompt Grounding and Scientific and Causal Correctness, while the automatic perceptual-quality proxy is nearly flat.Qualitative analysis further finds systematic mechanistic errors in leading models despite visually convincing outputs.
2 Related Work
Video-generation research has advanced from short, low-fidelity clips to high-fidelity, temporally coherent outputs through diffusion–transformer systems, while increasingly addressing physical-law adherence and commonsense reasoning. Existing benchmarks mainly evaluate perceptual, motion, compositional, interaction, and prompt-following capabilities rather than domain-specific scientific mechanisms.
- Video-Generation Benchmarks: Sci-VBench is positioned against prior benchmarks by incorporating expert involvement and reusable per-example evaluation guides or rubrics as comparison dimensions.Table 1 defines “Expert” for domain-expert participation and “Eval Spec” for released evaluation guides or rubrics.
- Video Generation: Video-generation systems have progressed from short, low-fidelity clips to high-fidelity, temporally coherent outputs through diffusion models and large-scale transformers.Recent work also increasingly targets adherence to physical laws and commonsense reasoning.
- Video-Generation Benchmarks: Existing video-generation benchmarks primarily assess perceptual fidelity and motion quality, with some extending evaluation to compositionality, object interaction, and higher-level prompt following.These benchmarks are not designed to test faithfulness to domain-specific mechanisms through rubric-verifiable scientific evaluation.
3 Sci-VBench Benchmark
Sci-VBench contains 1,253 expert-authored examples spanning 60 subjects across four disciplines, designed to test knowledge-grounded, multi-step causal and spatiotemporal reasoning. Its benchmark construction combines expert annotation and review with reference guides and 1–5 scoring rubrics that make mechanism-level evaluation accessible beyond domain experts.
- Benchmark desiderata: Sci-VBench covers 1,253 examples across 60 subjects and four disciplines, emphasizing broad domain knowledge and expert-level causal reasoning.Every prompt is authored by a domain expert and targets phenomena requiring grounded scientific understanding and multi-step reasoning.
- Evaluation dimensions: The benchmark evaluates perceptual fidelity, prompt grounding, scientific and causal correctness, and spatiotemporal consistency.The latter dimensions isolate mechanism-level failures that generic quality and prompt-alignment criteria do not capture.
- Benchmark construction: Experts select textbook-grounded concepts with observable, mechanism-governed visual realizations and write minimal prompts that require models to infer the underlying mechanism.Prompts specify the initial setup and intervention or objective while omitting the expected mechanistic trajectory and key phenomena.
- Evaluation specification: Each prompt receives a high-level reference guide and a detailed 1–5 rubric that define expected causal transitions, observable evidence, partial correctness, and failure conditions.This externalizes the expert knowledge needed for evaluation so that non-experts or MLLM judges can apply the specification.
- Quality control and splits: Every example undergoes independent domain-expert review for clarity, visual testability, and consistency among the prompt, reference guide, and rubric.The benchmark also provides a 150-example testmini split for rapid, cost-constrained evaluation; open-source models use both splits, whereas proprietary systems use testmini only.
4 Sci-VBench Evaluation Protocol
Sci-VBench uses expert 1–5 ratings as reference labels and a scalable automated protocol combining VBench metrics with rubric-conditioned MLLM judging. Reliability analyses compare these evaluations with non-expert and MLLM-as-Judge ratings, finding that evaluation specifications and larger judge models improve agreement with experts.
- Human Evaluation: Expert annotators provide the primary reference labels by scoring every generated video from 1–5 on all four evaluation dimensions after discipline-specific calibration.Each annotator receives the prompt, generated video, and per-example evaluation specification, and maps observable cues to rubric scores.
- Automated Evaluation: The automated protocol scales evaluation while preserving the per-example specifications: VBench Vision Tools report Low-level Perceptual Fidelity, and a rubric-conditioned judge scores the other three dimensions.The MLLM-as-Judge evaluates Prompt Grounding, Scientific and Causal Correctness, and Spatiotemporal Consistency using Qwen3.5-397B-A17B with native video input.
- Reliability Analysis: Reliability is measured against expert ratings using non-expert cohorts with different guidance conditions and MLLM judges supplied with evaluation specifications.The reliability analysis uses expert-scored videos and compares agreement across human and automated evaluators.
- Reliability Analysis: Non-expert agreement with experts is evaluated both with and without per-example evaluation specifications, testing whether rubric guidance improves reliability.The without-specification cohort sees only the prompt and video, whereas the with-specification cohort receives the reference guide and scoring rubric.
- Reliability Analysis: Qwen3.5-397B-A17B attains the highest overall correlations with expert ratings, while agreement broadly increases with MLLM evaluator scale.The comparison includes Qwen3.5-397B-A17B, Gemma-4-31B, and Qwen3.5-9B; every evaluator is assessed against expert ratings across the four dimensions.
5 Experiment
The experiment benchmarks 16 proprietary and open-source text-to-video models on Sci-VBench using verbatim prompts and finds that scientific and causal reasoning, rather than perceptual appearance, separates systems. Proprietary models lead on reasoning dimensions, while open-source models can match or exceed them on spatiotemporal consistency, and observed failures involve instruction adherence, scientific simulation, and temporal or visual quality.
- 5.1 Main Results: The benchmark evaluates 16 frontier text-to-video models—eight proprietary and eight open-source—using verbatim prompts under each model’s default configuration.Per-model versions, resolutions, frame rates, and clip durations are listed in Appendix A.3.
- 5.1 Main Results: Automatic perceptual quality is nearly flat across models, with VT ranging from 3.79 to 4.12, whereas mechanism-sensitive dimensions vary substantially.The results characterize current systems as differing more in mechanism than appearance.
- 5.1 Main Results: Testmini is a faithful low-cost proxy for the full benchmark, with open-source models’ per-model full-benchmark averages differing from testmini by at most 0.07.Table 4 reports testmini results for all 16 models, while complete full-benchmark per-dimension results appear in Appendix C.1.
- 5.1 Main Results: Proprietary models lead on scientific and causal reasoning, while open-source models lag substantially on SCC despite matching or exceeding proprietary systems on spatiotemporal consistency.MiniMax-H3 leads the open-source group with 1.63 on SCC, half the proprietary best, whereas Wan2.2-5B records the highest automatic SC score at 2.79.
- 5.2 Error Analysis: Observed errors fall into poor instruction adherence, inaccurate simulation of scientific principles, and deficiencies in temporal coherence and visual quality.These failures include missing prompt details, factually incorrect dynamics caused by prioritizing aesthetics, and objects changing illogically over time.
- 5.3 Prompting Ablation: An explicit-prompting ablation rewrites testmini prompts to make scenes and requested dynamics more visually concrete, then regenerates videos with Wan2.2-5B and HunyuanVideo-1.5 under unchanged settings.The gains are reported as consistently ordered across both models in Figure 4.
6 Conclusion
Sci-VBench evaluates text-to-video models as expert-domain world simulators through 1,253 expert-authored prompts spanning 60 subjects across four disciplines. Its benchmark emphasizes mechanistic reasoning and temporally coherent causal dynamics, revealing a gap between perceptual realism and scientific validity.
- Evaluation focus: Success on Sci-VBench depends on mechanistic reasoning and temporally coherent causal dynamics rather than surface-level realism.
- Key finding: Benchmarking 16 frontier proprietary and open-source models reveals a persistent gap between perceptual realism and scientific validity.
A Sci-VBench Dataset · A.1 Subject Selection
Sci-VBench evaluates knowledge- and reasoning-intensive scientific video generation using 1,253 expert-annotated examples across 60 subjects and four disciplines. Its rubric-based analysis finds substantial differences in scientific and causal correctness despite clustered perceptual-quality scores.
- A Sci-VBench Dataset: Sci-VBench establishes a rubric-based evaluation protocol for assessing generated scientific videos.
- A Sci-VBench Dataset: Non-expert human evaluators and MLLM-as-Judge systems achieve relatively high agreement with expert judgments under the protocol.
A.2 Data Quality Control … B.1 Low-level Video Quality: Definitions and Implementation Details (Adapted from VBench)
The benchmark audits evaluator consistency, documents model-generation settings and annotators, and evaluates low-level video quality through dimensions covering temporal coherence, motion, dynamics, aesthetics, and imaging fidelity.
- A.2 Data Quality Control: A targeted audit samples 200 examples, with a second same-subject annotator independently constructing each example’s full evaluation specification without seeing the first.The specifications include a high-level reference guide and anchored 1–5 rubrics for all dimensions.
- A.3 Video Generation Models: Table 6 documents the detailed settings of the evaluated video generation models.Videos use verbatim benchmark prompts and each model’s default configuration; supported durations are closest to 15 seconds, with some models capped at 5–10 seconds.
- A.4 Annotator Information: Annotator biographies are provided in Tables 7 and 8, while biographies of author-annotators are withheld to preserve anonymity.The tables cover 61 annotators involved in constructing Sci-VBench.
- B.1 Low-level Video Quality: Definitions and Implementation Details (Adapted from VBench): Subject and background consistency measure whether identities, attributes, layouts, lighting, and spatial structure remain stable across frames without unintended changes.Subject consistency emphasizes preserving the main subject’s identity and appearance, whereas background consistency addresses environmental continuity and flicker.
- B.1 Low-level Video Quality: Definitions and Implementation Details (Adapted from VBench): Motion Smoothness evaluates temporally continuous, physically plausible movement, while Dynamic Degree measures motion intensity and richness without judging correctness.The motion criterion penalizes jitter, abrupt jumps, and temporal artifacts; Dynamic Degree captures movement, deformation, and interaction rather than staticness.
- B.1 Low-level Video Quality: Definitions and Implementation Details (Adapted from VBench): Aesthetic Quality assesses perceptual appeal through composition, color harmony, lighting, balance, and stylistic coherence.It measures how pleasing and well-structured the video appears from a human perceptual perspective.
- B.1 Low-level Video Quality: Definitions and Implementation Details (Adapted from VBench): Imaging Quality measures frame-level visual fidelity using sharpness, resolution, noise, compression artifacts, blur, and rendering clarity.The metric reflects how clean and realistic generated images appear at the pixel level.
B.2 Comparing with Prior Video Scoring Methods · B.3 Evaluation Prompt Templates · C Additional Experimental Results
The rubric-conditioned MLLM-as-Judge aligns more closely with expert ratings than prior automatic video-scoring methods on key reasoning-centric dimensions. Its evaluation uses dimension-specific prompts with anchored 1–5 rubrics applied to each benchmark example.
- B.2 Comparing with Prior Video Scoring Methods: The comparison adopts VideoScore, which averages five released evaluator scores covering visual quality, temporal consistency, dynamic degree, text-to-video alignment, and factual consistency.
- B.2 Comparing with Prior Video Scoring Methods: The rubric-conditioned MLLM-as-Judge achieves the highest instance-level Pearson correlation among compared automatic methods on Prompt Grounding, Scientific and Causal Correctness, and Spatiotemporal Consistency.
- B.2 Comparing with Prior Video Scoring Methods: ETVA is comparatively competitive on Prompt Grounding and Scientific and Causal Correctness but degrades sharply on Low-level Perceptual Fidelity and Spatiotemporal Consistency.
- B.3 Evaluation Prompt Templates: Figure 6 presents the released evaluation specification: the verbatim generation prompt followed by anchored 1–5 rubrics for each judged dimension.
- B.3 Evaluation Prompt Templates: Figure 5 specifies a template that conditions MLLM-as-Judge on one evaluation dimension, with angle-bracketed fields filled separately for each example.
- B.3 Evaluation Prompt Templates: The rubric example uses the benchmark’s verbatim knee-jerk-reflex prompt and expert-authored evaluation anchors.
C.1 Full-Benchmark Results
This section presents complete automatic evaluation results for open-source models on all 1,253 Sci-VBench examples. Results cover each evaluation dimension and report corresponding overall averages for comparison.
- The evaluation reports results separately for each dimension, using metric definitions consistent with Table 4.
- Overall averages are included in the Full Avg. column of Table 4, while Table 9 sorts open-source models by its rightmost Avg. column.
- Table 9 covers the complete automatic evaluation results for open-source models across the full Sci-VBench benchmark of 1,253 examples.
C.2 Error Analysis
The error analysis identifies three major failure modes in knowledge-intensive text-to-video generation: poor instruction adherence, inaccurate scientific simulation, and deficiencies in temporal coherence and visual quality.
- Poor Adherence to Instructions: Models often miss fine-grained prompt details, such as failing to render a specified momentary push-button switch in the instructed setup.The example includes a single LED with a resistor while a hand repeatedly presses and releases the button.
- Inaccurate Simulation of Scientific Principles: Models frequently violate scientific principles by failing to reason from preconditions to scientifically plausible outcomes.Examples include the un-struck leg kicking during a knee-jerk reflex and a bag being squeezed beside the eye rather than from the side.
- Deficiencies in Temporal Coherence and Visual Quality: Temporal and visual defects include object permanence failures, discontinuous style shifts, coarse textures, and weak causal links between events.In one example, two balls released simultaneously on separate tracks merge into one during motion.