Source-linked AI summary
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Yuqi Tang, Yang Shi, Zhuoran Zhang, Qixun Wang, Xuehai Bai, Yue Ding, Ruizhe Chen, Bohan Zeng, Xinlong Chen, Xuanyu Zhu, Bozhou Li, Yuran Wang, Yifan Dai, Chengzhuo Tong, Xinyu Liu, Yiyan Ji, Yujie Wei, Yuhao Dong, Shilin Yan, Fengxiang Wang, Yi-Fan Zhang, Haotian Wang, Yuanxing Zhang, Pengfei Wan
TL;DR
It is unclear whether MLLMs can reliably perceive and reason about artifacts in AI-generated videos across diverse domains. Artifact-Bench addresses this gap with a hierarchical taxonomy and three complementary evaluation tasks, finding substantial weaknesses and misalignment with human realism preferences.
Problem
Existing benchmarks provide limited systematic evidence on whether MLLMs perceive and reason about AI-generated video artifacts beyond photorealistic, semantic, or isolated evaluation settings.
Method
Artifact-Bench organizes realism artifacts into a three-level taxonomy and evaluates MLLMs through authenticity classification, pairwise realism comparison, and fine-grained artifact identification.
Results
Many evaluated MLLMs show near-random or below-random performance on challenging tasks, with judgments often misaligned with human perceptual preferences.
Takeaways & Limitations
Current MLLMs remain unreliable for fine-grained realism assessment, artifact diagnosis, and use as evaluators or reward providers for video generative models.
Takeaways & Limitations
Resource constraints limit the number of human experts and the dataset scale, which can be expanded in future work.
Abstract
from arXiv · showhide
Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While Multimodal Large Language Models (MLLMs) show strong visual understanding capabilities, their ability to perceive and reason about such artifacts remains unclear. Existing benchmarks often lack systematic evaluation of artifact-aware perception and fine-grained diagnostic reasoning, especially across diverse AI-generated video domains beyond photorealistic content. To address this gap, we introduce Artifact-Bench, a comprehensive benchmark for evaluating MLLMs on AI-generated video artifact detection and analysis. We first establish a three-level hierarchical taxonomy of realism artifacts, covering photorealistic, animated, and CG-style videos. Based on this taxonomy, Artifact-Bench defines three complementary tasks: real vs. AI-generated video classification, pairwise realism comparison, and fine-grained artifact identification. Experiments on 19 leading MLLMs reveal substantial limitations in artifact perception and reasoning, with many models approaching random or even below-random performance in challenging settings. We further observe significant misalignment between MLLM judgments and human perceptual preferences, highlighting their limited reliability as general evaluators for AI-generated video realism.
1. Introduction
Artifact-Bench addresses the lack of systematic evaluation of MLLMs’ ability to perceive and reason about artifacts in AI-generated videos. It introduces a hierarchical artifact taxonomy and a multi-task benchmark, whose experiments expose substantial weaknesses in current models’ artifact-level perception and reasoning.
- Motivation: AI-generated videos still exhibit temporal inconsistencies, structural distortions, unnatural motion, and semantic incoherence that limit perceptual realism and reliable deployment.These imperfections may be subtle but remain fundamental barriers to realistic video generation.
- Motivation: Artifact-based detection offers a principled signal for distinguishing AI-generated videos from real videos because artifacts reflect intrinsic limitations of generation pipelines.The introduction motivates this direction for media authenticity, content moderation, and generative model evaluation.
- Research gap: Existing benchmarks emphasize semantic understanding or limited photorealistic scenarios, leaving unclear whether MLLMs use genuine artifact-aware perception rather than semantic priors or dataset biases.Although MLLMs process complex visual inputs and produce structured language, their artifact reasoning ability remains uncertain.
- Contributions: Artifact-Bench evaluates AIGC video detection, realism comparison, and artifact diagnosis across comprehensive scenarios, including CG and animation, using 30 evaluation aspects.The benchmark uses a multi-granularity progressive three-task system with difficulty levels.
- Contributions: Artifact-Bench establishes a three-level hierarchical taxonomy organizing AIGC artifacts from coarse visual abnormalities to fine-grained temporal and structural inconsistencies.The taxonomy is derived from systematic analysis of artifact characteristics, causes, and perceptual manifestations.
- Findings: Many MLLMs achieve near-random or below-random performance on certain tasks, revealing severe limitations in artifact-level perception, reasoning, and human-aligned realism understanding.These findings indicate that artifact-aware perception remains far from solved despite strong general vision-language capabilities.
2. Related Work
MLLMs have demonstrated strong visual understanding, temporal processing, and multimodal reasoning, enabling video applications and motivating their use in AI-generated video detection and realism assessment. Related benchmarks evaluate video quality through pairwise preference comparisons or diagnostic question answering.
- MLLMs for Video Understanding: MLLMs support video applications including visual question answering, captioning, and optical character recognition through temporal information processing.Their broader capabilities also include complex visual reasoning for sophisticated real-world scenarios.
- MLLMs for Video Understanding: Recent studies apply MLLMs to automated detection and realism assessment of AI-generated videos, including BusterX++ and Skyra.
- Artifact Assessment Benchmarks: UVE-Bench uses pairwise comparison scoring across fine-grained dimensions with human preference annotations, whereas VF-Eval frames evaluation as diagnostic question answering.
3. Artifact-Bench
Artifact-Bench organizes AI-generated video realism artifacts into a hierarchical, diagnostic taxonomy and evaluates MLLMs through three complementary tasks spanning recognition, comparison, and fine-grained explanation. Its taxonomy covers diverse video styles with 30 actionable artifact labels and supports multi-label annotation for co-occurring failures.
- Artifact taxonomy: Artifact-Bench’s taxonomy has three hierarchical tiers, progressing from broad artifact domains to fine-grained diagnostic labels.The taxonomy is designed to remain comprehensive, interpretable, and actionable for human annotation and model evaluation.
- Artifact taxonomy: The top-level domains are Surface Artifacts, Structural Defects, and Temporal-Semantic Violations.They respectively capture local visual defects, object-and-scene organization failures, and higher-level failures requiring cross-frame and semantic reasoning.
- Artifact taxonomy: The finest tier contains 30 concrete, visually observable artifact types, including Texture Inconsistency, Irreversibility Violation, and Cross-Shot Coherence.These fine-grained types form the operational label space for artifact-oriented evaluation.
- Artifact taxonomy: The taxonomy supports multi-label annotations because videos may contain co-occurring artifacts and failures spanning multiple analysis levels.A single example can involve both structural deformation and temporal inconsistency.
- Evaluation tasks: Artifact-Bench defines three complementary tasks: real-versus-AI classification, pairwise realism comparison, and fine-grained artifact identification.The tasks progressively assess authenticity recognition, relative realism judgment, and explanation of why an AI-generated video appears unrealistic.
4. Experiments
Experiments on 19 MLLMs show substantial limitations in detecting and diagnosing AI-generated video artifacts, especially on fine-grained artifact identification. Performance remains poorly aligned with human perceptual preferences, limiting MLLMs’ reliability for realism evaluation.
- Evaluation Setup: 19 MLLMs, spanning proprietary, open-source general-purpose, and specialized video-detection models, are evaluated on Artifact-Bench.The evaluation includes 2 proprietary, 14 open-source general-purpose, and 3 open-source specialized models.
- Overall Performance: Most MLLMs fail to consistently surpass the approximately 50% random-guessing baseline, particularly at higher difficulty levels.Even Gemini 3.1 Pro achieves only an overall score of 47.5 on Artifact-Bench.
- Task Difficulty: All models achieve less than 10% average accuracy on AID, which requires selecting multiple valid artifact categories from six candidates.AID is substantially more challenging than binary RVAC and pairwise PVRC because multiple answers may be correct.
- Failure Analysis: Fine-grained and temporal-spatial perception are critical bottlenecks because artifacts may occupy tiny regions or emerge across multiple frames.Examples include a paddle penetrating a boat hull and a football changing from two balls to one and back to two.
- Failure Analysis: Larger models and explicit reasoning do not necessarily improve artifact detection, with InternVL3.5-38B performing comparably to its 8B counterpart and some thinking variants underperforming instruction-tuned variants.These results indicate that artifact detection requires capabilities beyond general semantic understanding and chain-of-thought reasoning.
- Human Alignment: MLLM artifact perception substantially mismatches human perceptual preferences as realism ambiguity increases, limiting their reliability as general-purpose evaluators.The misalignment may also hinder their use as reward providers or automated judges for optimizing video generative models.
5. Conclusion
Artifact-Bench evaluates whether MLLMs can detect and diagnose artifacts in AI-generated videos through a three-level taxonomy and three complementary tasks. Experiments show that current MLLMs still struggle with artifact-level perception and reasoning.
- Benchmark objective: Artifact-Bench evaluates MLLMs’ ability to detect and diagnose artifacts in AI-generated videos.The benchmark targets artifact-aware video understanding rather than authenticity recognition alone.
- Evaluation design: A three-level artifact taxonomy and three complementary tasks provide systematic evaluation from coarse-grained authenticity recognition to fine-grained artifact identification.The evaluation spans multiple levels of artifact analysis and task granularity.
- Main finding: Current MLLMs still struggle with artifact-level perception and reasoning.The conclusion identifies persistent limitations in models’ ability to perceive and reason about video artifacts.
A. Experiment Details … A.3. Answer Extraction Prompt
Artifact-Bench standardizes inference with fixed video sampling and model-specific decoding, then evaluates three artifact-aware tasks through explicit prompts and deterministic answer extraction. The prompts distinguish AIGC artifacts from non-AIGC visual styles, compare realism by artifact severity, and support multi-label artifact identification.
- A.1. Experimental Setup: All models use FPS=5 video sampling, with frame resizing applied when context limits or high-resolution inputs make inference necessary.This setup is intended to ensure feasible inference across evaluated models.
- A.1. Experimental Setup: Decoding uses each model’s official recommended settings when available; otherwise, greedy decoding is the default.Gemini 3.1 Pro uses temperature = 1.0 and thinking_level = "high".
- A. Experiment Details: The benchmark provides prompt templates for reproducible evaluation across its three tasks.The templates cover RVAC, PVRC, and AID.
- A.2. Evaluation Prompt: RVAC classifies videos by AIGC-specific artifact patterns, excluding traditional CG, animation, game footage, and rendering-related artifacts as evidence of AIGC.Relevant artifacts include temporal inconsistency, structural distortions, unnatural motion, semantic incoherence, and abnormal appearance or texture.
- A.2. Evaluation Prompt: PVRC selects the more realistic of two AIGC videos using the severity, frequency, and perceptibility of artifacts rather than style or aesthetic preference.The video with fewer, less noticeable, or less severe artifacts is considered more realistic.
- A.2. Evaluation Prompt: AID asks models to select all clearly observable AIGC-specific artifacts from lettered options and return the selected letters.Multiple options may be selected, with letters separated by commas.
- A.3. Answer Extraction Prompt: Gemini 3.1 Pro extracts only final answers from potentially lengthy model responses, ignoring reasoning, explanations, self-corrections, and think blocks.For RVAC, valid outputs are exactly "yes" or "no"; unclear conclusions become "Invalid".
- A.3. Answer Extraction Prompt: For PVRC and AID, extraction enforces task-specific final-answer formats and returns "Invalid" when no clear final answer is present.PVRC permits only "<Video A>" or "<Video B>", while AID outputs valid selected option letters, comma-separated when multiple.
B. Benchmark Details · B.1. Representative Examples from Artifact-Bench
Artifact-Bench uses representative examples to convey the characteristics of its tasks, presenting two examples for each task in Figures 6, 7, and 8.
- B.1. Representative Examples from Artifact-Bench: Artifact-Bench presents representative examples to illustrate the characteristics of its tasks.The examples are intended to comprehensively convey task characteristics.
- B.1. Representative Examples from Artifact-Bench: Each task is accompanied by two representative examples.This provides paired examples for every task discussed in the benchmark.
- B.1. Representative Examples from Artifact-Bench: The examples are organized across Figures 6, 7, and 8.These figures collectively present the examples for the benchmark’s tasks.
- B.1. Representative Examples from Artifact-Bench: The examples focus on conveying task characteristics rather than reporting experimental results.The passage frames the figures as illustrative material for understanding the tasks.
- B.1. Representative Examples from Artifact-Bench: The figure-based presentation covers the tasks defined in Artifact-Bench.Two examples are presented for each task in the benchmark.
- B.1. Representative Examples from Artifact-Bench: Together, the examples provide a visual overview of how Artifact-Bench tasks are represented.This follows the stated goal of comprehensively conveying the characteristics of the tasks.
B.2. Task Distribution
Table 3 reports the distribution of QA pairs across difficulty levels for each Artifact-Bench task.
- Task Distribution: Table 3 shows the number of QA pairs at each difficulty level for every Artifact-Bench task.
C. Limitations
Artifact-Bench is limited by the number of human experts and the current dataset scale. Future work will expand video diversity, artifact coverage, and expert annotations to support more comprehensive and reliable evaluation.
- C. Limitations: The benchmark’s human-expert coverage and dataset scale remain limited by resource constraints.The authors identify both factors as areas for expansion.
- C. Limitations: Future work will enlarge the benchmark with more diverse video sources, artifact types, and expert annotations.These expansions aim to enable more comprehensive and reliable evaluation of artifact-aware video understanding.
D. Compute Resources
The experiments used a distributed setup of four identical machines, each with 8 NVIDIA H800 GPUs and 1000 GiB of system memory. The main results require no additional compute beyond the reported experiments, excluding preliminary runs.
- D. Compute Resources: Experiments ran on four identical machines, each equipped with 8 NVIDIA H800 GPUs and 1000 GiB of system memory.No additional compute beyond the reported experiments, excluding preliminary runs, is required to reproduce the main results.