Source-linked AI summary

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Junhao Cheng, Yuying Ge, Teng Wang, Yixiao Ge, Jing Liao, Ying Shan

arXiv:2505.21374v1cs.CV

TL;DR

Existing video benchmarks largely assess perception and isolated visual cues, leaving complex multi-clue video reasoning insufficiently evaluated. Video-Holmes addresses this gap with seven active-seeking tasks built from 270 annotated suspense films and 1,837 questions, finding substantial integration difficulties in current MLLMs, with Gemini-2.5-Pro reaching 45% accuracy.

  • Problem

    Existing video benchmarks primarily evaluate visual perception and grounding through explicit prompts or isolated visual cues rather than complex reasoning requiring multiple clues.

  • Method

    Video-Holmes uses 270 manually annotated suspense short films to construct seven tasks requiring models to actively locate and connect clues across video segments, while analyzing their reasoning processes.

  • Results

    Gemini-2.5-Pro achieves 45% accuracy and most models score below 40%, while models generally perceive visual content but struggle to integrate information and miss critical clues.

  • Takeaways & Limitations

    Video-Holmes serves as a benchmark for complex multimodal reasoning and highlights ongoing challenges in making MLLMs reason more like humans.

Abstract

from arXiv · show

Recent advances in CoT reasoning and RL post-training have been reported to enhance video reasoning capabilities of MLLMs. This progress naturally raises a question: can these models perform complex video reasoning in a manner comparable to human experts? However, existing video benchmarks primarily evaluate visual perception and grounding abilities, with questions that can be answered based on explicit prompts or isolated visual cues. Such benchmarks do not fully capture the intricacies of real-world reasoning, where humans must actively search for, integrate, and analyze multiple clues before reaching a conclusion. To address this issue, we present Video-Holmes, a benchmark inspired by the reasoning process of Sherlock Holmes, designed to evaluate the complex video reasoning capabilities of MLLMs. Video-Holmes consists of 1,837 questions derived from 270 manually annotated suspense short films, which spans seven carefully designed tasks. Each task is constructed by first identifying key events and causal relationships within films, and then designing questions that require models to actively locate and connect multiple relevant visual clues scattered across different video segments. Our comprehensive evaluation of state-of-the-art MLLMs reveals that, while these models generally excel at visual perception, they encounter substantial difficulties with integrating information and often miss critical clues. For example, the best-performing model, Gemini-2.5-Pro, achieves an accuracy of only 45%, with most models scoring below 40%. We aim that Video-Holmes can serve as a "Holmes-test" for multimodal reasoning, motivating models to reason more like humans and emphasizing the ongoing challenges in this field. The benchmark is released in https://github.com/TencentARC/Video-Holmes.

1 Introduction

Video-Holmes addresses a gap in video benchmarks that emphasize perception and isolated cues rather than multi-clue reasoning. It introduces a benchmark requiring active clue integration and finds that current MLLMs often miss critical information.

  • Existing video benchmarks mainly test visual perception and grounding through explicit prompts or isolated visual cues.
  • Humans often reason by actively searching for, integrating, and analyzing multiple clues before reaching a conclusion.
  • The benchmark also analyzes model reasoning processes to examine factors associated with correct and incorrect answers.
  • Gemini-2.5-Pro reaches only 45% accuracy, while most evaluated models score below 40%.
  • Video-Holmes contains 270 manually annotated suspense short films and 1,837 challenging questions across seven tasks requiring active multi-clue reasoning.

2 Related Works

Prior video benchmarks range from basic question answering and action recognition to reasoning over short and long clips. The reviewed literature characterizes these tasks as generally straightforward rather than requiring complex reasoning.

  • Video MLLM methods commonly treat videos as image sequences so language models can perform video understanding and reasoning.
  • Early benchmarks such as MSRVTT-QA, ActivityNet-QA, and NExT-QA evaluate action recognition and video question answering.
  • MMBench, TempCompass, and MVBench assess reasoning over short video clips, while LongVideoBench and Video-MME evaluate longer sequences.
  • The related benchmarks are described as generally straightforward and not requiring complex reasoning.

3 Video-Holmes

Video-Holmes is constructed from annotated suspense short films through task design, question generation, and reasoning-process analysis. Its tasks require models to find and connect clues distributed across video segments.

  • The benchmark pipeline comprises video collection and annotation, task definition, and question-answer-explanation generation.
  • Video Collection and Annotation: Suspense short films provide compact narratives containing hints, plot twists, and supernatural elements for complex reasoning evaluation.
  • Video Collection and Annotation: Annotations include segmented plot descriptions, key character relationships, and reasoning shots with timestamps, visual clues, and inferred conclusions.
  • Task Definition: Video-Holmes defines seven reasoning tasks that require actively locating and connecting relevant clues across different video segments.
  • Question-Answer Generation: DeepSeek-R1 generates questions from formatted manual annotations, while generated explanations support analysis of model reasoning processes.
  • After verification and annotation, the dataset contains 270 videos and 1,837 question-answer pairs.

4 Experiments

Experiments evaluate mainstream MLLMs on Video-Holmes and analyze whether their answers reflect valid reasoning. Results show substantial difficulty with the benchmark’s multi-clue reasoning demands, despite gains from thinking strategies.

  • Main Results: 45% overall accuracy is achieved by Gemini-2.5-Pro, while most models remain below 40% on Video-Holmes.Qwen2.5-VL (7B) reaches 27.8%, substantially below its performance on other video reasoning benchmarks.
  • Main Results: 12.5% performance gain is demonstrated by Gemini-2.0-Flash-Thinking over Gemini-2.0-Flash.Thinking strategies improve models over their vanilla versions, indicating that the benchmark distinguishes reasoning abilities.
  • Main Results: Below 40% accuracy is achieved by most models on each of the seven reasoning tasks.Performance remains relatively even across tasks, suggesting that all task types pose substantial challenges.
  • Reasoning Process Analysis: Around 60% of errors stem from reasoning errors involving logical comprehension of multiple visual clues, compared with approximately 35% from omitted critical visual information.The analysis categorizes errors into visual perception, visual omission, reasoning, and answer-selection types.
  • Reasoning Process Analysis: Over 80% of correct responses across most models are based on valid reasoning, measured by Think Right Answer Right (TRAR).The benchmark separately identifies cases where reasoning is wrong but the answer is right, or reasoning is right but the answer is wrong.
  • Analytical Study: Increasing input frames generally improves performance without substantial gains, while audio especially helps social reasoning tasks.CoT prompting benefits stronger closed-source models but can reduce accuracy for weaker open-source models; human annotations yield around 90% accuracy in text-only experiments.

5 Conclusion and Discussion

Video-Holmes evaluates complex video reasoning through 1,837 questions from 270 manually annotated suspense short films across seven tasks. State-of-the-art MLLMs often perceive visual content well but struggle to integrate information and miss critical clues.

  • Video-Holmes comprises 1,837 questions derived from 270 manually annotated suspense short films across seven reasoning tasks.The tasks require models to locate and connect visual clues scattered across different video segments.
  • The benchmark analyzes model reasoning processes and the factors associated with correct and incorrect answers.
  • MLLMs generally excel at visual perception but struggle to integrate information and frequently miss critical clues.
  • Video-Holmes is intended as a “Holmes-test” for multimodal reasoning that emphasizes ongoing challenges in the field.

A Model Implementation Details

The benchmark-generation procedure supplies structured film information and produces seven reasoning task types, while generated questions must support uniquely answerable six-option multiple-choice evaluation.

  • Inputs: The generation prompt uses segmented plot descriptions, character relationships, reasoning shots, and supernatural elements as film inputs.These inputs organize chronology, relationships, visual hints, and non-realistic phenomena for question construction.
  • Seven reasoning tasks: Social Reasoning infers implicit character relationships through clothing matches, interaction topology, and identity links across time.
  • Seven reasoning tasks: Intention and Motive Chaining infers psychological states or non-explicit intentions from actions, expressions, and environmental clues.
  • Seven reasoning tasks: Temporal Causal Inference connects non-continuous events by inferring causal mechanisms from camera language and multimodal clues.
  • Seven reasoning tasks: Timeline Analysis restores chronology from five key events, while Multimodal Hint Reasoning interprets camera, object, sound, and picture cues.
  • Seven reasoning tasks: Physical Anomaly Reasoning examines supernatural rules and implied meanings, whereas Core Theme Inference derives themes from plot, dialogue, or symbols.Physical Anomaly Reasoning creates two questions per supernatural element when such elements are provided.
  • Question format: Each question must provide six options with one objectively unique correct answer and distractors that overlap closely without expressing the same meaning.The prompt also requires the question type, question, answer, and explanation.

Model Evaluation Prompt (Thinking)

The thinking evaluation prompt asks the model to solve a video multiple-choice question by exposing reasoning before returning its final answer.

  • The model is instructed to reason about a video single-choice question and provide the reasoning inside <think> tags.
  • The prompt presents the question and options before requesting the model’s answer.
  • The final answer must appear inside <answer> tags.

Model Evaluation Prompt (Without Thinking)

The no-thinking evaluation prompt asks the model to answer a video multiple-choice question while returning only a tagged final choice.

  • The model is instructed to answer a single-choice video question without an explicit reasoning step.
  • The prompt provides the question and options before requesting the answer.
  • The final answer choice must be enclosed in <answer> tags.

Reasoning Process Analysis Prompt (Thinking Wrong)

The analysis prompt classifies why multimodal models answer video-reasoning questions incorrectly by comparing their thoughts with the film plot and ground-truth explanation. A parallel prompt evaluates whether correct answers are supported by aligned reasoning.

  • Reasoning Process Analysis Prompt (Thinking Wrong): The prompt compares a model’s thought process and incorrect answer with the film plot and correct explanation.It asks for the most critical cause of the error.
  • Reasoning Process Analysis Prompt (Thinking Wrong): Visual Perception Error means extracting incorrect visual information that leads to the wrong answer.The example contrasts perceiving a knife with perceiving a gun.
  • Reasoning Process Analysis Prompt (Thinking Wrong): Visual Omission Error means failing to extract key visual information, including important objects and events.The omitted information is identified as a cause of incorrect answers.
  • Reasoning Process Analysis Prompt (Thinking Wrong): Reasoning Error refers to mistakes in the reasoning process, such as misinterpreting implications.The prompt lists it alongside visual perception and omission errors.
  • Reasoning Process Analysis Prompt (Thinking Wrong): The prompt requires the model to provide a brief judgment reason inside <Reason></Reason>.This specifies the output format for the error analysis.
  • Reasoning Process Analysis Prompt (Thinking Right): The parallel prompt compares a model’s thought process and correct answer with the ground-truth explanation and film plot.It determines which reasoning-and-answer pattern the response belongs to.
  • Reasoning Process Analysis Prompt (Thinking Right): Think Wrong Answer Right describes reasoning that deviates from the ground truth while still reaching the correct answer.This category separates answer correctness from reasoning alignment.

C Key Statistics of Video-Holmes

Video-Holmes uses diverse suspense short films and balances its designed tasks while allowing task presence to depend on video content. Figures 5–12 illustrate annotation, question generation, model answers, and reasoning analysis.

  • C Key Statistics of Video-Holmes: Nine subkeywords, including Anim, Comic, Detective, Future, Horror, Social, Supernatural, and Thriller, support diverse suspense-film retrieval.The statistics discussion says these keywords are used when searching for suspense short films.
  • C Key Statistics of Video-Holmes: The nine designed tasks are relatively balanced, with more MHR tasks because one video may contain multiple human-annotated reasoning shots.PAR tasks are absent from videos without supernatural phenomena.
  • C Key Statistics of Video-Holmes: Figures 5–12 show human annotations, generated questions and explanations, model answers, and reasoning-process analyses.Figure 5 presents annotations, while Figures 6–12 present the other materials.

E Broader Impact

The broader-impact materials frame Video-Holmes as a rigorous evaluation resource while warning that its suspense-film content may distress some audiences. Examples show how questions probe relationships, intentions, causes, chronology, implications, supernatural rules, and themes.

  • E Broader Impact: Video-Holmes may affect complex-video reasoning research by providing a rigorous and comprehensive evaluation benchmark.This impact claim is stated as a potential contribution of the benchmark’s development and release.
  • E Broader Impact: Suspense short films may contain horror or thriller elements that could distress or be inappropriate for certain audiences.Responsible use includes content warnings and attention to the intended audience.
  • E Broader Impact: The examples span social relationships, character intentions, direct causes, event chronology, action implications, supernatural-power rules, and core themes.These examples illustrate the range of question types used in Video-Holmes.
  • E Broader Impact: The example questions require answers about relationships and intentions using character interactions and preceding events.The clown–David relationship and the clown’s retaliatory intention are represented as separate questions.
  • E Broader Impact: Other examples ask models to connect a character’s death to mockery, order multiple events chronologically, and infer an action’s implication.These questions respectively target causation, temporal ordering, and interpretation.
  • E Broader Impact: The benchmark also includes questions about the rules governing supernatural powers and the video’s core theme.These examples extend beyond isolated visual description toward broader interpretation.
Loading 2505.21374v1…