Source-linked AI summary
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
TL;DR
Existing video-generation benchmarks rarely assess whether instructed outcomes are achieved while preserving task-relevant semantic grounding in reference images. This paper introduces SemComp-Data and the VLM-based SemComp-Bench, finding that outcome achievement remains difficult even as generation reliability can be high.
Problem
Existing benchmarks rarely evaluate instructed outcome achievement jointly with high-level semantic grounding in reference images.
Method
The paper constructs SemComp-Data from full-context videos and introduces SemComp-Bench, which uses VLM-based binary judgments for Outcome Achievement and Generation Reliability.
Results
The best OA Score is 37.8%, while Seedance 2.0 achieves the highest GR Score of 91.8%, showing that generation reliability does not necessarily correspond to outcome achievement.
Takeaways & Limitations
Achieving intended outcomes while maintaining task-relevant semantic grounding remains challenging for representative video-generation models.
Takeaways & Limitations
The benchmark’s validity safeguards do not assess the completeness or procedural correctness of intermediate task steps.
Abstract
from arXiv · showhide
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding. Semantic grounding characterizes the correspondence between the reference image and the generated outcome in terms of high-level semantics relevant to the task. Evaluation focuses on the generated outcome and requires neither the presentation of a complete sequence of intermediate task steps nor conventional appearance consistency with the reference image. To support systematic evaluation, we construct SemComp-Data, an evaluation dataset covering six domains. Each instance comprises a reference image, a detailed instruction, a brief instruction, and an outcome-centric video clip. A scalable four-stage curation pipeline converts raw videos into standardized SemComp-Data instances. We further introduce SemComp-Bench, an evaluation protocol that uses a vision-language model (VLM) to answer structured binary questions. SemComp-Bench reports the OA Score and the GR Score for Outcome Achievement and Generation Reliability, respectively. Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.
1 Introduction
The paper formulates Semantic Task Completion Video Generation, requiring instructed outcomes while preserving task-relevant semantic relationships with reference images. It introduces SemComp-Data and SemComp-Bench to curate and evaluate this task through evidence-grounded VLM judgments.
- Dataset: Existing benchmarks rarely assess outcome achievement jointly with high-level semantic grounding, motivating the construction of SemComp-Data from full-context real-world videos.Each instance is an image-text-video triplet pairing the same reference and outcome with brief and detailed instructions.
- Dataset: SemComp-Data uses a scalable curation pipeline comprising Candidate Filtering, State Mining, Video Extension, and Instruction Structuring.The pipeline screens raw videos, localizes and verifies reference–outcome frames, and structures paired instructions.
- Task Formulation: Semantic Task Completion Video Generation requires generated videos to achieve instructed outcomes while preserving task-relevant semantic relationships with reference images.Grounding preserves relevant relationships and reference attributes while allowing unrelated attributes to change.
- Evaluation: SemComp-Bench measures Outcome Achievement and Generation Reliability using interpretable, evidence-grounded binary judgments from a vision-language model.These dimensions are reported as the OA Score and GR Score; OA evaluates outcome realization and semantic grounding, while GR provides a complementary measure.
2 Related Work
Existing video-generation benchmarks evaluate visual fidelity, temporal coherence, prompt adherence, identity, motion, physical plausibility, causal consistency, and commonsense constraints. Meanwhile, current open- and closed-source models increasingly support high-quality, instruction-following generation with diverse conditioning and control capabilities.
- Video Generation Benchmarks: Video-generation benchmarks assess visual fidelity, temporal coherence, prompt adherence, subject identity, intended motion, physical plausibility, causal consistency, and commonsense constraints.Together, these benchmarks cover generation quality, identity, motion, and physical plausibility.
- Video Generation Models: Current open-source systems support text- and image-conditioned generation, while LTX emphasizes efficiency.Examples include Wan2.2, CogVideoX, HunyuanVideo, and Pyramid Flow.
- Video Generation Models: Closed-source systems support multimodal conditioning, multi-shot narratives, camera control, and editing.The passage names Sora, Veo, MovieGen, Runway, Kling, Seedance, Hailuo, and Pika.
3 SemComp-Data Construction
SemComp-Data is constructed from Koala-36M full-context videos as structured reference–instruction–outcome triplets. A four-stage curation pipeline filters candidates, mines reliable states, extracts outcome-centric clips, and generates aligned instructions with task-relevant grounding constraints.
- Triplet formulation: Each instance pairs a reference frame, brief and detailed instructions, and an outcome-centric clip from the same task instance.The triplet is defined as x_i = (r_i, I_i, o_i), with the instruction pair describing the same intended outcome at different specificity levels.
- Candidate Filtering: Candidate Filtering removes narration-dependent videos using title keywords and creates video abstracts by uniformly sampled frames arranged as mosaics.Examples of excluded lexical patterns include “talk show,” “interview,” and “news.”
- State Mining: State Mining uses category-specific state definitions, VLM-based timestamp localization, and quality checking to extract representative reference and outcome frames.The method focuses on visual evidence for the two states without modeling intermediate processes or resolving ambiguous segment boundaries.
- Video Extension: Video Extension anchors shot-detected, same-scene segments to the verified outcome timestamp and merges consistent neighboring segments into an outcome-centric clip.The procedure follows the shot detection and same-scene merging approach of Panda-70M.
- Instruction Structuring: Instruction Structuring produces brief and detailed instructions by normalizing relations, selecting preserved attributes, and composing outcome descriptions.Brief instructions are constrained to at most 30 words, while detailed instructions combine alignment constraints with completed-outcome characteristics such as background, lighting, colors, shapes, and component details.
4 SemComp-Bench Evaluation
SemComp-Bench evaluates generated videos with structured binary questions across Outcome Achievement and Generation Reliability. It combines conjunctive outcome criteria with averaged reliability criteria to provide interpretable assessment and failure diagnosis.
- Evaluation dimensions: SemComp-Bench organizes binary criteria into Outcome Achievement for outcome realization and reference grounding, and Generation Reliability for visual, physical, and temporal quality.Each criterion receives a binary pass score, with the full question set provided in the supplementary material.
- Outcome Achievement: Outcome Achievement evaluates realization, semantic grounding, grounded entity consistency, and global visual continuity on 27 uniformly sampled frames.The first two criteria use the reference image and instruction, while the latter two use only the sampled sequence.
- Outcome Achievement: The OA Score is conjunctive: a sample passes only when it satisfies all four Outcome Achievement criteria.Failure-oriented entity inconsistency and global discontinuity questions map “No” responses to passing scores, and intermediate-step completeness is not required.
- Generation Reliability: Generation Reliability independently evaluates physical plausibility, visual clarity, artifact-free rendering, within-scene spatiotemporal coherence, and text and interface integrity.These criteria target specific sources of unreliability, including physical violations, unclear content, synthetic artifacts, and rendering instability.
- Generation Reliability: The GR Score averages the five binary reliability criteria for each sample and then averages those sample-level scores across the dataset.All five questions are failure-oriented, so “No” responses are mapped to passing scores.
5 Experiments
Experiments evaluate SemComp-Bench on a curated, domain-balanced subset using standardized video generation and VLM scoring. Results show strong differences across models, reference conditioning, and instruction specificity, while outcome achievement and reliable temporal coherence remain challenging.
- Dataset and Evaluation Setup: Approximately 20K Koala-36M videos yield 1,273 structured evaluation instances, while SemComp-Core contains 60 domain-balanced instances with 10 sampled per domain.Each SemComp-Core instance includes detailed and brief instruction variants sharing the same reference image and outcome target.
- Dataset and Evaluation Setup: Each instance produces one 720p video, 27 uniformly sampled frames, and scores from three independent VLM calls when seed control is available.Open-source and closed-source models use their respective default inference settings, with a fixed seed whenever explicit control exists.
- Outcome Achievement: 37.8%: HunyuanVideo-1.5-720P-I2V achieves the highest OA Score, followed by Wan2.2-I2V-A14B at 28.3%.HunyuanVideo-1.5-720P-I2V’s performance reflects relatively balanced pass rates across the four OA criteria.
- Generation Reliability: 91.8%: Seedance 2.0 achieves the highest GR Score, while Wan2.2-I2V-A14B leads evaluated open-source models at 89.0%.Within-scene spatiotemporal coherence is the primary bottleneck, with pass rates ranging from 0.328 to 0.739.
- Conditioning and Instructions: I2V variants consistently outperform T2V counterparts across all three model families, mainly through stronger semantic grounding, grounded entity consistency, and global visual continuity.Outcome-realization pass rates remain broadly comparable and favor T2V in two of the three model families.
- Conditioning and Instructions: Within T2V, detailed instructions consistently achieve higher OA Scores than brief instructions, improving outcome realization and semantic grounding but making coherent generation more difficult.Brief instructions often improve grounded entity consistency and global visual continuity, revealing a trade-off between specificity and generation difficulty.
6 Conclusion
The paper introduces Semantic Task Completion Video Generation, requiring instructed outcomes while preserving semantic grounding in reference contexts. It also develops SemComp-Data through a scalable four-stage curation pipeline using paired references, instructions, and outcome-centric clips.
- Semantic Task Completion Video Generation requires generated videos to realize an instructed outcome while maintaining semantic grounding in the reference context.
- SemComp-Data is developed through a scalable four-stage curation pipeline.
- Each SemComp-Data instance comprises a reference image, paired brief and detailed instructions, and an outcome-centric video clip.The paired image and clip are extracted from the same full-context video.
Supplementary Material … D Alignment and Attributes
The supplementary material documents SemComp-Data and SemComp-Bench, covering dataset organization, curation, evaluation, and additional analyses. It defines the filtering, taxonomy, state-mining, alignment, and preservation-attribute procedures used to construct and evaluate instances.
- Supplementary Material: The supplementary material covers title filtering, domain and category taxonomy, reference and outcome states, alignment types, preservation attributes, evaluation prompts, dataset statistics, and generated-video comparisons.
- Supplementary Material: The Code folder provides the complete curation pipeline and stage-specific prompts, while the Data folder contains metadata, source-video URLs, and reconstruction timestamps.
- B Domain and Category Taxonomy: SemComp-Data contains six broad domains and 21 fine-grained categories, distinguishing general application scenarios from task-specific reference-to-outcome patterns.
- C Reference and Outcome State Definitions: Category-specific state definitions localize representative reference and outcome frames for State Mining.Reference states capture initial visual evidence needed for task grounding, whereas outcome states capture visually verifiable completed results.
- D Alignment and Attributes: Instruction Structuring assigns one of four alignment types and selects instance-specific preservation attributes from a 17-option vocabulary.Alignment types identify the principal form of semantic grounding and do not require every associated attribute to remain unchanged.
E SemComp-Bench Evaluation Prompts
SemComp-Bench evaluates each generated video with structured binary VLM prompts over 27 uniformly sampled frames, using criterion-specific pass rules and visual evidence. Its prompts separate Outcome Achievement criteria from independently assessed Generation Reliability criteria, covering task fulfillment, semantic grounding, continuity, and rendering quality.
- Evaluation procedure: 27 frames are uniformly sampled in temporal order, and the evaluator answers each binary question with yes or no plus visual evidence.The common system prompt restricts each decision to yes or no; evidence uses provided frame indices or the 1-based sampled order when indices are absent.
- Outcome Achievement: OA uses positive questions that pass on yes and failure-oriented questions that pass on no.The two positive OA questions pass on yes, while the remaining OA questions pass on no.
- Outcome Achievement: The four OA criteria separate completed-state visibility, reference–instruction grounding, grounded-entity consistency, and global visual continuity.Outcome Realization checks the instructed completed state at a coarse semantic level; Semantic Grounding checks task-relevant correspondence; the continuity criteria assess entity drift and abrupt whole-frame changes.
- Generation Reliability: GR comprises five failure-oriented questions covering physical plausibility, visual clarity, artifact-free rendering, within-scene spatiotemporal coherence, and text and interface integrity.A no response denotes a pass, and the evaluation examines visible motion, contact, interaction, material behavior, recognizability, and scene consistency across the ordered sequence.
- Generation Reliability: Generation Reliability is assessed independently of task completion and semantic grounding.Its physical-plausibility question asks whether any visibly implausible content, motion, contact, interaction, or event occurs at any time.
F Dataset Statistics and Instances
The curation pipeline produces 1,273 SemComp-Data instances, including diverse task categories represented through reference images, brief instructions, and outcome-centric clips. Outcome-centric clips are extracted by selecting the event containing the verified outcome timestamp after shot partitioning and temporal merging.
- 1,273 SemComp-Data instances are produced by the curation pipeline.
- Figure S6 adds 22 instances spanning diverse task categories, each pairing a reference image and brief instruction with three temporally ordered outcome-clip frames.
- Outcome-centric clip extraction selects the event containing the verified outcome timestamp after partitioning videos into shots and merging adjacent visually consistent shots.The notation of intended semantic transitions does not require reproducing intermediate procedures.
G Qualitative Results
Additional SemComp-Core comparisons across four domains use matched reference conditions and instructions to compare seven I2V models through temporally ordered outputs. These sequences reveal complementary differences in outcome achievement and generation reliability that terminal frames alone may miss.
- Additional comparisons: Figure S7 compares seven I2V models on four SemComp-Core instances spanning Sculpting, Artwork Creation, Restoration, and Food Transformation.All models receive the same reference condition and detailed instruction in each comparison.
- Additional comparisons: Temporally ordered frame sequences expose complementary failure modes that are not fully captured by evaluating only the terminal frame.The comparisons illustrate differences in both outcome achievement and generation reliability.
H Result Variability
The paper quantifies evaluation variability using three independent VLM calls per generated video, reporting sample standard deviations for aggregate scores and criterion-level results across three supplementary tables. All deviations are reported in percentage points and follow the main-paper score definitions.
- Evaluation protocol: Each generated video is evaluated by three independent VLM calls made at different times, with sample standard deviations computed from the corresponding run-level scores.The variability analysis uses three evaluation runs.
- Score definitions: OA Score standard deviations use the conjunctive definition requiring a sample to pass all four OA criteria.This matches the aggregate-score definition used for the main-paper results.
- Score definitions: GR Score standard deviations are computed from the run-level mean of five GR criterion pass rates.All entries are measured in percentage points.
- Supplementary variability tables: Table S6 reports three-run sample standard deviations for SemComp-Core Outcome Achievement under detailed instructions.Values are reported in percentage points; †HY denotes HunyuanVideo.
- Supplementary variability tables: Tables S7 and S8 report three-run sample standard deviations for SemComp-Core Generation Reliability and the Outcome Achievement conditioning study, respectively.Both tables report values in percentage points and identify †HY as HunyuanVideo.