Source-linked AI summary

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

Yunhong Lu, Yanhong Zeng, Haobo Li, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Jiapeng Zhu, Hengyuan Cao, Zhipeng Zhang, Xing Zhu, Yujun Shen, Min Zhang

arXiv:2512.04678v2cs.CV

TL;DR

Streaming video generation requires low latency without sacrificing visual and dynamic fidelity, but existing attention-sink methods can overemphasize initial frames and weaken motion. Reward Forcing combines EMA-Sink’s continuously updated historical context with reward-weighted distillation, achieving state-of-the-art performance at 23.1 FPS on one H100 GPU.

  • Problem

    Streaming generation must maintain low latency and high visual-dynamic fidelity, while static attention sinks can produce initial-frame bias and diminished motion.

  • Method

    Reward Forcing combines EMA-Sink, which updates fixed-size sink tokens with evicted history, and Re-DMD, which weights distillation toward high-reward dynamic samples.

  • Results

    Reward Forcing achieves state-of-the-art video quality at 23.1 FPS on a single H100 GPU.

  • Takeaways & Limitations

    The framework balances visual fidelity, dynamic motion, and long-term coherence for real-time streaming video generation.

  • Takeaways & Limitations

    Reward optimization may emphasize temporal coherence over other VBench dimensions, so reward gains may not translate proportionally into VBench score gains.

Abstract

from arXiv · show

Efficient streaming video generation is critical for simulating interactive and dynamic worlds. Existing methods distill few-step video diffusion models with sliding window attention, using initial frames as sink tokens to maintain attention performance and reduce error accumulation. However, video frames become overly dependent on these static tokens, resulting in copied initial frames and diminished motion dynamics. To address this, we introduce Reward Forcing, a novel framework with two key designs. First, we propose EMA-Sink, which maintains fixed-size tokens initialized from initial frames and continuously updated by fusing evicted tokens via exponential moving average as they exit the sliding window. Without additional computation cost, EMA-Sink tokens capture both long-term context and recent dynamics, preventing initial frame copying while maintaining long-horizon consistency. Second, to better distill motion dynamics from teacher models, we propose a novel Rewarded Distribution Matching Distillation (Re-DMD). Vanilla distribution matching treats every training sample equally, limiting the model's ability to prioritize dynamic content. Instead, Re-DMD biases the model's output distribution toward high-reward regions by prioritizing samples with greater dynamics rated by a vision-language model. Re-DMD significantly enhances motion quality while preserving data fidelity. We include both quantitative and qualitative experiments to show that Reward Forcing achieves state-of-the-art performance on standard benchmarks while enabling high-quality streaming video generation at 23.1 FPS on a single H100 GPU.

1. Introduction

Reward Forcing addresses the latency and fidelity challenges of streaming video generation with EMA-Sink for long-term context and Re-DMD for motion-aware distillation.

  • Streaming video generation must combine low latency with high visual and dynamic fidelity for interactive, extended-horizon applications.
  • Autoregressive students enable real-time inference through sliding-window attention, but recursively conditioned outputs can accumulate errors over time.
  • Attention sinks reduce long-horizon drift by retaining initial tokens, but their static content can cause frame copying, diminished motion, and stiff transitions.
  • EMA-Sink: EMA-Sink updates fixed-size initial sink tokens with evicted history through exponential moving averages, preserving global context and recent dynamics without extra cost.
  • Rewarded distribution matching distillation: Re-DMD weights distribution-matching gradients using vision-language-model rewards, prioritizing samples with stronger motion while preserving data fidelity.
  • Reward Forcing is reported to achieve state-of-the-art video quality at 23.1 FPS on a single H100 GPU.

4. Experiments

Experiments evaluate Reward Forcing on short and long videos, component ablations, attention efficiency, and interactive generation. The results show strong video quality, dynamism, temporal consistency, and real-time speed.

  • Short video generation: 84.13 overall VBench score on 5-second clips surpasses all existing baselines.The method also uses the smallest attention window and achieves the fastest inference speed among compared methods.
  • Short video generation: 23.1 FPS provides real-time generation, with 47.14× and 1.36× speedups over SkyReels-V2 and Self Forcing, respectively.
  • Long video generation: 81.41 total score on 60-second videos surpasses LongLive’s 79.53, while the dynamic metric reaches 66.95 with an 88.38% boost in dynamic amplitude.
  • Ablation studies: Removing Re-DMD drops the dynamic score from 64.06 to 43.75, while removing EMA-Sink reduces motion smoothness from 98.91 to 98.64 and dynamic score from 43.75 to 35.15.
  • Ablation studies: Reducing α to 0.9 raises motion smoothness from 98.96 to 99.09 but increases drift from 2.52 to 3.23.
  • Analysis: Inference FPS is inversely proportional to attention-window size, identifying window size as a key efficiency factor.
  • Interactive video generation: EMA-Sink supports prompt changes during generation by enabling seamless transitions while maintaining high temporal consistency.

5. Conclusion

The conclusion presents Reward Forcing as a solution to motion stagnation in efficient streaming video generation. It combines EMA-Sink and Re-DMD to balance visual fidelity, dynamic motion, and real-time operation.

  • Reward Forcing addresses motion stagnation using EMA-Sink for dynamically maintained context and Re-DMD for high-reward motion samples.
  • Reward Forcing achieves state-of-the-art performance on standard benchmarks while enabling high-quality real-time streaming video generation.

S1: More Video Results

Supplementary videos evaluate long-horizon and interactive generation. They report preserved visual fidelity, stronger motion dynamics, and prompt-controlled event insertion during streaming.

  • More video results: The supplementary videos are compressed to approximately 40% of their original file size without significant quality degradation.
  • More video results: Reward Forcing preserves high visual fidelity and exhibits superior motion dynamics over ultra-long horizons in Scene Navigation and Object Motion prompts.
  • Interactive videos: Switching prompts and resetting the cross-attention cache allows new events to be introduced into an ongoing video.

S2: User Studies

A 20-participant user study compares Reward Forcing with three baseline methods using randomized video labels and 4-point ratings. Reward Forcing receives the highest scores across all reported criteria.

  • Experimental setup: 1,600 evaluations compare CausVid, Self-Forcing, LongLive, and Reward Forcing across 20 participants and 20 video groups.
  • Evaluation protocol: Participants rate long-range temporal consistency, dynamic complexity, and overall preference on a 1–4 Likert scale.
  • Results and analysis: Reward Forcing achieves the highest scores: 3.60 for Temporal Consistency, 3.72 for Dynamic Complexity, and 3.75 for Overall Preference.

S3: More Quantitative Results and Details

The paper evaluates video quality, semantic alignment, motion, and long-horizon consistency using VBench, user ratings, and vision-language-model scoring. These evaluations report strong short-video performance and analyze quality drift over extended sequences.

  • Long-horizon quality: Long-video quality drift is computed across 30 two-second clips per one-minute video using imaging-quality-score variability.Lower drift scores correlate with more consistent visual fidelity throughout long sequences.
  • Vision-language evaluation: Qwen3-VL evaluates text alignment, dynamics, and visual quality on a 1–5 scale.The evaluation template defines separate criteria for semantic matching, motion, and technical execution.
  • VBench evaluation: VBench Quality Score averages subject consistency, background consistency, temporal flickering, motion smoothness, aesthetics, imaging quality, and dynamic degree.The Semantic Score covers object, action, color, spatial, scene, style, and consistency dimensions.

S4: More Implementation details

The implementation uses flow matching with a shifted time schedule and a four-step sampling procedure for few-step video diffusion.

  • Noise schedule: The time-step shift uses factor k = 5 in the flow-matching implementation.The shifted timestep is applied in the forward sampling process.
  • Noise schedule: The forward process interpolates data x and Gaussian noise ϵ using the shifted timestep t′.ϵ is drawn from N(0, I), with t ranging from 0 to 1000.
  • Model parameterization: The data prediction model combines the noise term with the velocity prediction vθ through preconditioning coefficients.The model is written as Gθ(x, t, c) = cskip · ϵ − cout · vθ(cin · xt, cnoise(t′), c).
  • Sampling: Few-step diffusion sampling uses a uniform four-step schedule with time steps 1000, 750, 500, and 250.The preconditioning coefficients cskip, cin, and cout are all 1, while cnoise(t) = t.

S5: Further Related Works

The related work surveys video diffusion architectures and preference-alignment methods, including transformer-based generation, direct preference optimization, and reinforcement-learning approaches.

  • Video diffusion models: Video diffusion models evolved from UNet backbones toward Diffusion Transformers for improved spatio-temporal modeling and scalability.Representative systems include Sora, Hunyuan-Video, and Open-Sora.
  • Reinforcement learning for video models: Preference-alignment research applies reinforcement learning and direct preference optimization to objectives that better reflect human judgments.Examples include VideoDPO and VisionReward for temporal consistency and multi-objective preferences.

S6: Discussion and Future Work

The discussion emphasizes Reward Forcing’s general-purpose integration while identifying reward-evaluation misalignment and limitations in current video reward models. Future work focuses on richer, multi-dimensional reward modeling.

  • Generalizability: Reward Forcing is designed as a plug-and-play method compatible with various video generation architectures.The claimed practical benefit is adoption with minimal modifications to existing pipelines.
  • Limitations: Reward improvements may not translate proportionally to VBench gains when the reward function emphasizes some quality dimensions over others.The paper specifically contrasts temporal coherence with aesthetic qualities and VBench’s broader criteria.
  • Limitations: Current video reward models may miss long-range temporal dependencies, subtle artifacts such as frame jitter, and complex semantic attributes.Their subjective-annotation datasets may also incompletely represent video-quality dimensions.
  • Future work: Future directions include multi-objective, hierarchical, human-in-the-loop, domain-adaptive, and physics- or semantics-aware reward models.These directions aim to capture quality across dimensions, temporal scales, content types, and real-world dynamics.

S7: Border Social Impact

Reward Forcing offers efficiency benefits but raises risks involving misuse, bias, copyright, consent, and synthetic-content governance.

  • Reduced computational requirements could broaden access to video synthesis for educators, small organizations, and resource-limited researchers.
  • Faster, more accessible generation may lower barriers to deepfakes, misinformation, and identity fraud.
  • Reward-based prioritization may amplify vision-language-model biases and marginalize underrepresented groups or activities.
  • Deployment should address copyright, consent, and synthetic-content risks through watermarking, provenance tracking, and detection tools.
Loading 2512.04678v2…