Source-linked AI summary

LPM 1.0: Video-based Character Performance Model

Ailing Zeng, Casper Yang, Chauncey Ge, Eddie Zhang, Garvey Xu, Gavin Lin, Gilbert Gu, Jeremy Pi, Leo Li, Mingyi Shi, Shawn Wang, Sheng Bi, Steven Tang, Thorn Hang, Tobey Guo, Vincent Li, Xin Tong, Yikang Li, Yuchen Sun, Yue Zhao, Yuhan Lu, Yuwei Li, Zane Zhang, Zeshi Yang, Zi Ye

arXiv:2604.07823v2cs.CVcs.AIcs.MM

TL;DR

Existing video models struggle to combine expressive performance, real-time inference, and long-horizon identity stability, especially in conversation. LPM 1.0 addresses this with a multimodal full-stack system and benchmark, achieving consistent human preference while preserving real-time causal generation. Its scope remains limited to relatively simple, single-person, camera-facing interactions with weak broader-world grounding.

  • Problem

    Existing video models do not jointly achieve high expressiveness, real-time inference, and long-horizon identity stability in conversational character performance.

  • Method

    LPM 1.0 combines curated human-centric multimodal data, a multimodally conditioned 17B Base LPM, causal distillation into Online LPM, and LPM-Bench evaluation.

  • Results

    Base LPM and Online LPM are consistently preferred over state-of-the-art models in human evaluations, with matched-resolution judgments finding them indistinguishable in 42–88% of cases.

  • Takeaways & Limitations

    LPM 1.0 demonstrates that high-quality full-duplex conversational performance can be practical under deployable latency and stability constraints.

  • Takeaways & Limitations

    The system remains limited to mostly single, camera-facing characters and is not yet required to support long-horizon discourse memory, multi-party coordination, dynamic environments, or strong arbitrary-viewpoint 3D consistency.

Abstract

from arXiv · show

Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Learning such performance from video is a promising alternative to traditional 3D pipelines. However, existing video models struggle to jointly achieve high expressiveness, real-time inference, and long-horizon identity stability, a tension we call the performance trilemma. Conversation is the most comprehensive performance scenario, as characters simultaneously speak, listen, react, and emote while maintaining identity over time. To address this, we present LPM 1.0 (Large Performance Model), focusing on single-person full-duplex audio-visual conversational performance. Concretely, we build a multimodal human-centric dataset through strict filtering, speaking-listening audio-video pairing, performance understanding, and identity-aware multi-reference extraction; train a 17B-parameter Diffusion Transformer (Base LPM) for highly controllable, identity-consistent performance through multimodal conditioning; and distill it into a causal streaming generator (Online LPM) for low-latency, infinite-length interaction. At inference, given a character image with identity-aware references, LPM 1.0 generates listening videos from user audio and speaking videos from synthesized audio, with text prompts for motion control, all at real-time speed with identity-stable, infinite-length generation. LPM 1.0 thus serves as a visual engine for conversational agents, live streaming characters, and game NPCs. To systematically evaluate this setting, we propose LPM-Bench, the first benchmark for interactive character performance. LPM 1.0 achieves state-of-the-art results across all evaluated dimensions while maintaining real-time inference.

1. Introduction

LPM 1.0 frames conversational character performance as a systems problem spanning expressive behavior, real-time inference, and long-horizon identity stability. It combines curated multimodal data, controllable video generation, streaming distillation, and a benchmark for interactive performance.

  • Motivation: Learning performance directly from video could amortize authoring across identities, viewpoints, conversational contexts, emotions, and acting styles.The approach is motivated as an alternative to conventional pipelines that require proportional increases in assets, capture data, and expert labor.
  • Motivation: Conversation simultaneously demands expressive speaking and listening, real-time turn-taking, and stable identity across indefinite interaction horizons.These requirements jointly define the performance trilemma addressed by LPM 1.0.
  • Approach: LPM 1.0 addresses the trilemma with coordinated data construction, multimodal conditioning, and causal low-latency deployment.Its dataset pipelines target quality filtering, speaking-listening pairing, and identity-aware reference construction for full-duplex conversational generation.
  • Evaluation: LPM-Bench provides 1,000 test cases across speaking, listening, conversation, diverse motion, and character generalization with multimodal inputs.Its diversity spans appearance, performance, and audio, including long-form audio for extended-horizon consistency evaluation.
  • Evaluation: Base LPM and Online LPM are consistently preferred in human evaluations, while matched-resolution comparisons find them indistinguishable in 42–88% of cases across dimensions and scenarios.Base LPM is preferred over Kling-Avatar-2 and OmniHuman-1.5 by 64.3% and 42.5%, while Online LPM is preferred over LiveAvatar and SoulX by 82.5% and 64.1%.
  • Approach: The system introduces Base LPM, a 17B bidirectional Diffusion Transformer, and Online LPM, a causal streaming generator for real-time, infinite-length synthesis.The contributions also include multimodal training, identity-preserving references, and autoregressive distillation from the base model.

2. Data

LPM constructs a scalable, conversation-oriented multimodal dataset by filtering diverse videos, pairing speaking and listening behavior, and extracting identity-aware references. Its processing pipeline combines quality control, per-frame conversational labeling, semantic verification, and multi-granularity identity conditioning.

  • Dataset requirements: The dataset targets broad visual and motion coverage alongside speaking, listening, and multi-turn conversational behavior.It spans diverse shot scales, body motions, expressions, emotional states, and appearances while capturing both sides of interaction.
  • Data filtering: Raw videos are converted into standardized single-shot clips through scene segmentation, person detection, quality filtering, and content auditing.Filtering removes editing artifacts, visual defects, authenticity problems, and framing or composition issues.
  • Conversational processing: Conversational clips receive per-character, per-frame speak, listen, or idle labels to represent verbal production and non-verbal responses.The pipeline retains both single-person and multi-person clips, tracks people, crops person-centric views, and separates character-conditioned audio.
  • Conversational processing: 89.75% frame-level accuracy on Domain 1 and 87.63% on Domain 2 are achieved for speak-versus-listen classification after re-ranking and idle merging.Speak and listen recalls are balanced in Domain 1, while Domain 2 shows lower listen recall amid greater acoustic variability.
  • Semantic verification: 78.37 overall F1 is achieved by fine-tuned Qwen3-Omni, a +7.90 absolute gain over the Gemini baseline of 70.47.The largest gains occur for silence and speak, while confusion between listen_dialogue and conversation remains a residual weakness.
  • Identity-aware references: Three complementary reference types—global appearance, multi-view body, and facial expression images—provide an identity prior for consistent appearance and expressions.The references cover one to four body viewpoints and one to eight expressive states, addressing the limitations of single-image conditioning.

3. Base LPM

Base LPM uses a multimodal DiT architecture that jointly conditions video generation on visual references, text, and separate speaking and listening audio. Identity-aware reference tokens and temporal-extension training support expressive, identity-consistent generation over long sequences, while interleaved audio pathways distinguish speaking from listening motion.

  • Architecture: Base LPM predicts video noise from noisy video latents, text, speaking audio, listening audio, and identity reference images.Its DiT blocks combine spatiotemporal self-attention with multimodal cross-attention before VAE decoding.
  • Conversational Audio Conditioning: Speaking and listening audio use separate, interleaved conditioning pathways because they drive qualitatively different motion patterns.Speaking primarily produces high-frequency lip and rhythm movements, whereas listening correlates with slower posture and facial-expression changes.
  • Temporally-Aligned Audio-Video Attention: Speaking audio receives a local temporal attention window for precise lip synchronization, while listening audio uses a larger window for semantically informed reactions.The larger listening window supports responses such as subtle nods and gaze shifts to user audio over longer time scales.
  • Identity Multi-reference Conditioning: Identity references become spatial tokens in the video self-attention sequence, enabling fine-grained bidirectional interaction between video content and reference images.RoPE offsets distinguish reference types and support variable combinations of expression and body-view references.
  • Identity Multi-reference Conditioning: Multi-view and expression references provide geometric and facial-deformation priors, while persistent reference tokens anchor identity during long-video denoising and sliding-window inference.These mechanisms target head turns, profile views, natural speech-driven facial motion, and prevention of identity drift.
  • Training for Long-Horizon Generation: Temporal-extension training addresses chunk-wise inference mismatch by providing a global reference and randomly dropping 2-5 initial ground-truth video latents.This exposes the model to continuation windows beginning with causal latents rather than only ground-truth initial frames.

4. Online LPM

Online LPM converts offline conversational performance into causal, low-latency generation by addressing streaming-control mismatch and autoregressive error accumulation. Its backbone–refiner design separates temporal stabilization from high-fidelity detail recovery, with staged distillation and rollout-aware training.

  • Motivation: Base LPM is not directly deployable for interactive applications because offline continuation does not provide causal, low-latency, unbounded generation under incrementally changing controls.Online deployment lacks future context, and small rollout errors can accumulate over time.
  • Streaming control: Streaming audio and evolving text create a train–inference mismatch because online controllers repeatedly process truncated history instead of globally encoded context.Audio is especially sensitive because it drives lip motion, co-speech dynamics, and temporal synchronization.
  • Online architecture: The online generator uses a causal backbone for stable temporal anchoring and a causal refiner for recovering clean, high-frequency facial, appearance, and motion details.The backbone uses noisy-history KV caches, while the refiner uses clean-history KV caches under chunk-wise causal attention.
  • Training: DMD training progresses from supervised ODE initialization to off-policy and on-policy backbone distillation, followed by one-step refiner training on backbone-generated trajectories.This curriculum progressively reduces optimization difficulty and narrows the training–inference gap.
  • Training: On-policy training samples states from the backbone’s own autoregressive rollout distribution, matching the states encountered during online inference.Off-policy stages instead use teacher-derived latent states.

5. Infrastructure

The infrastructure builds Online LPM into a persistent runtime for continuous, responsive, long-horizon video generation. Chunked pipelining, efficient inference, state-aware updates, and synchronized audio handling support real-time deployment.

  • Runtime requirements: Online LPM is designed as a stateful service that continuously produces temporally coherent video while responding to streaming user inputs and control updates.Its target applications include conversational agents, live-streaming characters, and game NPCs.
  • Runtime requirements: The runtime targets low-latency generation, state-adaptive responsiveness, and long-horizon consistency with bounded memory cost.These requirements cover real-time streaming, prompt reactions to interruptions, and preservation of identity and motion over extended sessions.
  • Pipelined inference: Fixed 1-second chunks at 24 fps flow through overlapping Generator, Refiner, and VAE stages, reducing effective latency while maintaining throughput.The Generator and Refiner each take approximately 700 ms, while the VAE requires 180 ms.
  • Full-duplex integration: An external audio2audio module can provide streaming audio for live full-duplex dialogue, including through end-to-end or ASR–LLM–TTS systems.The Figure 7 workflow illustrates this broader runtime using audio conditions aligned to the streaming timeline.
  • Inference optimization: Fused kernels, efficient attention, and torch.compile reduce 1-NFE inference to approximately 0.35 s per chunk on a single GPU.The system tracks per-chunk latency, time-to-first-response, and steady-state throughput for deployment readiness.
  • System optimization: State splitting, boundary-aligned updates, and controlled lookahead handle interruptions and role transitions without disrupting visual continuity.Updated audio representations are aligned with the active chunk timeline to synchronize speech and facial motion.

6. Evaluation

LPM-Bench evaluates conversational performance across functional modes, human preferences, online deployment scenarios, and identity-reference ablations. Base and Online LPM show strong preference and stability advantages, while long action sequences and subtle listening behavior remain challenges.

  • Benchmark: LPM-Bench tests conversational performance with functional and generalization layers spanning dedicated modes, diverse identities, motions, and multimodal inputs.The benchmark targets behavioral expressiveness, conversational naturalness, and long-horizon character consistency rather than only general visual quality.
  • Base Model Evaluation: Base LPM supports minute-length generation with stable fidelity, precise text control, and consistent identity, while both baselines are limited to approximately 30 seconds.Short-form evaluation also reports natural, contextually appropriate acting and emotionally precise delivery.
  • Base Model Evaluation: Base LPM is preferred overall by 64.3% over Kling-Avatar-2 and 42.5% over OmniHuman-1.5.The strongest reported advantages are identity consistency at 58.5% and text controllability at 55.7%.
  • Base Model Evaluation: Temporally ordered action execution is the central remaining controllability bottleneck, with 78% of failures concentrated in multi-paragraph action sequences.
  • Online Model Evaluation: Online LPM is preferred over LiveAvatar overall by 82.5% and over SoulX by 64.1%, with SoulX retaining an identity-consistency advantage.Online LPM is favored over SoulX on motion dynamics, text controllability, and audio-video synchronization.
  • Online Model Evaluation: Online LPM preserves Base LPM capabilities while trading some subtle listening expressiveness for smoother speaking and stronger long-horizon conversational identity stability.In conversation, identity consistency favors Online LPM 48.0% to 10.0%; in listening, Base LPM leads motion dynamics 40.0% to 12.0%.
  • Ablation Study: Emotion references preserve person-specific facial expression cues, while view references preserve structural and pose-dependent appearance consistency.Their combination supports more complete identity preservation during expressive facial behavior and substantial body motion.

7. Summary and Discussion

LPM 1.0 frames conversational character performance as a systems problem spanning data, multimodal conditioning, generation, streaming, and stabilization. The resulting system makes high-quality full-duplex performance practical under deployable latency and stability constraints, but its current scope remains limited.

  • Summary and Discussion: LPM 1.0 addresses the performance trilemma through coordinated design across data construction, multimodal control, base-model training, and online inference.The paper treats expressiveness, real-time inference, and long-horizon stability as a system-level problem rather than an architectural challenge alone.
  • Summary and Discussion: The system demonstrates that high-quality full-duplex conversational performance can be practical under deployable latency and stability constraints.
  • Limitations: Current interaction is mostly limited to a single, camera-facing character and is weakly grounded in broader social, physical, and narrative structure.The system does not yet require long-horizon discourse memory, multi-party coordination, dynamic-environment behavior, or strong arbitrary-viewpoint 3D consistency.
  • Future Work: Future extensions require discourse memory, persona persistence, multi-party turn-taking, and grounding in scene geometry, objects, and contact.The paper presents LPM 1.0 as an initial systems-level step toward more unified actor models.

8. Safety, Security, and Responsibility

LPM 1.0 recognizes that realistic human-video generation raises identity, misuse, fairness, representation, and transparency concerns. The paper describes synthetic demonstration data, provenance and detection measures, input filtering, and responsible-deployment considerations.

  • Identity Protection and Consent: Reference-conditioned human-like video generation raises concerns about unauthorized identity replication, non-consensual likeness use, and personality-rights violations.
  • Misuse Prevention and Technical Safeguards: The system uses invisible watermarking, provenance metadata, companion detection models, and input-level safety filtering to address misuse risks.The stated misuse risks include non-consensual deepfakes, fraud, disinformation, and social engineering attacks.
  • Responsible Deployment and Social Impact: Responsible deployment requires attention to fairness, representation, and transparency while pursuing applications such as AI companions, education, and content creation.
Loading 2604.07823v2…