Source-linked AI summary
WildActor: Unconstrained Identity-Preserving Video Generation
Qin Guo, Tianyu Yang, Xuanhua He, Fei Shen, Yong Zhang, Zhuoliang Kang, Xiaoming Wei, Dan Xu
TL;DR
Existing human video generators struggle to preserve full-body identity across changing viewpoints and motions without restricting movement. WildActor combines the Actor-18M dataset with identity-preserving attention and viewpoint-adaptive sampling, outperforming prior methods on Actor-Bench.
Problem
Human video generation still lacks reliable full-body identity consistency across changing shots, viewpoints, and motions, with existing methods often face-centric or pose-locked.
Method
WildActor uses Actor-18M and combines Asymmetric Identity-Preserving Attention with Viewpoint-Adaptive Monte Carlo Sampling for any-view conditioned human video generation.
Results
WildActor consistently outperforms prior methods on Actor-Bench in identity preservation, semantic alignment, narrative coherence, and contextual generalization under viewpoint and motion variation.
Takeaways & Limitations
WildActor preserves full-body identity across dynamic shots, substantial motions, and large viewpoint transitions in unconstrained human video generation.
Takeaways & Limitations
The current implementation focuses on single-person videos and does not yet address multi-person scenarios with complex interactions.
Abstract
from arXiv · showhide
Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level consistency, or produce copy-paste artifacts where subjects appear rigid due to pose locking. We present Actor-18M, a large-scale human video dataset designed to capture identity consistency under unconstrained viewpoints and environments. Actor-18M comprises 1.6M videos with 18M corresponding human images, covering both arbitrary views and canonical three-view representations. Leveraging Actor-18M, we propose WildActor, a framework for any-view conditioned human video generation. We introduce an Asymmetric Identity-Preserving Attention mechanism coupled with a Viewpoint-Adaptive Monte Carlo Sampling strategy that iteratively re-weights reference conditions by marginal utility for balanced manifold coverage. Evaluated on the proposed Actor-Bench, WildActor consistently preserves body identity under diverse shot compositions, large viewpoint transitions, and substantial motions, surpassing existing methods in these challenging settings.
1. Introduction
Maintaining strictly invariant full-body identity across shots, viewpoints, and motions remains an open challenge because existing approaches can overemphasize facial cues or use naive reference injection. WildActor addresses this with Actor-18M, a large-scale dataset, and a framework combining asymmetric identity-preserving attention with viewpoint-adaptive sampling for any-view human video generation.
- Motivation: Strictly invariant identity across shots, viewpoints, and motions remains an open challenge in video generation.The introduction frames physical permanence as an anchor of visual storytelling.
- Limitations: Existing methods are predominantly face-centric or rely on naive injection, causing facial cues to be overemphasized while body identity is ignored.Face-centric methods can produce a “floating head” effect in which the body hallucinates.
- Dataset: Actor-18M comprises 1.6M videos and 18M corresponding human images spanning diverse viewpoints, environments, and motions.Multiple reference images of each subject within videos support learning identity-consistent representations under unconstrained conditions.
- Method: WILDACTOR combines Asymmetric Identity-Preserving Attention with Viewpoint-Adaptive Monte Carlo Sampling for robust any-view conditioning.The attention mechanism lets video tokens query identity cues while isolating reference tokens from noisy backbone features.
- Evaluation: On Actor-Bench, WILDACTOR consistently outperforms prior methods in narrative coherence and contextual generalization under large viewpoint and motion variations.Figure 1 also demonstrates stable human-centric narrative and multi-shot generation across substantial changes in environment, viewpoint, and action.
2. Related Work
Related work has advanced from U-Net architectures toward DiT-based video generation and foundation models with strong temporal coherence and photorealism, but strict subject identity control remains challenging. Existing human-centric datasets also inadequately capture full-body motion and appearance, limiting robust human video generation.
- Identity-Preserving Video Generation: Video generation has shifted from U-Net architectures to DiT, while Sora, Kling, and Wan demonstrate strong temporal coherence and photorealism.These advances have not resolved strict subject identity control.
- Identity-Preserving Video Generation: Strict subject identity control remains a significant challenge in modern video-generation models.
- Human-Centric Video Datasets: Talking-head datasets such as CelebV-HQ and TalkingFace-Wild focus on facial dynamics but lack body motion and appearance.
- Human-Centric Video Datasets: HumanVid primarily supports pose-driven generation, reflecting limitations in existing datasets for robust human-centric video generation.
3. The Proposed Actor-18M Dataset
Actor-18M is a large-scale human-video dataset built through identity filtering and augmentation across viewpoints, environments, attributes, and canonical views. Its three subsets address pose locking, background and lighting overfitting, and incomplete identity coverage while mitigating frontal-view bias.
- Dataset construction: Actor-18M contains 1.6M identity-consistent single-person videos and over 18M reference images with dense supervision across environments, viewpoints, and motions.A two-stage pipeline combines sparse-frame facial similarity filtering with dense point tracking and clip-similarity verification.
- Dataset subsets: Actor-18M-A synthesizes face and body references from six viewing angles, verifying identity and clothing consistency to counter pose locking from correlated viewpoints.The subset uses a multi-angle image editing model and Multimodal LLM verification.
- Dataset subsets: Actor-18M-B diversifies environments, expressions, lighting, and motions through attribute-conditioned editing while preserving subject identity.Its attribute pool covers 200 environments, 8 expressions, 10 lighting conditions, and 30 motions.
- Dataset subsets: Actor-18M-C provides canonical front, side, and back identity anchors generated from subjects visible across all three viewpoints.Pose estimation and MLLM refinement identify suitable subjects before reference-based character-sheet generation.
- Distribution balancing: Actor-18M-A shifts body-view distribution from 63.1% frontal bodies to 70.5% side views, while profile faces exceed 69% of generated samples versus 42% originally.Actor-18M-B and -C additionally expand attribute diversity and provide canonical anchors.
4. The Proposed Method
WildActor performs any-view-conditioned human video generation by conditioning a latent video DiT on textual prompts plus facial and body references, while preserving identity across viewpoints and temporal dynamics. Its method combines asymmetric identity injection, token-type positional separation, and viewpoint-adaptive reference sampling.
- Framework: WildActor conditions a latent video DiT on a textual prompt together with facial and body reference images for identity-preserving video generation.The goal is to maintain identity fidelity across viewpoints while preserving temporal coherence.
- Asymmetric Identity-Preserving Attention: AIPA isolates reference identity cues from noisy video latents by allowing reference tokens to inform video tokens without reciprocal contamination.It combines reference-only LoRA with asymmetric attention flow to reduce identity leakage and pose-locking artifacts.
- Identity-Aware 3D RoPE: I-RoPE assigns distinct temporal and spatial coordinates to facial, body, and video tokens to disambiguate static identity from temporal motion.Reference tokens use temporal offsets T + ∆f and T + ∆b, with ∆f = 4 and ∆b = 128, and shifted spatial indices beginning at (Hmax, Wmax).
- Viewpoint-Adaptive Monte Carlo Sampling: Viewpoint-Adaptive Monte Carlo Sampling dynamically suppresses angularly redundant candidates, biasing training toward complementary viewpoints and more uniform identity-manifold coverage.The scheme is designed to make multi-view references diverse and informative rather than repeatedly sampling similar views.
5. Experiments
Actor-Bench evaluates identity consistency and contextual generalization across 75 subjects and three conditioning settings. WildActor outperforms baselines qualitatively and quantitatively, while ablations show benefits from viewpoint-adaptive sampling, AIPA, and I-RoPE.
- Evaluation Setup: Actor-Bench contains 75 subjects evenly divided across canonical three-view, arbitrary-viewpoint, and in-the-wild conditioning settings.Each subject includes manually verified canonical three-view references for reliable evaluation.
- Evaluation Axes: Sequential narrative evaluation uses three consecutive prompts forming coherent storylines, while contextual generalization tests diverse environments, viewpoints, and motions.Both axes assess identity consistency beyond isolated clips.
- Qualitative Results: WildActor maintains consistent identity and coherent temporal progression, whereas Qwen-Image-Edit + I2V shows clip discontinuities and naive T2V →I2V accumulates identity drift.Identity references are incorporated throughout autoregressive generation to reduce long-form error accumulation.
- Quantitative Results: 0.952 body consistency is achieved by WildActor versus 0.905 for Vidu Q2 under contextual generalization, while WildActor also attains the highest VLM-level score.Vidu Q2 obtains higher face identity scores, attributed to appearance copying that favors facial similarity over structural flexibility.
- Dataset and Sampling Ablation: 0.952 average consistency is reached with Viewpoint-Adaptive sampling, improving over raw crops and random sampling by learning viewpoint-agnostic robustness.Raw-Crop degrades on side and back views because of viewpoint imbalance, while Random Sampling mitigates this issue.
- Model Ablation: Replacing AIPA with Full-Attn reduces semantic adherence, while removing I-RoPE sharply lowers body consistency by confusing reference and video features.AIPA improves facial identity preservation and instruction adherence; I-RoPE supports structural coherence and motion quality.
6. Conclusion
The paper introduces Actor-18M to address data limitations in identity-consistent human video generation and presents WILDACTOR for any-view conditioned generation. WILDACTOR preserves full-body identity across dynamic shots, large viewpoint transitions, and substantial motions.
- Contributions: Actor-18M is a large-scale human-centric video dataset providing diverse, complementary identity references across unconstrained environments, viewpoints, and motions.It addresses a key data limitation in identity-consistent human video generation.
- Contributions: WILDACTOR is a framework for any-view conditioned human video generation built on Actor-18M.The framework is designed for identity preservation under unconstrained video conditions.
- Contributions: WILDACTOR consistently preserves full-body identity across dynamic shots, large viewpoint transitions, and substantial motions.These conditions represent challenging settings for identity-consistent human video generation.
A. More Results
Additional qualitative results are provided in the supplementary material’s index.html.
- Additional qualitative results are available in the supplementary material’s index.html.
B. Implementation Details of Dataset Construction
This section describes the engineering pipeline used to construct Actor-18M, including model specifications, hyperparameters, and protocols for filtering, annotation, and generative augmentation.
- Pipeline scope: The Actor-18M construction pipeline covers data filtering, annotation, and generative augmentation.These components are presented as part of a comprehensive engineering pipeline.
- Implementation specifications: The implementation details specify the models and hyperparameter settings used in dataset construction.The section also details the algorithmic protocols supporting the pipeline.
B.1. Data Processing Pipeline · B.2. Generative Augmentation Pipeline
The data pipeline filters 1.6M videos for identity and motion consistency, then extracts face and body annotations. Generative augmentation creates viewpoint-transformed, attribute-edited, and canonical multi-view identity representations with verification at each stage.
- B.1. Data Processing Pipeline: 1.6M high-quality videos are selected from raw sources through a cascaded coarse-to-fine filtering strategy.The pipeline is designed to preserve identity consistency and annotation precision at massive scale.
- B.1. Data Processing Pipeline: Videos are first sampled at 1 fps and discarded when ArcFace similarity between the first and subsequent frames averages below 0.4.This coarse stage removes obvious identity shifts and false detections.
- B.1. Data Processing Pipeline: Passing videos are up-sampled to 8 fps, tracked with CoTracker, screened for occlusions and cuts, and retained only when CLIP consistency exceeds 0.45.The fine-grained stage evaluates motion and appearance stability.
- B.1. Data Processing Pipeline: RetinaFace and BiSeNet provide facial landmarks, bounding boxes, and fine-grained face segmentation, while YOLO-World and SAM2 produce temporal-consistent whole-body masks.The annotation pipeline targets precise pixel-level face and body regions.
- B.2. Generative Augmentation Pipeline: Subset A crops clear face and body references, synthesizes six viewpoints with Qwen-Image-Edit-Multiple-Angles, and keeps only views passing Qwen3-VL-32B consistency verification.The six viewpoints are Front, Left-Side, Right-Side, Back, Top-Down, and Bottom-Up.
- B.2. Generative Augmentation Pipeline: Subset B samples combinations from 200 environments, 8 expressions, 10 lighting conditions, and 30 motions to generate editing instructions and verified identity-preserving edits.Qwen3-VL-32B creates instructions, Qwen-Image-Edit performs editing, and CLIP-I checks the identity embedding.
- B.2. Generative Augmentation Pipeline: Subset C filters videos showing Front, Side, and Back views, selects three peak-visibility anchor frames, arranges them into orthogonal character sheets, and audits artifacts and clothing consistency.DWPose identifies suitable clips; Nano-Banana performs arrangement; Qwen3-VL-32B conducts the final audit.
C. Evaluation Metrics Details · C.1. Body Consistency · C.2. VLM-Level Semantic Alignment
The evaluation uses Gemini-3-Pro to measure frame-level Body Consistency and video-level Semantic Alignment. Body Consistency matches viewpoints against references, while Semantic Alignment judges holistic adherence to textual prompts.
- C. Evaluation Metrics Details: Gemini-3-Pro serves as the automated evaluator for identity consistency and semantic alignment.Body Consistency is evaluated at the frame level, whereas Semantic Alignment is assessed at the video level.
- C.1. Body Consistency: The viewpoint-aware identity pipeline addresses facial-metric limitations by evaluating full-body consistency under arbitrary viewpoints.Traditional metrics such as ArcFace primarily focus on facial features and are sensitive to viewpoint changes.
- C.1. Body Consistency: For each generated frame, Gemini-3-Pro estimates the subject’s dominant viewpoint.The evaluation uses ground-truth references covering Front, Side, and Back viewpoints.
- C.1. Body Consistency: Each frame is paired with the ground-truth reference sharing its closest viewpoint before identity comparison.This matching step aligns generated frames with viewpoint-corresponding references.
- C.1. Body Consistency: The VLM assigns sbody_t ∈{0, 1}, where 1 denotes consistent identity and 0 denotes inconsistency.The final Body Consistency score averages positive scores over the total number of evaluation frames T.
- C.2. VLM-Level Semantic Alignment: Semantic Alignment verifies whether generated videos follow complex prompts describing actions, environments, and lighting.The assessment uses the entire video clip together with its instructions, enabling holistic visual-content verification.
- C.2. VLM-Level Semantic Alignment: A binary Svideo ∈{0, 1} is assigned per video, and the final score averages these values across the benchmark.Because assessment is video-wise, semantic deviations occurring throughout a generated video are penalized; N denotes the total number of benchmark videos.
D. Limitations
WildActor currently focuses on single-person video generation, while multi-person scenarios with complex interactions remain an open challenge requiring explicit identity disentanglement.
- D. Limitations: The current implementation targets single-person video generation and does not yet handle multiple subjects simultaneously.Multi-person generation requires disentangling identity features within the attention mechanism to prevent attribute mixing between characters.
- D. Limitations: Extending identity control to multi-person scenarios with complex interactions remains a challenging direction for future research.The proposed modules, including AIPA, demonstrate robust identity control, but this capability has not yet been extended to interacting subjects.