Source-linked AI summary

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

Hao Sun, Hao Yan, Mengting Chen, Quanjian Song, Yu Li, Juan Cao, Jinsong Lan, Xiaoyong Zhu, Bo Zheng, Sheng Tang

arXiv:2606.26092v2cs.CV

TL;DR

Existing video virtual try-on methods are tied to input camera trajectories, limiting novel-view garment inspection. TryOnCrafter introduces a camera-controllable framework built around a renderable 4D try-on proxy and proxy-anchored DiT, and evaluations show it significantly outperforms existing baselines in structural consistency and garment identity under unconstrained camera movements.

  • Problem

    Existing video virtual try-on methods remain confined to input camera trajectories, limiting view-consistent garment synthesis and background alignment under arbitrary camera motion.

  • Method

    TryOnCrafter combines a renderable 4D try-on proxy that decouples subject and environment with a proxy-anchored DiT for camera-controllable video synthesis.

  • Results

    TryOnCrafter significantly outperforms existing baselines in preserving structural consistency and garment identity under unconstrained camera movements.

  • Takeaways & Limitations

    The editable 4D representation supports downstream applications including human relocalization and 360-degree orbital viewing.

  • Takeaways & Limitations

    Extreme viewpoint transitions can cause parallax and ambiguity, while iterative DiT denoising incurs high inference costs that hinder real-time trajectory editing.

Abstract

from arXiv · show

While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms remains fundamentally constrained by a passive dependency on source camera trajectories, failing to accommodate the requisite interactive freedom for omnidirectional viewpoint exploration. To address this limitation, we define a pioneering research frontier: Camera-controllable Video Virtual Try-on (CaM-VVT). Unlike conventional VVT, CaM-VVT not only necessitates viewpoint-agnostic texture hallucination but also strict structural synchronization between non-rigid human dynamics and background contexts under arbitrary, unconstrained camera movements. To tackle these challenges, we present TryOnCrafter, the first unified DiT-based framework specifically architected for the CaM-VVT task. Departing from implicit pixel-space manipulation, we introduce a Renderable 4D Try-on Proxy that explicitly decouples the human subject from the environment. This is achieved by distilling high-fidelity 2D try-on priors into a clothed 3DGS-based avatar, which is subsequently animated via SMPL-X sequences and metric-aligned into a reconstructed background point cloud. This proxy establishes a robust structural foundation with superior texture density and motion integrity. Our Proxy-Anchored Video DiT leverages this robust structural foundation as a primary geometric anchor, ensuring that the synthesized photorealistic videos are strictly constrained by prescribed trajectories and physically plausible deformations. Benefiting from the inherent editability of the 4D proxy, TryOnCrafter facilitates diverse downstream applications, including human relocalization, ``bullet time'' effects, and $360$-degree orbital viewing.

1 Introduction

Existing video virtual try-on methods are constrained by the input camera trajectory, motivating Camera-controllable Video Virtual Try-On (CaM-VVT) for free novel-view inspection. TryOnCrafter addresses this task with a unified DiT framework centered on a Renderable 4D Try-on Proxy and supports evaluation and interactive applications.

  • Motivation: Existing VVT methods remain confined to fixed monocular-video camera trajectories, limiting free inspection of garments from novel viewpoints.Users may want to assess details such as side profiles and back appearance.
  • Challenges: Sequentially combining video virtual try-on with V2V camera control is suboptimal because errors accumulate and garment geometry and deformation remain explicitly unmodeled.Texture inconsistencies can be amplified when camera-control models process out-of-distribution try-on results.
  • TryOnCrafter: TryOnCrafter is a unified, end-to-end DiT framework for CaM-VVT that decouples the human subject from the environment through a Renderable 4D Try-on Proxy.The proxy distills high-quality 2D image try-on priors into a clothed 3DGS-based avatar.
  • Evaluation and Applications: CaM-VVTBench provides a comprehensive evaluation protocol, while TryOnCrafter supports human relocalization, “bullet time” effects, and 360◦ orbital viewing.Experiments report improved preservation of structural consistency and garment identity relative to existing baselines.
  • TryOnCrafter: The Proxy-Anchored Video DiT establishes robust structural grounding and motion integrity for re-rendering under unconstrained camera movements.This component is presented as part of the paper’s re-rendering-based paradigm.

2 Related Work · 3 Method

TryOnCrafter defines Camera-controllable Video Virtual Try-on and addresses arbitrary camera trajectories with a renderable 4D proxy plus proxy-anchored video diffusion generation. Its pipeline reconstructs and aligns the scene, builds and animates a clothed 3DGS avatar, and combines rendered structural priors with reference and semantic conditioning.

  • 2 Related Work: Existing video virtual try-on methods target dynamic garment transfer with spatio-temporal consistency, whereas TryOnCrafter introduces camera-controllable synthesis across arbitrary camera trajectories.The method aims to preserve intricate garment textures and robust temporal coherence under unconstrained viewpoints.
  • 3 Method: The framework comprises 4D Try-on Proxy Construction and Proxy-Anchored Video Generation, using a metric-aligned world-space scene that supports interactive rendering across arbitrary camera trajectories.The proxy explicitly combines a dressed human avatar with persistent background context.
  • 3.1 4D Try-on Proxy Construction: Dense depth, confidence, and camera parameters reconstruct a global point cloud, which SAM 2 partitions into human and background components for target-garment replacement within a unified metric-aligned world space.The human component is replaced by a high-fidelity 3DGS avatar, while the background point cloud is preserved for seamless rendering.
  • Scene Reconstruction and Anchor-based Alignment: A similarity transformation anchored by the reconstructed human point cloud maps local-camera SMPL-X motion into world space, while confidence-aware alignment prioritizes reliable regions such as the torso.This resolves monocular scale ambiguity and precisely localizes the 3DGS avatar, maintaining temporal and spatial coherence.
  • Canonical 3DGS-based Avatar Generation: An optimal keyframe is converted into a target-garment reference image and distilled into a canonical 3DGS avatar capturing personalized identity and intricate apparel geometry, then animated with LBS-driven aligned SMPL-X poses.The posed avatar preserves fine-grained garment details and structural alignment with the background context.
  • 3.2 Proxy-Anchored Try-on Video Generation: The Proxy-Anchored Video DiT uses a rendered proxy video as a pixel-aligned spatio-temporal prior encoding geometry, poses, appearance, and the prescribed camera trajectory to constrain physically plausible synthesis.This design addresses structural distortion and identity drift under radical motion.
  • Cross-view Reference Adapter (CRA) · Multi-Modal Semantic Conditioning: CRA transfers fine-grained identity and background details through weight-sharing attention, while CLIP garment tokens and UmT5 attribute tokens provide semantic grounding that prevents drift in unobserved or occluded viewpoints.CRA complements proxy geometry with source features and preserves high-fidelity appearance across unobserved viewpoints with minimal computational overhead.

4 Experiments

Experiments evaluate TryOnCrafter on conventional VVT and the new CaM-VVT setting using quantitative, qualitative, and ablation studies. The results show improved video fidelity, texture preservation, structural robustness, and trajectory-aligned synthesis, while confirming the importance of explicit proxy anchoring.

  • Quantitative Comparison: On ViViD, TryOnCrafter establishes new SOTA performance across paired and unpaired settings, improving VFIDI and VFIDR while retaining competitive structural similarity and superior LPIPS.Paired evaluation additionally uses SSIM and LPIPS, while both paired and unpaired evaluations report VFIDI and VFIDR.
  • Qualitative Comparison: Qualitatively, TryOnCrafter improves texture transfer and color preservation, maintains intricate patterns and vibrancy, and reconstructs complex silhouettes through geometry-aware rendered priors.The comparison specifically contrasts TryOnCrafter with CatV2TON and Magic-Tryon, including full-length trousers that Magic-Tryon fails to resolve structurally.
  • Camera-controllable Video Virtual Try-On Benchmark: For CaM-VVT, TryOnCrafter is compared against Magic-Tryon combined with TrajectoryCrafter or ReCaMaster, covering explicit point-cloud and implicit parameter-injection camera control.These coupled baselines are established because end-to-end CaM-VVT models are absent.
  • Camera-controllable Video Virtual Try-On Benchmark: Under unconstrained trajectories, TryOnCrafter synthesizes viewpoint-agnostic fine-grained textures, realistic trajectory-aligned cloth deformations, and structurally robust, high-fidelity appearances.The passage contrasts this behavior with baselines that fail to preserve texture consistency or physically plausible clothing motion during violent maneuvers.
  • Ablation on 4D Try-on Proxy: Ablations show that removing the clothed avatar, background point cloud, or anchor-based alignment causes motion stochasticity, structural artifacts, or restricted scene consistency, while removing rendered video sharply reduces performance.The study also finds garment-semantic removal marginally beneficial in simple paired settings but severely degrading in unpaired scenarios.

5 Applications

The explicit 4D Try-on Proxy decouples human subjects, backgrounds, and camera trajectories within a unified world space. This enables generative applications such as arbitrary spatial edits to motion sequences for human relocalization.

  • Applications: The explicit 4D Try-on Proxy decouples human subjects, backgrounds, and camera trajectories within a unified world space Sw.This unified representation supports applications beyond standard video synthesis.
  • Human Relocalization: Human relocalization enables arbitrary spatial edits to motion sequences by integrating the subject and background into Sw.The passage presents this as a geometry-grounded application of the framework.

6 Limitation

TryOnCrafter remains limited by structural alignment failures under extreme viewpoint transitions and radical articulated motions, as well as high inference costs that hinder real-time trajectory editing.

  • Geometric alignment: Extreme viewpoint transitions introduce parallax and ambiguity, challenging structural alignment between the subject and background.The framework relies on a parametric 4D Try-on Proxy for geometric guidance, but spatial shifts can exacerbate estimation inaccuracies.
  • Geometric alignment: Radical articulated motions can produce misaligned hand poses and subtle geometric inconsistencies through SMPL-X inaccuracies.These failures occur occasionally during challenging motions.
  • Inference efficiency: Iterative denoising in the DiT-based framework incurs high inference costs, hindering real-time interaction for trajectory edits.The limitation affects interactive trajectory editing rather than the proxy’s underlying editability.

7 Conclusion

TryOnCrafter pioneers Camera-controllable Video Virtual Try-on (CaM-VVT) to overcome existing methods’ dependence on source camera trajectories. It combines a Renderable 4D Try-on Proxy with a Proxy-anchored DiT for explicit subject-environment decoupling and high-fidelity synthesis.

  • Conclusion: TryOnCrafter pioneers CaM-VVT to overcome the trajectory dependency of existing video virtual try-on paradigms.The framework targets camera-controllable video virtual try-on.
  • Conclusion: The Renderable 4D Try-on Proxy explicitly decouples the human subject from the environment.This proxy is one of TryOnCrafter’s two synergistic components.
  • Conclusion: The Proxy-anchored DiT uses a multi-tiered conditioning hierarchy for high-fidelity synthesis.The generative process is anchored to the 4D proxy, supporting physically plausible garment deformations.

Supplementary Materials · S1 Overview · S2 Training Scheme and Dataset Construction

The supplementary materials provide implementation details, analyses, and qualitative results, including the training scheme and dataset construction. TryOnCrafter uses progressive two-stage training with stage-specific monocular and synthetic multi-view data to support camera-controllable virtual try-on.

  • S1 Overview: The supplementary materials cover implementation details, analyses, qualitative results, keyframe sampling, hybrid rendering, benchmarking, efficiency, robustness, comparisons, and additional qualitative results.These topics are organized across Sections S1–S9.
  • Training Scheme: TryOnCrafter adopts a progressive two-stage training strategy for high-fidelity and spatially consistent multi-view virtual try-on.The strategy is summarized in Table S1.
  • Training Scheme: Stage 1 initializes from Wan2.1-I2V and performs monocular video virtual try-on pre-training for high-fidelity texture transfer and spatiotemporal coherence.A coarse-to-fine resolution schedule progressively increases input resolution to preserve visual details and capture intricate garment textures.
  • Training Scheme: Stage 2 introduces camera control and full-resolution fine-tuning to enforce cross-view spatial consistency and geometry-consistent rendering.The model learns view-aware garment deformation for dynamic and unconstrained camera trajectories.
  • Training Dataset Construction: Because synchronized multi-view try-on data are scarce, the two training stages use different dataset constructions.This design addresses the acquisition challenge for CaM-VVT training.
  • Training Dataset Construction: Stage 1 uses monocular ViViD videos and curated in-the-wild CaM-VVTBench sequences, while ViViD experiments restrict training to ViViD for controlled comparison.The combined data improve robustness across subjects, garments, poses, and environments.
  • Training Dataset Construction: Stage 2 constructs a large-scale synthetic multi-view try-on dataset by dual-reprojecting CaM-VVTBench monocular videos with randomized 2D/3D masking.Each synthesized pair includes a rendered proxy, as illustrated in Fig. S1.

S3 Keyframe Sampling Strategy · S4 Hybrid Rendering Strategy

The method selects a camera-aligned canonical frame from SMPL-X geometry, then combines 3D Gaussian Splatting for the human avatar with PyTorch3D background rendering. Depth-buffer fusion resolves occlusions and produces geometrically consistent, temporally coherent proxy videos for generative refinement.

  • S3 Keyframe Sampling Strategy: For each frame, the torso forward vector is estimated as the normalized cross product of the shoulder-to-shoulder and neck-to-pelvis joint vectors.The vectors are derived from SMPL-X joints.
  • S3 Keyframe Sampling Strategy: The canonical frame maximizes a viewing score measuring alignment between the torso forward vector and the camera optical axis.The score is defined by the normalized dot product between f_t and −z_c.
  • S3 Keyframe Sampling Strategy: This sampling favors frames in which the human body is well aligned with the camera, providing a stable reference for proxy construction and try-on synthesis.The selected frame serves as the canonical reference for subsequent processing.
  • S4 Hybrid Rendering Strategy: The reconstructed 4D try-on proxy supports video synthesis under both the input camera trajectory and novel camera trajectories.This capability is implemented through a hybrid rendering strategy.
  • S4 Hybrid Rendering Strategy: The human avatar Gp is rendered with a 3D Gaussian Splatting pipeline, while the background point cloud Pb is rendered using the PyTorch3D rasterizer.The two components use separate rendering pipelines before compositing.
  • S4 Hybrid Rendering Strategy: Rendered human and background outputs are fused through depth-buffer comparison to handle mutual occlusions.Db and Dh denote the rendered depths of the background and human avatar, respectively.
  • S4 Hybrid Rendering Strategy: The hybrid rendering scheme produces geometrically consistent and temporally coherent video proxies that guide subsequent generative refinement.The resulting proxies provide spatial-temporal guidance for the refinement stage.

S5 Benchmark Scale and Baseline Comparison · S6 Efficiency and Reusability Analysis · S7 Sensitivity and Robustness Analysis

CaM-VVTBench offers trajectory-aware evaluation across indoor and outdoor scenes at a moderate scale comparable to existing video virtual try-on benchmarks. TryOnCrafter’s proxy is reusable and efficient for trajectory changes, while robustness tests show stable performance under substantial upstream degradation, with reflection preservation remaining a limitation.

  • S5 Benchmark Scale and Baseline Comparison: CaM-VVTBench has moderate scale because trajectory annotation is labor-intensive, yet remains comparable to existing video virtual try-on benchmarks.It covers indoor and outdoor scenes and provides camera-trajectory-aware evaluation samples.
  • S5 Benchmark Scale and Baseline Comparison: CaM-VVTBench uniquely combines indoor and outdoor scenes with camera-trajectory-aware evaluation samples for camera-controllable video virtual try-on.
  • S6 Efficiency and Reusability Analysis: 1360.0s and 68.6GB are required for 720P video generation, versus 38.9s and 59.1GB for full proxy construction and rendering on one 80GB A100.The 14B DiT backbone, rather than proxy construction, is identified as the main computational bottleneck.
  • S6 Efficiency and Reusability Analysis: 14.3s and 20.3GB are required to re-render the constructed proxy for a changed camera trajectory, substantially lowering the marginal cost of multiple trajectories for one input.Point-cloud reconstruction, segmentation, SMPL-X estimation, separation, sampling, deformation, and rendering include independent, decoupled, or batchable stages.
  • S6 Efficiency and Reusability Analysis: TryOnCrafter’s generation cost is comparable to Wan2.1-I2V and much faster than two-stage baselines.The comparison is reported in Table S4 using runtime and peak memory.
  • S7 Sensitivity and Robustness Analysis: Despite incomplete reconstruction, occlusion, SMPL-X estimation errors, and motion blur, TryOnCrafter preserves coherent human motion and stable garment structure.These foreground stress tests show stability under imperfect upstream reconstruction.
  • S7 Sensitivity and Robustness Analysis: 10×/100× point-cloud downsampling and 25%/75% random point dropping still yield reasonable background consistency because the proxy supplies coarse geometric guidance rather than a strict pixel-level target.Moderate erosion and dilation of SAM2 masks are also tolerated, supporting robustness to imperfect in-the-wild segmentation.
  • S7 Sensitivity and Robustness Analysis: Unchanged reflections may appear in mirror-like scenes because residual reflection cues in the masked source video can preserve the original reflection under a new try-on condition.Robustness is supported by structured proxy guidance and training augmentation with dual reprojection and randomized 2D/3D masking.

S8 More Comparisons with SOTA Methods · S9 More Results of TryOnCrafter

TryOnCrafter is qualitatively compared with Magic-TryOn-based methods under challenging camera trajectories, where it demonstrates faithful garment details, stable identity preservation, and plausible visual quality. Additional synthesized and dynamic results are provided through Fig. S5 and the demo video.

  • S8 More Comparisons with SOTA Methods: TryOnCrafter is qualitatively compared with Magic-TryOn-based methods in Fig. S4.The comparison targets challenging camera trajectories.
  • S8 More Comparisons with SOTA Methods: TryOnCrafter produces more faithful garment details than the compared Magic-TryOn-based methods.This finding is reported for challenging camera trajectories.
  • S8 More Comparisons with SOTA Methods: TryOnCrafter provides more stable identity preservation under challenging camera trajectories.The result comes from the additional qualitative comparisons in Fig. S4.
  • S8 More Comparisons with SOTA Methods: TryOnCrafter achieves more plausible visual quality under challenging camera trajectories.This qualitative result is shown in comparisons with Magic-TryOn-based methods.
  • S9 More Results of TryOnCrafter: Fig. S5 presents more synthesized results of TryOnCrafter.These results extend the qualitative evidence reported for the framework.
  • S9 More Results of TryOnCrafter: The demo video provides additional dynamic results of TryOnCrafter.The passage directs readers to the demo video for these results.
Loading 2606.26092v2…