Source-linked AI summary

Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks

Hailong Guo, Bohan Zeng, Yiren Song, Wentao Zhang, Chuang Zhang, Jiaming Liu

arXiv:2501.15891v2cs.CV

TL;DR

Virtual try-on is constrained by scarce paired garment-model data, limited generalization and quality, and reliance on restrictive conditions or task-specific methods. Any2AnyTryon addresses these issues with the LAION-Garment dataset, a unified mask-free conditional framework, and adaptive position embeddings. Experiments report flexible, controllable, and high-quality generation that outperforms baseline and state-of-the-art methods across supported try-on tasks.

  • Problem

    Scarce paired garment-model data and task-specific methods limit generalization, quality, and user-friendly mask-free virtual try-on generation.

  • Method

    Any2AnyTryon combines the LAION-Garment dataset, a unified conditional generation architecture, clean condition latents, and adaptive position embeddings for mask-free VTON.

  • Results

    Experiments report that Any2AnyTryon outperforms baseline and state-of-the-art methods while producing higher-quality, realistic, and garment-aligned try-on images.

  • Takeaways & Limitations

    Any2AnyTryon supports flexible, controllable generation across garment reconstruction, model-free virtual try-on, and other diverse VTON tasks.

Abstract

from arXiv · show

Image-based virtual try-on (VTON) aims to generate a virtual try-on result by transferring an input garment onto a target person's image. However, the scarcity of paired garment-model data makes it challenging for existing methods to achieve high generalization and quality in VTON. Also, it limits the ability to generate mask-free try-ons. To tackle the data scarcity problem, approaches such as Stable Garment and MMTryon use a synthetic data strategy, effectively increasing the amount of paired data on the model side. However, existing methods are typically limited to performing specific try-on tasks and lack user-friendliness. To enhance the generalization and controllability of VTON generation, we propose Any2AnyTryon, which can generate try-on results based on different textual instructions and model garment images to meet various needs, eliminating the reliance on masks, poses, or other conditions. Specifically, we first construct the virtual try-on dataset LAION-Garment, the largest known open-source garment try-on dataset. Then, we introduce adaptive position embedding, which enables the model to generate satisfactory outfitted model images or garment images based on input images of different sizes and categories, significantly enhancing the generalization and controllability of VTON generation. In our experiments, we demonstrate the effectiveness of our Any2AnyTryon and compare it with existing methods. The results show that Any2AnyTryon enables flexible, controllable, and high-quality image-based virtual try-on generation. https://logn-2024.github.io/Any2anyTryonProjectPage

1 Introduction

Any2AnyTryon addresses limited task coverage, restrictive input conditions, and insufficient garment diversity in virtual try-on. It combines a large garment-model dataset with a unified architecture and adaptive position embeddings to support flexible, high-quality generation.

  • Existing VTON methods often support only specific try-on tasks and impose strict limitations on user-provided input conditions.
  • Any2AnyTryon introduces a user-friendly, mask-free framework that simultaneously supports multiple try-on generation tasks based on user instructions.The framework models virtual try-on, model generation, and garment generation within one conditional generation system.
  • LAION-Garment provides a large and diverse garment-model dataset with textual editing instructions and both in-the-wild and in-the-shop model images.The dataset is intended to provide sufficient training samples for VTON models.
  • Adaptive Position Embedding adjusts position embeddings using textual prompts and image conditions, supporting flexible, non-fixed conditional inputs.The architecture places conditions in a clean latent format within the same representation space as the target latent.
  • Extensive experiments and evaluations report that Any2AnyTryon outperforms other baseline methods in performance.

2 Related Work

Related work has improved image generation and virtual try-on through diffusion, transformer, adapter, and reference-image conditioning methods. However, existing approaches remain limited in their conditions and task scope, motivating Any2AnyTryon’s use of complex instructions and variable image inputs.

  • Latent diffusion models and large-scale transformer architectures have improved the quality and efficiency of generative tasks.DiT is presented as a cutting-edge transformer-based model family, with FLUX.1 building on flow-based and DiT-based generation.
  • Reference-based image generation uses image conditions to produce customized outputs with consistency in appearance, character identity, or style.Adapter structures, ControlNet, and ReferenceNet are described as mechanisms for conditioning generation on reference images.
  • Existing reference-based methods are limited to specific conditions and single-task scenarios.
  • Virtual try-on methods have evolved from warping-and-aggregation pipelines toward approaches based on powerful text-to-image models.Earlier systems used TPS or flow-based warping followed by aggregation with GANs, diffusion models, or related generative models.

3 Method

Any2AnyTryon unifies multiple mask-free virtual clothing tasks by conditioning generation on variable image inputs and textual instructions. Its dataset construction and adaptive position embedding support spatially aligned, high-quality outputs across diverse scenarios.

  • Any2AnyTryon Framework: Any2AnyTryon models virtual try-on as conditional generation across multiple images, integrating virtual try-on, model generation, and garment generation.The framework supports user instructions together with outfitted-model and garment images.
  • LAION-Garment Dataset Collection: LAION-Garment expands VTON training data by integrating public datasets, internet-crawled image pairs, synthetic inpainting triples, and quality filtering.The construction uses VITON-HD, DressCode, DeepFashion2, and LRVS-Fashion, with GPT-4o selecting authentic, high-quality, posture-consistent triples.
  • Model Architecture: Image conditions and target noisy images are concatenated along pixel dimensions, allowing variable quantities and resolutions of model and garment inputs.The concatenated representation contains noised latents, image-condition tokens, and user-instruction tokens.
  • Adaptive Position Embedding: Adaptive Position Embedding locates each image condition using a three-channel encoding for region identity, height, and width.This design supports pixel-level spatial alignment between input conditions and generated images, including consistency in non-edited regions.
  • Training Design: Clean condition latents replace noisy condition latents because noisy conditions caused undesirable changes in background regions.The model incorporates this condition representation alongside adaptive position embeddings and conditional flow matching.

4 Experiment

Experiments evaluate Any2AnyTryon across garment reconstruction, model-free virtual try-on, mask-free try-on, layered try-on, and ablations. The reported results show stronger reconstruction fidelity, image quality, garment alignment, and flexible generation than comparison methods.

  • Implementation Details: The unified model is trained across variable image sizes and multiple tasks, with try-on-in-layer data introduced during second-stage finetuning.Training uses 512x384, 768x576, and 512x384 resolutions for different models and tasks.
  • Garment Reconstruction: Any2AnyTryon reconstructs garments with higher quality and more accurate appearance matching than TryOffDiff.The comparison evaluates flattened garments reconstructed from images of models wearing the target garments.
  • Model-free Virtual Try-on: Any2AnyTryon significantly outperforms baseline methods on model-free virtual try-on while preserving garment appearance and generating high-quality outfitted model images at varied resolutions.The comparison uses VITONHD and metrics including DiffSim and FFA.
  • Virtual Try-on: Any2AnyTryon outperforms existing mask-free virtual try-on methods, particularly on FID and KID, while improving realism and alignment with input garments.Compared methods include GP-VTON, OOTD, IDM-VTON, CatVTON, and FitDiT.
  • Layered Try-on: In layered try-on, Any2AnyTryon better preserves the input model’s appearance and aligns the worn garment more accurately than DiOr.The comparison uses examples from DiOr’s original paper.
  • Ablation Study: The ablation study compares the full method against noisy image-condition inputs and a version without adaptive position embedding.These variants test the image-condition addition strategy and adaptive position embedding.

5 Conclusion

The conclusion presents Any2AnyTryon as a unified, mask-free framework for diverse virtual try-on scenarios. It reports improved performance across garment reconstruction, model-free try-on, and layered try-on, with high-fidelity and realistic outfitted images.

  • Conclusion: Any2AnyTryon provides a unified, mask-free solution for diverse virtual try-on scenarios.The framework uses LAION-Garment, adaptive position embedding, and clean condition latents.
  • Conclusion: Experiments report superior performance against state-of-the-art methods across garment reconstruction, model-free virtual try-on, and layered try-on.The reported outcomes include high-fidelity, realistic outfitted images under complex conditions.

A Validation Metrics

The validation uses complementary full-reference and perceptual metrics to assess reconstructed-garment alignment, generated-image quality, and fidelity.

  • Validation Metrics: Garment reconstruction uses SSIM, MS-SSIM, and CW-SSIM for alignment, plus LPIPS, FID, CLIP-FID, KID, and DISTS for image quality and fidelity.The metrics combine full-reference alignment measures with perceptual and distributional quality measures.

A.2 Virtual Try-on Genetation

The appendix evaluates model-free VTON and VTON generation with task-specific metrics and supplements the results with additional visualizations. These evaluations measure garment consistency, image faithfulness, and generation quality under paired and unpaired settings.

  • Model-free Virtual Try-on: Model-free VTON is evaluated with MP-LPIPS, CLIP-I, DiffSim, and FFA to measure consistency between garments and generated outfitted models.The evaluation follows MagiClothing and adds DiffSim and FFA as recent benchmarks.
  • VTON Generation: Paired VTON datasets use LPIPS, SSIM, FID, and KID, while unpaired datasets use FID and KID to assess generation quality.The metric choice reflects whether ground-truth images are available.

B.1 Model-free Virtual Try-on

Any2AnyTryon consistently generates high-fidelity model-free virtual try-on results for upper garments, lower garments, and complete outfit changes using the instruction “model in the shop.”

  • Any2AnyTryon consistently achieves high-fidelity model-free virtual try-on for upper, lower, and overall outfit changes.The results use the user instruction “model in the shop.”

B.2 Virtual Try-on

Any2AnyTryon produces rational, high-quality outfitted results across varied models and garments in the Virtual Try-on task. Across six models and six distinct-style garments, it generates 36 results.

  • 36 rational, high-quality outfitted results are generated from six models and six distinct-style garments.The combinations cover varied model-garment pairs in the Virtual Try-on task.
  • Any2AnyTryon stably realizes impressive virtual try-on generation.

B.3 Garment Reconstruction

Any2AnyTryon preserves input-garment details better than TryOffDiff and remains effective on challenging in-the-wild inputs. It also produces rational, high-quality results for virtual try-on in layers in complex scenes.

  • Garment Reconstruction: Any2AnyTryon preserves input models’ garment details better than TryOffDiff in garment reconstruction.
  • Garment Reconstruction: Any2AnyTryon generates impressive garment results for challenging in-the-wild model images.
  • Virtual Try-on in Layers: In complex in-the-wild scenarios, Any2AnyTryon produces rational, high-quality outfitted images for virtual try-on in layers.
Loading 2501.15891v2…