Source-linked AI summary

Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items

Mengting Chen, Zhengrui Chen, Yongchao Du, Zuan Gao, Taihang Hu, Jinsong Lan, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Mingzhou Zhang, Jun Zheng, Qinye Zhou, Xiaoyong Zhu, Bo Zheng

arXiv:2604.19748v3cs.CV

TL;DR

Existing virtual try-on methods remain limited for challenging real-world images, detailed garment preservation, and complex multi-garment coordination. Tstars-Tryon 1.0 addresses these demands through an integrated commercial-scale pipeline spanning architecture, data, training, and inference optimization. The authors report robust, realistic, flexible, and efficient performance, including stability on complex multi-garment tasks.

  • Problem

    Existing virtual try-on methods face limitations in robustness, garment-detail preservation, and flexible multi-image, cross-category generation for commercial-grade applications.

  • Method

    Tstars-Tryon 1.0 combines a full-stack pipeline with scalable data curation, unified MMDiT processing, specialized training, and inference optimization.

  • Results

    Tstars-Tryon 1.0 maintains stability and achieves the highest overall results among tested models on complex multi-garment try-on while preserving garment fidelity and structural logic.

  • Takeaways & Limitations

    The system supports robust multi-garment coordination across diverse conditions and provides a benchmark designed to assess virtual try-on under commercial standards.

  • Takeaways & Limitations

    Existing academic benchmarks use homogeneous backgrounds and restricted garment categories, limiting their reflection of real-world deployment conditions.

Abstract

from arXiv · show

Recent advances in image generation and editing have opened new opportunities for virtual try-on. However, existing methods still struggle to meet complex real-world demands. We present Tstars-Tryon 1.0, a commercial-scale virtual try-on system that is robust, realistic, versatile, and highly efficient. First, our system maintains a high success rate across challenging cases like extreme poses, severe illumination variations, motion blur, and other in-the-wild conditions. Second, it delivers highly photorealistic results with fine-grained details, faithfully preserving garment texture, material properties, and structural characteristics, while largely avoiding common AI-generated artifacts. Third, beyond apparel try-on, our model supports flexible multi-image composition (up to 6 reference images) across 8 fashion categories, with coordinated control over person identity and background. Fourth, to overcome the latency bottlenecks of commercial deployment, our system is heavily optimized for inference speed, delivering near real-time generation for a seamless user experience. These capabilities are enabled by an integrated system design spanning end-to-end model architecture, a scalable data engine, robust infrastructure, and a multi-stage training paradigm. Extensive evaluation and large-scale product deployment demonstrate that Tstars-Tryon1.0 achieves leading overall performance. To support future research, we also release a comprehensive benchmark. The model has been deployed at an industrial scale on the Taobao App, serving millions of users with tens of millions of requests.

1 Introduction

Tstars-Tryon 1.0 targets commercial virtual try-on by combining robustness, fidelity, flexible multi-item editing, and efficient inference in an integrated system. It is evaluated with a business-oriented benchmark and reported to outperform leading models across key dimensions.

  • Commercial virtual try-on must handle arbitrary user photos, preserve garment details, support multi-item styling, and generate results near real time.
  • Tstars-Tryon 1.0 reformulates the full-stack pipeline across data curation, model architecture, training strategies, and inference optimization.The system uses a scalable data engine, unified MMDiT architecture, progressive training strategies, prompt enhancement, and distilled inference.
  • 3.92 seconds single-garment and 6.74 seconds multi-garment try-on are achieved using a streamlined 5B-parameter model with CFG and step distillation.The multi-garment result uses five reference images on average and is reported without compromising visual fidelity.
  • The system supports diverse poses, extreme lighting, and complex garment combinations while preserving identity and intricate clothing textures.
  • Multi-item generation spans eight fashion categories and supports non-photorealistic inputs such as digital humans and anime characters.

2 Tstars-VTON Benchmark

Tstars-VTON Benchmark is designed for commercial-standard evaluation by covering realistic multi-item try-on conditions, diverse data, and human-aligned quality assessment. Its construction combines broad collection, refinement and anonymization, physically plausible pairing, and VLM-based scoring.

  • Benchmark scope: 1780 paired samples span 5 garment categories, 3 accessory categories, 465 fine-grained subcategories, and 1–6 layered try-on items.The benchmark explicitly includes multi-garment layering, complex backgrounds, and diverse human poses.
  • Limitations of academic benchmarks: Existing academic benchmarks underrepresent real deployment by using homogeneous backgrounds, limited garment categories, and mainly basic tops, bottoms, and dresses.The paper contrasts these constraints with broader e-commerce and daily-life fashion settings that include outerwear, footwear, bags, and accessories.
  • Benchmark principles: Tstars-VTON supports 1–6 item free combinations, including layered garments and accessories, to capture real-world dressing complexity.This expands evaluation beyond the single-garment focus of most prior benchmarks.
  • Data construction: The benchmark collects diverse Internet and e-commerce data and uses tag-based retrieval, VLM refinement, manual checking, and controllable sampling across fine-grained attributes.The data process addresses source-distribution bias and supports structured coverage of model and garment characteristics.
  • Data construction: Privacy protection replaces open-source portrait identities with matched licensed surrogate faces through face swapping, followed by automated filtering and human inspection.Surrogates are matched using attributes such as skin tone, gender, and age, with failed cases iteratively corrected.
  • Data construction: A three-stage pipeline combines broad collection, quality filtering and anonymization, and physically and semantically constrained try-on pairing.The pairing strategy maximizes matching diversity while enforcing rules such as gender compatibility, layering logic, and unique image utilization.
  • Evaluation paradigm: A VLM-driven protocol scores try-on quality on four semantically distinct dimensions using 1–10 Likert scales.The protocol uses separate image inputs for garment-aware evaluation and assesses the relationship between target items and the subject, including identity consistency and garment fidelity.

3 Evaluation Results

Tstars-Tryon 1.0 achieves strong results on commercial and academic evaluations, maintaining stability as try-on complexity increases. Human evaluations likewise favor it over leading commercial models, especially for multi-garment tasks.

  • Quantitative Results: Tstars-Tryon 1.0 is evaluated on the Tstars-VTON Benchmark across single- and multi-garment scenarios designed to reflect complex real-world conditions.The benchmark emphasizes pose diversity, background complexity, and high-fidelity texture requirements.
  • Quantitative Results: Tstars-Tryon 1.0 consistently achieves superior or competitive performance across garment fidelity, identity consistency, and background preservation against leading proprietary models.The reported advantage is especially clear in capturing intricate textures, material drapes, and fine patterns.
  • Quantitative Results: General-purpose editing models often collapse in multi-garment tasks by omitting items, mishandling layering, or degrading identity and overall structure.These failures become more severe as the number of visual conditions increases.
  • Evaluation Caveat: GPT-Image-1.5 and GPT-Image-2 failed to generate 168 and 134 test cases, respectively, and reported metrics exclude those missing instances.The failures are attributed to platform restrictions.
  • Quantitative Results: Tstars-Tryon 1.0 maintains stability and achieves the highest overall results in multi-garment evaluation while preserving individual items and physical and structural logic.The system uses a unified MMDiT architecture and specialized multi-garment training pipeline.
  • Academic Benchmarks: On VITON-HD and DressCode, Tstars-Tryon 1.0 achieves state-of-the-art quantitative results without training data from either benchmark, indicating zero-shot generalization to unseen distributions.The comparison uses the unpaired setting, which is presented as more representative of practical try-on generalization.
  • Human Evaluation: In human evaluation, Tstars-Tryon 1.0 wins 41.1% against Nano Banana Pro, 41.9% against GPT-Image-2, and 54.4% against Seedream5 lite in overall preference.The corresponding competitor preference rates are 17.3%, 15.5%, and 9.0%, respectively.
  • Human Evaluation: As garment count rises from one to five, win rates increase from 33.6% to 54.8% against Nano Banana Pro, from 36.4% to 50.0% against GPT-Image-2, and from 46.1% to 70.2% against Seedream5 lite.The widening gap indicates stronger relative preference in higher-complexity scenarios.

4 Demonstrations

The demonstrations show Tstars-Tryon 1.0 handling single- and multi-garment try-on across difficult poses, lighting, backgrounds, identities, and subject types. It also supports layered outfits, accessories, holistic outfit transfer, and out-of-domain subjects.

  • Single-Garment Try-On: Tstars-Tryon 1.0 maps flat-lay garments onto crouching and swinging poses while preserving material textures and lighting.The examples emphasize robustness across varied perspectives and input conditions.
  • Multi-Garment Try-On: The system composes coordinated multi-item outfits with accessories across users and environments while retaining poses, body types, and complex backgrounds.Examples include jackets, shorts, caps, sneakers, hats, and crossbody bags.
  • Instruction Following: Tstars-Tryon 1.0 follows complex garment and layering instructions, including keeping an outer coat open to reveal the inner layer.The demonstrated instruction combines garments, footwear, bags, and headwear while preserving pose and background.
  • Multi-Condition Try-On: The model handles low-light neon scenes, satin and lace materials, unconventional orientations, and lying-down perspectives without losing global visual consistency.These examples target heterogeneous lighting and challenging geometric configurations.
  • OOTD Swap: Holistic OOTD swapping transfers complete ensembles between different individuals while adapting garments to different body types and challenging postures.The transferred ensemble includes tops, bottoms, outerwear, footwear, and accessories.
  • Out-of-Domain Subjects: The system extends try-on beyond standard human photography to stylized 3D characters with non-standard body proportions.The demonstration describes semantic generalization to out-of-domain subjects.

5 Industrial-Scale Deployment

Tstars-Tryon 1.0 is deployed in Taobao’s consumer-facing AI Try-On service, supporting apparel try-on and user-composed multi-garment outfits. The deployment has reached millions of users and tens of millions of requests, demonstrating large-scale operational use.

  • Consumer Service: Taobao users can initiate AI Try-On from the shopping assistant or product detail pages after uploading a personal portrait.The service supports a wide range of apparel and freely composed DIY multi-garment outfits.
  • Service Capabilities: The demonstrations include versatile multi-item synthesis under heterogeneous lighting, unconventional perspectives, and multi-subject interactions.These capabilities complement the consumer service’s support for multi-garment composition.
  • Production Scale: The system has served several million users and fulfilled tens of millions of try-on requests in production.The paper presents this scale as evidence of robustness, scalability, and commercial viability under industrial workloads.
  • Commercialization: The authors report that deployment addresses the trade-off between C-end serving cost and generation quality, moving virtual try-on from a research prototype to a consumer-facing product.A broader rollout is planned for the Taobao user base.
Loading 2604.19748v3…