Source-linked AI summary

Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories

Junyao Hu, Zhongwei Cheng, Waikeung Wong, Xingxing Zou

arXiv:2603.14153v1cs.CV

TL;DR

Existing VTON datasets and systems remain limited for full outfits involving multiple garments, accessories, fine-grained categories, layering, and styling. Garments2Look introduces an 80K-pair multimodal outfit-level dataset built from real and synthesized data with structured annotations, then evaluates adapted VTON and image-editing baselines. The evaluation finds that current methods struggle with seamless complete-outfit try-on and correct layering and styling.

  • Problem

    Existing VTON datasets lack simultaneous support for diverse categories, accessories, multi-garment layering, and coherent outfit-level composition.

  • Method

    Garments2Look combines real and synthesized multi-reference garment-to-look pairs with item, layering, dressing-technique, and styling annotations, then benchmarks adapted VTON and editing models.

  • Results

    Current methods struggle to generate complete outfits seamlessly and to infer correct layering and styling, producing misalignment and artifacts.

  • Takeaways & Limitations

    Outfit-level VTON introduces challenges beyond single-item try-on, especially for variable-length outfits, structural consistency, and coordinated styling.

Abstract

from arXiv · show

Virtual try-on (VTON) has advanced single-garment visualization, yet real-world fashion centers on full outfits with multiple garments, accessories, fine-grained categories, layering, and diverse styling, remaining beyond current VTON systems. Existing datasets are category-limited and lack outfit diversity. We introduce Garments2Look, the first large-scale multimodal dataset for outfit-level VTON, comprising 80K many-garments-to-one-look pairs across 40 major categories and 300+ fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (Average 4.48), a model image wearing the outfit, and detailed item and try-on textual annotations. To balance authenticity and diversity, we propose a synthesis pipeline. It involves heuristically constructing outfit lists before generating try-on results, with the entire process subjected to strict automated filtering and human validation to ensure data quality. To probe task difficulty, we adapt SOTA VTON methods and general-purpose image editing models to establish baselines. Results show current methods struggle to try on complete outfits seamlessly and to infer correct layering and styling, leading to misalignment and artifacts.

1. Introduction

Outfit-level VTON must handle multiple garments, accessories, fine-grained categories, layering, and styling, but existing datasets and methods do not support these requirements comprehensively. Garments2Look addresses this gap with a multimodal dataset and a structured outfit-level VTON task.

  • Current VTON research has not comprehensively addressed multiple items, layered garments, fine-grained categories, and styling techniques.
  • The dataset comparison emphasizes outfit-level data paired with diverse clothing and accessories, including layering and styling information.
  • Existing datasets primarily target single garments, omit accessories and coordination annotations, or provide limited category diversity despite some multi-reference support.
  • Outfit-level VTON must represent variable layering, occlusion, and dressing techniques such as draping or cinching garments.
  • Garments2Look introduces an open-source multimodal dataset with 80K item-model pairs and a task using multiple reference items, matching relationships, and structured annotations.

2. Related Work

Prior image VTON datasets expanded resolution, garment coverage, or accessory categories, but generally remained limited in scope. Garments2Look instead targets outfit-level composition with diverse references, high-resolution outputs, and layering and styling annotations.

  • Earlier datasets variously improve resolution, expand garment types, or add accessory categories, while some address only specific multi-garment challenges.
  • Garments2Look supports outfit-level reference input, high-quality real-world data, approximately 1M-pixel outputs, and textual annotations for items, outfits, layering, and styling.
  • Multi-reference image-generation research commonly uses open-vocabulary models for layouts or segmentation and develops strategies to reduce copy-paste artifacts.
  • The dataset construction follows the broader use of advanced editing models for synthesis and filtering but specializes in full-outfit consistency, layering order, and styling techniques.

3. Garments2Look Dataset

Garments2Look combines real and synthesized outfit data through collection, synthesis, filtering, and evaluation, targeting broad category coverage and richly annotated outfit-level VTON pairs. Its pipeline constructs plausible outfits, retrieves items, generates look images, and applies quality screening.

  • Construction pipeline: The construction process comprises data collection, data synthesis, data filtering, and data evaluation for outfit-level VTON.
  • Data collection: The dataset combines naturally paired gold-standard data with unpaired garment and outfit sources to balance try-on fidelity against scale and diversity.
  • Outfit synthesis: Outfit synthesis heuristically constructs lists from style knowledge, user context, language-model generation, image retrieval, and re-weighted item sampling.
  • Outfit synthesis: The retrieval stage queries relevant category items and increases selection opportunities for historically underused items to improve corpus utilization and distribution uniformity.
  • Look synthesis: Look synthesis arranges outfit items in an OOTD grid for image generation, while visual-language models provide richer textual descriptions.
  • Filtering and evaluation: The dataset defines 40 primary clothing and accessory categories with 300+ fine-grained subcategories and applies expert-assisted screening to item images, outfit lists, and garment-look pairs.
  • Filtering and evaluation: The final 80K outfit-level pairs maintain an approximately 1:1 real-to-synthetic ratio and cover varied genders, outfit lengths, layering lengths, categories, combinations, and textual annotations.

4. Experiments

Experiments show that Garments2Look is difficult for current VTON and editing models, especially as outfit cardinality, layering, accessories, and styling complexity increase. Structured textual annotations improve generation quality and controllability, while existing methods still exhibit omissions, distortions, artifacts, and styling errors.

  • Experimental setup: Current models struggle with complete outfit try-on, including accessories, layering order, styling techniques, and fine-grained garment details.The dataset is designed to expose these limitations through qualitative and quantitative evaluation of VTON and general-purpose editing models.
  • Q1. How many items can be worn simultaneously?: When outfits exceed 4 items, fixed-label VTON models often drop garments or render only the outer layer, whereas editing models handle variable-length outfit lists more robustly.The results suggest that training-time input modality constrains outfit composition capacity.
  • Q2. How consistent are look images with references?: More references universally degrade consistency through accessory shape distortion, altered textures, color deviations, and fusion of independent garments.A holistic OOTD reference often outperforms disparate single-garment references because it preserves contextual co-occurrence and implicit relations.
  • Q3. How good is the overall try-on effect?: Layered and styled outfits remain problematic, producing missing details, incorrect garment geometry, distorted textures, inaccurate lengths, and tidy tucked outfits despite prompts for non-standard styling.These errors become more pronounced as the number of reference items grows.
  • Q4. Why do results differ across benchmarks?: Editing models outperform VTON models on Garments2Look because its broader categories and more complex multi-garment styling exceed the fixed category and single-layer focus of many VTON systems.Garments2Look is characterized as a substantial yet tractable challenge for future research.
  • 4.3. What new insights does it bring?: Adding outfit-level and increasingly fine-grained textual annotations yields measurable gains across evaluation metrics and improves style and pose controllability.The findings support combining structured textual guidance with visual features for outfit-level VTON.

5. Conclusions

Garments2Look addresses the dataset gap in outfit-level VTON by providing multimodal structural annotations and evaluating current methods. The evaluation reveals substantial shortcomings in generating complete outfit-level results.

  • Garments2Look provides 80K high-fidelity outfit-level VTON pairs with structural annotations for layering, accessories, and styling.
  • Existing datasets lack support for multi-garment layering, accessory integration, and fine-grained layering and styling annotations.
  • State-of-the-art methods show substantial shortcomings when generating complete outfit-level try-on results.
  • Future work will develop task-specific metrics and models that fuse visual and textual cues for precise end-to-end outfit-level VTON.

Supplementary Material

The supplementary material presents example data from Garments2Look, illustrating the dataset’s outfit-level samples.

  • Figure S1 shows example data from the Garments2Look outfit-level VTON dataset.

A.1. Visual Samples in Garments2Look

Garments2Look is designed around complex, diverse outfit combinations that reflect real-world virtual try-on scenarios. It combines visual coverage of garments, accessories, layering, and styling with detailed textual annotations reviewed by fashion experts.

  • The dataset contains high-quality, high-precision, and highly diverse samples of complex complete clothing combinations.
  • Its visual coverage ranges from basic items and accessories to single-layer and multi-layer outfits with multiple styling options.
  • Text annotations describe item categories, fine-grained attributes, overall outfits, wearing order, styling, body shape, and posture.
  • The dataset is intended to support deeper research and broader applications in virtual try-on.

A.2. Data Division

The supplementary material describes Garments2Look’s evaluation split and synthesis-related annotation procedures. Sampling uses inverse-frequency weighting, while look-image annotations and layering examples are generated and reviewed to represent diverse outfit structures.

  • Data Division: Approximately 1K stratified samples form the test set, while the remaining samples constitute training data.
  • Data Division: Table S1 presents the division of Garments2Look.
  • Outfit Synthesis: Inverse-frequency weighting encourages selection of historically underrepresented items during outfit sampling.
  • Text Annotation: Look-image annotations cover overall appearance, item wearing and layering, styling techniques, model attributes, pose, and background.
  • Layering & Styling: Figure S2 illustrates how the same garment can appear under different layering arrangements.

B.3. Data Filtering

The dataset metadata includes a manually corrected category structure spanning 40 primary fashion categories.

  • 40 primary categories are listed, covering garments, footwear, jewelry, accessories, and other fashion items.

C.1. Split Results

Split evaluations compare methods across Garments2Look subsets and examine reference inputs and limited fine-tuning. Nano Banana remains strong across splits, while skeleton references add little and fine-tuning improves QIE-2509.

  • Split evaluation: Nano Banana series maintain SOTA performance across both Garments2Look test subsets, outperforming all baselines.
  • Reference ablation: Adding a skeleton image as a third reference did not significantly improve image-editing metrics and slightly decreased performance.The extra pose reference may be difficult for general image-editing models to disambiguate from garment references.
  • Fine-tuning: Fine-tuning Qwen-Image-Edit-2509 on 10K randomly sampled training examples improved quantitative, qualitative, and user-study results.The user study focused on layering quality, including occlusion rationality and layering accuracy.

C.5. More Examples and Results

Supplementary figures provide additional dataset examples and model comparisons, illustrating the annotations, task difficulty, and dataset utility for outfit-level VTON.

  • Additional data examples: Figures S5–S9 show sample data from multiple datasets with rich annotations for Garments2Look.
  • Additional comparisons: Figures S10–S14 present additional generated results from SOTA VTON, general-purpose editing, and fine-tuned models.
  • Interpretation: The supplementary comparisons illustrate outfit-level VTON performance, the inherent difficulty of the task, and the dataset’s utility.
  • Prompt guidance: Table S4 provides the prompt template used for minimalist style guidance.
Loading 2603.14153v1…