Source-linked AI summary

Beyond Global Realism: Virtual Try-On Evaluation and Optimization with Dimension-wise Garment Fidelity Assessment

Kaidong Zhang, Yukang Ding, Xiaoyu Liu, Ying Chen

arXiv:2608.29804v1cs.CV

TL;DR

VTON evaluation needs to measure garment fidelity across attributes that global image metrics miss. DAT provides a seven-dimensional assessment model trained with staged supervision and imbalance-aware learning, and it also supplies adaptive rewards for Qwen-Image-Edit optimization. The 8B model outperforms proprietary judges on balanced accuracy, SROCC, and PLCC while improving reward-aligned garment preservation.

  • Problem

    Global image metrics struggle to capture the multi-dimensional, fine-grained consistency between generated and reference garments in virtual try-on.

  • Method

    DAT predicts discrete attributes across seven garment-fidelity dimensions, using two-stage supervision, weighted cross-entropy, and adaptive dimension-wise rewards for reinforcement learning.

  • Results

    DAT outperforms Gemini-3.1, Qwen3.7-plus, and GPT-5.5 on balanced accuracy, SROCC, and PLCC while using 8B parameters and improves reward-aligned garment preservation.

  • Takeaways & Limitations

    A domain-specialized, interpretable assessment model can evaluate garment fidelity and serve as an optimization signal for more faithful virtual try-on generation.

Abstract

from arXiv · show

Virtual try-on (VTON) requires not only realistic generation but also faithful preservation of garment characteristics. However, existing evaluation metrics such as PSNR, SSIM, KID and FID struggle to measure the consistency between the generated and reference garments, particularly in capturing the multi-dimensional characteristics of garment fidelity. To address this, we propose DAT: a Dimension-wise Assessment framework for virtual Try-on, which decomposes garment consistency into seven interpretable dimensions: silhouette, color, neckline and sleeve shape, major decoration and structure, material texture, fine-detail fidelity, and logo preservation, each formulated as a discrete attribute-level prediction task. To train this specialized assessment model, we adopt a two-stage learning paradigm comprising large-scale weak supervision on 50K samples, followed by refinement on 10K higher-quality annotations obtained via multi-model voting. Furthermore, we employ weighted cross-entropy loss to mitigate the severe label imbalance inherent across evaluation dimensions. Beyond its role as an evaluation framework, the assessment model can be integrated into reinforcement learning optimization of Qwen-Image-Edit for VTON, where dimension-wise rewards are adaptively aggregated to emphasize under-optimized aspects during training. Experimental results show that our method (8B parameters) achieves state-of-the-art performance in terms of balanced accuracy, SROCC, and PLCC, outperforming strong proprietary models such as Gemini-3.1, Qwen3.7-plus, and GPT-5.5, while also serving as an effective optimization signal for reward-guided VTON generation

Introduction

VTON evaluation must assess garment fidelity beyond global visual realism because consistency spans multiple attributes. DAT addresses this gap with interpretable dimension-wise assessment, imbalance-aware training, and reward-guided optimization that improves garment-preserving generation.

  • Global image metrics overlook fine-grained garment attributes, even though deviations can reduce the commercial reliability of virtual try-on results.
  • DAT evaluates garment consistency across seven interpretable dimensions using discrete attribute-level prediction tasks, with an additional not-applicable category for absent attributes.The dimensions cover silhouette, color, neckline and sleeve shape, major decoration and structure, material texture, fine-detail fidelity, and logo preservation.
  • DAT training combines 50K weakly labeled samples, 10K multi-model-voted refinement samples, and weighted cross-entropy to address annotation noise and label imbalance.The weak labels come from Gemini-3.1, while refinement uses voting among Gemini-3.1, Qwen3.7-plus, and GPT-5.5.
  • Adaptive aggregation of DAT’s dimension-wise rewards emphasizes under-optimized garment aspects during reinforcement learning optimization of Qwen-Image-Edit.This replaces simple reward averaging with dynamically adjusted dimension weights.
  • DAT outperforms proprietary judges on balanced accuracy, SROCC, and PLCC while using an 8B-parameter model, and reward-guided optimization further improves garment preservation.The compared judges include Gemini-3.1, Qwen3.7-plus, and GPT-5.5.

Related Work

Prior VTON methods and metrics emphasize generation and global consistency, whereas fine-grained garment identity assessment remains less developed. DAT complements this work by focusing specifically on localized garment-consistency evaluation and reward modeling.

  • Early VTON methods transfer garments using pose estimation, human parsing, geometric warping, and garment-to-body alignment.
  • Existing image metrics capture global consistency but often overlook localized discrepancies and fine-grained garment identity.
  • Reward modeling and preference learning commonly derive rewards from pairwise comparisons, rankings, or model-generated annotations across generative and multimodal tasks.

Method

DAT formulates VTON assessment as seven-dimensional attribute prediction rather than holistic scoring, then trains it with staged supervision and imbalance-aware optimization. The resulting model also supports adaptive reward-guided reinforcement learning for garment-preserving generation.

  • Assessment formulation: DAT predicts garment consistency across seven dimensions: silhouette, color, neckline and sleeve shape, major decoration and structure, material texture, fine-detail fidelity, and logo preservation.Fine-detail fidelity and logo preservation additionally support a not-applicable category when the attribute is absent.
  • Assessment training: The training framework addresses the tension between annotation scale and fidelity through weak supervision followed by consensus-based refinement.Stage 1 uses broad, noisy annotations for initialization; Stage 2 uses multi-model voting to improve reliability and calibration.
  • Imbalance-aware optimization: Weighted cross-entropy handles skewed score distributions without resampling, which could distort the joint distributions across the seven simultaneously labeled dimensions.Inverse-frequency weights emphasize rare categories, while clipping limits unstable updates; supervision is applied only to valid labels.
  • Reward-guided optimization: DAT supplies dimension-wise rewards for reinforcement learning optimization of Qwen-Image-Edit through adaptive aggregation that emphasizes under-optimized dimensions.The reward-guided pipeline uses DAT as its reward model and dynamically adjusts dimension weights during training.
  • Assessment results: DAT outperforms the baseline and proprietary judges across assessment metrics, with nearly 20% improvement over Qwen3-VL-8B while retaining an 8B-scale architecture.It performs best across all metrics on out-of-domain sources and maintains more balanced dimension-wise performance.
  • Optimization results: Reward-guided optimization yields an average 0.35 gain in seven-dimensional garment-consistency scores and improves performance consistently across dimensions in a blind human study.The study compares base and reward-guided Qwen-Image-Edit results across overall similarity and seven garment-fidelity aspects.

Conclusion

The paper introduces a fine-grained framework for assessing garment consistency in virtual try-on and reports improved evaluation and optimization outcomes. DAT outperforms proprietary judges and improves garment preservation when used for reinforcement learning.

  • DAT decomposes garment fidelity into seven interpretable dimensions, enabling more informative attribute-level evaluation than conventional metrics.
  • DAT outperforms Gemini-3.1, Qwen3.7-plus, and GPT-5.5 in binary accuracy and correlation with human annotations.
  • Using DAT as a reward signal for reinforcement learning optimization of Qwen-Image-Edit improves garment preservation and produces more faithful try-on results.

Supplementary Material: Beyond Global Realism: Virtual Try-On Evaluation and

The supplementary material details DAT’s seven-dimensional protocol, data preparation, and human-comparison rationale. The protocol uses five-point dimension ratings with N/A handling and supports fine-grained diagnosis of garment fidelity.

  • Assessment protocol: DAT compares source and generated garments across silhouette, color, neckline and sleeve shape, decoration and structure, material texture, fine details, and logos.
  • Assessment protocol: The protocol distinguishes core identity, style-defining structure, material appearance, small semantic details, and commercially important logos.
  • Assessment protocol: Each dimension receives a 1–5 consistency rating, while non-applicable dimensions are marked N/A and excluded from dimension-specific analysis.
  • Assessment protocol: Fine-detail and logo preservation are assessed only when relevant non-logo elements or clearly visible logos are present.
  • Human evaluation: The resulting benchmark reveals which garment-fidelity aspects current virtual try-on methods preserve well and which remain challenging.
  • Data preparation: The dataset filters corrupted, low-resolution, blurry, multi-product, and multi-person images before constructing synthetic garment-person pairs.
  • Data preparation: Synthetic try-on results are generated from randomly paired garment and person images using multiple image-editing models and automatically generated instructions.
  • Human evaluation: Pairwise human judgments provide fine-grained comparisons of which system better preserves identity-defining garment attributes.
Loading 2608.29804v1…