Source-linked AI summary

Beyond the Mirror: Balancing Interaction Modality and Avatar Fidelity in Public 3D Virtual Try-On Systems

Yueqian Guo, Tianzhao Li, Xin Lv

arXiv:2608.23345v1cs.HC

TL;DR

Public VTON systems face physical fatigue from mid-air interaction and social inhibition from being observed. The paper develops a real-time 3D avatar apparatus and evaluates these barriers through two studies, finding that latency shapes fatigue while avatar fidelity trades off social comfort and trust. The authors conclude that adaptive fidelity can balance privacy, immersion, and commercial trust within the studied public-interaction context.

  • Problem

    Public VTON must address both mid-air interaction fatigue and social inhibition, while existing work has rarely examined latency and avatar appearance together in public spaces.

  • Method

    The paper uses a real-time 3D avatar system with markerless motion capture, programmable latency control, and independently swappable avatar fidelity in two empirical studies.

  • Results

    Study 1 found that low-latency gestures eliminated physical fatigue while providing better hygiene and higher immersion than touch; social acceptability remained lower at M = 3.4 than touch at M = 3.8.

  • Takeaways & Limitations

    Adaptive fidelity is proposed to balance user privacy and commercial trust in public interactive systems.

  • Takeaways & Limitations

    Both studies used simulated laboratory environments, where authentic bystander effects may be more intense in crowded retail settings.

Abstract

from arXiv · show

Virtual Try-On (VTON) systems deployed on large public displays face a dual barrier: the physical strain of mid-air interaction and the social inhibition caused by public self-consciousness. This paper presents a real-time 3D avatar system integrating markerless motion capture with dynamic visual fidelity control to investigate and mitigate both barriers. Through a dual-study empirical evaluation, we first decoupled physical fatigue from gesture interaction ($N=20$), demonstrating that interaction fatigue is primarily driven by visuomotor latency rather than the physical act of gesturing; our optimized low-latency gesture pipeline achieved usability comparable to touchscreens while delivering superior immersion and hygiene. Building on these insights, our second study ($N=25$) investigated the "avatar fidelity paradox" via a $2 \times 2$ factorial design manipulating interaction modality (gestures vs. touch) and visual fidelity (photorealistic MetaHuman vs. stylized mannequin). Results reveal that while high fidelity and mid-air gestures independently maximize virtual embodiment ($p < .05$), their combination elicits the highest social awkwardness. Crucially, low-fidelity avatars serve as a "psychological mask" that alleviates public embarrassment during expressive gestures, while mid-air gestures simultaneously act as a compensatory mechanism to preserve perceived try-on trust despite reduced visual realism. Finally, we propose a context-aware fidelity framework to balance privacy, immersion, and commercial trust in public spatial interactions.

I. INTRODUCTION

Public VTON systems must address both the physical fatigue of mid-air interaction and the social inhibition of being observed. This paper uses a dual-study real-time 3D avatar system to examine latency, interaction modality, and avatar appearance as joint contributors to these barriers.

  • Motivation: Public mid-air VTON interaction combines physical fatigue with social awkwardness in semi-open retail settings.Users stand before large displays while potentially watched by passersby.
  • Physical barrier: Visuomotor mismatch from system latency is proposed as a major source of gesture-related fatigue and cognitive load.Latency disrupts synchrony between physical movement and avatar feedback.
  • Social barrier: The avatar fidelity paradox describes realism that can increase product trust while also amplifying public self-consciousness.A realistic avatar may make users feel more exposed during visible gestures.
  • Research approach: The dual-study design compares touch, low-latency gestures, and high-latency gestures, then varies avatar fidelity to examine social inhibition and try-on trust.Study 1 targets physical fatigue and immersion; Study 2 targets appearance-related social effects.
  • Contribution: The paper proposes adaptive fidelity as a design guideline for balancing privacy and commercial trust in public interactive systems.This addresses the paper’s stated physical and social barriers together.

B. Visuomotor Synchrony and Physical Fatigue

The paper frames visuomotor synchrony as central to physical fatigue and develops a controlled real-time apparatus to vary latency independently. It also connects this physical barrier to the remaining social challenge of public avatar exposure.

  • Visuomotor synchrony: Visuomotor synchrony is the precise temporal match between a user’s movement and the avatar’s visual feedback.The paper identifies this match as a prerequisite for agency and full-body illusion.
  • Physical fatigue: Latency-induced sensory mismatch is linked to broken embodiment, increased cognitive load, and exacerbated perceived fatigue.Study 1 addresses the stated gap concerning prolonged mid-air gestures in public retail contexts.
  • Remaining barrier: Reducing physical fatigue does not remove the separate social inhibition associated with using gestures in public.The paper therefore treats physical and identity exposure as distinct barriers.
  • System apparatus: The apparatus uses markerless pose estimation to map camera-derived 3D skeletal data onto a digital human avatar in real time.MediaPipe extracts skeletal data from a single RGB camera feed.
  • Latency control: The Motion Codec includes a programmable delay buffer that injects a precise 0.5-second delay without changing frame rate or rendering quality.This isolates visuomotor latency as the experimental manipulation.

C. Modality and Fidelity Control (Targeting the Social Barrier)

Study 2’s apparatus supports independent control of interaction modality and avatar fidelity while preserving motion, interface state, and clothing physics. Its personalized digital twin and generic mannequin operationalize identity exposure versus psychological masking.

  • Experimental control: The apparatus supports a 2×2 matrix by swapping avatars while holding motion capture, UI state, and clothing physics constant.This isolates interaction modality and visual fidelity as experimental factors.
  • Avatar conditions: High-fidelity avatars personalize gender and body proportions to create a digital twin matching the current user.This establishes identity exposure for the high-fidelity condition.
  • Avatar conditions: Low-fidelity avatars remove personalized facial features and skin textures, rendering a generic faceless mannequin while preserving motion tracking and cloth simulation.The mannequin functions as the paper’s psychological mask.
  • Study 1: Study 1 was designed to evaluate visuomotor latency’s effects on physical fatigue and usability using a controlled public-display apparatus.The study’s primary objective was empirical evaluation of the proposed mid-air gesture framework.
  • Study 1: Study 1 hypothesized that high latency would increase fatigue and reduce usability, while optimized gestures would match touch comfort and improve perceived hygiene.Conditions A, B, and C represent low-latency gestures, touch, and high-latency gestures.

C. Experimental Design and Procedure

Study 1 uses a counterbalanced within-subjects comparison of low-latency gestures, touch, and gestures with an injected 0.5-second delay. Participants complete standardized VTON tasks while the system records time, errors, and user feedback.

  • Design: The within-subjects design compares three modalities: low-latency gestures, touch-based input, and high-latency gestures.The high-latency condition adds a 0.5-second visuomotor delay.
  • Design: Condition order was counterbalanced to mitigate learning and fatigue effects.Each participant experienced all three conditions.
  • Procedure: Participants browsed and tried on three garments, searched for a specified item, and coordinated a complementary outfit.These tasks covered browsing, targeted search, and outfit coordination.
  • Measures: The evaluation collected objective performance data and subjective user feedback across the interaction modalities.The system automatically logged task completion time and operation error rate.
  • Measures: Task completion time measures total task duration, while operation error rate counts incorrect gestures or failed touch registrations.The error definition differs by interaction condition.

2) Subjective Usability and Ergonomic Metrics:

Study 1 evaluated usability, ease of use, fatigue, hygiene, immersion, and social acceptability across touch and gesture conditions using questionnaires and performance measures. Low-latency gestures were broadly viable, while high latency degraded performance and subjective experience.

  • The questionnaire measured SUS, perceived ease of use and naturalness, physical fatigue, hygiene, immersion, and social acceptability.
  • Cronbach’s Alpha was 0.809 and KMO was 0.824, indicating good internal consistency and suitability for analysis.
  • Condition B (Touch) was fastest at M = 171.40 seconds, while Condition A (Low-Latency Gesture) reached M = 234.50 seconds.
  • Condition B produced the fewest operation errors at M = 2.45; Condition A generated 4.35 more errors on average, while Condition C generated 10.35 more.

1) Objective Performance: Time and Errors:

Objective and subjective results show that low-latency gestures reduce the physical costs associated with mid-air interaction while improving immersion and hygiene. However, social acceptability remained lower than with touch, leaving a distinct public-use barrier.

  • Low-latency gestures and touch both scored highly for SUS and perceived ease of use, and both exceeded high-latency gestures.
  • Condition A produced the highest immersion because low-delay body tracking preserved an intuitive mirror-like experience, unlike Condition C.
  • Both low-latency gestures and touch produced low physical-fatigue scores with no significant difference, whereas high-latency gestures produced significantly higher fatigue.
  • Touchless conditions scored higher on perceived hygiene than touch, whose display showed visible fingerprints after repeated use.
  • Social acceptability was lower for low-latency gestures than touch, with M = 3.4 versus M = 3.8, because visible arm gestures induced self-consciousness.
  • Resolving latency and fatigue did not remove the psychological pressure of performing gestures publicly, motivating Study 2’s fidelity manipulation.

V. STUDY 2: THE INTERPLAY OF INTERACTION MODALITY AND AVATAR FIDELITY ON SOCIAL INHIBITION

Study 2 addressed the remaining social barrier by examining how interaction modality and avatar fidelity jointly shape public embarrassment. It framed avatar realism as a potential trust benefit and social liability.

  • Study 2 followed Study 1’s physical-barrier findings and focused on public embarrassment as the unresolved social barrier.
  • The study manipulated physical exposure through interaction modality and identity exposure through avatar fidelity.
  • The high-fidelity avatar was hypothesized to produce stronger virtual embodiment and try-on trust regardless of modality.
  • The gesture-plus-high-fidelity combination was hypothesized to produce the highest social embarrassment through simultaneous physical and identity exposure.
  • Replacing the high-fidelity avatar with a low-fidelity avatar was hypothesized to reduce social inhibition during mid-air gestures.
  • Twenty-five participants imagined using the system in a crowded shopping-mall atrium while completing standardized try-on tasks.

C. Experimental Design

Study 2 used a counterbalanced within-subjects 2 × 2 design crossing interaction modality with avatar visual fidelity. The design measured embodiment, social inhibition, and perceived try-on trust with repeated-measures analyses.

  • The four conditions crossed mid-air gesture versus direct touch with high-fidelity versus low-fidelity avatars.
  • Condition order was counterbalanced using a Latin Square design to reduce order effects.
  • High-fidelity avatars were personalized photorealistic MetaHumans, whereas low-fidelity avatars were featureless mannequins preserving clothing-simulation geometry.
  • The conditions were High-Fidelity + Gesture, Low-Fidelity + Gesture, High-Fidelity + Touch, and Low-Fidelity + Touch.
  • A 9-item, 7-point Likert questionnaire assessed virtual embodiment, social inhibition, and perceived try-on trust.
  • Two-way repeated-measures ANOVAs evaluated modality and fidelity effects for each construct, with sphericity inherently satisfied because each factor had two levels.

1) Virtual Embodiment (VE):

Virtual embodiment increased independently with both mid-air gestures and high-fidelity avatars, while their interaction was not significant. Gestures strengthened agency, and photorealistic avatars enhanced self-identification.

  • F(1, 24) = 8.74, p = .007: Mid-air gestures elicited significantly stronger virtual embodiment than direct touch.
  • F(1, 24) = 5.47, p = .028: High-fidelity avatars yielded higher virtual embodiment scores.
  • No significant interaction effect indicated that modality and fidelity contributed independently to virtual embodiment.
  • Gestures provided agency through their performative interaction, while the photorealistic MetaHuman avatar enhanced self-identification.
  • The highest virtual presence resulted from combining visual realism with kinematic synchrony in spatial systems.

2) Try-on Trust: Gestures as a Compensatory Mechanism:

Gesture interaction partially preserved try-on trust when visual fidelity was reduced, while high-fidelity avatars paired with gestures produced the highest descriptive social awkwardness. Qualitative accounts connect these patterns to agency and anonymity.

  • p = .074: Low-fidelity avatars reduced try-on trust more under direct touch than under mid-air gestures.
  • Mid-air gestures mitigated the trust deficit associated with low visual fidelity, suggesting kinematic feedback can compensate for reduced photorealism.
  • The High-Fidelity + Gesture condition produced the highest descriptive levels of social awkwardness, despite non-significant social-inhibition effects.
  • Replacing a personalized avatar with a low-fidelity mannequin descriptively reduced embarrassment during mid-air gestures.
  • Interviews linked high-fidelity gesture awkwardness to over-exposure and low-fidelity gesture trust to instant, controllable movement.
  • Study 2 jointly manipulated interaction modality and avatar realism to examine physical and identity exposure in public spatial interfaces.

A. Key Findings Synthesized

Across the two studies, low latency addressed physical fatigue, while embodiment and social acceptability required balancing interaction agency with avatar fidelity. The authors therefore recommend context-aware adaptation, while noting laboratory and task-complexity limits.

  • Key Findings Synthesized: System latency, rather than holding arms aloft, was identified as the main driver of perceived physical fatigue.
  • Key Findings Synthesized: The optimized low-latency gesture system achieved usability comparable to a mature touch interface while offering superior immersion and hygiene.
  • Key Findings Synthesized: Visual realism and kinematic synchrony independently and additively boosted virtual embodiment.
  • Implications for Public Spatial Interaction Design: Minimizing processing delay is presented as necessary for preventing cognitive mismatch, reducing fatigue, and establishing agency.
  • Implications for Public Spatial Interaction Design: Touchless gesture interaction offers practical hygiene benefits and supports a magic-mirror illusion with real-time 3D tracking.
  • Implications for Public Spatial Interaction Design: A context-aware fidelity strategy uses high-fidelity avatars privately and low-fidelity psychological masks in highly public, high-exposure settings.
  • Limitations and Future Work: The studies were conducted in simulated laboratory settings, where authentic bystander pressure may be less intense than in crowded retail environments.
  • Limitations and Future Work: Because experimental tasks were relatively straightforward, modality differences might widen during more complex, prolonged interactions.
Loading 2608.23345v1…