Source-linked AI summary

UiAs: User-Independent 3D Facial Anti-Spoofing via Multi-modal Wireless Signals

Zhiwei chen, Lebin Lyu, Yimo Zhang, Dingyu Zhong, Yijie Li, Yichao Chen, Dian Ding, Jiguo Yu, Xiaosong Zhang, Yongzhao Zhang

arXiv:2608.29084v1cs.CR

TL;DR

High-fidelity 3D spoofing attacks challenge face authentication because wireless liveness cues remain entangled with user-dependent facial geometry. UIAS combines mmWave and acoustic sensing with cross-modal suppression and skin-anchored contrastive learning, achieving 93.25% accuracy for unseen users without user-specific physical-signal enrollment.

  • Problem

    Wireless liveness cues are entangled with user-dependent facial geometry, limiting cross-user generalization and deployment scalability.

  • Method

    UIAS combines co-located mmWave and acoustic sensing to suppress shared geometry variations while preserving complementary liveness cues, then uses skin-anchored contrastive learning for material variations.

  • Results

    93.25% detection accuracy was achieved under unseen-user evaluations with real 3D spoofing masks and multi-material conditions, outperforming baselines.

  • Takeaways & Limitations

    UIAS enables liveness detection without user-specific physical-signal enrollment under the evaluated unseen-user settings.

  • Takeaways & Limitations

    UIAS may be affected by wireless signal injection and sensor distribution shifts across manufacturers or models in large-scale deployment.

Abstract

from arXiv · show

Face authentication is widely deployed in security-sensitive applications, while increasingly realistic 3D spoofing attacks pose growing threats. High-fidelity 3D masks can reproduce facial appearance and geometry but cannot replicate the intrinsic physical responses of living tissue, which can be actively probed by wireless signals. However, the resulting liveness cues captured by wireless signals are entangled with user-dependent facial geometry, limiting cross-user generalization. We present UiAs, a multimodal user-independent 3D facial anti-spoofing system using electromagnetic (mmWave) and mechanical (acoustic) waves. The two modalities share similar user-dependent geometric variations, allowing UiAs to suppress them through cross-modal subtraction while preserving modality-specific liveness cues. Their complementary physical responses further improve live/spoof discrimination. In practical deployments, multiple materials (e.g., skin, hair, eyeglasses, or face coverings) may also bias liveness representations, while spoofing materials are diverse and open-ended. UiAs addresses both through skin-anchored contrastive learning. We evaluate UiAs with real 3D spoofing attacks, which achieves 93.25\% accuracy for unseen users without user-specific physical-signal enrollment.

1 Introduction

UIAS targets user-independent 3D face anti-spoofing by combining electromagnetic and acoustic sensing to suppress shared geometry variations while preserving liveness cues. Skin-anchored contrastive learning further addresses material interference, achieving 93.25% accuracy under unseen-user evaluations.

  • High-fidelity 3D masks reproduce realistic skin appearance and facial geometry, motivating anti-spoofing methods that capture intrinsic physical responses.
  • Wireless liveness cues remain entangled with user-dependent facial geometry, biasing representations toward training identities and limiting cross-user generalization.
  • UIAS combines co-located mmWave and acoustic sensing because both modalities share geometry-related variations but capture complementary physical properties.
  • Cross-modal feature extraction suppresses facial-geometry variations while preserving modality-specific liveness cues; orthogonality and reconstruction constraints reduce overlap and prevent feature collapse.
  • Skin-anchored contrastive learning mitigates interference from hair, eyeglasses, face coverings, and diverse spoofing materials by aligning representations with clean genuine-skin anchors.
  • 93.25% accuracy was achieved under unseen-user evaluations with real 3D spoofing masks and multi-material conditions, outperforming single-modal and conventional multimodal baselines.

2 Background

High-fidelity 3D masks have become a practical threat because they can reproduce facial appearance and structure and bypass face-authentication systems. Active electromagnetic and acoustic probing instead measures material-dependent physical responses, motivating complementary multimodal sensing.

  • High-fidelity 3D masks reproduce both the appearance and facial structure of target users, and their fabrication has become increasingly accessible.
  • Conventional visual, depth, and thermal sensing may be weakened because high-fidelity masks reproduce facial appearance and contours and can absorb body heat.
  • Active wireless probing characterizes physical responses by transmitting known signals and analyzing reflected waveforms.
  • Electromagnetic and acoustic signals respond to different properties, including complex permittivity and acoustic impedance.
  • UIAS combines mmWave and acoustic sensing as complementary modalities, although both remain affected by the same user-dependent facial geometry.

3 Preliminary Study

Preliminary experiments show that high-fidelity 3D masks threaten visual authentication, while active wireless sensing captures liveness cues that fail to generalize across users because user-dependent geometry remains entangled. This motivates suppressing shared geometric information through cross-modal comparison.

  • 3.1 High-Fidelity 3D Mask Threat: 36.67%–83.33% ASRs, averaging 64.17%, let high-fidelity 3D masks bypass evaluated commercial face-unlock systems, whereas conventional 2D attacks fail.The masks also bypassed all evaluated open-source face-recognition systems, while conventional 2D attacks failed against most of them.
  • 3.2 User-dependent Bias in Active Probing: Both mmWave and acoustic representations, as well as naive fusion, separate live faces from high-fidelity masks under the same-user setting.These results indicate that both modalities capture liveness cues from the physical properties of the presented medium.
  • 3.2 User-dependent Bias in Active Probing: Around 50% accuracy on unseen users shows that user-dependent facial geometry severely limits cross-user liveness detection, despite multimodal fusion.Feature distributions shift across users and live and mask representations increasingly overlap because geometry changes propagation paths and local reflection conditions.
  • 3.2 User-dependent Bias in Active Probing: Scaling the dataset is impractical because broad generalization requires diverse users and customized high-fidelity masks costing approximately $2,000 each.Participant recruitment and user-specific mask fabrication therefore introduce substantial cost and effort.
  • 3.4 Motivation: Co-located mmWave and acoustic modalities encode consistent user-dependent variations, making their shared information suitable for suppression through cross-modal comparison.The task-level model separates user-specific geometry, modality-specific physical response, and environmental noise, with UIAS targeting liveness information while suppressing geometry.

4 Threat Model

The threat model considers an attacker presenting a high-fidelity 3D mask through a normal face-authentication interface. UIAS operates as the liveness component, using actively probed wireless responses independently of visual information.

  • 4 Threat Model: The attacker attempts to impersonate a legitimate user by presenting a high-fidelity 3D spoofing mask through the normal authentication interface.Successful end-to-end impersonation requires both passing liveness detection and matching the target identity in face recognition.
  • 4 Threat Model: UIAS determines liveness from actively probed mmWave and acoustic responses after face detection, without relying on visual information.It can therefore serve as a plug-in liveness module independent of camera configuration and illumination conditions.

5 Methodology

UIAS combines mmWave and acoustic sensing with cross-modal processing to suppress user-dependent facial geometry while retaining liveness cues. Skin-anchored contrastive projection then improves robustness to material variations and supports live/spoof classification.

  • Bimodal Facial Representation Construction: UIAS converts synchronized mmWave and acoustic reflections into modality-specific facial representations after reducing environmental multipath.The mmWave branch produces range–azimuth and 2D-AoA responses, while the acoustic branch retains a local time-of-flight window around the strongest CIR peak.
  • User-Dependent Information Suppression: The two modalities encode shared user-dependent facial geometry alongside modality-specific physical-response liveness information.This shared geometry is treated as nuisance information because it can bias classifiers toward identities seen during training.
  • User-Dependent Information Suppression: UIAS aligns latent semantics and calibrates feature energy before subtraction, making geometry-related components approximately comparable across modalities.Semantic alignment addresses coordinate mismatch, while energy calibration reduces scale imbalance and preserves modality-specific differences.
  • Skin-Anchored Contrastive Liveness Projection: Skin-anchored contrastive projection suppresses material-induced drift by learning stable genuine-skin characteristics instead of modeling individual spoofing materials.The method uses a population-level skin anchor because genuine-skin representations are consistent across users, whereas spoof-material representations are heterogeneous.
  • Skin-Anchored Contrastive Liveness Projection: Stage II adapts the projection head to multi-material presentations while updating the skin anchor with an exponential moving average.This optimization recovers clean-skin-related liveness characteristics from material-affected representations without explicitly modeling each spoofing material.
  • Skin-Anchored Contrastive Liveness Projection: The overall training strategy yields a liveness representation for user-independent facial anti-spoofing without user-specific physical-signal enrollment.UIAS is evaluated with COTS mmWave and acoustic sensors under variations in identity, 3D spoofing media, and multi-material interference.

6 Evaluation

UIAS is evaluated under registered and unseen-user settings, using multimodal sensing, controlled robustness tests, and ablations. It achieves strong liveness classification and substantially better cross-user generalization when alignment, subtraction, and contrastive projection are retained.

  • Experimental setup: The evaluation uses COTS mmWave and acoustic sensors, seven-fold unseen-user cross-validation, and metrics including ACC, FAR, FRR, EER, AUC, and F1.The system uses a TI IWR1843BOOST radar and co-located acoustic sensing.
  • Registered-user evaluation: UIAS achieves 99.38% ACC, 0.67% EER, 0.9996 AUC, and 99.38% F1 for registered users.
  • Unseen-user evaluation: 93.25% ACC, 4.80% FAR, 8.70% FRR, 0.9695 AUC, and 93.12% F1 are achieved for unseen users.UIAS outperforms MiniFAS, AFace, MID, and mmFAS in unseen-user accuracy.
  • Cross-modal subtraction: 79.90% user-classification accuracy for kG versus 20.48% for kL indicates that subtraction suppresses user-dependent information in the liveness representation.Random guessing for seven users is 14.29%.
  • Ablation study: Removing alignment reduces unseen-user ACC from 93.25% to 59.25%, while removing subtraction reduces it to 53.25%.Without contrastive projection, ACC drops to 64.50% and FAR rises from 4.80% to 57.20%.
  • Robustness: UIAS achieves at least 90.7% accuracy for face coverings, 91.8% for eyeglasses, and 96.6% for hairstyles.The skin-anchored projection maps material-affected representations into a skin-referenced liveness space.

7 Related Work

Related work spans vision-based 2D anti-spoofing and active RF or acoustic sensing. The paper positions UIAS against methods whose physical-response representations remain entangled with user-dependent facial geometry.

  • 2D face anti-spoofing: Vision-based methods detect 2D presentation attacks using texture, motion, reflection patterns, depth, infrared, and ultrasonic sensing.
  • Robustness dimensions: The evaluation examines sensing range, head orientation, eyeglasses, hairstyles, and other deployment factors as robustness dimensions.The supplied figure passages identify these factors without reporting their outcomes.
  • Wireless anti-spoofing: RF and acoustic systems capture richer physical evidence than RGB but jointly encode liveness cues and user-dependent facial geometry.Existing systems may rely on user-specific enrollment or template-based verification.

8 Discussion and Limitations

UIAS is designed as a plug-in verification module and tolerates common materials such as eyeglasses and face coverings. The discussion also identifies deployment boundaries involving attacks, device variation, and multi-target scenes.

  • Integration: UIAS can serve as a plug-in verification module for existing face-recognition pipelines without modifying the recognition model or authentication circuitry.Compact mmWave and acoustic components in some smart devices could potentially be reused.
  • Multi-material presentations: UIAS tolerates common materials such as eyeglasses and face coverings during facial anti-spoofing.The paper notes that high-security scenarios may still require these materials to be removed manually.
  • Deployment caveats: Signal injection within sensing bands could affect liveness decisions, while heterogeneous sensors may introduce cross-device distribution shifts.Signal-level injection defenses and cross-device calibration are outside the paper’s scope.
  • Scope boundary: UIAS currently targets individual detection, where one approaching user dominates the captured response, leaving crowded multi-target scenarios outside its current scope.

9 Conclusion

UIAS uses co-located mmWave and acoustic sensing to suppress shared user-dependent variations while preserving liveness-discriminative physical responses. Experiments on 14 subjects achieve 93.25% detection accuracy under unseen-user settings.

  • UIAS achieves 93.25% detection accuracy and outperforms baselines under unseen-user settings.The experiments use 14 subjects.
  • UIAS combines co-located mmWave and acoustic sensing to suppress cross-modal shared user-dependent variations while preserving liveness-discriminative physical responses.

Ethical Considerations

The study collected face-related sensing measurements from human participants who provided informed consent. Participants spanned three self-identified racial groups.

  • Human participants and informed consent: Participants voluntarily provided informed consent before data collection.They were informed about the study’s purpose and procedure.
  • Human participants and informed consent: The study collected mmWave, acoustic, and camera information related to facial sensing.Camera information was used to define the facial sensing region.
  • Human participants and informed consent: Participants represented Asian, White, and Black self-identified racial groups.

Open Science

The authors provide an anonymized artifact to support reproducibility, while withholding participant-derived biometric data for privacy protection. The artifact includes the implementations needed to reproduce the paper’s main processing, modeling, training, evaluation, and demonstration components.

  • Artifact availability: An anonymized artifact is available to support reproducibility and future research.
  • Artifact availability: The artifact includes implementations for signal processing, model architecture, training objectives, evaluation, and the demonstration video.
  • Data availability: Raw and processed participant data are not released because they contain biometric information derived from facial characteristics.Participant-derived data are excluded from the artifact for privacy protection.
Loading 2608.29084v1…