Source-linked AI summary
Biometric Face Presentation Attack Detection with Multi-Channel Convolutional Neural Network
Anjith George, Zohreh Mostaani, David Geissenbuhler, Olegs Nikisins, Andre Anjos, Sebastien Marcel
TL;DR
Face PAD remains unreliable against sophisticated attacks, particularly when only visible-spectrum data are used. The paper proposes MC-CNN, which combines aligned color, depth, thermal, and infrared information through transfer learning, and introduces the diverse WMCA dataset. The framework improves over baselines, while requiring aligned multi-sensor capture and potentially reduced performance when channels are unavailable.
Problem
Visible-spectrum PAD remains challenging for sophisticated attacks such as 3D masks, and existing PAD methods may generalize poorly to unseen attacks.
Method
MC-CNN combines multiple aligned channels using a pretrained LightCNN, retraining low-level features while retaining high-level layers.
Results
The proposed MC-CNN outperforms selected baselines by a big margin in the grandest evaluation protocol.
Takeaways & Limitations
The WMCA dataset broadens face-PAD evaluation with diverse 2D and 3D attacks and multiple sensing channels.
Takeaways & Limitations
The framework requires spatially and temporally aligned channels and cannot use data recorded independently across multiple sessions.
Abstract
from arXiv · showhide
Face recognition is a mainstream biometric authentication method. However, vulnerability to presentation attacks (a.k.a spoofing) limits its usability in unsupervised applications. Even though there are many methods available for tackling presentation attacks (PA), most of them fail to detect sophisticated attacks such as silicone masks. As the quality of presentation attack instruments improves over time, achieving reliable PA detection with visual spectra alone remains very challenging. We argue that analysis in multiple channels might help to address this issue. In this context, we propose a multi-channel Convolutional Neural Network based approach for presentation attack detection (PAD). We also introduce the new Wide Multi-Channel presentation Attack (WMCA) database for face PAD which contains a wide variety of 2D and 3D presentation attacks for both impersonation and obfuscation attacks. Data from different channels such as color, depth, near-infrared and thermal are available to advance the research in face PAD. The proposed method was compared with feature-based approaches and found to outperform the baselines achieving an ACER of 0.3% on the introduced dataset. The database and the software to reproduce the results are made available publicly.
I. INTRODUCTION
Face presentation attacks limit reliable face-recognition deployment, especially as sophisticated 3D attacks become harder to detect using visible-spectrum information alone. The paper proposes multi-channel PAD and introduces the WMCA database to support broader attack evaluation.
- Presentation attacks can fool face-recognition systems and limit their reliable use in unsupervised applications.
- Attacks include impersonation, obfuscation, and presentation attack instruments such as printed photos, replayed videos, and manufactured 3D masks.
- Most existing PAD research focuses on print and replay attacks using visible-spectrum cues such as color, texture, motion, and physiological information.
- Visible-spectrum information alone may be insufficient for sophisticated attacks and generalization to unseen presentation attack instruments.
- The proposed MC-CNN combines multi-channel information while retraining only low-level LightCNN features and retaining high-level pretrained layers.
- The WMCA database contains aligned color, depth, thermal, and infrared data spanning diverse 2D and 3D attacks.
II. RELATED WORK
Prior face-PAD research has developed handcrafted, CNN, transfer-learning, and auxiliary-supervision methods, but much of it remains focused on 2D attacks. Limited training data and poor generalization to unseen or cross-dataset attacks motivate alternative approaches.
- Most reviewed PAD methods address 2D print and replay attacks, often treating PAD as binary classification.
- Feature-based approaches use image quality, color, texture, motion, physiological, and local-pattern cues to distinguish attacks.
- CNN-based PAD methods include 3D spatiotemporal models, deep CNNs with SVM classifiers, and auxiliary supervision for depth or rPPG estimation.
- Transfer learning is commonly used because PAD datasets are often too small to train deep networks from scratch.
- Auxiliary-task approaches have been reported to improve generalization capability in PAD.
- FASNet used pretrained VGG16 features and reported HTERs of 0% on 3DMAD and 1.20% on Replay-Attack.
C. Multi-channel based approaches and datasets for face PAD
Multi-channel PAD research uses complementary sensing information to address limitations of visible-spectrum methods, but earlier datasets and acquisition setups have constrained attack diversity or channel fusion.
- Visible-spectrum PAD becomes harder as capture devices, printers, and 3D mask manufacturing improve.
- Multi-channel countermeasures have been proposed as a way to improve robustness against diverse spoofing attempts.
- Prior studies combined visible, thermal, infrared, depth, multispectral, or other channels using feature, score, or patch-level fusion.
- Some prior multi-channel data were recorded independently, preventing spatial or temporal fusion across channels.
- Thermal and depth channels were reported as useful for detecting 3D masks and 2D attacks, respectively.
- Earlier multi-channel datasets often contained limited presentation attack instrument variety.
D. Discussions
The discussion motivates multi-channel PAD for increasingly realistic attacks and describes MC-CNN as a transfer-learning framework that jointly represents multiple channels while adapting few parameters. Its design depends on aligned inputs and may require reduced performance or additional sensing hardware in deployment.
- D. Discussions: Realistic 3D masks and partial attacks remain difficult for PAD, while visible-spectrum-only discrimination becomes harder as attack instruments improve.
- D. Discussions: The proposed MC-CNN uses a joint representation from multiple channels with transfer learning from a pretrained face-recognition network.
- D. Discussions: The framework preprocesses color and non-RGB channels through face alignment, with non-RGB processing requiring spatial and temporal alignment to color.
- D. Discussions: The approach adapts a minimal parameter set by retraining lower-level features while sharing higher-level features across channels.
- D. Discussions: The network concatenates 256-dimensional channel embeddings and adds two fully connected layers for PAD classification.
- D. Discussions: Binary cross entropy trains the model using attack and bona fide ground-truth labels.
IV. THE WIDE MULTI-CHANNEL PRESENTATION ATTACK DATABASE
WMCA is a 72-identity database of short bonafide and presentation-attack videos captured across multiple sensing channels. It provides aligned RGB, depth, NIR, and thermal data for face PAD research.
- WMCA contains short video recordings from 72 different identities, covering bonafide presentations and presentation attacks.
- The acquisition system records RGB, depth, Near-Infrared, and thermal streams, with RGB-D and NIR supplied by Intel RealSense SR300 and thermal data by Seek Thermal Compact PRO.
- The database is designed to provide high-quality information across visual, infrared, and 3D sensing domains.
1) Intel RealSense SR300 sensor:
WMCA combines consumer RGB-D and thermal sensors in a calibrated, aligned acquisition setup and records diverse bonafide and attack sessions. The resulting database contains 1,679 presentations across multiple attack categories and channels.
- Sensors: The setup combines an Intel RealSense SR300 RGB-D sensor with a Seek Thermal Compact PRO thermal camera.
- Camera integration and calibration: A rigid optical mounting frame controls sensor orientation, while a thermally visible checkerboard and custom marker extraction provide cross-channel alignment.
- Data collection procedure: Data were collected across seven sessions over five months under varied backgrounds and illumination, with ten seconds captured per bonafide or attack presentation.
- Presentation attacks: The database includes diverse attacks spanning glasses, prints, replays, fake heads, rigid masks, flexible silicone masks, and paper masks.
- Database statistics: 1,679 total presentations comprise 347 bonafide samples and 1,332 attacks, recorded across the four channels.
V. EXPERIMENTS
The experiments evaluate the proposed system on all four WMCA channels under both seen and unseen attack scenarios.
- All four channels from the Intel RealSense SR300 and Seek Thermal Compact PRO are used in the experiments.
- Experiments evaluate performance in both “seen” and “unseen” attack scenarios.
- In the “seen” protocol, all PAI types occur in the train, development, and testing subsets while client identities remain disjoint across folds.
A. Protocols
WMCA defines seen and leave-one-out unseen protocols, samples aligned four-channel videos into frame-level decisions, and evaluates PAD using standardized error metrics and ACER. Feature-based channel baselines use IQM or LBP representations with logistic regression and score fusion.
- Data sampling: Each biometric sample uses uniformly sampled frames from all four spatially and temporally aligned channels, with one PAD score produced per frame.
- Protocols: The grandtest protocol distributes attack categories across train, development, and evaluation partitions to emulate seen attacks, while seven leave-one-out protocols evaluate unseen attacks.
- Evaluation metrics: Thresholds are selected on development data at BPCER = 1%, then APCER and BPCER are reported on the test set using ISO/IEC 30107-3 metrics.
- Evaluation metrics: ACER summarizes performance as the average of APCER and BPCER, and is reported for both development and test sets.
- Baselines: Feature-based baselines use IQMs for RGB, spatially enhanced LBP histograms for non-RGB channels, logistic regression classifiers, and mean score fusion.
2) RDWT-Haralick-SVM baseline:
The baseline experiments show that multi-channel fusion improves PAD accuracy, but the best baseline remains inadequate for critical deployment. MC-CNN is designed to use multi-channel information more efficiently than these baselines.
- MC-CNN architecture: MC-CNN combines a LightCNN face-recognition subnetwork with channel-specific embeddings and added fully connected layers.The architecture extends the 29-layer LightCNN across four channels, concatenating embeddings while sharing unretrained layers.
- Baseline performance: Thermal and infrared provide the most discriminative information among individual channels, while score fusion improves feature-based baseline accuracy.RDWT-Haralick-SVM in infrared achieves the best individual-channel accuracy, and fusion improves both feature-based baselines.
- Baseline performance: FASNet outperforms IQM and RDWT-Haralick features in the color channel but is difficult to extend directly to additional channels.Its ImageNet-based three-channel design and fine-tuning of only the final fully connected layers limit straightforward multi-channel extension.
- Baseline performance: Multiple-channel fusion boosts PAD performance, yet the best baseline systems remain inadequate for deployment in critical scenarios.The lower accuracy of fusion baselines motivates methods that use multi-channel information more efficiently.
- MC-CNN results: MC-CNN outperforms selected baselines by a large margin by learning a joint representation of complementary multi-channel information.Transfer learning from a pretrained face-recognition model supports deep multi-channel PAD while adapting only a minimal set of layers.
- MC-CNN results: At a BPCER threshold of 1%, MC-CNN classifies attacks perfectly except for the glasses attack.The glasses case is discussed separately as a source of performance degradation.
B. Generalization to unseen attacks
MC-CNN generalizes well to most attacks excluded from training, but glasses attacks remain difficult because their appearance resembles bonafide subjects wearing medical glasses. The layer-adaptation study identifies a small adapted subset as preferable to adapting all layers.
- B. Generalization to unseen attacks: MC-CNN performs well on most unseen attacks evaluated in protocols that systematically exclude one attack from training.Each protocol trains on the remaining attacks and evaluates on the held-out attack with bonafide samples.
- B. Generalization to unseen attacks: The ROC in Fig. 7 compares MC-CNN with baseline methods on WMCA grandtest protocol evaluation set 5.The figure reports receiver operating characteristic curves for the proposed and baseline systems.
- B. Generalization to unseen attacks: Fig. 8 presents four-channel preprocessed data for bonafide subjects with glasses and for the funny-eyes-glasses attack.The bonafide examples occupy the first row and the attack examples the second row.
- B. Generalization to unseen attacks: Glasses attacks produce very poor performance for both MC-CNN and baselines because they resemble bonafide medical glasses across most channels.Because glasses attacks are absent from training, they are classified as bonafide and reduce performance.
- B. Generalization to unseen attacks: Partial attacks are especially difficult in variable face regions, including eyes affected by prescription glasses and lower chins affected by facial hair.Bonafide variability in these regions can make corresponding partial attacks harder to detect.
- Layer adaptation: Adapting lower layers initially improves performance, but adapting more layers eventually degrades it, consistent with overfitting from learning too many parameters.The best configuration adapts a minimal layer set and achieves an ACER of 0.3%.
2) Experiments with different combinations of channels:
Channel-combination experiments show that all four channels perform best, while a three-channel subset remains effective when thermal data is unavailable. Single-channel results are treated as ablations rather than optimized MC-CNN configurations.
- 2) Experiments with different combinations of channels: The experiments evaluate individual channels and channel combinations to identify useful PAD inputs and assess deployment when some channels are unavailable.Color, depth, and infrared come from Intel RealSense SR300, while thermal is obtained separately.
- 2) Experiments with different combinations of channels: The proposed architecture retains its general design while concatenating embeddings from whichever selected channels are used in the final fully connected layers.Training and testing follow the grandtest protocol procedure.
- 2) Experiments with different combinations of channels: The all-four-channel system achieves the best performance with an ACER of 0.3%.The four-channel configuration uses grayscale, depth, infrared, and thermal inputs.
- 2) Experiments with different combinations of channels: The CDI combination achieves an ACER of 1.04% using three channels from the same Intel RealSense device.This configuration indicates a potentially useful subset when all channels are unavailable.
- 2) Experiments with different combinations of channels: Among individual channels, thermal achieves the best performance with an ACER of 1.85%.These individual-channel experiments are ablations, and the network is not optimized for single-channel use.
D. Discussions
The proposed multi-channel system outperformed selected feature-based baselines, with transfer learning supporting training under limited data. Across channel, cross-database, and unseen-attack experiments, complementary channels improved performance, while deployment depends on aligned sensing and available hardware.
- Performance: The proposed algorithm surpasses selected feature-based baselines, and transfer learning from a face-recognition network is effective with limited training data.The framework reuses representations learned for face recognition rather than training the deep multi-channel CNN from scratch.
- Performance: All four channels achieved the best performance in both channel-comparison and cross-database experiments.The authors attribute the improvement to complementary information added by the channels.
- Unseen attacks: The binary-trained framework generalized well to most unseen attacks when their properties could be learned from other presentation types.Training with sufficient varieties of complex presentation attack instruments may support reasonable performance on simpler unseen attacks.
- Unseen attacks: The penultimate-layer representation can support one-class classifiers or anomaly detectors for detecting unseen attacks.This provides an alternative use of the learned representation beyond binary classification.
- Deployment limitations: The framework requires spatially and temporally aligned channels, with synchronized sensing needed for supervised temporal alignment.Small temporal misalignments are tolerated, but data collected across multiple unsynchronized sessions cannot be used directly.
- Deployment limitations: Missing sensors can be accommodated by retraining on available channels, but performance is reduced when not all channels are deployed.The framework can also be extended with additional channel types.
- Overall conclusion: Visible-spectrum-only algorithms perform poorly on the diverse 2D and 3D attacks in the introduced dataset, whereas adding multiple channels improves results greatly.The evaluations include unseen-attack protocols intended to reflect scenarios where deployment encounters attacks absent from training.