Source-linked AI summary
Face Morphing Attack Generation & Detection: A Comprehensive Survey
Sushma Venkatesh, Raghavendra Ramachandra, Kiran Raja, Christoph Busch
TL;DR
Face morphing can undermine border-control face recognition by allowing one morphed passport image to verify against two contributors. This survey synthesizes morph generation, MAD methods, datasets, evaluation, and open challenges; reported benchmarks show detection remains difficult, with D-MAD outperforming S-MAD in the discussed evaluations. The field remains constrained by limited varied public datasets and unresolved effects of covariates such as aging, gender, ethnicity, and image quality.
Problem
Morphed face images can compromise border-control FRS by being verified against both contributing subjects, creating a need for reliable attack detection.
Method
The survey systematically reviews morph generation techniques, MAD taxonomies, datasets, benchmarking outcomes, vulnerability assessments, metrics, and open challenges.
Results
Detection remains challenging; D-MAD methods show more robust benchmark performance than S-MAD methods, while feature-difference D-MAD outperforms face de-morphing in the reported evaluation.
Takeaways & Limitations
Reliable morph-attack detection requires evaluation across varied morph-generation methods and image conditions, with feature-difference D-MAD showing a practical advantage in the discussed benchmarks.
Abstract
from arXiv · showhide
The vulnerability of Face Recognition System (FRS) to various kind of attacks (both direct and in-direct attacks) and face morphing attacks has received a great interest from the biometric community. The goal of a morphing attack is to subvert the FRS at Automatic Border Control (ABC) gates by presenting the Electronic Machine Readable Travel Document (eMRTD) or e-passport that is obtained based on the morphed face image. Since the application process for the e-passport in the majority countries requires a passport photo to be presented by the applicant, a malicious actor and the accomplice can generate the morphed face image and to obtain the e-passport. An e-passport with a morphed face images can be used by both the malicious actor and the accomplice to cross the border as the morphed face image can be verified against both of them. This can result in a significant threat as a malicious actor can cross the border without revealing the track of his/her criminal background while the details of accomplice are recorded in the log of the access control system. This survey aims to present a systematic overview of the progress made in the area of face morphing in terms of both morph generation and morph detection. In this paper, we describe and illustrate various aspects of face morphing attacks, including different techniques for generating morphed face images but also the state-of-the-art regarding Morph Attack Detection (MAD) algorithms based on a stringent taxonomy and finally the availability of public databases, which allow to benchmark new MAD algorithms in a reproducible manner. The outcomes of competitions/benchmarking, vulnerability assessments and performance evaluation metrics are also provided in a comprehensive manner. Furthermore, we discuss the open challenges and potential future works that need to be addressed in this evolving field of biometrics.
I. INTRODUCTION
Face morphing threatens border-control FRS because a morphed passport image can resemble both contributors and enable unauthorized travel. The survey motivates systematic study of morph generation, detection, benchmarking, and unresolved challenges.
- I. INTRODUCTION: Face recognition is widely used for secure access control, including border control comparisons against passport or visa records.This broad deployment makes attacks against FRS relevant to identity verification systems.
- I. INTRODUCTION: Morphed passport images can let a malicious person pass border control while the accomplice’s identity remains recorded.The attack exploits facial resemblance to the applicant and the genuine passport issuance process.
- I. INTRODUCTION: Digital renewal and visa-upload processes create additional opportunities to submit morphed facial images without trusted supervision.The survey specifically identifies online facial-image submission as a concern.
- I. INTRODUCTION: The survey covers morph generation, MAD algorithms, public databases, evaluation metrics, competitions, vulnerability assessments, and future challenges.Its scope spans both attack creation and detection research.
- I. INTRODUCTION: A morphed image may be verified against both contributing subjects with a high FRS similarity score.This is the central mechanism by which morphing undermines identity verification.
III. FACE MORPH ATTACK GENERATION
Face morphs are generated through landmark-based or deep-learning approaches, with landmark methods warping corresponding facial points and practical tools requiring varying post-processing effort.
- III. FACE MORPH ATTACK GENERATION: Morph generation is broadly classified into landmark-based and deep-learning-based approaches.The taxonomy distinguishes traditional landmark processing from newer learned synthesis methods.
- Landmark Based Morph Generation: Landmark-based methods warp corresponding facial points, such as the eyes, nose, and mouth, toward averaged positions.Reported deformation procedures include free-form deformation, moving least squares, mass-spring, and Bayesian approaches.
- Landmark Based Morph Generation: Open-source landmark tools can generate morphs but require substantial post-processing to remove artefacts, whereas commercial tools support large-scale generation with reasonable effort.The comparison concerns practical generation effort rather than detection performance.
- III. FACE MORPH ATTACK GENERATION: The survey organizes generation methods by their advantages and limitations in a dedicated comparison table.The table is presented as a synthesis of face morph generation methods.
B. Deep Learning Based Morph Generation
Deep-learning morph generation, including GAN-based synthesis, expands the attack space alongside landmark and splicing methods, while available datasets vary in scale, modality, openness, and realism.
- B. Deep Learning Based Morph Generation: GAN-based methods synthesize morphed faces by sampling two facial images in a deep-learning latent space.MorGAN uses encoders, decoders, and a discriminator to generate 64 × 64 images.
- B. Deep Learning Based Morph Generation: Morph databases span landmark, automatic, splicing, print-scan, aging, and deep-learning-generated images, but their availability and standards differ.The survey catalogs both public and private resources used for MAD and vulnerability evaluation.
- B. Deep Learning Based Morph Generation: Public and private morph databases are summarized to support benchmarking of morph attack detection methods.The survey explicitly presents a database comparison table covering both categories.
- B. Deep Learning Based Morph Generation: The Bologna-SOTAMD dataset contains 5,748 morphed and 1,396 bona fide face images from subjects across varied locations, ethnicity, gender, and age.It supports recent public competition and benchmarking under a sequestered protocol.
A. Discussion
Public reproducibility remains constrained because most morphing datasets are unavailable for redistribution under data-protection and licensing conditions.
- A. Discussion: Most morphing datasets are private because data-protection regulations and licensing conditions restrict redistribution.This limits direct comparison with previously published morphing-detection methods.
V. HUMAN PERCEPTION AND MORPHED FACE DETECTION
Human observers often struggle to detect morphed faces, although expertise and training improve detection. This motivates automatic MAD methods, including single-image approaches with varied feature families.
- Human perception: Skilled and unskilled observers both frequently miss morphed faces, while substantial training can improve detection performance.High-quality morphs are especially difficult for human observers to detect.
- Automatic MAD: Automatic MAD is organized into single-image and differential-image categories to address limitations of human morph detection.The supplied section emphasizes single-image MAD and identifies differential-image MAD as the second major category.
- S-MAD: Single-image MAD detects a suspected morph from one submitted face image in passport enrolment, without requiring a trusted comparison image.This setting applies to physical or digital passport-image submission.
- S-MAD: S-MAD methods use texture, quality, residual-noise, deep-learning, or hybrid features, with hybrid approaches generally outperforming single-mode methods.Texture methods include LBP, LPQ, and BSIF; hybrid methods combine multiple extractors, scores, or classifiers.
- S-MAD limitations: Texture-based S-MAD reports strong digital-image accuracy but limited generalizability across image quality, sensors, and print-scan processes.Micro-texture methods have shown reasonable performance on both digital and print-scan S-MAD.
B. Differential Image Based MAD (D-MAD)
D-MAD determines whether a suspected passport image is morphed by comparing it with a trusted live or reference image. Its main approaches use feature differences or demorphing, with deep CNN features reporting the best performance among feature-difference methods.
- D-MAD overview: D-MAD compares a suspected passport morph with a trusted live or reference image, making it suitable for automated border-crossing scenarios.The trusted image may be captured at an ABC gate.
- D-MAD taxonomy: D-MAD comprises feature-difference methods and demorphing methods that attempt to detect morphs using paired-image information.Feature-difference methods compare representations from the suspected and trusted images, whereas demorphing reverses the morphing process.
- Feature-difference D-MAD: Deep CNN features have indicated the best performance among reported feature-difference D-MAD approaches.Most existing studies use digital images, while recent work also explores print-scan datasets with improved results.
VII. PERFORMANCE METRICS
The survey describes vulnerability metrics for assessing whether morphed images can deceive face-recognition systems. These metrics use comparison scores, thresholds, contributing subjects, and repeated probe attempts to quantify attack strength.
- Vulnerability assessment: Vulnerability analysis measures whether generated morphs can be verified against all contributing subjects at the face-recognition threshold.Most studies set the threshold to correspond to an FMR of 0.1%, following FRONTEX guidance.
- Vulnerability assessment: The vulnerability plot’s Q-III contains morphs verified against both contributing subjects and therefore represents the highest threat quadrant.Q-I indicates verification against neither subject, while Q-II and Q-IV indicate verification against only one subject.
- MMPMR: MMPMR measures the proportion of morphed images verified with their contributing images.Its formulation aggregates successful verification across contributing subjects and compares scores with a predefined threshold.
- FMMPMR: FMMPMR extends MMPMR by requiring verification against both contributing subjects while accounting for multiple probe attempts.The metric uses a threshold corresponding to FMR = 0.1% and is presented as a realistic measure of morph attack strength.
B. MAD Performance Metrics
MAD performance is evaluated as binary attack-versus-bona-fide classification using standardized error rates. Because APCER and BPCER cannot be optimized jointly, evaluations typically fix one rate and report the other, with DET curves supporting operating-point comparisons.
- Standard metrics: MAD robustness is benchmarked using ISO/IEC 30107-3 performance metrics for binary classification of attack and bona fide images.The survey focuses on metrics widely used to report MAD performance.
- Standard metrics: APCER is the proportion of attack samples incorrectly classified as bona fide images.
- Standard metrics: BPCER is the proportion of bona fide images incorrectly classified as attack samples.
- Operating points: Because APCER and BPCER cannot be optimized jointly, studies commonly fix APCER at 1%, 5%, or 10% and report the dependent metric.A DET curve compares algorithms at selected operating points.
- Operating points: At APCER of 5% or 10%, MAD Algorithm 3 is preferred over the two other algorithms in the cited benchmark.
C. Joint Evaluation of MAD and Vulnerability
AMPMR evaluates vulnerability when morph attack detection is integrated with face recognition. It captures whether morphed images both pass enrolment and match contributor probes despite MAD.
- AMPMR measures vulnerability in an integrated FRS–MAD pipeline by requiring a morph to pass enrolment and match contributor probe images.
- The metric aggregates recognition scores for morphed images against contributor probes while incorporating MAD scores.N denotes the total number of morph images; M_i denotes contributors to morph i.
- Higher AMPMR values indicate greater vulnerability to morphing attacks despite MAD processing.
VIII. PUBLIC EVALUATION AND BENCHMARKING
Public benchmarks provide shared datasets, protocols, and computational environments for evaluating morph attack detection. FRVT MORPH results show that reliable operational performance remains unmet and depends on morph quality and detection design.
- Public FRVT and Bologna-SOTAMD benchmarks provide common datasets, evaluation protocols, and computational environments for reproducible MAD assessment.
- The FRVT MORPH dataset covers low-quality, automatically generated, and high-quality morphs to reflect varied attack-generation conditions.
- None of the evaluated algorithms reliably met the FRONTEX operational requirement, leaving face morphing detection challenging.
- Morph quality directly affects both S-MAD and D-MAD performance; hybrid features led S-MAD, while ArcFace latent feature differences led D-MAD.
B. Bologna-SOTAMD: Differential Morph Attack Detection
The Bologna-SOTAMD benchmarks show that differential MAD remains insufficient for operational requirements, while feature-difference methods outperform face de-morphing. Across the broader field, unknown generation conditions, subject selection, face covariates, datasets, and metrics remain major challenges.
- B. Bologna-SOTAMD: Differential Morph Attack Detection: D-EER was 3.36% for both digital and print-scan data, yet submitted D-MAD techniques remained insufficient for FRONTEX operational requirements.Feature-difference methods performed better than face de-morphing methods in the benchmark.
- C. Bologna-SOTAMD: Single Morph Attack Detection: S-MAD baseline D-EER reached 37.10% on print-scan and 38.99% on digital morphed images, underscoring detection difficulty.The benchmark used passport-like images and morphs generated with commercial and open-source software.
- D. Discussion on Public Evaluation: S-MAD performance is severely degraded relative to D-MAD, while hybrid S-MAD features improve generalisability across morph-generation methods.D-MAD feature-difference methods showed more robust performance across both benchmarks.
- IX. OPEN CHALLENGES: Large, varied public datasets remain unavailable because realistic print/scan generation is costly and sharing is restricted by licensing, privacy, and GDPR concerns.Existing public benchmarks primarily test submitted algorithms rather than supporting broader MAD development.
- IX. OPEN CHALLENGES: Unknown printers, scanners, and morph-generation sources degrade MAD performance, especially for learning-based S-MAD systems.Print/scan variation changes data quality, while known-data decision policies limit generalisation to unseen sources.
- IX. OPEN CHALLENGES: Open evaluation priorities include look-alike subject selection, face covariates, standardised vulnerability metrics, and user-convenient MAD systems.These gaps cover subject pairing, aging and image conditions, harmonised evaluation, and deployment with minimal operator or applicant intervention.
X. CONCLUSION
The survey reviews face morph generation and detection, while emphasizing that robust generalization remains difficult because public databases lack sufficient variation across morph-generation techniques.
- Robust generalization of morph attack detection remains difficult because large public databases with varied morph-generation techniques are limited.The survey identifies this database challenge as a central obstacle to developing robust detection methods.
- The survey catalogs advances in multiple face morph generation techniques and summarizes corresponding morphing attack detection methods.
- The paper reports performance metrics and discusses challenges to guide future research on detecting morphed face images.