Source-linked AI summary
Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set
Michael Zang, Haiyu Wu, Mrinal Sharma, Kevin W. Bowyer
TL;DR
Face recognition for MZ twins remains difficult, and existing datasets provide limited metadata about distinguishing marks and mirror asymmetry. The paper analyzes CTTS-80, evaluates current matchers on twin image pairs, tests whether they encode skin marks, and examines generative AI for imagined-twin training images; it concludes that twin recognition remains an open problem and that generative-AI construction has an unresolved fundamental limitation.
Problem
MZ-twin recognition is challenging, while existing twin datasets lack sufficient scale and metadata for distinguishing skin marks and possible mirror asymmetry.
Method
The paper analyzes CTTS-80, evaluates face-verification matchers using cross-validation, investigates skin-mark encoding, and explores Grok, ChatGPT, and Gemini for imagined-twin images.
Results
CTTS accuracy is lower than that of test sets emphasizing age, pose, facial hair, or illumination differences, and current matchers show negligible embedding effects from visible skin marks.
Takeaways & Limitations
CTTS enables comparisons of face matchers on MZ-twin discrimination while focusing analysis on skin marks and asymmetry.
Takeaways & Limitations
The facial differences between real MZ twins that would produce well-separated positive- and negative-pair distance distributions are not yet known, and flip-averaged embeddings may obscure skin-mark location.
Abstract
from arXiv · showhide
Past literature on face recognition for monozygotic (("identical") twins points to facial marks and mirror asymmetry as possible directions for improved accuracy of twins recognition. The Celeb Twins Test Set (CTTS) contains web-scraped image pairs for 80 sets of celebrity twins. It is the only twins test set with meta-data for twins with distinguishing skin marks and possible mirror asymmetry. CTTS is organized in the manner of face verification test sets such as LFW, CALFW, CPLFW, CFP-FP, and AgeDB-30. Current deep CNN matchers can achieve over 76% accuracy in classifying CTTS same-person / different-person image pairs. We show that current matchers do not make use of skin marks, or asymmetry, and discuss reasons for this. Finally, we discuss the feasibility of using generative AI tools such as Grok, ChatGPT and Gemini to create images of imagined monozygotic twins as a means to increase representation of twins in face recognition training sets.
1. Introduction
MZ twins remain difficult for face recognition, while progress is constrained by limited experimental datasets and missing metadata about distinguishing marks and mirror asymmetry. This paper analyzes CTTS-80, a larger celebrity-twin test set designed to address those gaps.
- Motivation: A 0.0001 false match-rate threshold for non-twins produced 0.98–0.99 false match rates for most MZ-twin algorithms.The lowest reported algorithmic false match rate was 0.475, meaning one twin was accepted as the other about half the time.
- Motivation: Limited twin datasets hinder automated face-recognition research for both training and testing.The ND-Twins test set was introduced to address this limitation, but its images come from the ND Twins Days dataset.
- Contribution: CTTS-80 contains 21,120 image pairs for 80 celebrity-twin sets, over 3.5 times the 6,000 pairs in several established test sets.More than half of its twin sets have distinguishing skin marks, and seven are identified as potentially mirror twins through opposite handedness.
- Contribution: Unlike ND-Twins and ND Twins Days, CTTS-80 includes metadata for distinguishing skin marks and possible mirror asymmetry.Other face datasets may contain skin-mark metadata, but they are not focused on twins.
- Scope: The paper evaluates CTTS, examines whether current matchers use skin marks and asymmetry, and considers generative AI for creating imagined-twin training images.The paper’s sections cover biological background, human and automated twin recognition, CTTS evaluation, matcher analysis, and generative-AI feasibility.
2. MZ Twins: Biological Background
MZ twins share an embryo origin and closely related DNA, but their facial appearance is not literally identical. Differences can arise through skin marks, asymmetry, aging, and other developmental or environmental factors, while mirror-twin status is only imperfectly observable.
- Overview: MZ twins may acquire small facial differences despite originating from one embryo and sharing the same initial DNA.Possible distinguishing elements include moles, freckles, scars, and mirrored asymmetry.
- Mirror twins: “Mirror twins” are an MZ-twin subset in which the handedness of some physical features is reversed.Possible reversals include handedness, hair-whorl direction, and left-right dental patterns; facial asymmetry may or may not be useful for distinguishing a pair.
- Mirror twins: About 1 in 4 MZ twins are estimated to be mirror twins, but no DNA test determines mirror-twin status.Mirror twins generally result from later embryo splitting, around 7 to 10 days after conception, after left-right development begins.
- Distinguishing features: Moles and freckles are not generally evidence of mirrored locations; they can distinguish MZ twins whether or not they are mirror twins.Mirror twins may or may not have birth marks in mirrored positions, while non-mirror MZ twins generally do not have marks in the same location.
- Mirror twins: Opposite handedness identifies some mirror twins, but not all mirror twins have opposite handedness.One study estimates that approximately 50% to 75% of mirror twins are opposite-handed, and CTTS uses known cases to categorize some twins.
- Aging: Facial differences between MZ twins tend to increase with age because aging effects can diverge through environmental conditions and health habits.The NIST study noted lower average similarity between MZ-twin faces with increased age.
3. Human Ability to Distinguish Twins
Human performance in distinguishing MZ twins varies by task, observer, and viewing conditions. Facial marks emerge as an important cue, while training can improve discrimination and face-recognition ability is partly influenced by genetics.
- Human experiments: 93% mean human accuracy was reported for judging whether frontal image pairs showed the same person or MZ twins.Across 25 subjects and 100 pairs, accuracy ranged from 78% to 100%; participants cited features including moles and scars.
- Human experiments: A 30 ms viewing task found twins and close friends were no more accurate at identifying one twin than the other.The experiment used cropped face images from 10 MZ-twin sets and required categorization as self, twin, or close friend.
- Human–algorithm comparison: A ResNet101-based matcher exceeded the accuracy of nearly all tested human participants across the reported image-pair conditions.The comparison used frontal, 45°, and 90° pose combinations for same-person, twin-impostor, and general-impostor pairs.
- Facial marks: Facial marks were studied as moles, freckles, spots, birthmarks, scars, and related features, with manually and automatically annotated marks supporting twin recognition.Reported equal error rates were about 24%–35% for manual annotations and about 27% for automatic annotations.
- Individual differences: Face-recognition ability appears strongly influenced by genetics, with MZ-twin CFMT-score correlation of 0.70 versus 0.29 for DZ twins.The study compared 164 MZ twin pairs with 125 same-gender DZ pairs.
- Synthesis: The literature identifies facial marks as important cues, finds no explicit human-study discussion of asymmetry, and reports substantial variation and trainability in face-recognition ability.These observations summarize the reviewed evidence on human twin discrimination.
4. Automated Face Recognition of MZ Twins
Automated MZ-twin recognition has used holistic, local, landmark, skin-mark, asymmetry, and fused feature approaches, but results are difficult to compare across studies. Existing evaluations report substantial performance variation, limited dataset availability, and persistent false matches between twins.
- Literature caveat: A claimed ND Twins Days experiment was questioned because published figures appear copied from other works, including one source that was not cited.The critique concerns the provenance of figures and the claimed dataset use.
- Approaches: Automated approaches have included feature-, score-, and decision-level fusion, LFDA with regional Gabor features, LBP and HOG methods, landmark-based recognition, and deep CNN matchers.Reported studies also include facial-mark representations, asymmetry features, aging-related features, and comparisons involving the WV Twins Days dataset.
- Deep CNN studies: 61% overall accuracy was reported on one twin dataset, with 77% accuracy on positive pairs and 45% on negative pairs, unchanged by fine-tuning.The dataset used 31 twin pairs for fine-tuning, and its availability to other researchers was not reported.
- Large-scale evaluation: The NIST evaluation found that algorithms had high false matches and could not distinguish identical and fraternal same-sex twins as different people.Its twins false match rate measured the fraction of MZ-twin similarity scores exceeding a threshold set at 1-in-10,000 false match rate on non-twins.
- Skin-mark methods: Automatically detected skin marks yielded 96% accuracy on 319 images from 74 twin sets.The approach used a pre-trained network to detect mark types and locations, then represented them in a feature vector.
- Evidence gaps: Previous work suggests facial marks are useful and asymmetry may help, but the asymmetry literature remains limited.The field primarily relies on the ND Twins Days dataset, which is more than 15 years old and too small for effective deep CNN training.
- Comparability: Accuracy comparisons across papers are difficult because studies use different datasets, covariates, and matchers.The reviewed work spans multiple datasets and experimental designs, including the FGFV dataset’s correlation between image source and pair category.
5. The Celeb Twins Test Set
CTTS is a web-scraped, twin-set-disjoint face-verification test set covering 80 celebrity twin pairs, with metadata for distinguishing skin marks and mirror-twin indicators. Using 10-fold cross-validation, current matchers show substantial variation across twin sets and lower accuracy on twins than on standard face-verification benchmarks.
- Dataset construction: CTTS-80 contains 80 celebrity twin sets, with 43 sets having distinguishing skin marks and 7 showing left-right handedness consistent with mirror twins.The set includes 36 male and 44 female twin sets spanning births from the 1940s through the early 2000s.
- Dataset construction: Each twin set contributes 264 balanced positive and negative image pairs, evaluated using twin-set-disjoint 10-fold cross-validation.Thresholds are selected on nine folds and applied to the held-out tenth fold, preventing a twin set's pairs from influencing its own threshold.
- Results: CTTS-80 and ND-Twins have lower accuracy than traditional face-verification sets, while higher standard deviations make firm matcher and training-set comparisons difficult.Accuracy is more stable across matchers and training sets for CTTS than for ND-Twins, likely because CTTS contains more image pairs.
- Caveats: Potential overlap between CTTS identities and web-scraped matcher training sets is unlikely to have a significant overall effect because overlapping twins represent a very small fraction of training data.The paper nevertheless notes that individual twins appearing in training data could in principle receive higher accuracy.
- Results: 7 twin sets scored 50–60% accuracy, 19 scored 60–70%, 26 scored 70–80%, 11 scored 80–90%, and 17 scored 90–100%.The wide range reflects differences in positive-versus-negative distance-distribution separation and how well the cross-validation threshold generalizes to each twin set.
- Results: Well-separated examples generally use approximately frontal faces with little occlusion, while poorly separated examples include caps, off-frontal pose, and atypical expressions.Figure 3 reports accuracy above 96% despite a threshold lower than ideal for its displayed distributions.
6. Do Current Matchers Encode Skin Marks?
Current matchers largely ignore visible skin marks, and standard flip-based training and scoring may suppress asymmetry-related information. Removing horizontal-flip augmentation produced mixed effects across mirror-twin sets rather than a consistent improvement.
- Skin marks: Skin marks are often distinguishing in CTTS, but pose, cosmetics, resolution, and blur can make them invisible or ambiguous at 112x112 resolution.Over half of CTTS twin sets have potentially distinguishing skin marks.
- Skin marks: Editing out Tamera Mowry’s skin mark did not significantly change the distance distributions against Tia Mowry images.The experiment edited only the skin-mark region before matching.
- Skin marks: Original-versus-erased comparisons produced nearly overlapping positive and negative distributions, with the 12 mixed-mark pairs clustered near zero distance.This provides direct evidence that the visible mark has almost no effect on the computed embedding.
- Skin marks: The same negligible skin-mark effect appeared for the Hassan and Kaczynski twins, indicating it is a general property of the model rather than one twin set.Original and erased images yielded essentially the same pair distributions and zero difference for the same image.
- Training and scoring: Horizontal-flip augmentation can teach networks not to use skin-mark position or asymmetry, while averaging original and flipped embeddings may further obscure location information.The paper also notes that twins are rare in typical training data, likely limiting learning of twin-related cues.
- Mirror asymmetry: Across seven known mirror-twin sets, removing flip augmentation increased overlap for three sets, left one unchanged, and decreased overlap for three.The largest decreases occurred for two sets, suggesting that only a fraction of mirror twins may have useful facial asymmetry.
7. Twins Training Images from Generative AI?
The paper tests whether generative AI can create useful images of imagined monozygotic twins for face-recognition training. Grok, ChatGPT, and Gemini fail to reproduce the identity-preserving variation and separated distance distributions observed for real twins.
- Motivation: 20% or more of training identities targeted to a condition produced noticeable accuracy increases in related recognition studies, but the needed fraction for MZ twins remains unknown.WebFace-scale training sets would therefore imply tens to hundreds of thousands of twin sets under this rule of thumb.
- Method: The study asks whether generative AI can create non-existent MZ-twin images suitable for augmenting face-matcher training data.The authors used the same prompt sequence with Grok v3, ChatGPT v5.4, and Gemini 3.1 Pro.
- Grok: Grok produced pose, expression, hairstyle, and age variation, but some images were implausibly non-human and same-person images were not identity-preserving.Its same-person and different-person distance distributions were effectively wholly overlapped.
- ChatGPT: ChatGPT images appeared plausibly human and showed hairstyle and expression variation, yet same-person and different-person distance distributions were effectively wholly overlapped.The images lacked the pose and age variation evident in the Grok images.
- Gemini: Gemini repeatedly stated that it could not produce images satisfying the prompt and recommended using an established labeled dataset of real human twins.Its responses described the requested generation as not technically possible.
- Discussion: The initial experiments indicate that generative AI lacks explicit information about the natural facial differences that distinguish real MZ twins.The paper identifies this missing representation as a fundamental problem for generating useful imagined-twin training data.
8. Conclusions and Discussion
Distinguishing monozygotic twins remains challenging, but CTTS enables direct evaluation of face matchers on this problem. Existing matchers exceed chance performance while apparently not relying on facial marks or asymmetry, leaving the discriminative features unresolved.
- Conclusions and Discussion: All 12 matcher versions in Table 1 exceed 70% accuracy on CTTS, although distinguishing MZ twins remains a current open problem.If twin faces were literally identical, positive and negative distance distributions would not separate and expected accuracy would be 50%.
- Conclusions and Discussion: CTTS supports research on MZ-twin recognition with explicit attention to skin marks and asymmetry.Its accuracy is lower than on test sets emphasizing age, pose, hairstyle, or illumination differences.