Source-linked AI summary
The MegaFace Benchmark: 1 Million Faces for Recognition at Scale
Ira Kemelmacher-Shlizerman, Steve Seitz, Daniel Miller, Evan Brossard
TL;DR
Face recognition systems that perform near perfectly on small benchmarks must be evaluated for accurate identification at planetary scale. MegaFace introduces a public million-face benchmark with identification and verification tests, showing that performance differences, age variation, and pose effects become more pronounced at scale.
Problem
Planetary-scale applications require finding true matches among billions of people with negligible false positives, but existing benchmarks and datasets contain too few people for this setting.
Method
MegaFace provides a publicly available benchmark evaluating face recognition with up to one million unconstrained distractor photos, using identification and verification scenarios.
Results
35-75% identification rates with 1M distractors were achieved by algorithms exceeding 95% on LFW, while Joint Bayes and LBP dropped below 10%; larger pose and age differences remained challenging.
Takeaways & Limitations
Testing at scale exposes performance differences among algorithms that appear similar on smaller benchmarks and reveals that age and pose variation remain challenging for recognition.
Takeaways & Limitations
Datasets containing 8M and 500M people were not publicly available and were used only for training, not testing.
Abstract
from arXiv · showhide
Recent face recognition experiments on a major benchmark LFW show stunning performance--a number of algorithms achieve near to perfect score, surpassing human recognition rates. In this paper, we advocate evaluations at the million scale (LFW includes only 13K photos of 5K people). To this end, we have assembled the MegaFace dataset and created the first MegaFace challenge. Our dataset includes One Million photos that capture more than 690K different individuals. The challenge evaluates performance of algorithms with increasing numbers of distractors (going from 10 to 1M) in the gallery set. We present both identification and verification performance, evaluate performance with respect to pose and a person's age, and compare as a function of training data size (number of photos and people). We report results of state of the art and baseline algorithms. Our key observations are that testing at the million scale reveals big performance differences (of algorithms that perform similarly well on smaller scale) and that age invariant recognition as well as pose are still challenging for most. The MegaFace dataset, baseline code, and evaluation scripts, are all publicly released for further experimentations at: megaface.cs.washington.edu.
1. Introduction
MegaFace addresses the gap between near-perfect small-scale benchmark results and the demands of recognition with millions of unrelated identities. It introduces a public, unconstrained benchmark that evaluates algorithms across scale, training data, age, and pose.
- MegaFace evaluates face recognition with up to 1M distractors, or people absent from the test set.
- MegaFace uses Flickr’s publicly releasable Creative Commons photos because celebrity datasets contain too few unique individuals.
- Algorithms above 95% on LFW achieve only 35–75% identification with 1M distractors, while Joint Bayes and LBP fall below 10%.
- Larger training sets tend to improve performance at scale, although FaceN compares favorably with FaceNet on FaceScrub despite less training data.
- Age gaps and child probes are more challenging, with FG-NET showing lower performance and a larger decline than FaceScrub as distractors increase.
- Recognition drops with greater pose variation between matching probe and gallery images, and the effect is more significant at scale.
2. Related Work
Prior face recognition datasets progressed toward unconstrained imagery but remained limited in public identity diversity or lacked large-scale testing. MegaFace builds on this gap while extending evaluation to age-invariant recognition.
- Earlier benchmarks moved from controlled imagery toward unconstrained conditions, but collecting photos of many individuals remained difficult.
- NIST reported 90% recognition on 1.6M controlled subjects, but those results were not representative of photos in the wild.
- LFW contains 13K photos of 5K people and became a major unconstrained benchmark, despite its limited identity count.
- Massive private training datasets are unavailable publicly and were used for training rather than testing.
- Public CASIA-WebFace contains 500K photos of 10K celebrities but has no associated benchmark and is used for training rather than testing.
- MegaFace augments FG-NET with 1M distractors because most modern recognition algorithms had not been evaluated for age invariance.
3. Assembling MegaFace
MegaFace was assembled from public Flickr imagery to provide a broad, unconstrained distractor set with close to a million likely unique faces. Its statistics span identity diversity, geography, pose, group composition, and resolution.
- The dataset was designed as a public, licensing-unrestricted collection of unconstrained photos with close to 1M unique identities.
- Identity diversity was increased by sampling from 500K Flickr user IDs and treating multiple faces in one photo as likely different identities.
- 1,296,079 downloaded faces yielded 690,572 faces with a high probability of representing unique individuals after stricter detection and blur removal.
- The collection includes worldwide GPS locations, varied Flickr tags, diverse people and poses, and group photographs.
- More than 197K faces have yaw angles beyond ±40°, exceeding the typical less-than-±30° range of unconstrained face datasets.
- More than 50% of MegaFace photos, totaling 514K, exceed 40 pixels of interocular distance, corresponding to LFW’s approximately 100x100 face size.
4. The MegaFace Challenge
The MegaFace Challenge tests recognition against galleries containing up to 1M distractor faces using both identification and verification protocols. It evaluates rank-based retrieval and false-accept versus false-reject tradeoffs across FaceScrub and FG-NET probes.
- MegaFace galleries contain up to 1M faces of unknown people, while FaceScrub and FG-NET provide the probe identities.
- Recognition scenarios: Identification rank-orders gallery images for each probe and reports the probability that a correct image appears within rank K using CMC curves.
- Recognition scenarios: Verification classifies image pairs as same or different identities and reports ROC curves over 4B negative pairs.
- The challenge emphasizes large distractor counts and evaluates both identification and verification rather than relying only on LFW-style verification.
- Probe sets: FaceScrub contributes 100K photos of 530 celebrities, while FG-NET contributes 975 photos of 82 people spanning substantial age ranges.
- Evaluation and baselines: Participants computed features on MegaFace and were evaluated with L2 distance without training on FaceScrub or FG-NET.
5. Results
The MegaFace challenge evaluates recognition across identification and verification tasks as gallery size, training data, age, and pose vary. At million-distractor scale, performance drops substantially, with larger training sets generally helping and age and pose differences remaining difficult.
- More than 100 groups registered, but results from 5 groups that uploaded all features by the deadline are presented.
- Participating algorithms: FaceNet uses more than 500M photos of 10M people, while FaceN’s large model uses more than 18M photos of 200K people.
- Verification results: At low false accept rates, FaceScrub verification performance drops by about 40% on average as scale increases, while FaceNet and FaceN drop by only about 15%.For MegaFace, FAR of 10^-5 or 10^-6 is meaningful, compared with 1%-5% FAR typically implied by LFW equal-error-rate reporting.
- Verification results: For FGNET, verification performance drops by about 60% for everyone but FaceNet, which achieves impressive performance across the board.
- Identification results: Identification rates drop for all algorithms as gallery size increases, and FGNET at scale reveals a dramatic performance gap except for FaceNet.The curves suggest rates will be even lower beyond 1M distractors, such as at 100M.
- Training set size: Algorithms trained on data larger than 500K photos and 20K people generally perform better than others.
- Age: Age differences reduce identification performance, with small gallery-probe age gaps performing better and adults matched more accurately than children at scale.The age analysis compares 1K and 1M distractors across algorithms.
- Pose: Recognition accuracy depends strongly on pose; similar and more frontal poses are easier to recognize, with variation becoming more dramatic at 1M distractors.
6. Discussion
MegaFace shows that face-recognition performance degrades with a large gallery, exposes differences between algorithms, and makes age variation especially challenging. The benchmark is intended as a step toward recognition across billions of people.
- Performance degrades as the gallery grows, even when the probe set remains fixed.
- Million-scale testing distinguishes algorithms that appear similarly strong at smaller scales.
- Age differences between probe and gallery remain especially challenging for recognition.
- MegaFace provides a million-face benchmark as a first step toward evaluating recognition with billions of people.
7. Supplementary material
The supplementary material tests result stability across three random MegaFace gallery subsets and presents identification and verification results across distractor scales and probe sets.
- Three random MegaFace subsets were evaluated for each distractor size from 10 to 1M.All algorithms were rerun on each subset for both FaceScrub and FGNET probe sets.
- Set #1 is the random gallery subset presented in the main paper.
- The supplementary figures compare verification performance at 1M and 10K distractors across the two probe sets.Figures 12–14 focus on performance at low false accept rates.
- The supplementary figures compare identification at 1M and 10K distractors and report rank-10 performance for both probe sets.