Source-linked AI summary

UFPR-PEs: A Brazilian Face Recognition Benchmark with Self-Declared Race/Color Labels

Alexandre Diano, Bernardo Biesseck, Gabriel Polo, Vinicius Gregorio, Laura Lopes, Diego Addan, David Menotti

arXiv:2608.30688v1cs.CV

TL;DR

Face-recognition fairness evaluation needs evidence that covers both demographic reliability and uncontrolled visual conditions. UFPR-PEs addresses this by combining official self-declared Brazilian race/color labels with compressed public videos and difficulty-stratified evaluation, finding that performance and subgroup gaps vary with image quality.

  • Problem

    Face-recognition benchmarks need to assess demographic reliability and robustness beyond curated images, including Brazilian demographic categories not represented in common schemas.

  • Method

    The paper constructs UFPR-PEs from elected politicians’ public videos, official self-declared race/color records, and explicit visual-difficulty stratification.

  • Results

    Recognition performance is high across race/color groups in easy and medium subsets but drops substantially for all groups in hard conditions, where relative group differences widen.

  • Takeaways & Limitations

    Demographic subgroup performance should be interpreted jointly with image quality, and benchmarks should preserve difficult samples for realistic fairness analysis.

  • Takeaways & Limitations

    The labels represent administrative self-declarations and may inherit inconsistencies from the institutional and political context of electoral registration.

Abstract

from arXiv · show

While face recognition systems are widely deployed, ensuring their demographic reliability and robustness under uncontrolled visual conditions remains a critical challenge. To bridge this gap, we present UFPR-PEs, a benchmark for face recognition bias evaluation using public videos of elected Brazilian politicians annotated with official self-declared race/color categories. The dataset adopts the Brazilian census taxonomy, including the parda category, which has no direct equivalent in the U.S.- or Europe-centric schemas commonly used in prior benchmarks. Our benchmark is built from compressed public video and preserves difficult samples so that performance can be analyzed under realistic conditions. We describe the construction pipeline, report dataset statistics, and evaluate face recognition performance across verification and (closed- and open-set) identification settings, including subgroup analysis by race/color and difficulty level. The results show that recognition performance varies substantially with image quality, and that subgroup gaps must be interpreted jointly with visual difficulty rather than in isolation. Overall, UFPR-PEs provides a reproducible and demographically grounded setting for studying face recognition bias under challenging public video conditions.

I. Introduction

UFPR-PEs introduces a Brazilian benchmark for face-recognition bias evaluation using official self-declared demographic records and difficult public-video imagery. It preserves uncontrolled conditions and supports analysis of demographic disparities across visual difficulty levels.

  • Benchmark motivation and contribution: The benchmark targets demographic reliability and robustness because recognition errors can vary across groups beyond curated image settings.This makes benchmark design central to fair face-recognition evaluation.
  • Benchmark motivation and contribution: UFPR-PEs uses public videos of elected Brazilian politicians annotated with official self-declared race/color labels.The labels come from official records rather than third-party annotations or automatic labeling.
  • Benchmark motivation and contribution: UFPR-PEs includes Brazil’s parda category, which has no direct equivalent in common U.S.- or Europe-centric benchmark taxonomies.The dataset follows Brazilian race/color classifications grounded in the national demographic context.
  • Benchmark motivation and contribution: The benchmark preserves compressed, uncontrolled, and difficult video samples instead of filtering them out.Its protocol is designed to analyze demographic disparities across difficulty strata.
  • Benchmark motivation and contribution: The paper presents a reproducible setting for studying bias under realistic public-video conditions and stratified visual difficulty.The paper reports construction, statistics, and fairness evaluations across recognition operating conditions.

II. Related Work

The related-work section situates UFPR-PEs within the evolution of face-recognition fairness benchmarks and frames the design limitations addressed by the paper.

  • Related work: Prior fairness benchmarks differ in demographic coverage, annotation methodology, visual conditions, and construction methodology.The section introduces these dimensions as the basis for reviewing representative datasets.

A. Bias Benchmarks

Existing face-recognition fairness benchmarks provide important subgroup evaluations but often rely on limited labels, curated imagery, or narrow demographic schemas. UFPR-PEs addresses these gaps with official Brazilian demographic records and challenging public-election videos.

  • Prior benchmarks: RFW, BFW, DemogPairs, and BUPT datasets established race-aware, balanced, pair-based, and large-scale subgroup evaluation approaches.These benchmarks also examined demographic imbalance and training composition effects.
  • Brazilian demographic context: Brazilian race/color categories include parda, whose semantics are not preserved by direct English translations or ancestry-based schemas.The category reflects social and administrative self-declaration rather than a genealogical claim.
  • Prior benchmarks: Prior benchmarks often use annotator or classifier labels, curated still images, narrow geographic coverage, and evaluation settings that may not reflect public video.These limitations motivate alternative annotation sources and visual conditions.
  • UFPR-PEs design: UFPR-PEs uses election-related videos with gallery portraits, query frames, distractor identities, and manual verification of extracted images.The benchmark evaluates challenging conditions including illumination, compression, blur, pose, expression, and occlusion.
  • UFPR-PEs design: UFPR-PEs trims source videos to 30-second segments and retains difficult samples for realistic robustness and fairness evaluation.The source media typically ranges from 360p to 720p and uses H.264 encoding.

A. Dataset Statistics

UFPR-PEs combines self-declared demographic attributes with an automated-plus-manual video-to-identity pipeline and explicit cosine-similarity difficulty strata. These design choices support demographic analysis across progressively degraded probe frames.

  • Dataset attributes: All identities have self-declared race/color, gender, and age labels aligned with official Brazilian categories.The labels are grounded in legal and administrative records, and their distributions are summarized in Fig. 2.
  • Construction and annotation: RetinaFace detection and cosine-similarity gallery matching select candidate faces using thresholds of 0.3 and 0.2 during construction.The lower-threshold second pass is intended to recover challenging target faces missed initially.
  • Dataset statistics: Table II details the composition of the UFPR-PEs gallery and query sets.The table is the benchmark’s stated summary of dataset statistics.
  • Construction and annotation: Every extracted frame is manually reviewed to correct automatic selections and identify target faces missed or incorrectly discarded.Manual inspection compares the gallery portrait with all frames from each video.
  • Difficulty stratification: Easy probes have CS ≥0.60, medium probes have 0.30 < CS < 0.60, and hard probes have CS ≤0.30.The strata correspond respectively to high-quality frontal, ambiguous, and heavily degraded or profile faces.

C. Access Policy and Ethics

UFPR-PEs is restricted to academic purposes and access is available upon request, with ethics approval documented for its construction and use. The paper also states that publication rights for the data are limited.

  • C. Access Policy and Ethics: UFPR-PEs is restricted to academic purposes and available upon request.Its release policy and institutional access requirements are documented separately.
  • C. Access Policy and Ethics: The dataset construction was approved by the Federal University of Paraná ethics committee and registered in Plataforma Brasil.The cited process is CAAE 93419925.4.0000.0214.
  • C. Access Policy and Ethics: The authors obtained authorization to collect, handle, and analyze publicly available politician data but do not have rights to publicly release it.
  • C. Access Policy and Ethics: Individuals whose images appear in the paper provided written informed consent for publication.

IV. Results

The evaluation measures verification and identification performance across easy, medium, and hard subsets using a pretrained R50 model with ArcFace loss and cosine similarity. Severe visual degradation causes genuine and impostor similarity distributions to overlap substantially.

  • IV. Results: The study evaluates verification and identification performance across three difficulty subsets.The reported experiments include verification (1:1) and identification (1:N).
  • IV. Results: R50 produces 512-dimensional face embeddings from normalized 112×112 aligned face images for cosine-similarity comparisons.The model is a ResNet50 trained with ArcFace loss and pretrained on WebFace260M.
  • IV. Results: High-quality easy faces yield near-perfect genuine–impostor separability, while hard faces produce substantial distribution overlap.Severe occlusions, low resolution, and profile poses reduce genuine similarity scores in the hard subset.

A. Overall Performances

Overall recognition performance is high on easy and medium subsets but collapses under severe degradation in the hard subset. Verification and closed-set identification both show this difficulty-dependent pattern.

  • A. Overall Performances: Verification and closed-set identification performance are visualized with ROC and CMC curves across the three difficulty subsets.
  • A. Overall Performances: AUC is 1.000 for easy, 0.999 for medium, and 0.440 for hard verification subsets.The hard subset falls below chance level, indicating a collapse under severe degradation.
  • A. Overall Performances: 99.97%, 98.97%, and 19.10% rank-1 accuracy are achieved on easy, medium, and hard closed-set subsets, respectively.The corresponding rank-5 accuracy reaches only 41.49% for hard samples.
  • A. Overall Performances: The benchmark reports closed-set rank-1 performance by race/color and difficulty subset, including hard-set CMC curves by demographic group.

B. Bias by Race/Color Groups

Race/color differences are modest on easy and medium subsets but become more pronounced under hard visual conditions. In hard-set closed-set identification, the white group performs best, while yellow and indigenous groups have lower cumulative accuracy.

  • B. Bias by Race/Color Groups: Easy and medium closed-set subsets achieve very high rank-1 accuracy across all race/color groups.Performance approaches perfect recognition in these subsets.
  • B. Bias by Race/Color Groups: Hard-subset closed-set performance drops substantially across every race/color group.
  • B. Bias by Race/Color Groups: In hard-set CMC curves, white maintains the highest identification rates, followed by black and parda with similar trajectories.
  • B. Bias by Race/Color Groups: Yellow and indigenous categories show lower cumulative accuracy in hard closed-set identification.Increasing the candidate-list rank improves accuracy for all demographics, but group gaps persist through Rank-20.
  • B. Bias by Race/Color Groups: Open-set groups show moderate performance differences, with indigenous highest at rank-1 and yellow lowest.The performance gap widens in the hard subset across difficulty levels.

V. Limitations

UFPR-PEs has limitations involving its administrative demographic labels, politically shaped public-video composition, detector dependence, and evaluation of only one model family.

  • Self-declared race/color labels represent administrative classifications rather than fixed biological categories and may reflect electoral institutional and political contexts.The records reduce annotation noise and provide an official source, but may inherit inconsistencies from registration contexts.
  • Public political videos shape demographic and regional composition through online-material availability and election context, potentially leaving some groups underrepresented.Material quantity and quality can vary across regions and groups, affecting probe-set sizes and difficulty distributions even after stratified sampling.
  • Automated gallery search introduces residual detector dependence because faces it fails to localize never reach search or annotation review.The reported gaps are expected to underestimate rather than overstate race/color disparities because undetected faces concentrate in the most degraded observations.
  • The evaluation covers only R50 with ArcFace loss pretrained on WebFace260M, so observed demographic gaps may not generalize across models.Future evaluations are needed to assess whether disparities are model-specific or inherent to benchmark conditions.

VI. Conclusion

UFPR-PEs is a benchmark built from compressed public political video, official self-declared race/color labels, and explicit difficulty stratification. Results show that recognition gaps vary with image quality, widening in hard samples, so subgroup performance should be interpreted jointly with difficulty.

  • UFPR-PEs combines compressed public political video, TSE self-declared race/color labels, uncontrolled conditions, and explicit difficulty stratification.The benchmark is designed to study whether demographic gaps persist and vary across recognition operating conditions.
  • Open-set identification analysis examines R50+ArcFace performance by race/color and difficulty subsets on UFPR-PEs.
  • In easy and medium subsets, recognition performance is high across race/color groups, whereas hard-subset performance drops substantially for every group.
  • Relative differences between race/color groups widen in the hard subset, indicating that subgroup gaps are not constant across difficulty levels.
  • UFPR-PEs provides a reproducible, demographically grounded setting for studying face recognition bias under realistic public-video conditions, including Brazil’s parda category.Future work will extend evaluation to additional model families, alternative operating points, and different training-data compositions.
Loading 2608.30688v1…