Source-linked AI summary

BiosecurID: a multimodal biometric database

Julian Fierrez, Javier Galbally, Javier Ortega-Garcia, Manuel R Freire, Fernando Alonso-Fernandez, Daniel Ramos, Doroteo Torre Toledano, Joaquin Gonzalez-Rodriguez, Juan A Siguenza, Javier Garrido-Salas, E Anguiano, Guillermo Gonzalez-de-Rivera, Ricardo Ribalda, Marcos Faundez-Zanuy, JA Ortega, Valentín Cardeñoso-Payo, A Viloria, Carlos E Vivaracho, Q Isaac Moro, Juan J Igarza, J Sanchez, Inmaculada Hernaez, Carlos Orrite-Urunuela, Francisco Martinez-Contreras, Juan José Gracia-Roche

arXiv:2111.03472v1cs.CRcs.CVeess.IV

TL;DR

BiosecurID addresses the need for realistic, representative multimodal biometric data suitable for moving beyond controlled laboratory evaluation. The paper presents and characterizes a 400-subject, multisession database spanning many biometric traits and realistic acquisition conditions, while noting that difficult-to-detect acquisition errors may remain and require future updates.

  • Problem

    Realistic multimodal biometric data are needed to support valid biometric-system results beyond controlled laboratory conditions and avoid combining unrelated databases into “chimerical subjects.”

  • Method

    The paper constructs and describes a multisession database collected in realistic conditions, including diverse biometric traits, demographic metadata, attack samples, and skilled forgeries.

  • Results

    The resulting BiosecurID database contains 400 subjects, eight biometric-trait categories, four sessions over four months, and low-quality samples retained as a realistic benchmark feature.

  • Takeaways & Limitations

    BiosecurID supports research on recognition, temporal variability, sample quality, demographic effects, sensor interoperability, and template adaptation or update.

  • Takeaways & Limitations

    Some acquisition errors are difficult to detect even after validation and post-editing, so future updated database versions may be needed.

Abstract

from arXiv · show

A new multimodal biometric database, acquired in the framework of the BiosecurID project, is presented together with the description of the acquisition setup and protocol. The database includes eight unimodal biometric traits, namely: speech, iris, face (still images, videos of talking faces), handwritten signature and handwritten text (on-line dynamic signals, off-line scanned images), fingerprints (acquired with two different sensors), hand (palmprint, contour-geometry) and keystroking. The database comprises 400 subjects and presents features such as: realistic acquisition scenario, balanced gender and population distributions, availability of information about particular demographic groups (age, gender, handedness), acquisition of replay attacks for speech and keystroking, skilled forgeries for signatures, and compatibility with other existing databases. All these characteristics make it very useful in research and development of unimodal and multimodal biometric systems.

1 Introduction

BiosecurID addresses the need for realistic, sufficiently large multimodal biometric data that can support valid conclusions beyond controlled laboratory experiments. It presents a multisession database designed to avoid methodological problems associated with combining unrelated unimodal datasets.

  • Motivation: Realistic multimodal biometric data are needed to narrow the gap between laboratory error rates and practical authentication deployments.The motivation is to support valid inference from controlled experiments toward final applications.
  • Contribution: The BiosecurID project acquired a realistic, multisession database intended to represent potential biometric-application users and support valid results.The database was conducted by a consortium of six Spanish universities.
  • Methodological problem: BiosecurID helps address the use of “chimerical subjects,” created by combining separate unimodal databases under an independence assumption questioned in prior work.The paper characterizes this practice as a serious methodological flaw.
  • Design scope: The database combines many subjects, biometric traits, and temporally separated sessions with realistic acquisition, demographic information, attacks, forgeries, and database compatibility.These characteristics are presented as useful for developing and testing automatic recognition systems.
  • Paper structure: The paper describes related databases, BiosecurID acquisition and protocol, validation and post-processing, and potential database uses.This is the paper’s stated organizational structure.

2 Related works

Prior multimodal biometric databases vary in modalities, scale, acquisition settings, and session structure. The paper summarizes these alternatives using a common feature table to position BiosecurID among existing resources.

  • Existing databases: Existing databases include combinations of face, speech, fingerprints, iris, hand, handwriting, signature, and other biometric traits.Examples include BIOMET, MyIDEA, BIOSEC, BIOSECURE, and other datasets.
  • Existing databases: BIOSECURE extends earlier related databases toward approximately 1,000 subjects under multiple realistic acquisition conditions.Its scenarios include Internet, desktop, and mobile-device datasets.
  • Existing databases: The related datasets differ in subject counts, sessions, devices, and modalities, with some sharing subjects across the whole database.One described resource has around 1,000 subjects, while two others have about 700 users and roughly 400 common subjects.
  • Comparison framework: The comparison table consolidates the most relevant database features, treating palmprint and palm geometry as Hand and on-line and off-line signature as Signature.When session sizes differ, the table reports subjects common to all sessions.
  • Comparison framework: The table uses abbreviations for subject counts and modalities, including Face, Fingerprint, Hand, Handwriting, Iris, Keystroking, Signature, and Speech.The caption defines the notation used in the comparison.

3 The BiosecurID database

BiosecurID was collected across six realistic office-like sites and covers a broad set of biometric traits, demographic attributes, subjects, and acquisition times. Its multisession design captures temporal variability relevant to several traits.

  • Database scope: BiosecurID was collected at six sites in an uncontrolled office-like environment designed to simulate realistic conditions.The database emphasizes both environmental realism and a broad variety of biometric traits.
  • Database scope: The database includes speech, iris, face photographs and talking-face videos, on-line and off-line signature and handwriting, fingerprints, hand traits, and keystroking.Hand includes palmprint and contour geometry.
  • Scale and sessions: 400 subjects were acquired across six sites, with the site counts reported for UAM, UPC Mataro, UPC Terrasa, UVA, UPV/EHU, and UniZar.The site-specific counts are 130, 40, 35, 77, 52, and 66, respectively.
  • Scale and sessions: Four sessions were distributed across four months, representing within-session, between-weeks, and between-months temporal variability.The paper highlights this design for traits such as face, speech, handwriting, and signature.
  • Demographic information: Recruitment targeted balanced demographic groups, including specified age-band proportions and a forced gender balance.The reported recruitment process had difficulty obtaining older donors relative to younger participants.
  • Demographic information: The database stores age, gender, handedness, manual-worker status, and vision-aid information to support experiments involving demographic groups.Manual-worker status identifies contributors reporting work associated with eroded fingerprints.

4 Acquisition environment

Each acquisition site used a standardized kiosk with relaxed environmental guidance, while human supervision and centralized software supported protocol compliance and sample management. Manual verification addressed invalid data after acquisition.

  • Acquisition environment: Each site prepared an acquisition kiosk with neutral lighting, limited background noise, and seated contributors in non-revolving chairs.The relaxed conditions were intended to create realistic variability between sites.
  • Protocol control: A human operator instructed contributors during acquisition, but human and software errors still occurred.The protocol therefore required additional verification after collection.
  • Protocol control: A human expert manually verified every biometric sample and corrected or discarded non-valid data.The validation guidelines are described in a later section.
  • Acquisition environment: The setup included acquisition devices connected to a standard PC running software designed around the database protocol.Figure 1 illustrates an example site setup and some acquisition devices.
  • Protocol control: The acquisition software centralized device control, sample naming and storage, and database management to minimize acquisition errors.It coordinated the functioning and launching of the connected devices.

5 Acquisition protocol

BiosecurID used a centralized software-and-device setup to collect multimodal biometric data across four sessions, with consent, supervision, and detailed trait-specific protocols. The protocol included genuine samples and designed impostor data across speech, fingerprints, iris, hand, face, handwriting, signatures, and keystroking.

  • Physical traits: The protocol captured fingerprints with two sensors, four samples per subject, alternating index and middle fingers across both hands.The interleaving was designed to introduce intravariability among images of the same fingerprint.
  • Physical traits: Each iris and hand was sampled four times, alternating eyes for iris captures and alternating hands for hand images.Glasses were removed for iris acquisition, while the hand scanner was isolated from external illumination.
  • Face and handwriting: Face data included four frontal still images and a five-second talking-face video with synchronized audio while prohibiting moving background objects.The video recorded the subject saying an eight-digit PIN.
  • Face and handwriting: Handwriting was collected as common Spanish text, digits, and uppercase words using an inking pen, yielding both on-line signals and off-line scanned images.The protocol prohibited corrections or crossing out in the lowercase text.
  • Forgeries and keystroking: Signatures included four genuine samples per session and one forgery from each of three preceding subjects, with progressively skilled imitation scenarios.Keystroking captured repeated subject names and surnames plus the names of three other subjects as forgeries.

6 Validation process

Validation combined supervised reacquisition with expert review to reduce invalid samples while retaining low-quality data from the uncontrolled acquisition setting. Subjects and samples were completed, corrected, copied, or removed according to protocol-compliance and completeness rules.

  • Definitions: Invalid samples violated acquisition specifications, whereas low-quality samples were expected to perform poorly in automatic recognition systems.Examples included mislabeled fingers or wrong PINs for invalid samples, and noisy voice or blurred iris images for low-quality samples.
  • Validation objective: Low-quality samples were deliberately retained because they reflect the non-controlled scenario and support realistic evaluation of biometric systems.The validation process sought to reduce invalid samples, not reject low-quality samples.
  • Validation stages: Validation proceeded in two stages: supervised sample-by-sample checking and reacquisition, followed by expert review of all samples.The expert completed missing data, corrected invalid samples, or removed incomplete subjects.
  • Subject-level rules: Subjects missing any of the four sessions were removed from the database.Subjects with substantial missing or invalid biometric data in one or more sessions were also removed.
  • Subject-level rules: Subjects with fewer than approximately 10% missing or invalid genuine samples retained those subjects after the samples were copied from valid samples in the same session.The resulting identical genuine samples can be filtered before experiments because they are the only identical genuine samples.
  • Forgery handling: Missing or invalid forgeries in PIN, signature, or keystroking data were produced by the verifying expert.This rule applied specifically to forgery samples rather than genuine samples.
  • Residual errors: Some acquisition errors may remain undetected until database use, making future updated releases likely.The limitation persists despite careful acquisition and post-editing efforts.

7 Compatibility with other databases

BiosecurID was designed for interoperability with existing biometric databases, enabling combined experiments, larger subject pools, and long-term variability studies.

  • Compatibility with BIOSEC and BIOSECURE: Compatible sensors and protocols allow BiosecurID to be combined with BIOSEC and BIOSECURE for interoperability and long-term variability research.BIOSEC and BiosecurID share 37 subjects with about one year between acquisitions; BiosecurID and BIOSECURE share 29 subjects with a similar interval.
  • Compatibility with BIOSEC and BIOSECURE: Combining BiosecurID with BIOSEC produces a multimodal database of 650 subjects.
  • Compatibility with BIOSEC and BIOSECURE: BiosecurID is compatible with BIOSECURE for optical/thermal fingerprints, iris, and signature.
  • Compatibility with other databases: BiosecurID portions can also be combined with MyIDEA or MCYT to increase subject numbers, but without common subjects to add sessions.
  • Compatibility with other databases: Table 5 summarizes BiosecurID's main compatibilities with existing multimodal databases.

8 Potential uses of the database

BiosecurID supports research across its eight modalities and multimodal systems, including effects of time, sample quality, demographics, sensors, and attacks under realistic acquisition conditions.

  • Modalities and multimodal systems: The database supports research in any of its 8 modalities and in multibiometric systems combining them.
  • Temporal effects: Its multisession design enables short-, medium-, and long-term evaluation of time effects, including template adaptation and update.Long-term studies use common subjects across BIOSEC/BIOSECURE and BiosecurID.
  • Data quality and demographics: Realistic uncontrolled acquisition and retention of low-quality samples support studies of sample quality effects on multibiometric systems.
  • Data quality and demographics: Balanced age and gender distributions support research on age-related recognition rates, compensation, and gender-dependent performance.
  • Sensor interoperability: Multiple devices for fingerprint and speech enable evaluation of sensor interoperability and its effect on multibiometric systems.
  • Security evaluation: The database's size and number of modalities support evaluation of potential attacks against unimodal and multibiometric systems.

9 Conclusions

The paper presents BiosecurID as a large, multisession multimodal database collected under realistic conditions to address the shortage of public real-world biometric data and support recognition research.

  • Motivation: Large public multimodal databases acquired under real working conditions remain scarce for developing, testing, and evaluating biometric recognition systems.
  • Contribution: The paper describes BiosecurID's relevant features and places it within an overview of existing multimodal biometric databases.
  • Database scope: BiosecurID covers speech, iris, face, signature, handwriting, fingerprints, hand traits, and keystroking across 400 subjects.
  • Database scope: The database was captured in 4 sessions over a 4 month time span.
  • Availability: Distribution details are provided through the BiosecurID Multimodal Biometric Database website.
Loading 2111.03472v1…