Source-linked AI summary

Ear Recognition: More Than a Survey

Žiga Emeršič, Vitomir Štruc, Peter Peer

arXiv:1611.06203v2cs.CV

TL;DR

Ear recognition is an active field with promising contactless, nonintrusive acquisition, but unresolved challenges limit wider commercial deployment. This paper surveys descriptor-based 2D methods, introduces a publicly available unconstrained web dataset and toolbox, and shows that the dataset remains difficult for existing techniques.

  • Problem

    Open challenges in ear recognition remain insufficiently addressed despite active research, limiting the technology’s wider commercial deployment.

  • Method

    The paper surveys ear recognition approaches, compares reported techniques, and introduces a publicly available web-gathered dataset and toolbox of state-of-the-art methods.

  • Results

    The AWE dataset produced mean rank-1 rates from 39.8% to 49.6%, with POEM highest, while the best verification EER was around 30%.

  • Takeaways & Limitations

    The survey, dataset, and toolbox provide a public reference point for evaluating and advancing ear recognition research.

  • Takeaways & Limitations

    Ear detection remains unsolved, with no widely adopted solution and insufficient large-scale real-world datasets for training competitive detectors.

Abstract

from arXiv · show

Automatic identity recognition from ear images represents an active field of research within the biometric community. The ability to capture ear images from a distance and in a covert manner makes the technology an appealing choice for surveillance and security applications as well as other application domains. Significant contributions have been made in the field over recent years, but open research problems still remain and hinder a wider (commercial) deployment of the technology. This paper presents an overview of the field of automatic ear recognition (from 2D images) and focuses specifically on the most recent, descriptor-based methods proposed in this area. Open challenges are discussed and potential research directions are outlined with the goal of providing the reader with a point of reference for issues worth examining in the future. In addition to a comprehensive review on ear recognition technology, the paper also introduces a new, fully unconstrained dataset of ear images gathered from the web and a toolbox implementing several state-of-the-art techniques for ear recognition. The dataset and toolbox are meant to address some of the open issues in the field and are made publicly available to the research community.

I. Introduction

Ear recognition is an active biometric research area with appealing contactless, nonintrusive acquisition and distinctive ear features, but unresolved challenges limit broader deployment. This paper surveys recent methods and contributes a new in-the-wild dataset, an open-source toolbox, and reproducible comparative evaluation.

  • Motivation: Ear images can be acquired contactlessly and nonintrusively, including from profile photographs or video, without requiring the subject’s cooperation.
  • Motivation: Ear distinctiveness, including reported distinctions between identical twins, supports security applications and multimodal biometric use.
  • Open challenges: Despite active research and public datasets, commercial ear recognition remains limited, with the paper attributing this gap to unresolved open challenges.
  • Contributions: The paper surveys 2D ear recognition through 2015, emphasizing descriptor-based approaches and discussing open problems and challenges.
  • Contributions: The Annotated Web Ears dataset provides publicly available ear images collected from the web and presents a challenging, highly variable recognition setting.
  • Contributions: The open-source AWE toolbox implements state-of-the-art feature extraction and evaluation tools, while consistent protocols support direct, reproducible comparisons.

II. Ear Recognition Essentials

Ear structure varies across individuals and supports identity recognition, while 2D systems typically detect, segment, align, normalize, describe, and classify ear images. The paper organizes these approaches into geometric, holistic, local, and hybrid categories.

  • A. Ear Structure: Ear cartilage formations vary in shape, appearance, and relative position across people, providing features for identity recognition.
  • A. Ear Structure: Ear size increases with age while width remains relatively constant, but the effect of growth on automatic recognition remains unresolved.Available longitudinal datasets are not sufficiently long; initial studies used images captured less than a year apart.
  • C. Ear Recognition Approaches: 2D ear recognition systems detect a profile face, segment and align the ear, normalize illumination, extract features, and classify identity.
  • C. Ear Recognition Approaches: The paper divides 2D approaches into geometric, holistic, local, and hybrid methods according to their feature extraction strategies.Geometric methods model ear geometry; holistic methods encode global appearance; local methods describe local regions; hybrid methods combine categories or representations.
  • C. Ear Recognition Approaches: Table I reports rank-1 recognition rate, equal error rate, or verification rate, but results are not directly comparable because protocols usually differ.

1) Geometric Approaches:

Geometric approaches describe ears through contours, points, shapes, and spatial relationships, often using computationally simple edge-based processing. Their reliability is constrained by sensitivity to imaging conditions and difficult point localization.

  • 1) Geometric Approaches:: Geometric approaches are generally computationally simple and often use edge detection to derive geometry-related representations.This processing discards other image information that may contain discriminative value.
  • 1) Geometric Approaches:: Geometric methods exploit ear shape, characteristic points, contours, angles, and relationships among selected parts for recognition.
  • 1) Geometric Approaches:: Edge detectors are susceptible to illumination variation and noise, while exact ear-point localization can fail in poor-quality or occluded images.

2) Holistic Approaches:

Holistic methods represent global ear appearance, whereas local methods describe neighborhoods or dense local regions; hybrid methods combine representations to improve recognition. These choices trade global structure against robustness to occlusion and computational complexity.

  • 2) Holistic Approaches:: Force Field Transform methods model image pixels as Gaussian force sources, analyze the resulting force field, and use its properties for image similarity.
  • 2) Holistic Approaches:: Local approaches describe neighborhoods around image points rather than relying on anatomically meaningful point locations or interpoint relationships.
  • 2) Holistic Approaches:: Keypoint-based local descriptors support partial matching and can improve robustness to partial occlusions, but typically lose global ear-structure information.
  • 2) Holistic Approaches:: Dense local descriptors retain global image structure but generally sacrifice robustness to partial occlusions.
  • 2) Holistic Approaches:: Hybrid approaches combine categories or multiple representations to increase recognition performance, though their building blocks can make them computationally more complex.

III. Existing Datasets

Existing ear datasets differ in image variability, acquisition conditions, subject coverage, and annotations of factors such as pose, occlusion, accessories, gender, and ethnicity. Their development has moved toward more realistic imaging conditions.

  • III. Existing Datasets: Existing datasets vary in characteristics and in the image variability they present, as summarized comparatively in Table II.
  • III. Existing Datasets: The CP dataset contains 102 ears from 17 subjects captured under controlled conditions, with variability mainly from minor pose changes and identity.
  • III. Existing Datasets: IITD datasets contain 493 images from 125 subjects and 793 images from 221 subjects, captured indoors at varied lighting and approximately one profile angle.Pre-processing tightly crops, resizes, centers, and aligns the ears; left ears are mirrored.
  • III. Existing Datasets: Figure 4 traces dataset development over time and illustrates progression toward more realistic imaging conditions.
  • III. Existing Datasets: USTB provides four datasets spanning cropped ears and complete head profiles, with examples including 185 images from 60 volunteers and 308 images from 77 volunteers.

E. WPUT Ear Dataset

The paper reviews diverse ear-recognition datasets spanning cropped ears, profile-face images, video, 2D/3D data, and ear-prints, with substantial variation in acquisition conditions and annotations.

  • WPUT Ear Dataset: WPUT provides thousands of ear images with annotations covering demographic, pose, lighting, background, and occlusion factors.The downloadable collection contains 3348 images from 471 subjects, including duplicates and missing subject IDs.
  • WPUT Ear Dataset: FEARID differs from image datasets by containing ear-prints acquired with scanning hardware, without occlusions, variable angles, or illumination variation.Ear-prints therefore expose different sources of appearance variability than regular ear images.
  • WPUT Ear Dataset: Several datasets support recognition under pose, illumination, motion, or occlusion variation, including WVU, UBEAR, IITK, and PIE.UBEAR specifically includes moving subjects, while IITK and PIE organize images across fixed or varied head poses, illumination conditions, and expressions.

A. The AWE Dataset

The AWE dataset addresses the limited variability of existing ear datasets by collecting tightly cropped ear images from the web and supplying annotations and standardized evaluation protocols.

  • The AWE Dataset: AWE was designed because many existing datasets use controlled or posed images with limited, author-selected variability.The paper contrasts this setting with uncontrolled image collections used in face recognition.
  • The AWE Dataset: AWE images were collected from the web using person-specific image-search queries and a web crawler.The subject list primarily included actors, musicians, politicians, and similar publicly represented people.
  • The AWE Dataset: The dataset contains 1000 tightly cropped ear images from 100 subjects, with 10 images per subject and substantial variation in image quality and size.Image sizes range from 15 × 29 to 473 × 1022 pixels, with an average of 83 × 160 pixels.
  • The AWE Dataset: Each image is annotated for gender, ethnicity, accessories, occlusion, head pose, and head side, with the tragus location also marked.Head pitch, roll, and yaw labels were estimated from head pose.
  • The AWE Dataset: The proposed protocols cover identification and verification, using development/test partitioning, cross-validation, and standardized performance metrics.Identification uses CMC and rank-1 rates; verification uses ROC curves and EERs.

B. The AWE Toolbox

The AWE toolbox provides an open-source, configurable pipeline for reproducible ear-recognition experiments, covering data handling, feature extraction, matching or classification, and evaluation.

  • The AWE Toolbox: The toolbox addresses reproducibility problems caused by gaps between published methods and available implementations.Its open-source design also supports modification, sharing, and unrestricted use.
  • The AWE Toolbox: Its four major components handle data, feature extraction, distance calculation or classification, and result visualization.The complete pipeline can be configured in different combinations and automatically executed on a supplied dataset.
  • The AWE Toolbox: The data-handling component reads, normalizes, caches, and organizes samples by class, without restricting use to ear images.The toolbox can also support other biometric modalities when datasets follow the required directory structure.
  • The AWE Toolbox: The feature-extraction component applies preselected techniques and can export computed feature vectors for later experiments or external tools.The toolbox implements eight descriptor-based feature-extraction techniques.
  • The AWE Toolbox: Matching supports distance measures or user-trained classifiers, while evaluation generates ROC, CMC, score-distribution, rank-1, and EER outputs.Specific classifiers are not included, although classifiers available with Matlab can be used.

V. Experiments and Results

The experiments compare eight descriptor-based techniques under consistent protocols on aligned and non-aligned datasets, showing that binary-pattern methods generally outperform other local descriptors while alignment strongly affects performance.

  • Experiments and Results: Eight descriptor-based techniques were evaluated on IITD II and USTB II using the same reproducible experimental protocol.The experiments included identification and verification tasks with 5-fold cross-validation and CMC and ROC evaluation.
  • Experiments and Results: The comparison implements LBP, LPQ, BSIF, POEM, HOG, DSIFT, RILPQ, and Gabor features with cross-validated open hyper-parameters.Chi-square distance is used for histogram-based descriptors and cosine similarity for the remaining features.
  • Experiments and Results: Performance is significantly better for aligned images than for non-aligned images, consistent with pose variability being a major performance factor.The aligned images come from IITD II, whereas the non-aligned images come from USTB II.
  • Experiments and Results: Binary-pattern techniques perform similarly to one another, while HOG, DSIFT, and Gabor generally perform worse, especially in verification.This pattern appears on both aligned and non-aligned images.
  • Experiments and Results: The experiments conclude that binary-pattern methods should generally be preferred over other local methods, with within-group selection based on non-performance criteria.The reported results are comparable to prior literature, supporting the competitiveness of the toolbox implementations.
  • Experiments and Results: On AWE, cross-validation curves demonstrate both dataset difficulty and substantial remaining room for improvement in ear recognition.The descriptors show almost equal performance in the AWE experiments.

B. Comparative Assessment on the AWE Dataset

Experiments on the unconstrained AWE dataset show substantially lower recognition performance than on earlier datasets, underscoring the difficulty of uncontrolled ear images.

  • AWE experiments used eight feature-extraction techniques with the previous hyper-parameters and predefined experimental protocols.The methods were based on LBP, LPQ, RILPQ, BSIF, DSIFT, POEM, Gabor, and HOG features.
  • The development data were evaluated using 5-fold cross-validation, with CMC and ROC curves plus predefined performance metrics.The dataset provides experiment lists for each fold, and results are reported with means and standard deviations.
  • AWE performance was significantly lower than on the earlier datasets for both identification and verification.The comparison covers all assessed techniques and both evaluation tasks.
  • 39.8%–49.6% mean rank-1 recognition rates were obtained, with Gabor lowest and POEM highest.The remaining recognition rates were close and all above 40%.
  • Around 30% EER was the best verification performance, while POEM was again the top performer across the ROC curve.Verification performance among the evaluated methods was otherwise close.
  • The results demonstrate that AWE is difficult and that more research is needed for recognition from completely uncontrolled images.The table caption likewise characterizes the dataset as useful and difficult.

VI. Open Questions and Research Directions

The paper identifies unresolved problems in automatic ear recognition, especially detection, alignment, and variability handling, and outlines directions involving richer data, open tools, and 3D modeling.

  • Open Questions: The field still contains important open research questions, motivating a list of topics for future investigation.The paper frames these questions as areas that are less studied than in other biometric fields.
  • Ear Detection and Alignment: Ear detection remains unsolved, with no widely adopted solution for efficiently locating ears in real-world images or video.The paper calls for large-scale nonlaboratory datasets and publicly available open-source detection tools.
  • Ear Detection and Alignment: Alignment is a major robustness issue because pose variation affects ear recognition, although suitable techniques can address in-plane and moderate out-of-plane changes.The paper notes manual alignment, active-shape approaches, active-contour methods, and landmark-based alternatives.
  • Handling Ear Variability: Occlusions from hair, clothing, accessories, viewpoint changes, and pose are major error sources in ear recognition systems.These factors also affect detection and alignment, with errors propagated through the processing pipeline.
  • Handling Ear Variability: Existing occlusion methods can handle moderate cases, but their breaking points and tolerance limits remain unclear.The paper suggests contextual information and modeling the entire facial profile as possible directions.
  • Pose Variation: 3D data, generic 3D models, and analysis-by-synthesis could support future pose normalization without forcing images into predefined canonical forms.These approaches are presented as promising directions for handling pose variation.

C. Feature Learning

The paper points toward CNN-based feature learning and broader modalities while emphasizing unresolved questions about scale, symmetry, aging, and inheritance.

  • Feature Learning: CNN-based feature learning is identified as the next development step for ear description and potentially end-to-end system design.Earlier approaches progressed from geometric and holistic methods toward local and hybrid approaches.
  • Dataset Scale: Existing datasets contain at most a few hundred subjects, leaving scalability to large-scale identification unclear.The paper calls for experiments and mathematical models of ear individuality.
  • Other Modalities: Research should extend beyond 2D images to 3D ear data and ear-prints, including detection, segmentation, feature extraction, and modeling.These modalities are described as presenting similar research challenges.
  • Other Modalities: Heterogeneous ear recognition remains an important longer-term problem involving matching images captured with different imaging modalities.The paper identifies cross-modality matching as a future research direction.
  • Understanding Ear Recognition: Ear symmetry, aging, and inheritance are insufficiently understood and require further study in automatic recognition.Open questions include exploiting bilateral symmetry, assessing template aging, and determining whether inherited characteristics aid recognition or create errors.
  • Conclusion: The paper combines a survey, a publicly available web-gathered dataset, and a state-of-the-art toolbox to examine challenges and support future research.It reports that the dataset is a considerable challenge and that the toolbox may later include detection, alignment, normalization, descriptors, and feature learning.
Loading 1611.06203v2…