Source-linked AI summary
Survey on RGB, 3D, Thermal, and Multimodal Approaches for Facial Expression Recognition: History, Trends, and Affect-related Applications
Ciprian Corneanu, Marc Oliu, Jeffrey F. Cohn, Sergio Escalera
TL;DR
Interpreting facial expressions and their relation to affect remains challenging despite extensive research. This paper surveys RGB, 3D, thermal, and multimodal approaches through a comprehensive taxonomy, datasets, and benchmarked methods, concluding with trends and research questions. It reports that multimodal data can support a wider range of expressions than RGB alone, while identifying limitations in feature extraction and spontaneous-expression data.
Problem
Automatic facial-expression recognition has advanced substantially, but recognizing expressions in naturalistic settings and interpreting their affective meaning remain difficult.
Method
The paper synthesizes RGB, 3D, thermal, and multimodal facial-expression analysis through a taxonomy covering processing stages, datasets, influential methods, and affect-related applications.
Results
Multimodal approaches can represent a wider range of expressions than RGB alone.
Takeaways & Limitations
The survey provides a broad perspective on automatic facial-expression recognition, affect inference, current trends, important questions, and future research lines.
Takeaways & Limitations
Feature extraction remains constrained because thermal images are unsuitable for geometric information, while dynamic learned feature extraction has not been attempted.
Abstract
from arXiv · showhide
Facial expressions are an important way through which humans interact socially. Building a system capable of automatically recognizing facial expressions from images and video has been an intense field of study in recent years. Interpreting such expressions remains challenging and much research is needed about the way they relate to human affect. This paper presents a general overview of automatic RGB, 3D, thermal and multimodal facial expression analysis. We define a new taxonomy for the field, encompassing all steps from face detection to facial expression recognition, and describe and classify the state of the art methods accordingly. We also present the important datasets and the bench-marking of most influential methods. We conclude with a general discussion about trends, important questions and future lines of research.
1 INTRODUCTION
Facial expressions are central signals in human social communication, and automatic recognition has developed across behavioral science, neurology, and artificial intelligence. This survey organizes RGB, 3D, thermal, and multimodal approaches while emphasizing affect inference and recent trends.
- Facial expressions convey cues about people’s emotional states and contribute, alongside voice, language, hands, and posture, to social communication.
- Historical evolution: Modern facial-expression research grew from historical studies by Duchenne de Boulogne and Darwin, followed by influential work from Izard and Ekman.
- Related surveys: Earlier surveys addressed processing-oriented analysis, natural conditions, and 3D facial-expression recognition from different perspectives.
- Survey contribution: The survey defines a comprehensive taxonomy of automatic RGB, 3D, thermal, and multimodal computer-vision approaches for automatic facial-expression recognition.
- Survey contribution: It complements the taxonomy with affect inference, historical evolution, important trends, and a broader perspective emphasizing research since 2009.
2 INFERRING AFFECT FROM FES
Facial expressions can be described categorically or dimensionally and may serve emotional, communicative, physiological, and social functions. Automatic affect inference has broad applications, but subjective-experience evidence remains limited by labor-intensive measurement.
- Describing affect: Affect is commonly described either through distinct emotion categories or within dimensional spaces such as valence, activation, and control.
- Describing affect: Categorical approaches commonly use six primary emotions: happiness, sadness, fear, anger, disgust, and surprise.
- Describing affect: Dimensional descriptions can represent more complex emotions, but automatic systems often simplify them into limited categories or two-dimensional quadrants.
- Evolutionary perspective: Evolutionary accounts emphasize adaptive functions of expressions, while later models also describe their communicative and survival-related functions.
- Evidence limitations: Manual facial-expression annotation and facial EMG remain labor intensive, limiting replication of studies on subjective experience.
- Applications: Automatic affect inference supports adaptive environments, socially aware systems, e-learning, games, clinical monitoring, driving safety, and psychological-distress analysis.
3 A TAXONOMY FOR RECOGNIZING FES
The survey’s taxonomy organizes automatic facial-expression recognition around parametrization and recognition, then details modality-dependent processing, multimodal fusion, and labeled datasets. It applies across data modalities.
- Taxonomy structure: The taxonomy has two main components: parametrization, which defines expression coding schemes, and recognition, which discriminates among expressions.
- Parametrization: Parametrization distinguishes descriptive coding schemes based on facial surface properties from judgmental schemes based on latent emotions or affects.
- Recognition pipeline: Automatic facial analysis commonly localizes faces, registers fiducial points, extracts modality-dependent features, and recognizes categorical or continuous expressions.
- Recognition pipeline: Recognition methods may model temporal dynamics, depending on whether the output represents categorical expressions or a continuous space.
- Multimodal fusion: Multimodal systems add fusion at direct, early, late, or sequential stages when combining facial data with sources such as speech or physiology.
- Datasets: The survey characterizes labeled datasets by content, capture conditions, modality, and participant distributions.
3.1 Parameterization of FEs
Facial-expression parametrization separates descriptions of visible facial actions from judgments about underlying affect. These schemes differ because expressions and emotion labels do not have a one-to-one correspondence.
- Descriptive coding: Descriptive coding schemes represent what the face can do, including FACS and Face Animation Parameters.
- Descriptive coding: FACS describes facial expressions through anatomically based Action Units, each representing contraction of one or more facial muscles.
- Descriptive coding: FACS also specifies visual detection rules and temporal segments, including onset, apex, offset, and ordinal intensity.
- Judgmental and hybrid coding: Judgmental coding schemes describe expressions through latent emotions or affects, whereas hybrid schemes define emotion labels through specific observable signs.
- Judgmental and hybrid coding: A single emotion may produce multiple expressions, so facial actions and emotion labels do not have a one-to-one correspondence.
3.2 Recognition of FEs
The survey organizes automatic facial-expression recognition into a four-step pipeline and classifies modality-specific methods for localization, registration, feature extraction, and recognition. It also distinguishes descriptive and judgement coding schemes, feature types, and multimodal fusion strategies.
- An AFER system consists of face detection, face registration, feature extraction, and expression recognition.
- Recognition and fusion: FACS provides descriptive coding through Action Units, while judgement schemes represent inferred emotions; fusion approaches include early, late, and sequential fusion.Early fusion merges features and can exploit synchronous correlations, but increases feature dimensionality and over-fitting risk.
- Face localization: Face localization uses detection to obtain face bounding boxes or geometry, or segmentation to assign binary labels to pixels.
- Face localization: RGB face detection commonly uses Viola&Jones, while CNNs, SVMs, pose-specific detectors, and segmentation methods address alternative settings.Viola&Jones is fast but has problems with occlusions and large pose variations.
- Face registration: Registration aligns 2D landmarks or captured 3D geometry with a model, using methods including AAM, SDM, ICP, and deformable models.3D registration establishes geometric correspondence between captured geometry and a model.
- Feature extraction: Features are predesigned or learned, and global or local; appearance features use image intensity, whereas geometric features measure facial shape and deformation.Geometric features cannot be extracted from thermal data because dull facial features hinder precise landmark localization.
- Feature extraction: Examples include LBP-TOP for RGB, TDHFs and StaFs for thermal images, and optical flow, MHI, FFDs, deformation descriptors, and motion vectors for dynamics.LBP-TOP describes spatiotemporal information across three orthogonal planes in a volume formed by stacked frames.
3.3 FE datasets
The survey characterizes facial-expression datasets by content, capture modality, and participant properties, covering RGB, 3D, and thermal collections. Across datasets, variation includes posed versus spontaneous expressions, viewpoints, illumination, environments, labels, and participant diversity.
- Datasets are grouped by content, capture modality, and participants, with content including intentionality, labels, and static or dynamic format.Capture properties include lab context, perspective, illumination, and occlusions; participant statistics include age, gender, and ethnic diversity.
- RGB datasets: CK contains small, posed, primary-expression samples with limited demographic diversity, frontal views, and homogeneous illumination; CK+ adds 22% more posed samples and spontaneous expressions.
- RGB datasets: MMI adds profile views, broader FACS Action Unit coverage, and onset, apex, and offset labels, while Multi-PIE adds varied viewpoints and illumination.
- 3D datasets: 3D datasets include posed collections such as BU-3DFE, Bosphorus, and BU-4DFE, while BP4D uses authentic emotion-induction tasks to obtain spontaneous expressions.BP4D videos were annotated by experienced FACS coders and checked using self-report, FACS analysis, and human observer ratings.
- Thermal datasets: Thermal facial-expression datasets are few and also include RGB data; NVIE contains 215 subjects displaying six spontaneous and posed expressions.
4 HISTORICAL EVOLUTION AND CURRENT TRENDS
AFER progressed from early landmark tracking and posed-expression analysis toward dynamic, spontaneous, multimodal, and affect-related applications. Current trends include intensity estimation, modality fusion, and recognition of depression, personality, cognitive states, pain, fatigue, and other complex behaviors.
- Historical evolution: AFER began with landmark-motion tracking in 1978, then revived in the early 1990s before the CK dataset helped inaugurate modern AFER in 2000.Early progress was slowed by poor face detection, face registration, and limited computational power.
- Historical evolution: AFER expanded beyond posed primary expressions to spontaneous expressions, action units, pain, fatigue, frustration, depression severity, psychological distress, and cognitive states.These applications opened new territory for facial expression research.
- Historical evolution: Representations evolved from static RGB, 3D, and thermal features toward dynamic geometric and sequence-based appearance representations.Dynamic approaches tracked facial deformations across frames or extracted features directly from frame sequences.
- Estimating intensity of facial expressions: Intensity estimation became a major trend because expression intensity and timing can distinguish posed from spontaneous smiles and polite from embarrassed smiles.The FERA challenge added intensity estimation, supported by RGB and 3D datasets with spontaneous-expression intensity labels.
- Estimating intensity of facial expressions: 3D was not uniformly better than RGB for intensity estimation, but RGB–3D fusion significantly improved overall performance.RGB performed better on the upper face, whereas 3D performed better on the lower face; both modalities had greater difficulty with lower-face action units.
- AFER for detecting non-primary affective states: Affect-related applications increasingly use multimodal or direct dynamic representations to predict depression, dimensional affect, and personality.Personality studies found facial expressions more informative than basic visual activity measures but less effective than audio, especially prosodic cues.
- AFER in naturalistic environments: Naturalistic AFER remains difficult because of pose and illumination variation, spontaneous low-intensity expressions, multiple apexes, speech-related facial displays, and simultaneous expressions by multiple people.Methods therefore commonly extract dynamic appearance or learn representations from frame sequences.
5 DISCUSSION
The discussion organizes AFER around a modality-aware pipeline while highlighting advances in affective-state analysis, multimodal fusion, and broader expression modeling. It also identifies persistent representation and interpretation constraints across modalities and settings.
- Pipeline: AFER commonly proceeds through face detection, registration, feature extraction, and recognition, with multimodal systems adding a fusion step.Registration may be unnecessary for some global-feature or deep-learning methods, while fusion can be direct, early, late, or sequential.
- Feature extraction: Feature extraction uses handcrafted descriptors or learned representations, covering facial appearance, geometry, or both.The taxonomy distinguishes predesigned descriptors from CNN- and DBN-based methods that learn features implicitly with recognition.
- Feature extraction: Thermal images are unsuitable for geometric feature extraction because their captured images are dull, while RGB local static geometry may lack registration precision.The discussion also notes that dynamic extraction from learned features had not been attempted to the authors’ knowledge.
- Recognition: Recognition can classify categorical expressions or represent them continuously, with models optionally incorporating temporal dynamics.Most methods use multiclass classification, commonly targeting six basic emotions, although continuous expressive spaces are also possible.
- Expression modeling: AU-based coding supports a wider range of expressions and facilitates microexpression detection, whereas RGB alone makes broader expression recognition more difficult.The discussion links this expansion to limitations of datasets focused on a small set of primary emotions.
- Multimodal fusion: Recent work increasingly combines RGB, depth, thermal, audio, language, gestures, and physiological signals to enrich representations and improve emotion inference.The survey expects continued integration of visual and non-visual modalities, while warning that redundant modalities can make simple feature concatenation inefficient.