Source-linked AI summary

A Survey of Automatic Facial Micro-expression Analysis: Databases, Methods and Challenges

Yee-Hui Oh, John See, Anh Cat Le Ngo, Raphael Chung-Wei Phan, Vishnu Monn Baskaran

arXiv:1806.05781v1cs.CVcs.MM

TL;DR

Automatic facial micro-expression analysis addresses the challenge of computationally spotting and recognizing brief, subtle expressions in video, a field that remains relatively new despite established psychological study. The survey reviews databases, methods, processing stages, and task-specific challenges, concluding that data quality, spotting robustness, and cross-domain recognition remain important concerns.

  • Problem

    Automatic micro-expression analysis lacks comprehensive coverage despite challenges in spotting subtle, short expressions and limited spontaneous databases.

  • Method

    The paper surveys micro-expression databases, spotting and recognition methods, preprocessing stages, and challenges across the automatic analysis pipeline.

  • Results

    The survey reports that cross-database recognition performs more poorly than single-database evaluations, indicating a need for methods robust across domains.

  • Takeaways & Limitations

    Reliable automatic analysis requires attention to data collection, precise spotting, and robustness across differing databases and domains.

  • Takeaways & Limitations

    Spontaneous micro-expression data remain difficult to acquire and imbalanced across emotions and subjects.

Abstract

from arXiv · show

Over the last few years, automatic facial micro-expression analysis has garnered increasing attention from experts across different disciplines because of its potential applications in various fields such as clinical diagnosis, forensic investigation and security systems. Advances in computer algorithms and video acquisition technology have rendered machine analysis of facial micro-expressions possible today, in contrast to decades ago when it was primarily the domain of psychiatrists where analysis was largely manual. Indeed, although the study of facial micro-expressions is a well-established field in psychology, it is still relatively new from the computational perspective with many interesting problems. In this survey, we present a comprehensive review of state-of-the-art databases and methods for micro-expressions spotting and recognition. Individual stages involved in the automation of these tasks are also described and reviewed at length. In addition, we also deliberate on the challenges and future directions in this growing field of automatic facial micro-expression analysis.

1 INTRODUCTION

Facial micro-expressions are brief, subtle, involuntary expressions associated with concealed emotions, making them difficult to analyze automatically. This survey frames automatic analysis as an emerging computational field focused on spotting and recognition.

  • 1 INTRODUCTION: Micro-expressions are brief, subtle, involuntary facial expressions that can occur when genuine emotions are concealed.They typically last 1/25 to 1/5 of a second, with a generally accepted upper limit of 0.5 second.
  • 1 INTRODUCTION: Micro-expression analysis has potential applications in clinical diagnosis, business negotiation, forensic investigation, and security systems.The text also reports that human detection performance remained limited even with dedicated training.
  • 1 INTRODUCTION: Computational analysis has become feasible as computer algorithms and video-processing technology have advanced.Researchers have consequently moved beyond predominantly psychological and manual analysis toward automated video-based methods.
  • 1 INTRODUCTION: Automatic micro-expression analysis remains challenging because expressions are short and subtle, making their temporal occurrence difficult to spot in video.Recognition is also hindered by inadequate features and a lack of complete, spontaneous, dynamic databases.
  • 1 INTRODUCTION: The survey addresses the absence of a comprehensive review by covering databases, spotting, recognition, task stages, challenges, and future directions.It distinguishes spotting as locating micro-expressions in video from recognition as assigning emotion labels.

2 MICRO-EXPRESSION DATABASES

Micro-expression databases differ in whether expressions are posed or spontaneous, with spontaneous data being more relevant to genuine emotions but harder to collect. Existing datasets address frame rate, scale, camera diversity, and spotting context, yet important data limitations remain.

  • 2 MICRO-EXPRESSION DATABASES: Spontaneous micro-expressions are more relevant to genuine emotional states than posed expressions, but they are substantially harder to elicit and annotate.Posed expressions may not reproduce the appearance and timing of spontaneous micro-expressions.
  • 2 MICRO-EXPRESSION DATABASES: Earlier posed databases may not represent real-life spontaneous expressions, and their reported durations can exceed the generally accepted micro-expression range.SMIC annotations also rely on participant self-reports for action units and emotion classes.
  • 2 MICRO-EXPRESSION DATABASES: SMIC was extended with high-speed, visual, and near-infrared recordings, while extended datasets added surrounding non-micro frames to support spotting.CAS(ME)2 and SAMM were also developed as extensions for algorithm development.
  • 2 MICRO-EXPRESSION DATABASES: CASME II is described as the largest and most widely used database, containing 247 videos recorded with high-frame-rate cameras at 200 fps.It was developed to address CASME’s short videos and lower 60 fps capture rate.
  • 2 MICRO-EXPRESSION DATABASES: Database designs vary in annotation and context: SAMM uses expert emotion labels and neutral frames, whereas Silesian Deception emphasizes deception-related facial cues.SAMM includes about 200 neutral frames before and after each micro-movement and is described as culturally diverse.

3 SPOTTING OF FACIAL MICRO-EXPRESSIONS

Automatic micro-expression analysis separates spotting from recognition, with spotting locating the temporal interval of a micro-movement before higher-level classification. Spotting frameworks typically combine preprocessing, feature description, and detection.

  • 3 SPOTTING OF FACIAL MICRO-EXPRESSIONS: Automatic analysis comprises micro-expression spotting and recognition, where spotting detects temporal intervals and recognition assigns emotion classes.Accurate frame identification is presented as a prerequisite for higher-level recognition.
  • 3 SPOTTING OF FACIAL MICRO-EXPRESSIONS: Micro-expression spotting models temporal dynamics across onset, apex, offset, and neutral phases.The onset reflects increasing facial change, the apex the peak expression, and the offset relaxation.
  • 3 SPOTTING OF FACIAL MICRO-EXPRESSIONS: A general spotting framework consists of preprocessing, feature description, and final micro-expression detection.The survey discusses these stages in subsequent sections.

3.1 Pre-processing

Preprocessing prepares facial video for spotting by locating and tracking landmarks, registering faces, masking irrelevant motion, and retrieving facial regions. The appropriate choices depend on image conditions, geometric deformation, and which facial movements are informative.

  • 3.1 Pre-processing: Preprocessing for facial micro-expression spotting includes landmark detection and tracking, face registration, masking, and facial-region retrieval.These operations reduce nuisance variation and prepare inputs for subsequent feature and detection stages.
  • 3.1 Pre-processing: Facial landmark detectors such as ASM, DRMF, SCMS, Face++, and CLM are widely used, often detecting points once and fixing them across consecutive frames.This convention assumes micro-expression movements are too subtle to substantially alter landmark locations.
  • 3.1 Pre-processing: Face registration aligns reference and sensed images to reduce head translations and rotations, using area-based or feature-based approaches.The chosen method should match the assumed geometric deformation because no single registration method suits all degraded face images.
  • 3.1 Pre-processing: Masking can suppress nuisance motion such as eye blinks, but masking the eyes may remove a highly discriminative region for affect recognition.This creates a practical need to distinguish expression-related eye activity from irrelevant blinking.
  • 3.1 Pre-processing: Facial regions may be analyzed separately because concealed-emotion studies support treating upper and lower face areas independently.The eye region is reported as more salient than the whole face in later recognition work.

3.2 Facial Micro-expression Spotting

Facial micro-expression spotting detects the temporal interval of micro-movements and uses classifier-based or rule-based approaches. Methods range from frame classification and temporal thresholds to feature-difference, optical-flow, random-walk, and phase-based techniques, with reported performance varying across databases.

  • Spotting methods are broadly divided into classifier-based approaches and rule-based approaches using thresholds or heuristics.
  • Early works: Early methods classified frames into neutral, onset, apex, or offset phases, but this formulation was considered unsuitable for real-time spotting.
  • Movement spotting: Temporal methods locate micro-expressions by analyzing frame-to-frame changes, feature-difference magnitudes, or thresholded motion peaks.
  • Movement spotting: LBP consistently outperformed HOOF across several datasets, while MDMD achieved recall 0.32, precision 0.35, and F1-score 0.33 on CAS(ME)2.
  • Movement spotting: Alternative approaches use optical flow, random-walk probabilities, or Riesz-Pyramid phase variation to identify movement intervals and distinguish micro-expressions from eye movements.
  • Apex spotting: Apex spotting focuses on locating the most expressive frame, with optical flow described as the most all-rounded feature for direction- and movement-based detection.

3.3 Performance Metrics

Micro-expression spotting is evaluated as a binary detection problem by comparing detected peaks or sequences with ground-truth intervals. The survey describes ROC-based measures, apex-specific metrics, and a benchmark intended to standardize evaluation.

  • ROC analysis evaluates spotted peaks by comparing them with ground-truth labels to determine true and false detections.
  • TPR is the number of correctly spotted micro-expression frames divided by the total number of ground-truth micro-expression frames, while FPR uses incorrectly spotted frames over non-micro-expression frames.
  • The MESB benchmark uses sliding-window, multi-scale protocols and Intersection over Union to support fairer and more comprehensive spotting evaluation.
  • Mean Absolute Error measures how close estimated apex frames are to ground-truth apex frames.
  • Apex Spotting Rate measures whether an estimated apex frame falls within the annotated onset-to-offset interval, assigning 1 for an in-range frame and 0 otherwise.

4 FACIAL MICRO-EXPRESSION RECOGNITION

Automatic micro-expression recognition classifies brief facial videos into emotion classes using preprocessing, feature representations, and classifiers, but uneven durations, subtle motion, and imbalanced datasets complicate reliable analysis.

  • Recognition classifies micro-expression videos into emotion classes, but available datasets contain unevenly represented emotion categories.
  • Pre-processing: Preprocessing can include landmark detection, face registration, region retrieval, motion magnification, and temporal normalization before feature extraction and classification.
  • Pre-processing: DMDSP outperformed TIM on comparable features and classifiers by retaining significant temporal structures while removing irrelevant facial dynamics.
  • Facial Micro-expression Representations: Feature representations are broadly single-level or multi-level and may encode geometric, appearance, motion, or transformed-domain information.
  • Facial Micro-expression Representations: LBP-TOP is widely used as a baseline, while variants such as LBP-SIP reduce redundant information and can improve recognition performance.
  • Experimental Protocol & Performance Metrics: SVM is the most widely used classifier, while F1-score is preferable to accuracy for imbalanced datasets because accuracy can favor larger classes.

5 CHALLENGES

The survey identifies unresolved challenges in automatic micro-expression spotting and recognition, especially limited and biased data, unstable facial processing, real-world variability, and weak cross-database generalization.

  • Spontaneous micro-expression datasets remain limited and imbalanced across emotions and subjects, creating potential bias toward better-represented categories.
  • Existing datasets often lack ethnic diversity, FACS metadata, reliable inter-coder measures, and realistic conditions representative of real-life applications.
  • Micro-expression Spotting: Inaccurate landmark detection can destabilize face alignment and introduce head-motion or gaze noise that complicates correct micro-expression detection.
  • Micro-expression Spotting: Eye regions are both discriminative and vulnerable to blinks, so spotting systems must distinguish expressive eye movements from irrelevant motion and overlapping blinks.
  • Micro-expression Spotting: Onset and offset frames remain harder to locate than movement peaks, especially when facial movements continuously change in real-life settings.
  • Micro-expression Recognition: Recognition still faces low-intensity, short-duration signals, feature-selection trade-offs, scarce imbalanced data for deep learning, and weaker cross-database performance.

6 CONCLUSION

The survey identifies under-addressed issues that limit practical automatic micro-expression spotting and recognition, especially robustness to motion, viewpoint, irrelevant facial variation, limited data, and database shift.

  • More precise spotting must handle macro movements, varied head poses, and camera views to extend constrained systems toward real-time in-the-wild settings.
  • Defining onset and offset frames more firmly would enable short micro-expression sequences to be extracted from long videos for emotion classification.
  • Recognition systems must exclude alignment and head-rotation artifacts because micro-expressions are subtle and easily interfered with by irrelevant facial information.
  • The survey highlights limited data as a constraint on encoding subtle movements, even when feature representations are rich.
  • The survey calls for cross-database evaluation because single-database testing can give a false impression of method performance when databases lack diversity.

TABLES

The tables organize micro-expression databases and literature methods across spotting preprocessing, spotting studies, and recognition benchmarking, with some database and reporting caveats.

  • Table 1 catalogs posed and spontaneous micro-expression databases, including dataset size, recording properties, labels, and annotations where reported.
  • The database listings distinguish spontaneous resources such as YorkDDT, Silesian Deception, CAS(ME)2, and SAMM, while noting that some annotations or emotion classes are unavailable.
  • Table 2 surveys preprocessing techniques used in facial micro-expression spotting, while Table 3 compiles spotting and micro-movement studies.
  • Table 4 benchmarks facial micro-expression recognition methods, with reported results subject to differences such as omitted samples, emotion-class counts, combined databases, and computed values.

FIGURES

The figures provide sample facial-expression sequences from several micro-expression databases, illustrating varied emotions and micro-movement content.

  • Sample frames depict a ’Surprise’ sequence from Subject 1 in SMIC.
  • A ’Happiness’ sequence from Subject 6 illustrates sample frames from CASME II.
  • A ’Disgust’ sequence from Subject 15 provides sample frames from CAS(ME)2.
  • SAMM and MEVIEW examples show sequences containing micro-movements, including one marked with AU L12.
  • A CASME II video sequence illustrates the temporal order of onset, apex, and offset frames.
Loading 1806.05781v1…