Source-linked AI summary

Face Expression Recognition and Analysis: The State of the Art

Vinay Bettadapura

arXiv:1203.6722v1cs.CV

TL;DR

Automatic facial-expression recognition has advanced substantially, but the field still needs standardized, complete, and spontaneous-expression data. This tutorial-style survey synthesizes post-2001 developments, parameterization methods, applications, databases, classifiers, and challenges, concluding that robust spontaneous recognizers are expected to support real-time systems and emotion-sensitive human–computer interfaces.

  • Problem

    The field lacks standardized databases containing spontaneous expressions under varied conditions, while capturing and labeling suitable data remains difficult and expertise-intensive.

  • Method

    The paper provides a tutorial-style survey of post-2001 research covering applications, facial parameterization, expression features, databases, classifiers, and future challenges.

  • Results

    The survey reports substantial progress over the past decade and a shift from posed toward spontaneous-expression recognition.

  • Takeaways & Limitations

    The paper projects that robust spontaneous-expression recognizers will be developed and deployed in real-time systems and emotion-sensitive human–computer interfaces.

  • Takeaways & Limitations

    Available expression databases may contain incomplete temporal sequences, and eliciting authentic fear and anger through videos or films is challenging.

Abstract

from arXiv · show

The automatic recognition of facial expressions has been an active research topic since the early nineties. There have been several advances in the past few years in terms of face detection and tracking, feature extraction mechanisms and the techniques used for expression classification. This paper surveys some of the published work since 2001 till date. The paper presents a time-line view of the advances made in this field, the applications of automatic face expression recognizers, the characteristics of an ideal system, the databases that have been used and the advances made in terms of their standardization and a detailed summary of the state of the art. The paper also discusses facial parameterization using FACS Action Units (AUs) and MPEG-4 Facial Animation Parameters (FAPs) and the recent advances in face detection, tracking and feature extraction methods. Notes have also been presented on emotions, expressions and facial features, discussion on the six prototypic expressions and the recent studies on expression classifiers. The paper ends with a note on the challenges and the future work. This paper has been written in a tutorial style with the intention of helping students and researchers who are new to this field.

1. Introduction

Facial-expression research spans historical studies of expression and physiognomy, Darwin’s classification of expressions, and the later development of automatic recognition. Since the 1990s, automatic facial-expression recognition has become an active field, with this survey concentrating mainly on work published from 2001 onward.

  • Historical background: Facial-expression research has historical roots in studies of appearance, facial movements, and expression categories extending from antiquity through the nineteenth century.Darwin established general principles of expression, grouped expressions into categories, and cataloged associated facial deformations.
  • Historical background: Darwin grouped expressions into categories including low spirits, joy, reflection, contempt, surprise, and self-attention.These categories include related states such as grief, love, determination, disgust, fear, shame, and modesty.
  • Development of automatic recognition: Paul Ekman’s work since the 1970s strongly influenced the development of modern automatic facial-expression recognizers.The survey devotes substantial discussion to Ekman’s work and its impact on present-day systems.
  • Development of automatic recognition: Suwa et al. took an early step toward automatic recognition in 1978 by analyzing facial expressions in image sequences with twenty tracking points.Research in this direction was not pursued extensively until the early 1990s.
  • Scope of the survey: Comprehensive surveys covered published automatic-recognition work from 1990 to 2001, so this paper focuses mainly on publications from 2001 onward.Earlier work is directed to the surveys by Pantic and Rothkrantz and by Fasel and Luttin.

2. Applications

Automatic facial-expression recognition supports robotics, affect-sensitive human–computer interaction, and applications across communication, behavioral, entertainment, medical, safety, and educational domains. Demonstrated systems include expression-mirroring characters and robotic platforms, while future uses depend on greater real-time robustness.

  • Primary applications: Robotics and affect-sensitive human–computer interaction are presented as the two main application areas for automatic facial-expression recognition.Robots increasingly interacting with humans require improved ability to understand human moods and emotions, while affective computing uses expression recognition to build responsive interfaces.
  • Additional applications: Expression recognition has also been applied or proposed for telecommunications, behavioral science, video games, animation, psychiatry, automobile safety, music and television systems, and educational software.These examples extend beyond the two principal application areas identified by the paper.
  • Demonstrated systems: Practical real-time systems have included an expression-mirroring animated character and deployments on Sony’s Aibo Robot and ATR’s RoboVie.The paper also describes EmotiChat as another application of expression recognition.
  • Future applications: The paper anticipates additional innovative applications as expression-recognition systems become more real-time and robust.The statement is presented as a future direction rather than as an already demonstrated capability.

3. Facial Parameterization

Facial parameterization represents facial behavior through standardized action or animation parameters. The survey presents FACS AUs and MPEG-4 FAPs as the two principal parameterization approaches, connecting muscle-related facial actions, feature points, and expression representations.

  • FACS: FACS represents facial behavior through Action Units that correspond to changes caused by individual facial muscles or muscle groups.The system was developed to address limitations and lack of consensus in earlier observer-based and disparate measurement approaches.
  • FACS: FACS expressions can combine additive or non-additive AUs; for example, fear is represented by AUs 1, 2, and 26.Additive AUs appear independently, whereas non-additive AUs modify one another’s appearance.
  • Recognition challenges: Recognizing AUs remains challenging, including for profile views and certain actions, although automatic systems such as AFA recognize six upper-face and ten lower-face AUs.The survey also notes that comprehensive parameterization is needed for difficult questions about facial variation across cultural and economic backgrounds.
  • MPEG-4 FAPs: MPEG-4 FAPs provide a standardized parameter set for facial animation and are closely related to FACS AUs.The standard defines a neutral face model, 84 feature points, and parameters covering facial actions, head motion, tongue, eyes, and mouth control.

FAP No. FAP Name

MPEG-4 normalizes facial animation parameters against a neutral face and feature-based units, while related work maps FACS AUs to FAPs for expression analysis and synthesis. The section also highlights automatic extraction algorithms used with recognition systems.

  • FAP normalization: FAP values are normalized with Facial Animation Parameter Units defined as fractions of distances between key facial features.This allows FAP-described movements to adapt to face models of different sizes and shapes.
  • Expression parameters: MPEG-4 defines six primary facial expressions, including joy, anger, sadness, fear, disgust, and surprise, within its expression parameter group.The expression parameter is distinct from visemes, which are used for speech-related studies.
  • AU–FAP mapping: Research has mapped FACS AUs to MPEG-4 FAPs to connect facial-expression analysis with facial-expression synthesis.The cited work also derives models intended to bring these disciplines together.
  • Automatic extraction: Automatic FAP extraction algorithms include GVF snakes, parabolic templates, combination methods, and active contours, with HMMs used in expression recognition.The cited evaluation found GVF snakes more sensitive to random noise and reflections than the alternatives.

4. Emotions, Expressions and Features

The section surveys how emotions and facial expressions are categorized, elicited, recognized, and distinguished across basic, non-basic, posed, spontaneous, and temporal forms. It also emphasizes facial features, occlusion, cultural context, and individual differences as central recognition considerations.

  • Emotions and expressions: Ekman and Friesen’s cross-cultural studies established six prototypic expressions: happiness, sadness, anger, surprise, disgust, and fear.These expressions became the main target of automatic recognition systems, although social context influences how emotions are displayed.
  • Culture and context: Facial expression interpretation is broadly cross-cultural, whereas emotional display through facial changes depends on social context.American and Japanese viewers showed similar expressions to eliciting videos, but Japanese viewers suppressed displays more in the presence of authority.
  • Expression categories: Automatic systems increasingly extend recognition beyond six basic expressions by identifying individual Action Units and temporal segments.Temporal analysis distinguishes onset, apex, and offset, enabling recognition of a broader range of expressions.
  • Posed and spontaneous expressions: Posed expressions differ from spontaneous expressions in appearance, timing, and temporal characteristics, while many database examples are exaggerated and artificial.Genuine expressions are usually subtler, motivating systems that can recognize spontaneous behavior.
  • Facial features and occlusion: Mouth and eyebrows carry substantial expression information, but the mouth can be more informative under occlusion.Recognition accuracy was 78% with only the mouth visible versus 50% with only the eyebrows visible; surprise, joy, and disgust also achieved 100%, 93.4%, and 97.3%.
  • Individual differences: Individual differences in skin texture, facial hair, and eye opening across demographic groups complicate expression recognition.These variations are identified as an additional facial-feature consideration for recognition systems.

5. Characteristics of a Good System

The paper defines an ideal facial expression recognition system as fully automatic, real-time, person-independent, unobtrusive, and robust across conditions. It also argues that existing research addresses separate requirements, leaving integration toward an ideal system unresolved.

  • Automation and operation: A good system should operate automatically in real time on both images and video feeds without requiring preprocessing.The system is expected to function directly on incoming visual data.
  • Recognition scope: It should recognize spontaneous and non-prototypic expressions, potentially through recognition of the full set of facial Action Units.The desired scope extends beyond the six prototypic expressions.
  • Generalization: It should be person-independent and work across cultures, skin colors, and ages, including infants, adults, and older people.The requirement explicitly spans demographic variation.
  • Robustness: It should remain invariant to facial hair, glasses, makeup, lighting conditions, image resolution, and moderate occlusion.These conditions represent practical sources of visual variation and obstruction.
  • Capture conditions: It should recognize expressions from frontal, profile, and intermediate viewing angles while remaining unobtrusive.The target setting is not restricted to frontal, controlled capture.
  • Open challenge: Existing research groups have addressed different requirements separately, but the paper calls for integrating these ideas into more ideal systems.The stated challenge is combining robustness, spontaneity, occlusion handling, and other capabilities rather than solving only one aspect.

6. Face Detection, Tracking and Feature Extraction

The paper organizes automatic expression recognition into face detection and tracking, feature extraction, and expression classification, focusing here on the first two modules. It surveys model-based, statistical, appearance-based, and probabilistic tracking approaches used in recent systems.

  • System modules: Automatic facial expression recognition systems comprise face detection and tracking, feature extraction, and expression classification modules.This section covers detection and tracking plus feature extraction; classifiers are discussed separately.
  • Detection and tracking: Face detection localizes a face in an image, whereas face tracking follows it across frames in a video sequence.The distinction defines the first processing stage for image- and video-based analysis.
  • Tracking approaches: Major approaches include Kanade-Lucas-Tomasi, statistical detection, AdaBoost-based Viola–Jones detection, PBVD, Candide models, Ratio Templates, and PersonSpotter.These methods span early feature tracking, rapid frontal-face detection, 3-D models, geometric templates, and model graphs.
  • Model-based tracking: The PBVD tracker uses a generic 3-D wireframe face model associated with 16 Bezier volumes for real-time tracking.The same model can also support facial-motion analysis and computer animation.
  • Model-based tracking: Candide face models use triangular meshes, with Candide-3 identified as the model currently used by most researchers.The paper illustrates Candide-1, Candide-2, and Candide-3 as successive model variants.
  • PersonSpotter: PersonSpotter detects moving regions through difference images, applies skin and convex detectors, and uses a face model graph to suppress background.Its tracking system was demonstrated as robust against considerable background motion.
  • Probabilistic tracking: Particle filters are widely used because they handle noise, occlusion, clutter, and uncertainty, with later variants adding motion prediction, mean shift, or AdaBoost.These extensions address fast movement, degeneracy, and multiple-target detection or tracking.

7. Databases

Facial-expression databases are central to benchmarking, yet standardization remains difficult because expressions vary between posed and spontaneous settings and across recording conditions. Recent databases improve coverage, but labeling, temporal completeness, illumination, occlusion, and authentic spontaneous-expression capture remain unresolved challenges.

  • Database standardization: A standardized expression-recognition database remains an open problem despite significant progress compared with earlier surveys.A common database would simplify comparison and benchmarking, but must satisfy diverse requirements.
  • Database standardization: Posed and spontaneous expressions differ in characteristics, temporal dynamics, and timing, requiring standardized data focused on spontaneous expressions under varied conditions.The desired database should include images and videos at different resolutions, lighting conditions, occlusions, and head rotations.
  • Available databases: Many established databases contain only posed expressions and are unsuitable for spontaneous-expression recognition.The Cohn-Kanade, AR, CMU PIE, and JAFFE databases are cited in this context.
  • Available databases: The MMI database advances standardization by combining posed and spontaneous expressions with profile-view data in a freely searchable and downloadable web database.It went online in February 2009.
  • Annotation: Database labeling traditionally relies on expert observers, AU coders, or participants’ self-reports, making annotation time-consuming and expertise-dependent.Labels are applied after data capture and augmented with metadata.
  • Limitations: Researchers have encountered incomplete temporal sequences and missing databases for varied illumination or facial occlusion, limiting specific evaluations.Complete onset, apex, and offset patterns are needed for temporal studies.
  • Future direction: New public databases include spontaneous expressions, frontal and profile views, 3D data, occlusion, and lighting variation, supporting broader future research.The paper still describes creating one database serving everyone’s needs as very difficult.

8. State of the Art

The survey organizes 19 expression-recognition studies published from 2001 onward in chronological tables, alongside reported system performances and sample sizes. Its state-of-the-art overview spans diverse classifiers and system configurations rather than a single unified benchmark.

  • Surveyed systems: The survey selects 19 important and diverse papers and presents them chronologically from 2001 in a summary table.The detailed descriptions are abbreviated in favor of the tabulated overview.
  • Method diversity: The surveyed methods use a broad range of classifiers, including ANN, kNN, HMM, NB, TAN, SSS, SVM, MLP, and related variants.The abbreviation list also includes several discriminant, correlation, and multi-stream methods.
  • Reported evaluations: The state-of-the-art material includes system performance summaries for Bourel et al. and Cohen et al.The supplied passages identify the referenced tables but do not provide their numerical entries.

9. Classifiers

Expression-recognition classifiers range from static probabilistic models to dynamic sequence models and modular or fused systems. Reported findings emphasize distribution choice, data-label availability, temporal structure, feature condensation, and robustness to occlusion as important design considerations.

  • Classifier landscape: The classification module is the final stage after face detection and feature extraction, and recent work evaluates many alternative classifiers.The survey treats classifier selection as a major research area.
  • Classifier landscape: Cohen et al. studied static NB, TAN, and SSS classifiers alongside dynamic HMM and ML-HMM classifiers.Their work compares models that do not explicitly model sequences with models designed for temporal patterns.
  • Probabilistic classifiers: Cauchy-distribution NB performed better than Gaussian-distribution NB, while TAN outperformed NB in Cohen et al.’s experiments.The paper notes that facial motions are highly correlated with displayed emotions, challenging NB’s independence assumption.
  • Static versus dynamic models: Dynamic classifiers are suggested for person-dependent tests, whereas static classifiers are suggested for person-independent tests because dynamic models are sensitive to temporal and appearance changes.The recommendation is attributed to Cohen et al.’s reported scenario-dependent behavior.
  • Semi-supervised learning: SSS outperformed NB and TAN when unlabeled data was added, although NB and TAN performed well with labeled training data.This result concerns semi-supervised learning with labeled and unlabeled databases.
  • Feature selection: AdaBoost-based feature selection produced an improved representation for SVM training and improved classification performance.The cited work used AdaBoost to speed feature selection.
  • Robustness to occlusion: Bourel et al. used localized modular classifiers with data fusion so occlusion in one facial region need not disable the other regional classifiers.The design replaces one monolithic classifier with several local classifiers.
  • Classifier evaluation: NB and kNN were stable in Sebe et al.’s evaluation, with voting algorithms producing no significant performance improvement.The study reported results for 14 classifiers, including bagging and boosting variants.

10. The 6 Prototypic Expressions

The six prototypic expressions are happiness, sadness, anger, surprise, disgust, and fear. Surveyed studies highlight recurring confusions, uneven elicitation difficulty, and expression-specific effects of occlusion.

  • The six prototypic expressions are happiness, sadness, anger, surprise, disgust, and fear, which are comparatively easiest to recognize.
  • Confusions: Automatic systems commonly confuse anger with disgust, while fear is often confused with happiness or anger and sadness with anger.
  • Elicitation: Emotion-inducing films and clips commonly elicit expressions, but fear and especially anger are difficult to evoke naturally.
  • Occlusion: Mouth occlusion reduces recognition accuracy for anger, fear, happiness, and sadness, whereas eye and brow occlusion affects disgust and surprise.

11. Challenges and Future Work

Future work centers on authentic spontaneous-expression data and more robust, fully automatic recognition across difficult emotions, labels, populations, timing, viewpoints, and fine-grained expressions.

  • Spontaneous data: Authentic spontaneous-expression datasets remain difficult to capture because awareness of recording can make expressions unnatural.Semi-authentic databases such as MMI offer a more practical alternative and include posed and spontaneous expressions in a searchable, downloadable web resource.
  • Elicitation: Emotion-inducing videos make happiness and amusement easier to elicit than fear and anger, leaving alternative elicitation methods as an open need.
  • Data annotation: Data labeling is time-consuming and potentially error-prone because it requires observer or Action Unit coding expertise.Semi-supervised learning is identified as one way to use both labeled and unlabeled data.
  • Expression scope: Recognizing spontaneous non-basic expressions remains more challenging than recognizing spontaneous basic expressions and is still open.
  • Robustness: Recognition accuracy differs across expressions, with anger often confused with disgust, motivating efforts toward more uniform recognition.
  • Robustness: Systems must become robust to facial and expression differences across cultures and age groups.
  • Temporal dynamics: Temporal dynamics may help distinguish posed from spontaneous expressions because onset, apex, and offset timings vary.
  • Viewpoint variation: Recognition from intermediate head angles between frontal and profile views remains largely unaddressed.

12. Conclusion

The paper surveys recent advances in automatic facial-expression recognition and its associated components in a tutorial format. It concludes that systems have improved while the field shifts toward spontaneous expressions and prospective real-time, emotion-sensitive interfaces.

  • The survey covers expression-recognition history, applications, facial parameterization, expressions and features, ideal-system characteristics, detectors, trackers, databases, classifiers, prototypic expressions, and future challenges.
  • The paper is designed to introduce recent advances accessibly to newcomers without prior background in the field.
  • Face expression recognition systems have improved over the past decade, with research shifting from posed toward spontaneous expression recognition.
  • The paper anticipates robust spontaneous recognizers being deployed in real-time systems and emotion-sensitive human-computer interfaces.
Loading 1203.6722v1…