Source-linked AI summary
A Systematic Review on Affective Computing: Emotion Models, Databases, and Recent Advances
Yan Wang, Wei Song, Wei Tao, Antonio Liotta, Dawei Yang, Xinlei Li, Shuyong Gao, Yixuan Sun, Weifeng Ge, Wei Zhang, Wenqiang Zhang
TL;DR
Affective computing must recognize human affect from physical information and physiological signals, each offering practical benefits but also important limitations. This review systematically taxonomizes unimodal and multimodal methods, emotion models, databases, architectures, and performances, and identifies baseline databases, multimodal fusion, and unsupervised learning as key future directions.
Problem
Affective computing seeks reliable emotion and sentiment recognition, but physical expressions may conceal inner emotion while physiological signals are difficult to acquire practically.
Method
The review synthesizes more than 380 research papers and organizes affective computing into unimodal affect recognition and multimodal affective analysis across emotion models, databases, and ML- or DL-based methods.
Results
The review provides a taxonomy of recent unimodal and multimodal affective-computing advances, distinguishing traditional ML-based methods from DL-based models that improve feature representation and classification.
Takeaways & Limitations
Future progress requires extended multimodal baseline databases, improved fusion strategies, and further exploration of zero/few-shot or unsupervised learning.
Takeaways & Limitations
Affective computing databases differ substantially in size, quality, and collection conditions, while limited physical-physiological databases constrain video-physiological emotion recognition.
Abstract
from arXiv · showhide
Affective computing plays a key role in human-computer interactions, entertainment, teaching, safe driving, and multimedia integration. Major breakthroughs have been made recently in the areas of affective computing (i.e., emotion recognition and sentiment analysis). Affective computing is realized based on unimodal or multimodal data, primarily consisting of physical information (e.g., textual, audio, and visual data) and physiological signals (e.g., EEG and ECG signals). Physical-based affect recognition caters to more researchers due to multiple public databases. However, it is hard to reveal one's inner emotion hidden purposely from facial expressions, audio tones, body gestures, etc. Physiological signals can generate more precise and reliable emotional results; yet, the difficulty in acquiring physiological signals also hinders their practical application. Thus, the fusion of physical information and physiological signals can provide useful features of emotional states and lead to higher accuracy. Instead of focusing on one specific field of affective analysis, we systematically review recent advances in the affective computing, and taxonomize unimodal affect recognition as well as multimodal affective analysis. Firstly, we introduce two typical emotion models followed by commonly used databases for affective computing. Next, we survey and taxonomize state-of-the-art unimodal affect recognition and multimodal affective analysis in terms of their detailed architectures and performances. Finally, we discuss some important aspects on affective computing and their applications and conclude this review with an indication of the most promising future directions, such as the establishment of baseline dataset, fusion strategies for multimodal affective analysis, and unsupervised learning models.
1. Introduction
Affective computing encompasses emotion recognition and sentiment analysis, using physical, physiological, unimodal, and multimodal data. This review addresses limitations in prior surveys by organizing the field broadly and clarifying methods, databases, and performance.
- Scope and modalities: Affective computing covers emotion recognition and sentiment analysis across textual, audio, visual, and physiological data.Emotion recognition commonly targets visual, speech, and physiological signals, while sentiment analysis focuses on textual evaluations and opinion mining.
- Motivation for multimodality: Physical signals are plentiful and accessible, but physiological signals can better capture subtle or concealed emotional states despite acquisition challenges.The review therefore motivates combining physical and physiological modalities for complex affective analysis.
- Motivation for multimodality: Multimodal affective analysis combines complementary textual, auditory, visual, and physiological information to characterize complex human affect.The review identifies modality selection and fusion strategy as key components of multimodal analysis.
- Review gap: Prior reviews often took specialist perspectives and omitted broad coverage of deep-learning or multimodal methods and clear performance implications.The present review aims to cover different aspects of affective computing through methods, results, discussion, and future work.
- Review scope and contributions: The review categorizes affective computing into unimodal affect recognition and multimodal affective analysis while surveying more than 380 papers from the past 20 years.It also organizes benchmark databases by modality and summarizes representative methods and quantitative performance.
2. Related works
Recent reviews cover physical, physiological, and multimodal affective computing, but gaps remain in comprehensive coverage of recent deep-learning advances and comparative analysis across unimodal and multimodal settings.
- Review coverage: Table 2 compares recent reviews by emotion model, databases, modalities, fusion, methods, and quantitative evaluation.The proposed review is presented as systematically covering all listed aspects.
- Physical-based reviews: Existing physical-based reviews primarily address visual, textual, and audio modalities, including facial expressions, micro-expressions, body gestures, sentiment, and speech.These reviews commonly examine deep-learning methods and modality-specific recognition tasks.
- Physiological-based reviews: Physiological reviews have examined machine-learning methods in discrete and dimensional emotion spaces, while relatively few have addressed deep-learning methods.Reviewed topics include physiological signals, nonlinear EEG indexes, feature representation, classification, and performance.
- Multimodal reviews: Recent multimodal reviews increasingly combine physical and physiological signals through feature representations, fusion strategies, recognition methods, databases, and applications.The surveyed modalities include audio, visual, textual, EEG, and other physiological signals.
- Remaining gaps: Earlier reviews have not fully elaborated comparisons in unimodal and multimodal affective analysis, which the proposed review identifies as an objective.The review also targets recent research advances in deep-learning-based affective computing.
3. Emotion models
Affective computing uses discrete and dimensional emotion models to represent emotional states. Discrete models assign emotions to categories, whereas dimensional models locate them in continuous spaces defined by affective dimensions.
- Model families: Affective computing commonly uses two generic emotion-model families: discrete emotion models and dimensional emotion models.No unanimously accepted emotion model exists across the multidisciplinary study of emotion.
- Discrete emotion model: Discrete models define emotions as limited categories, including Ekman’s six basic emotions and Plutchik’s eight-emotion wheel.Ekman’s categories include anger, disgust, fear, happy, sad, and surprise; Plutchik’s wheel also represents relationships and intensity.
- Discrete emotion model: Plutchik’s componential wheel arranges stronger emotions toward the centre and weaker emotions toward the extremes according to relative intensity.The model also encodes relationships such as oppositions and possible emotional developments.
- Dimensional emotion model: The PAD model represents emotion in three dimensions: Pleasure, Arousal, and Dominance.These dimensions correspond respectively to joy or distress, physiological activity or alertness, and influence over or by the surrounding environment.
- Dimensional emotion model: The Valence-Arousal circumplex represents complex emotions in a continuous two-dimensional space organized by pleasantness and activation.Its quadrants associate combinations of valence and arousal with emotions such as happy, sad, and angry.
4. Databases for affective computing
Affective-computing databases are organized by modality into textual, speech/audio, visual, physiological, and multimodal collections, whose properties influence model and architecture design. The review surveys representative datasets spanning posed and naturalistic expressions, body gestures, physiological recordings, and combined signals.
- Database taxonomy: Affective-computing databases are classified as textual, speech/audio, visual, physiological, or multimodal according to their data modalities.These database properties influence affective-computing model design and network architecture.
- Textual databases: Textual databases contain word-, sentence-, or document-level data annotated with emotion or sentiment labels.MDS contains more than 100,000 product-review sentences, while IMDB provides 25,000 training and 25,000 testing movie reviews.
- Speech databases: Speech databases include non-spontaneous performances and spontaneous recordings, with Emo-DB representing an early actor-performed resource.Emo-DB contains about 500 utterances spoken by 10 actors.
- Physiological databases: Physiological databases use signals such as EEG, RESP, and ECG that are less affected by social masking than physical emotion signals.Task-driven resources include DSdRD, which records multiple signals from 24 volunteers during real-world driving tasks.
- Multimodal databases: Multimodal databases combine multiple physical modalities or physical and physiological signals to represent affect expressed through coordinated channels.IEMOCAP combines face, head, hand, and speech information with both categorical and continuous emotion annotations.
5. Unimodal affect recognition
The review organizes unimodal affect-recognition methods by affect modality, separating physical signals from physiological signals. Physical modalities include text, audio, and visual data, while physiological modalities include EEG and ECG.
- Method taxonomy: Unimodal affect-recognition methods are summarized by physical modalities and physiological modalities.Physical modalities include textual, audio, and visual signals; physiological modalities include EEG and ECG.
5.1 Textual sentiment analysis
Textual sentiment analysis has progressed from knowledge-based and statistical feature engineering toward hybrid and deep learning models. Representative approaches use lexicons, classifiers, CNNs, RNNs, attention, and combined architectures to model sentiment at multiple granularities.
- Traditional methods: Traditional textual sentiment analysis relies on feature engineering to identify useful sentiment-related features in large user-generated datasets.Knowledge-based methods use emotional vocabularies and linguistic rules, while statistical methods use annotated datasets and classifiers.
- Traditional methods: Knowledge-based methods use lexicons and linguistic rules, whereas statistical methods train classifiers from annotated data and prior statistics or posterior probabilities.Lexicon-based systems can classify word polarity but perform poorly without linguistic rules.
- Hybrid methods: Hybrid methods combine knowledge and statistical models to handle contextual polarity and difficult neutral or ambiguous expressions.One hybrid method achieved an average accuracy of 87.13% across Amazon, IMDb, and Yelp.
- Deep learning methods: Deep textual sentiment analysis includes CNN, RNN, ConvNet-RNN, attention-based, graph-based, and adversarial architectures.These approaches target document-, sentence-, aspect-, or word-level sentiment analysis and learn representations with varying contextual and local information.
- Deep learning methods: Combining ConvNets and RNNs exploits local feature extraction and long-sequence modeling for sentiment analysis.Parallel CNN-BiLSTM designs are used to capture both feature types.
5.2 Audio emotion recognition
Audio emotion recognition detects emotion from speech using either engineered acoustic features with conventional classifiers or end-to-end deep models. Recent systems incorporate temporal modeling, attention, CNN-RNN combinations, and adversarial learning to address emotional context and distribution differences.
- Overview: Audio emotion recognition processes speech signals to detect embedded emotions using machine-learning or deep-learning systems.The review summarizes representative methods by feature representation, classifier, database, and performance.
- ML-based SER: Traditional speech emotion recognition learns acoustic feature representations and applies classifiers such as HMM, GMM, SVM, RF, or ANN.Prosodic, spectral, voice-quality, and other acoustic features may be fused, and model-feature combinations produce different performance levels.
- ML-based SER: Prosodic features encode emotional speech through fundamental frequency, energy, and duration, while voice quality includes jitter, shimmer, and harmonics.Feature selection and fusion achieved 83.10% accuracy on the Berlin Database in one speaker-independent system.
- DL-based SER: Deep speech emotion recognition automatically learns discriminative representations with CNNs, RNNs, hybrid networks, and attention mechanisms.CNNs process spectrograms, while RNN variants capture temporal information and emotional states across sequences.
- DL-based SER: Attention mechanisms emphasize emotionally salient speech regions and can be paired with silence removal to suppress irrelevant frames.CNN-RNN systems combine frequency and temporal dependencies, while adversarial models address mismatched training and test distributions.
- DL-based SER: Generative models can augment speech emotion recognition by producing synthetic emotional features from unlabeled source data.CycleGAN is used to generate feature vectors representing target emotions.
5.3 Visual emotion recognition
Visual emotion recognition covers facial expressions and body gestures, with FER organized by image dynamics, expression duration, feature types, and learning architecture. Methods progress from hand-crafted geometric or appearance features to deep models that capture spatial, temporal, and salient expressive information.
- Visual emotion recognition primarily comprises facial expression recognition and body gesture emotion recognition.
- Facial expression recognition: FER is divided into static and dynamic systems, and into macro-FER and micro-FER according to expression duration and intensity.
- Facial expression recognition: ML-based FER uses hand-crafted geometry-based features, appearance-based features, feature fusion, or feature selection.
- Facial expression recognition: Geometry-based FER represents facial shapes and component positions, while appearance-based FER analyzes spatial or spatiotemporal facial information.
- Facial expression recognition: Feature fusion combines geometry and appearance information to enhance FER robustness, whereas feature selection removes redundant attributes from high-dimensional representations.
- Deep learning for FER: Deep FER is categorized into ConvNet, ConvNet-RNN, and adversarial learning according to network architecture.
- Deep learning for FER: ConvNets address small-database overfitting through transformed learning or loss functions, while attention mechanisms emphasize distinctive expression-related regions.
- Deep learning for FER: ConvNet-RNN systems combine CNN-derived spatial characteristics with RNN or LSTM-derived temporal characteristics for dynamic facial sequences.
5.4 Physiological-based emotion recognition
Physiological emotion recognition addresses the tendency of physical expressions to be forged by using signals that directly reflect emotional changes. The review covers EEG and ECG pipelines, including feature extraction, reduction or selection, classification, and deep representation learning.
- Physical affect cues can be forged, whereas physiological signals directly reflect emotional changes and provide more objective emotion information.
- Physiological emotion-recognition systems stimulate emotions, record signals, extract and reduce features, and train recognition models.
- EEG-based emotion recognition: EEG is frequently used because it measures brain-activity changes directly and offers high temporal resolution for real-time emotional-state monitoring.
- EEG-based emotion recognition: ML-based EEG recognition designs time-domain, frequency-domain, and time–frequency-domain features, followed by dimensionality reduction or selection and classification.
- EEG-based emotion recognition: Common ML-based EEG classifiers include SVM and its variations.
- EEG-based emotion recognition: DL-based EEG methods jointly learn features and classify emotions with CNNs, RBMs, graph neural networks, or dense convolutional networks.
- ECG-based emotion recognition: ECG detects emotion-related waveform transformations, and ML pipelines use preprocessing, R-wave detection, windowing, feature extraction, normalization, and selection.
- ECG-based emotion recognition: 79.51% accuracy with 0.13ms classification time was reported by fusing linear, nonlinear, time-domain, and time-frequency ECG features on BioVid Emo DB.
6. Multimodal affective analysis
Multimodal affective analysis fuses physical and physiological signals to seek more accurate and comprehensive emotion understanding. The review classifies approaches by modality combinations and fusion levels, reporting gains from multimodal and decision-level strategies in several settings.
- Multimodal affective analysis integrates multiple unimodal signals to obtain more accurate results and a more comprehensive understanding than unimodal recognition.
- The review classifies multimodal analysis by multi-physical, multi-physiological, and physical-physiological modality combinations, alongside four fusion strategies.
- Fusion strategies: Feature-level fusion combines modality features before classification, while decision-level fusion combines independently generated decision vectors.
- Multi-physical modality fusion: Visual-audio recognition generally outperforms visual-only or audio-only recognition, reflecting the complementary nature of these communication cues.
- Multi-physical modality fusion: Factorized bilinear pooling achieved 62.48% recognition accuracy on AFEW for visual-audio emotion recognition.
- Text-audio emotion recognition: Text-audio fusion addresses the difficulty of performing speech emotion recognition and textual sentiment analysis separately by combining linguistic and speech clues.
- Text-audio emotion recognition: Decision-level fusion reached 69.2% four-class accuracy on IEMOCAP, compared with 55.4% for feature-level fusion in one comparison.
- Multi-physiological modality fusion: Multi-physiological fusion achieved overall highest accuracies of 99.0% on AMIGOS and 90.8% on DREAMER using majority voting across classifiers.
7. Discussions
The review examines how signals, modality combinations, fusion strategies, learning paradigms, databases, and applications shape affective computing. Visual signals are widely used, while physiological signals offer objective outcomes but remain harder to obtain.
- Effects of different signals: Visual signals are the most widely used unimodal modality because they are easier to capture and provide helpful emotional information.
- Effects of different signals: Textual affective analysis achieves the highest reported accuracy, whereas audio-based recognition is more susceptible to noise than visual recognition.
- Effects of different signals: Physiological emotion recognition remains attractive because wearable-sensor signals produce objective and reliable outcomes despite being difficult to obtain.
- Modality combinations and fusion strategies: Multimodal affective analysis combines physical, physiological, or both signal types using feature-level, decision-level, hybrid-level, or model-level fusion.
- Modality combinations and fusion strategies: Feature-level fusion is more common and can be strongly affected by mismatched feature time scales and metric levels, while decision-level fusion is easier but ignores cross-modal feature relevance.
- ML-based and DL-based models: Deep learning generally outperforms machine learning through learned feature representations, although it has not had a huge impact on physiological emotion recognition relative to machine learning.
- Potential factors: Database size, quality, and collection conditions differ substantially, with laboratory-collected body-gesture datasets often containing only several hundred samples and limited categories.
8. Conclusion and new developments
The review synthesizes more than 400 papers into a taxonomy of emotion models, databases, unimodal recognition, multimodal analysis, learning methods, and applications. It concludes that stronger benchmarks, fusion strategies, and learning methods are needed for robust affective computing in challenging scenes.
- Review scope: The review surveys more than 400 papers and organizes affective computing around emotion models, metrics, benchmark databases, and computational models.
- Taxonomy: Recent advances are grouped into unimodal affect recognition and multimodal affective analysis, each further divided into machine-learning and deep-learning approaches.
- Discussion scope: The review discusses signal effects, modality combinations, fusion strategies, learning models, database and metric factors, and real-life applications.
- Open challenges: Only a few robust and effective algorithms currently predict emotion and recognize sentiment under diverse and challenging scenes.
- Future directions: Future work should develop extended multimodal databases covering spontaneous and non-spontaneous scenarios with both discrete and dimensional annotations.
- Future directions: Zero-shot, few-shot, and unsupervised methods warrant further study for affective analysis under limited or biased databases.