Source-linked AI summary

Affective Image Content Analysis: Two Decades Review and New Perspectives

Sicheng Zhao, Xingxu Yao, Jufeng Yang, Guoli Jia, Guiguang Ding, Tat-Seng Chua, Björn W. Schuller, Kurt Keutzer

arXiv:2106.16125v1cs.CVcs.AIcs.MM

TL;DR

AICA must address the mismatch between visual features and induced emotions, differences among viewers, and imperfect or missing labels. This survey reviews two decades of emotion representations, datasets, features, learning methods, and applications, then identifies open directions for more robust analysis. It concludes that an effective, efficient, and robust AICA algorithm for unconstrained conditions has yet to be designed.

  • Problem

    AICA faces the affective gap, perception subjectivity, and label noise or absence, while available datasets also exhibit bias and limited scale or label quality.

  • Method

    The survey synthesizes emotion models, datasets, handcrafted and deep features, learning methods, applications, and emerging research directions.

  • Results

    The survey finds that recent deep learning-based AICA methods have made remarkable progress, but robust performance under unconstrained conditions remains unresolved.

  • Takeaways & Limitations

    Future AICA research should address image content and context understanding, group emotion clustering, and viewer-image interaction.

  • Takeaways & Limitations

    Existing datasets are either small and well labeled or large with automatically obtained labels whose quality cannot be guaranteed.

Abstract

from arXiv · show

Images can convey rich semantics and induce various emotions in viewers. Recently, with the rapid advancement of emotional intelligence and the explosive growth of visual data, extensive research efforts have been dedicated to affective image content analysis (AICA). In this survey, we will comprehensively review the development of AICA in the recent two decades, especially focusing on the state-of-the-art methods with respect to three main challenges -- the affective gap, perception subjectivity, and label noise and absence. We begin with an introduction to the key emotion representation models that have been widely employed in AICA and description of available datasets for performing evaluation with quantitative comparison of label noise and dataset bias. We then summarize and compare the representative approaches on (1) emotion feature extraction, including both handcrafted and deep features, (2) learning methods on dominant emotion recognition, personalized emotion prediction, emotion distribution learning, and learning from noisy data or few labels, and (3) AICA based applications. Finally, we discuss some challenges and promising research directions in the future, such as image content and context understanding, group emotion clustering, and viewer-image interaction.

1 INTRODUCTION

Affective image content analysis (AICA) infers emotions induced by images, extending beyond objective visual semantics to support emotional understanding and applications. The survey organizes two decades of work around the affective gap, perception subjectivity, dataset bias, and methods for extracting features, learning emotions, and applying AICA.

  • Motivation: AICA analyzes emotions induced by images rather than only perceptual properties such as objects or segments.It can help evaluate psychological health, discover affective anomalies, and prevent extreme behaviors.
  • Challenges: The affective gap arises because low-level visual features may not represent high-level emotions: similar objects can evoke different feelings, while different content can evoke similar ones.The survey reviews handcrafted features, deep features, and image regions as ways to bridge this gap.
  • Challenges: Context can change the emotion elicited by the same image, including detailed scene context and accompanying text.The survey therefore treats context as complementary to visual feature extraction.
  • Challenges: Perception subjectivity means that viewers may assign different emotions to the same image because of personal and contextual factors.AICA addresses this through personalized emotion prediction and multiple-label or distribution-based representations.
  • Challenges: Dataset bias and domain shift make direct transfer between differently styled or sourced datasets unreliable.The survey discusses bias, label noise, and label absence alongside methods for learning from noisy or few labels.
  • Survey scope: The survey reviews emotion models, datasets, feature extraction, learning methods, applications, and future research directions in AICA.Its coverage includes dominant emotion recognition, personalized prediction, emotion distribution learning, and learning with noisy or few labels.

2 EMOTION MODELS FROM PSYCHOLOGY

The survey distinguishes categorical and dimensional emotion representations, while noting that multimedia emotions may be expected, perceived, or induced. It focuses on using established psychological models without discussing their differences or correlations in depth.

  • Emotion representation models: Psychologists mainly use categorical emotion states and dimensional emotion spaces to represent emotions in AICA.Categorical models assign discrete emotion categories, whereas dimensional models locate emotions in continuous spaces.
  • Categorical emotion states: Categorical models range from binary polarity to fine-grained categories and hierarchical schemes.One hierarchy uses positive and negative at level 1, six categories at level 2, and 25 fine-grained categories at level 3.
  • Dimensional emotion space: Dimensional models represent emotions in continuous two-dimensional, three-dimensional, or higher-dimensional Cartesian spaces.Examples include valence-arousal-dominance and spaces augmented with intensity, novelty, or related dimensions.
  • Comparing CES and DES: Compared with categorical models, dimensional models capture finer-grained and more comprehensive emotional descriptions.The two representation families are related, and prior work studies transformations between them.
  • Emotion types: Multimedia emotion can be distinguished as expected, perceived, or induced emotion.Expected emotion is intended by the creator, perceived emotion is interpreted as expressed, and induced emotion is actually felt by the viewer.
  • Scope: The survey excludes detailed discussion of differences or correlations among emotion concepts and models.It instead treats findings from psychology and cognitive science as beneficial to AICA.

3 DATASETS

AICA datasets range from small psychology- and art-based collections to large web-crawled resources with diverse annotation schemes. The survey compares their label noise and dataset bias, finding substantial differences across image types, class distributions, and dataset cleanliness.

  • Dataset overview: AICA datasets expanded from small psychology or artistic collections to large-scale datasets crawled from online social networks.The survey summarizes these datasets in Table 3.
  • Psychology-based datasets: IAPS contains 1,182 natural-color images rated by about one hundred college students on valence, arousal, and dominance.
  • Emotion-labeled datasets: Emotion6 contains 1,980 images annotated by 15 participants with valence-arousal scores and discrete emotion distributions.
  • Large-scale datasets: FI contains 23,308 images receiving at least three annotator agreements after weakly labeled web images were filtered and assessed.The dataset was constructed from images collected using eight emotion keywords.
  • Personalized perception: IESN contains more than one million images from 11,347 Flickr users and includes metadata for studying personalized emotion perception.Its labels include expected emotions from uploaders and actual emotions from viewers.
  • Label noise: Event has the least label noise, whereas Abstract has the largest, partly because abstract images are difficult to understand distinctly.
  • Dataset bias: Dataset bias is larger across different image types, is asymmetric with cleaner targets, and is influenced by class-distribution similarity.The model trained on Comics obtains 0.522 accuracy on Abstract and 0.546 on GAPED, while Event has noise ratio ρm = 0.074.
  • Dataset bias: FI has the best generalization ability among the compared datasets, with its trained model exceeding 60% classification accuracy on all datasets.

4 EMOTION FEATURE EXTRACTION

AICA emotion feature extraction has progressed from handcrafted representations to deep features developed alongside convolutional neural networks.

  • AICA feature extraction studies investigate multiple representation types, including handcrafted features and emerging deep features based on CNNs.

4.1 Hand-crafted Features

Handcrafted AICA features span low-, mid-, and high-level representations, moving from visual statistics toward interpretable composition, semantic content, and affective concepts.

  • Low-level features: Low-level handcrafted features represent emotional content using visual statistics such as lines, color, texture, shape, and image transforms.Early low-level features often lacked reasonable interpretation.
  • Mid-level features: Mid-level features bridge low-level visual properties and high-level emotions through attributes, composition, and artistic principles.Examples include SUN attributes, figure-ground contrast, multiscale representations, and principles-of-art features.
  • Mid-level features: Principles-of-art features combine balance, emphasis, harmony, variety, gradation, and movement as an emotion representation.
  • Mid-level features: Roundness, angularity, and visual complexity provide effective mid-level representations using only one scalar per characteristic.
  • High-level features: High-level handcrafted features use semantic content, facial expressions, and viewer-attention-related elements that can directly evoke emotions.Facial expressions are commonly organized according to Ekman’s six basic emotions.
  • High-level features: SentiBank represents visual sentiment with 1,200 adjective-noun pair concepts derived from Plutchik’s emotion theory.

4.2 Deep Features

Deep AICA features progressed from global CNN representations toward multi-level and localized representations that target emotionally informative image regions. These approaches include transfer learning, weak-label denoising, multi-network fusion, attention, region proposals, and joint holistic-local modeling.

  • Global and Deep Features: Progressive CNN training removes noisy weakly labeled instances across iterations, improving robustness when transferring to small strongly labeled datasets.Instances with large polarity differences are retained for subsequent training rounds.
  • Multi-level Features: Multi-level feature methods combine complementary CNN outputs, including parallel semantic, aesthetic, and texture networks for richer emotional representation.These architectures address cases where high-level semantic features alone are insufficient, such as abstract paintings.
  • Local Features: Local-feature methods extract and aggregate patches or detected regions because emotionally informative content may occupy only important parts of an image.Approaches include multi-scale patch features, Fisher Vector aggregation, attention, object proposals, sentiment maps, and holistic-local feature fusion.
  • Local Features: Region-proposal methods concatenate multi-level ROI features for image emotion classification and achieved the best reported classification performance on several benchmark datasets.The comparison also evaluates local, global, and combined local-global features using average classification accuracy and rank.

4.3 Quantitative Feature Comparison

Feature comparisons across six datasets show that deep representations generally provide the strongest performance, while high-level semantic features usually exceed middle-level handcrafted features. Classifier-level averaging further indicates that representation quality depends on the feature type and classifier pairing.

  • Feature Comparison: Deep features obtain the best performance across the reported feature comparisons on six widely used datasets.The evaluated representations include handcrafted PAEF, Sun attributes, and SentiBank, plus MVSO and pretrained VGGNet-16 features.
  • Feature Comparison: High-level SentiBank features generally outperform middle-level PAEF features because adjective-noun-pair concepts map more directly to emotional semantics.The comparison averages results across classifiers for each feature type.

5 LEARNING METHODS FOR DIFFERENT TASKS

AICA learning methods address dominant emotion recognition, personalized perception, emotion distributions, and limited or noisy labels through increasingly specialized representations and learning objectives. The surveyed approaches incorporate multiple feature levels, emotional polarity, local regions, user context, domain adaptation, and few-shot transfer.

  • Overview: The survey organizes AICA learning methods into dominant emotion recognition, personalized emotion prediction, emotion distribution learning, and learning from noisy data or few labels.These tasks reflect different treatments of affective labels and perception variability.
  • Traditional Methods: Traditional methods classify handcrafted emotional features with SVMs, while abstract-art recognition uses multiple kernel learning and non-linear matrix completion to combine feature spaces.Multiple kernel learning adjusts feature weights automatically to capture different emotional patterns.
  • Learning-based Methods: Deep recognition methods model emotion through multi-level feature dependencies, polarity-aware metric learning, and local-global representations built from informative regions.Contrastive loss brings same-category features closer and separates different-category features, while salient-region models combine sub-images with entire images.
  • Learning-based Methods: Across backbones, RCA is more robust than WSCNet and PDANet, while recent specialized methods generally outperform DCNN and dataset-specific winners differ.RCA’s multi-layer representation fluctuates less across architectures; local-feature methods perform especially well on the small Twitter II dataset.
  • Personalized Emotion Prediction: Personalized emotion prediction combines visual content with social context, temporal emotion history, and location through rolling multi-task hypergraph learning.Hypergraph vertices represent a user, target image, and recent history image set, with corresponding target-, history-, and user-centric hyperedges.
  • Emotion Distribution and Limited Labels: Emotion distribution learning predicts graded emotion descriptions, while few-shot and zero-shot methods transfer from seen to unseen classes using visual features and semantic side information.The affective gap makes visual-to-semantic similarity difficult to compute for unseen classes.
  • Learning from Noisy Data or Few Labels: EmotionGAN addresses domain shift in emotion distribution learning by alternately optimizing adversarial, semantic-consistency, and regression losses.The model aims to generate an intermediate domain indistinguishable from target images while preserving source labels.

6 AICA BASED APPLICATIONS

AICA applications span social analysis, product and business intelligence, psychological health, psychology research, tourism, advertising, entertainment, and image-based dialogue. These applications use affective image understanding to interpret users, experiences, decisions, and cross-modal content.

  • Social and Business Applications: Social-media image emotion analysis can infer users’ emotions and attitudes toward events or products, while group-emotion analysis may help predict societal tendencies.The surveyed work connects visual emotion understanding with public-opinion analysis.
  • Social and Business Applications: Visual sentiment analysis supports product, service, and venue review understanding by incorporating user, item, visual, and textual information.Some approaches jointly classify visual and textual product-review content.
  • Psychological Health: Image-based emotion signals have been studied for detecting depression, anxiety, and stress from social-web profiles and posts.These applications target monitoring of users who repeatedly share negative information.
  • Psychology Research: Psychological image systems provide standardized affective stimuli and average elicited-emotion ratings for research on psychological status and emotion theories.IAPS is described as a database designed to evoke target emotions in people.
  • Commercial and Entertainment Applications: Affective image analysis is applied to advertising, tourism, and entertainment by relating visual emotion to consumer decisions, travel experiences, and cross-media retrieval.Reported examples include emotional advertising design, travel-photo analysis, and image-music retrieval.
  • Commercial and Entertainment Applications: Entertainment-oriented AICA includes comic-emotion analysis and image-based chatbot conversation rather than text-only interaction.Emotion is treated as an important element of comics and as a bridge toward image-driven dialogue.

7 FUTURE DIRECTIONS

Future AICA research should jointly improve image-content and contextual understanding while modeling viewer-specific factors. These directions address unresolved variation in how images evoke emotions across viewers and situations.

  • Image Content and Context Understanding: AICA needs more accurate image-content analysis because subtle objects, combinations, and settings can change the emotion evoked by an image.The paper highlights cases where flowers or laughter receive different emotional interpretations depending on surrounding context or the person involved.
  • Image Content and Context Understanding: Combining interpretable handcrafted features with deep features may improve performance, while using handcrafted features to guide interpretable deep representations remains open.The review notes that deep features generally outperform handcrafted ones, but their relationship to specific emotions is unclear.
  • Viewer Contextual and Prior Knowledge Modeling: Context modeling is important because the same viewer may experience different emotions for the same image under different climate, time, or social conditions.Probabilistic graph and hypergraph models are identified as feasible ways to represent correlations among contextual factors.
  • Viewer Contextual and Prior Knowledge Modeling: Viewer demographics and other prior knowledge can improve emotion classification, but social-network information may be inaccurate and requires noise filtering.Reported factors include gender, marital status, occupation, and personality-related information.

7.3 Learning from Noisy Data or Few Labels

Future work on sparse or shifted AICA data should improve few-shot and zero-shot representations while addressing increasingly complex domain-transfer settings. The paper also proposes moving beyond individual personalization toward group-level emotion prediction.

  • Few-shot or Zero-shot Learning: Few-shot and zero-shot AICA methods may lose information and fail to exploit data distributions, motivating representative-image selection and reliable synthesis for unseen classes.The paper identifies generative models as a potential basis for synthesizing samples from estimated distributions.
  • Domain Adaptation/Generalization: Domain adaptation research should determine the minimum labeled target data needed and address heterogeneous, open-set, and category-shift settings.These settings differ in label spaces, unknown classes, or category composition across source and target domains.
  • Domain Adaptation/Generalization: Domain generalization differs from adaptation because it trains without accessing target images, making source-domain diversity and domain randomization important directions.The target domain is treated as one of the randomized domains during training.
  • Group Emotion Clustering: Group-based prediction could occupy a useful middle ground between generic dominant-emotion recognition and highly specific personalized prediction.Users sharing tastes, interests, or backgrounds may respond similarly to the same image.
  • Group Emotion Clustering: Affective group-emotion recognition remains largely unexplored because existing group-emotion work mainly analyzes people attending social events rather than groups’ induced responses to images.The paper connects this direction to recommendation, where one member’s product interest may influence others in the same group.

7.5 Viewer-Image Interaction

Viewer-image interaction research can complement image-content analysis with viewers’ audiovisual or physiological responses. The paper presents joint modeling as a way to address open problems involving implicit emotion signals, missing data, and affective image generation.

  • Viewer-Image Interaction: Implicit emotional tagging records viewers’ audiovisual or physiological responses, but current methods mainly study videos rather than still images.Examples include facial expressions and electroencephalogram signals.
  • Viewer-Image Interaction: Jointly modeling image content and viewer responses may better bridge the affective gap and improve performance.The review frames this as a largely open research topic for image emotion analysis.
  • Viewer-Image Interaction: Missing or corrupted physiological signals require methods that can handle incomplete viewer-response data.The paper notes that some physiological signals may not be successfully captured in practice.
  • Viewer-Image Interaction: When physiological responses are unavailable, privileged-modality information used during training may still outperform image-only learning.This comparison is proposed for real-world applications where physiological signals are difficult to capture or are suppressed.
  • Viewer-Image Interaction: AICA applications may adjust image color, texture, or other features to shift an image’s evoked emotion distribution toward a target.The paper identifies GAN-based adversarial models as potentially suitable for generating affective images.

7.7 Efficient AICA Learning

Efficient AICA learning is needed because mobile and edge devices lack the computing, memory, and power resources that support conventional deep learning. Progress also depends on larger, higher-quality datasets that resolve the trade-off between small labeled collections and noisy automatic annotations.

  • Efficient Model Design: Edge devices constrain AICA because they may lack the computing capacity, memory, and power required by conventional deep-learning factors.The paper identifies efficient model design as a requirement for mobile deployment.
  • Efficient Model Design: AICA efficiency remains underexplored, and computer-vision efficiency methods could be adapted by incorporating AICA-specific properties such as emotion hierarchy.The paper also proposes online incremental learning for on-device models.
  • Benchmark Dataset Construction: Existing datasets trade off scale and label quality: small datasets lack training samples, while keyword-based large datasets have uncertain automatic annotations.The review uses IAPSa and IESN as examples of these two dataset types.
  • Benchmark Dataset Construction: A large, high-quality benchmark could collect personalized emotion judgments together with viewer backgrounds, contextual responses, social interactions, and spontaneous reactions.Crowdsourcing and online systems are proposed to obtain annotations from viewers with diverse backgrounds.

8 CONCLUSION

The survey synthesizes two decades of AICA research, comparing representative models, datasets, methods, applications, and emerging directions. Despite deep-learning progress, effective, efficient, and robust AICA under unconstrained conditions remains unresolved.

  • The survey reviews representative emotion models, datasets, feature-extraction methods, learning approaches, applications, open issues, and future directions.
  • Deep learning has produced remarkable progress in AICA, but robust performance under unconstrained conditions remains an open problem.
  • The authors identify brain science, psychological emotion measurement, and new machine-learning architectures as drivers of continued AICA research.
Loading 2106.16125v1…