Source-linked AI summary

From Affect to Complex Behavior: Advancing Multimodal Human-Centered AI at the 10th ABAW Workshop & Competition

Dimitrios Kollias, Panagiotis Tzirakis, Alan Cowen, Stefanos Zafeiriou, Irene Kotsia, Eric Granger, Marco Pedersoli, Simon Bacon, Jens Madsen, Soufiane Belharbi, Muhammad Haseeb Aslam, Chunchang Shao, Guanyu Hu

arXiv:2605.27451v1cs.CV

TL;DR

The workshop addresses the challenge of understanding affect and behavior in real-world, unconstrained environments, where multimodal and socially meaningful signals must be analyzed. It presents a competition and paper track covering affect estimation, behavior analysis, multimodal learning, evaluation, fairness, robustness, and deployment. Its conclusion emphasizes multimodal fusion, pretrained models, temporal reasoning, and deployment-aware design as important across these tasks.

  • Problem

    The workshop addresses limited understanding of human affect and behavior in real-world, unconstrained environments, including multimodal and temporally evolving signals.

  • Method

    The workshop combines a competition benchmarking affective and behavioral tasks with a paper track spanning multimodal learning, pose and behavior estimation, evaluation, fairness, robustness, and deployment.

  • Results

    The workshop’s results highlight multimodal fusion, strong pretrained models, temporal reasoning, and deployment-aware design across real-world affective and behavioral analysis tasks.

  • Takeaways & Limitations

    The 10th ABAW Workshop provides a platform for benchmarking, collaboration, and innovation in multimodal, robust, and human-centered AI.

  • Takeaways & Limitations

    The competition restricts refinement and methodology development to task-specific annotations, excluding related labels such as expression or action units in the valence-arousal challenge.

Abstract

from arXiv · show

The 10th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop and Competition, held at CVPR 2026, continues to advance research on modelling, analysis, understanding of human affect and behavior in real-world, unconstrained environments. The workshop maintains its dual structure, comprising both a competition and a paper track. The ABAW Competition introduces a diverse set of challenges targeting key aspects of affective and behavioral understanding, including continuous affect (valence-arousal) estimation, discrete affect (expression and action unit) recognition, as well as more complex behavior analysis tasks, such as emotional mimicry intensity estimation, ambivalence/hesitancy recognition and fine-grained violence detection. These challenges are built upon large-scale in-the-wild datasets, providing comprehensive benchmarks for state-of-the-art approaches. In parallel, the paper track presents a wide range of contributions spanning pose, motion & behavior estimation, affect modelling & multimodal learning, benchmarks, datasets & evaluation protocols, fairness, robustness & deployment. Overall, the 10th ABAW Workshop and Competition continues to serve as a key platform for benchmarking, collaboration and innovation, shaping the development of next-generation multimodal, human-centered AI systems.

Hume AI, USA

The listed contributors are affiliated with Google DeepMind, LIVIA, ILLS, ETS Montreal, and Hume AI in the USA and Canada.

  • Alan Cowen is affiliated with Google DeepMind in the USA.
  • Eric Granger, Marco Pedersoli, Soufiane Belharbi, and Muhammad Haseeb Aslam are affiliated with LIVIA, ILLS, and ETS Montreal in Canada.
  • Jens Madsen is affiliated with Hume AI in the USA.

1. Introduction

The introduction frames affect and behavior as multimodal, socially relevant phenomena and positions the 10th ABAW Workshop as a forum for advancing robust, human-centered analysis in-the-wild.

  • Affect and behavior shape communication, decisions, social interaction, and responses to the environment, motivating their study in human-centered AI.
  • Human affect and behavior are inherently multimodal, combining facial, bodily, vocal, linguistic, and sometimes physiological signals.
  • Multimodal analysis supports scientific study and applications including healthcare, education, robotics, human-computer interaction, and security.
  • ABAW has evolved from core affective tasks toward complex, temporally evolving, socially meaningful, and application-relevant behavior analysis.
  • The 10th ABAW Workshop spans pose, motion, behavior, affect modelling, multimodal learning, benchmarks, fairness, robustness, and deployment.

2. Workshop Overview

The workshop overview spans advances in pose and behavior estimation, affective multimodal learning, evaluation resources, fairness, robustness, and deployment.

  • Pose, Motion and Behavior Estimation: Pose and behavior papers model inter-person correlations, self-distillation, language-aligned pose representations, and enriched skeleton cues.
  • Affect Modelling and Multimodal Learning: Affect-modelling papers address lightweight affect dynamics, annotator uncertainty, reliability-aware fusion, and richer emotion embeddings.
  • Benchmarks, Datasets and Evaluation Protocols: Benchmark contributions provide grounded scene-graph annotations, temporally annotated laughter datasets, event-centric evaluation, and leave-one-dataset-out AU evaluation.
  • Fairness, Robustness and Deployment: Fairness and robustness contributions examine visual confounders in facial-expression assessment and topology-guided test-time adaptation under domain shift.

3. Competition Overview

The competition spans continuous affect estimation, discrete expression and action-unit recognition, and complex behavior challenges using in-the-wild datasets. Leading approaches consistently use multimodal inputs, pretrained encoders, temporal modeling, and adaptive or cross-modal fusion.

  • Valence-Arousal Estimation: The VA challenge predicts continuous valence and arousal per frame, with values ranging from -1 to 1.The dataset contains 594 videos, 2,993,081 annotated frames, and 584 subjects.
  • Valence-Arousal Estimation: The VA baseline uses ImageNet-pretrained ResNet-50 and achieves an average CCC of 0.22 on validation.
  • Valence-Arousal Estimation: The leading VA systems combine facial, behavioral, and audio cues with Transformer, Mamba, temporal, and reliability-aware fusion mechanisms.RAS uses GRADA, Qwen3-VL, and WavLM; EmoDX uses SAGE with stage-adaptive reliability-guided fusion; IMLAB adds semantic guidance and hierarchical cross-modal attention.
  • Multimodal Behavior Analysis: Top-performing AU and behavior systems use pretrained multimodal encoders, hierarchical alignment, long-range temporal modeling, and reliability-aware fusion.USTC-IAT-United combines DINOv2 and WavLM with audio-guided state-space updates, while other systems model complementary or conflicting visual, audio, and textual cues.
  • Fine-Grained Violence Detection: The violence-detection baseline combines ImageNet-pretrained ResNet-50 with a bidirectional LSTM and reaches a macro F1 score of 0.64 on validation.The winning HSEmotion system fuses visual and skeleton streams through bidirectional cross-attention before two-layer BiLSTM decoding.

4. Conclusion

The 10th ABAW Workshop and Competition advances multimodal, human-centered AI by expanding affect analysis toward richer, temporally grounded and socially meaningful behavior understanding. Its challenges and papers emphasize multimodal fusion, pretrained models, temporal reasoning and deployment-aware design for real-world tasks.

  • The workshop marks a shift from core affect recognition toward richer, multimodal, temporally grounded and socially meaningful human behavior analysis.
  • Six competition challenges and a diverse paper track provide a platform for benchmarking, collaboration and innovation.
  • The edition highlights multimodal fusion, strong pretrained models, temporal reasoning and deployment-aware design as important for real-world affective and behavioral analysis.
Loading 2605.27451v1…