Source-linked AI summary

MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions

Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat

arXiv:2609.11322v1cs.CLcs.CVcs.MM

TL;DR

Existing humour-recognition systems often reduce humour to binary classification and lack datasets capturing psychological styles, emotional undertones, and expressive variation. MultiHuSE addresses this gap with a multimodal dataset and benchmarks fusion models, which achieve 80.1% accuracy versus 77.4% for text-only models while revealing complementary audio and visual contributions.

  • Problem

    Most multimodal humour-recognition approaches focus on binary humour classification rather than psychological humour styles and their associated emotional expressions.

  • Method

    MultiHuSE contains over 2,400 recordings of fifty performers across four psychological humour styles and neutral content, with multiple performances of the same texts and performer-reported emotion annotations.

  • Results

    80.1% accuracy from exponential fusion exceeded 77.4% for text-only models, while cross-attention achieved 79.7%; ablations identified text as dominant but audio and visual cues as complementary.

  • Takeaways & Limitations

    MultiHuSE supports analysis of expressive diversity and computational investigation of relationships between humour styles and emotions, including connections to psychological theories.

  • Takeaways & Limitations

    Moderate inter-annotator agreement and recruitment constrained to London may limit representation of global humour expressions and introduce cultural bias.

Abstract

from arXiv · show

Computational recognition of verbal humour remains a challenging task, requiring an understanding of language, delivery style, emotions, and cultural context. Most existing approaches focus on binary classification and lack datasets that capture psychological dimensions of humour alongside variations in expression. We introduce MultiHuSE, a multimodal dataset comprising 2,407 high-definition videos of 50 demographically diverse actors performing 1,463 text samples across four psychological humour styles (affiliative, aggressive, self-enhancing, and self-deprecating), as well as neutral content. A subset is additionally annotated for underlying emotions. The dataset uniquely captures multiple actor interpretations of the same texts, enabling systematic analysis of expressive diversity. Baseline experiments show that multimodal fusion outperforms unimodal approaches (80.1% vs. 77.4% accuracy) in humour style classification, with particularly strong gains for affiliative humour (66% to 74%). While text provides the strongest individual signal, fusion models deliver meaningful improvements. We hope that MultiHuSE provides empirical support for psychological theories linking humour and emotion, while also opening new avenues for research in human communication, well-being, and AI-driven interaction. The dataset is available for academic use under an End-User Licence Agreement.

I. INTRODUCTION

Existing humour systems largely detect humour as a binary phenomenon, while MultiHuSE targets psychologically distinct styles, emotional expression, and performance variation in multimodal data.

  • Research Motivation: Humour styles have differing reported relationships with psychological well-being, motivating computational recognition of styles and associated emotional expressions.Affiliative and self-enhancing styles are generally linked to positive outcomes, whereas aggressive and self-deprecating styles can be detrimental to mental health.
  • Dataset Gaps and Research Motivation: MultiHuSE addresses datasets that typically overlook psychological humour styles, emotional undertones, and multiple interpretations of the same content.Prior resources include binary humour datasets, style-limited domains, or single performances per instance.
  • Dataset Contribution: The dataset contains over 2,400 recordings by fifty diverse performers across four humour categories and neutral content, with underlying emotion annotations.Its multiple actor interpretations of identical texts support analysis of expressive variation and multimodal humour perception.
  • Benchmark Contribution: 80% versus 77% F1-score shows multimodal fusion outperforming unimodal models in the benchmark experiments.The contribution statement presents this comparison as evidence that the dataset supports multimodal humour-style recognition.

A. Dataset Gaps and Research Motivation

MultiHuSE is designed to address binary, single-expression, and weakly theory-linked humour datasets through controlled, diverse recordings that preserve interpretative variation.

  • Dataset Gaps and Research Motivation: The study targets three gaps: binary classification, single-expression capture, and limited integration of psychological theories linking humour, emotion, and well-being.These gaps motivate a dataset at the intersection of computational analysis and psychological theory.
  • Dataset Overview: The dataset contains 2,407 high-definition recordings by fifty participants performing nearly 1,500 psychologically labelled text instances.The recordings, performances, and emotion annotations are novel contributions, while the text labels come from prior work.
  • Recording Setup: Recordings used a professional controlled environment with consistent lighting, backdrop, positioning, camera, and microphone conditions.The setup was intended to keep facial expressions and body gestures visible while reducing environmental variability.
  • Dataset Composition: The collection is near-balanced across self-enhancing, aggressive, neutral, self-deprecating, and affiliative categories.Category counts range from 426 affiliative videos to 529 aggressive videos.
  • Interpretative Diversity: 943 re-performed videos capture multiple interpretations of identical content, enabling direct comparison of expressive variation.The controlled design makes performer-specific delivery differences observable while holding the text constant.

B. Annotation and Performance Approach

MultiHuSE combines psychologically grounded text categories, interpretative freedom, and performer self-reports to connect humour delivery with underlying emotional states.

  • Annotation and Performance Approach: Actors performed 1,463 text instances spanning four humour styles and 334 neutral non-jokes using labels from an existing psychologically grounded corpus.The four styles are self-enhancing, self-deprecating, affiliative, and aggressive.
  • Performance Guidance and Interpretative Freedom: Participants were encouraged to deliver identical jokes across different emotional spectrums, explicitly allowing natural variation rather than correct interpretations.The provided emotion list included joy, anger, neutral, superiority, embarrassment, and empathy.
  • Emotion Annotation: Performers self-reported their emotional state immediately after each delivery, distinguishing internal experience from audience or observer perception.These reports were intended to support analysis of humour-embedded emotional states and psychological theory validation.
  • Emotion Annotation: Emotion agreement was evaluated on 687 paired original and re-performance instances, yielding 0.440 simple agreement and Cohen’s Kappa of 0.325.The kappa value indicates fair agreement above random chance, while only 687 of 943 re-performances had emotion labels.

C. Humour Styles-Emotion Relationships

Emotion–humour associations remained consistent across original and re-performed annotations, with distinct styles showing characteristic emotional pairings.

  • Humour Styles-Emotion Relationships: The consistency of these patterns across performers supports the reported robustness of emotion–humour relationships and their alignment with psychological theories.The authors present the relationships as potential support for computational recognition of emotional aspects of humour styles.
  • Humour Styles-Emotion Relationships: Neutral humour showed a strong association with neutral emotion in both original and re-performance annotations.The reported counts were 90 for original performances and 98 for re-performances.
  • Humour Styles-Emotion Relationships: Aggressive humour paired predominantly with anger and superiority, while self-deprecating humour aligned strongly with embarrassment across performances.Aggressive pairings were Anger 60 versus 74 and Superiority 42 versus 40; self-deprecating Embarrassment was 70 versus 55.
  • Humour Styles-Emotion Relationships: Affiliative humour primarily linked to joy, while self-enhancing humour co-occurred with empathy and joy.Affiliative Joy counts were 59 and 67; self-enhancing Empathy counts were 39 and 31, and Joy counts were 36 and 45.

D. Ethical Considerations and Data Access

MultiHuSE was collected under formal ethical oversight, with consent, compensation, opt-out provisions, and pseudonymisation alongside visible facial features. Access is limited to academic and research use under an End-User Licence Agreement, with documentation addressing offensive aggressive-humour content.

  • The study received ethics committee approval, and participants gave informed consent, received £50 compensation, and could decline uncomfortable content.
  • Facial features remained visible for research purposes, while personal identifiers were pseudonymised and participants were informed of this requirement.
  • The dataset is accessed through the project website under an End-User Licence Agreement permitting academic and research use only.
  • Documentation states that including potentially offensive aggressive-humour samples does not imply endorsement and restricts use to mental-health and computational-humour research.

IV. EXPERIMENTS

The experiments establish foundational baselines for MultiHuSE because direct comparison with existing datasets is not feasible. They evaluate humour-style and emotion recognition using multimodal features extracted by state-of-the-art encoders.

  • MultiHuSE baselines evaluate humour-style and emotion recognition across single and combined modalities, providing foundational benchmarks for future research.
  • State-of-the-art encoders extract multimodal features for the experimental evaluation.
  • BERT-base-uncased produces 768-dimensional text embeddings from the [CLS] token using sequences padded to 128 tokens.
  • Dasheng-0.6B extracts 1280-dimensional audio features intended to capture nuanced speech and sound information.
  • MC3-18 provides 768-dimensional spatiotemporal video features optimized for short clips and fine-grained motion representation.

B. Data Splits and Preprocessing

The evaluation uses stratified data splitting and compares unimodal, weighted-fusion, and cross-modal-attention approaches. These models combine modality-specific features either through performance-based prediction weighting or learned cross-modal relationships.

  • An 80:20 train-test split is stratified by humour style, with 5-fold stratified cross-validation applied to the training data.
  • The study evaluates XGBoost classifiers and a Cross-Modal Attention Transformer for humour-style and emotion classification.
  • Unimodal baselines train separate XGBoost models on text, audio, and visual encoder features.
  • Exponential weighted fusion combines modality-specific XGBoost predictions using α = 2 and performance-based dynamic weights.
  • Cross-modal attention learns relationships between modalities by weighting one modality's feature relevance when processing another.

E. Evaluation

Text is the strongest individual modality, but multimodal fusion improves overall humour-style classification and especially benefits affiliative humour. The results support semantic dominance with complementary expressive information from audio and video.

  • The evaluation reports weighted accuracy, precision, recall, and F1-score to measure classification performance.
  • 77.4% performance from text features exceeds audio at 58.5% and visual features at 40.0%, identifying language as the strongest individual signal.
  • 80.1% accuracy from exponential weighted fusion and 79.7% from cross-attention outperform unimodal baselines, although gains over text-only are +2.7% and +2.3%.
  • Affiliative humour improves from 66% to 74% F1 with exponential fusion, while neutral content reaches 85–90% F1.
  • Text remains dominant across fusions, whereas audio-visual-only features achieve 58.3%, indicating semantic information is enriched rather than replaced by multimodal expression.

B. Dataset Quality and Validation

MultiHuSE validation indicates reliable multimodal feature extraction and confirms that repeated performances preserve content-related similarity while exhibiting statistically significant expressive variation.

  • The validation was organized around feature extraction reliability, re-performance variation, and annotation coverage.
  • 1) Feature Extraction Reliability:: 100% feature extraction success was achieved across all modalities using modern pre-trained encoders.This provides complete modality coverage for multimodal fusion experiments.
  • 2) Re-performance Variation:: Within-joke facial-feature distances were significantly smaller than random-baseline distances, showing content consistency alongside expressive diversity.The analysis covered 1,872 re-performance videos corresponding to 927 unique jokes.

3) Annotation Coverage:

Humour-style labels cover the full video collection, while emotion annotations cover most but not all instances; these labels support psychologically informed applications described by the paper.

  • 3) Annotation Coverage:: All 2,407 videos have complete humour-style labels, while emotion annotations cover 2,150 instances, or 89.3% of the dataset.Emotion labels are available for all 1,463 original performances and 687 re-performances, representing 72.7% of re-performed instances.
  • 3) Annotation Coverage:: The partial emotion coverage resulted from introducing the re-performance annotation protocol partway through data collection.
  • The paper links humour-style differentiation to mental-health outcomes, with affiliative and self-enhancing styles associated with protection and aggressive and self-deprecating styles with increased risks.
  • Proposed applications include monitoring humour expressions during therapy and supporting clinicians in assessing coping mechanisms and therapeutic progress.
  • Additional applications include detecting harmful humour and enabling virtual assistants to distinguish supportive from harmful humour.

VII. CONCLUSION

MultiHuSE is a multimodal dataset built around psychologically grounded humour styles, repeated performances, and emotion recognition. Its validation reports strong extraction and fusion results, while its London-based sampling and moderate annotation agreement constrain generalization.

  • VII. CONCLUSION: MultiHuSE comprises over 2,400 high-definition recordings of fifty performers across four psychologically grounded humour categories, including 943 re-performed samples.The repeated performances enable direct comparison of expressive variation for identical textual content.
  • VII. CONCLUSION: 80.1% accuracy was achieved by exponential fusion versus 77.4% for text-only classification, while cross-attention reached 79.7%.Ablation results identify text as the strongest individual signal while showing complementary audio and visual contributions.
  • VII. CONCLUSION: Validation found 100% feature-extraction success and significant performance variation reflecting interpretative diversity, with p < 0.000001.
  • VIII. LIMITATIONS AND FUTURE DIRECTIONS: Moderate inter-annotator agreement (Cohen’s κ = 0.325) and recruitment limited to London constrain emotional-label consistency and global cultural representation.The authors note that the geographical limitation may reinforce Western humour perspectives and introduce cultural bias in models trained on the dataset.
  • VIII. LIMITATIONS AND FUTURE DIRECTIONS: The controlled dataset contains 2,407 videos and emphasizes quality and balance, including 79.4% humorous content evenly distributed across styles.The authors caution that direct comparison with existing datasets is challenging.
  • VIII. LIMITATIONS AND FUTURE DIRECTIONS: Future work will expand recruitment across countries and cultures, study longitudinal humour-style changes, and examine automatically detected versus human-perceived emotions.
Loading 2609.11322v1…