Source-linked AI summary

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin

arXiv:2607.19011v1cs.CLcs.AIcs.MM

TL;DR

Multimodal humor remains difficult because its intended meaning depends on implicit mechanisms, cultural knowledge, and communicative intent beyond literal perception. This survey organizes recognition, interpretation and reasoning, and generation, finding that interpretation remains the central bottleneck and that robust progress requires socially and rhetorically grounded reasoning.

  • Problem

    Multimodal humor requires inferring implicit mechanisms, cultural references, violated expectations, and communicative intent beyond observable visual and textual content.

  • Method

    The survey organizes literature, benchmarks, and modeling paradigms around recognition, interpretation and reasoning, and generation capabilities.

  • Results

    Interpretation-oriented tasks remain the central bottleneck, with current MLLMs consistently trailing humans when inferring intended humorous mechanisms.

  • Takeaways & Limitations

    Robust multimodal humor understanding and generation require socially and rhetorically grounded reasoning across recognition, interpretation, and generation.

  • Takeaways & Limitations

    Existing research and datasets focus mainly on Western, internet-centric visual humor, leaving many cultural traditions and non-mainstream media underexplored.

Abstract

from arXiv · show

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.

1 Introduction

Multimodal humor understanding requires interpreting how visual content and associated text jointly convey non-literal meaning, social targets, and communicative stance beyond explicit recognition. This survey organizes the literature by recognition, interpretation and reasoning, and generation while synthesizing benchmarks, evaluation protocols, modeling paradigms, and technical and socio-technical bottlenecks.

  • Problem setting: Multimodal humor understanding interprets image-based artifacts whose humorous, satirical, or ironic meaning emerges from visual–text interaction.The scope covers single-image and multi-panel artifacts; humor generation is treated as an emerging downstream frontier.
  • Why this gap matters: Literal recognition can miss the metaphor, social target, or punchline because humor requires reasoning about juxtaposition, background knowledge, and communicative stance.A model may identify depicted elements correctly while failing to recover what those elements jointly encode.
  • The gap in existing surveys: Existing surveys treat text-based humor, sarcasm detection, meme classification, or general multimodal reasoning separately without systematically addressing visual grounding and cross-modal incongruity.The survey is designed to fill this gap with dedicated treatment of visual humor understanding.
  • Survey contribution: The survey organizes the literature around recognition, interpretation and reasoning, and generation as progressively demanding capabilities.This capability-centric framing aligns background concepts, benchmark design, and modeling paradigms.
  • Survey contribution: The survey synthesizes datasets and evaluation protocols while identifying shortcut-prone evaluation, sparse mechanism-level annotation, weak cultural grounding, and safety and ownership concerns as bottlenecks.It emphasizes what current benchmarks do and do not measure about interpretive understanding.

2 Background

Multimodal humor comprises image–text artifacts whose meaning extends beyond literal depiction through visual–text coordination, non-literal mechanisms, and shared social knowledge. Its study spans static and sequential representations and requires capabilities progressing from recognition to interpretation and reasoning, then generation.

  • Definition: Multimodal humor intentionally conveys meaning beyond literal depiction through coordinated visual content and associated text in memes, cartoons, comic strips, and satirical images.
  • Interpretive challenge: Creative-media understanding requires perceptual recognition plus inference of implicit intent, shared social knowledge, and meaning distributed across visual–text interactions or comic panels.
  • Representation forms: StaVT uses a single image optionally paired with short text, whereas SeqVN uses an ordered multi-panel sequence requiring entity, event, causal, and cross-panel integration.
  • Expressive mechanisms: Recurring expressive mechanisms include incongruity, analogy and conceptual mapping, exaggeration and hyperbole, and narrative structure involving setup, payoff, and causal progression.
  • Core challenge: AI systems struggle because humor depends on latent metaphors, violated expectations, cultural references, implicit norms, and communicative intent rather than directly observable evidence.

3 Task Hierarchy: Recognition, Interpretation, and Generation

The survey organizes multimodal humor research into recognition, interpretation, and generation, pairing each capability with its most direct evaluation signal. This hierarchy clarifies that gains in label prediction do not automatically transfer to explanation or generation.

  • Task Hierarchy: The capability hierarchy comprises recognition, interpretation, and generation, with each level paired to its most direct evaluation signal.The distinction is useful even though recent benchmarks often mix levels and the organization is not purely chronological.
  • Recognition: Recognition tasks detect humorous or sarcastic effects and may extract roles or estimate intensity using classification, AUROC, F1, accuracy, or human-rating correlation.These tasks require reliable visual-text alignment but can reward recurrent surface patterns without genuine understanding.
  • Interpretation: Interpretation tasks explain why an artifact is humorous, distinguishing cue description from mechanism-grounded reasoning about conflicts, analogies, targets, or narrative steps.Evaluation should combine language quality with interpretive faithfulness, including whether models identify relevant cues, name the correct mechanism, and remain evidence-consistent.
  • Generation: Generation produces humor-consistent captions, punchlines, explanations, or comic continuations and is treated as an emerging downstream frontier requiring partial understanding of incongruity, target, tone, and setup.Because valid outputs are diverse, evaluation emphasizes human or rubric-based judgments and faithfulness to the source image, intended target, and rhetorical device.

4 Modeling Paradigms in the Large-Model Era

In the large-model era, multimodal humor modeling has shifted from handcrafted fusion toward alignment-driven representation learning, explicit reasoning, and evidence-grounded generation. These paradigms are complementary, but their integration remains ad hoc and requires joint evaluation across alignment, reasoning, and generation.

  • Recognition-oriented alignment: Recognition-oriented alignment maps visual and textual inputs into shared representations, enabling scalable classification while leaving humor mechanisms implicit.Models may predict labels without recovering violated expectations or social targets.
  • Interpretation-oriented reasoning: Interpretation-oriented methods add natural-language rationales through prompted, supervised, or theory-guided reasoning to explain why an artifact is humorous.Theory-guided decomposition commonly follows setup extraction, conflict detection, and resolution.
  • Interpretation-oriented reasoning: Reasoning methods trade off prompted chain-of-thought’s zero cost and fragility against rationale supervision’s strength, annotation expense, and risk of annotator-phrasing overfitting.External retrieval supplies missing social, cultural, or temporal information through evidence selected from an external corpus.
  • Generation and creative control: Generation conditions humorous output on creative controls such as intended mechanism, target concept, and tone, positioning generation as a downstream application and potential diagnostic of understanding.The paper notes that generation quality could probe interpretive competence more strongly than MCQ accuracy if suitable evaluation protocols existed.
  • Cross-paradigm integration: The strongest recent systems combine alignment, reasoning, retrieval, and generation components, but integration remains ad hoc and calls for joint benchmarks on the same artifacts.Open questions concern alignment–reasoning interaction, retrieval timing, and generation’s diagnostic role.

5 Datasets and Benchmarks

The benchmark landscape follows a capability hierarchy from standardized recognition tasks to interpretation, reasoning, and generation. Recognition resources are abundant but classification-oriented, whereas generation benchmarks remain sparse and less mature in evaluation.

  • Capability hierarchy: Benchmarks are organized by whether they test label prediction, explanation, cross-panel inference, or image-grounded generation.This capability hierarchy distinguishes surface classification from deeper understanding and controlled output.
  • Recognition Resources: Recognition benchmarks are the most abundant and standardized, covering humor presence, target type, dialogue act, and intensity in cartoons, image–text pairs, and comics.They support perception and alignment training but provide only indirect evidence of understanding the joke, target, or narrative conflict.
  • Interpretation and Reasoning Resources: Interpretation and reasoning benchmarks use explanations, rationales, multiple-choice questions, and cross-panel inference to test why content is funny and what contradictions or implicit targets drive it.Examples include Do Androids Laugh at Electric Sheep?, V-FLUTE, PixelHumor, and the YesBut series.
  • Generation Resources: Generation benchmarks are fewer and often repurpose understanding data, covering humorous captioning, joke completion, meme rewriting, and structured continuation.MEMECAP and OxfordTVG-HIC focus on caption generation, while Oogiri-GO and XMeCap extend toward completion, rewriting, and continuation.
  • Generation Resources: Generation benchmarks diagnose whether models can turn image interpretations into controlled novel outputs, but their evaluation protocols remain less mature than recognition and explanation evaluations.The survey characterizes these resources as sparse and their evaluation as still developing.

6 Cross-Benchmark Empirical Analysis

The cross-benchmark analysis evaluates recent MLLMs under each benchmark’s original protocol across seven humor-understanding benchmarks and twelve task settings. Recognition is generally stronger than interpretation and reasoning, but performance varies substantially by capability, model, and task.

  • Evaluation protocol: The study evaluates a broader set of recent MLLMs using original benchmark prompts and official splits when available, supporting within-benchmark comparisons and descriptive cross-benchmark analysis.Because prompts differ by benchmark, the results do not provide a strictly controlled comparison of task difficulty.
  • Benchmark organization: The analysis covers seven humor-understanding benchmarks and twelve task settings, grouping ExHVV, DarkHumor, HumorDB, and MangaUB under recognition and YesBut-v2, NYCC, and MemeQA under interpretation and reasoning.Recognition tasks identify roles, humor categories, or visual and structural elements, whereas interpretation tasks infer morals or select titles and captions.
  • Recognition: GPT-4o achieves 97.60% on background recognition, 99.21% on panel localization, 64.30% on next-panel inference, and 95.80% on onomatopoeia-scene recognition.These are the best scores on four of the five MangaUB subtasks, while DarkHumor remains difficult, with a best score of 66.43% for Qwen3-VL-8B-Thinking.
  • Interpretation and reasoning: Interpretation and reasoning show larger and more consistent gaps from human performance than recognition-oriented tasks.On YesBut-v2, Qwen3.5-27B reaches 84.72% for moral selection and 83.28% for title selection, versus human scores of 91.30% and 97.50%.
  • Model comparison: No single model dominates all humor capabilities: GPT-4o leads NYCC and most MangaUB subtasks, while Qwen3.5-27B leads both YesBut-v2 subtasks and HumorDB.Qwen3.5-35B-A3B achieves the highest ExHVV and MemeQA scores, whereas Qwen3-VL-8B-Thinking performs best on DarkHumor.
  • Reasoning variants: Reasoning variants help selectively: Qwen3-VL-8B-Thinking improves DarkHumor detection from 49.43% to 66.43% and ExHVV from 76.24% to 78.64%, but decreases several interpretation-task scores.The reported decreases include YesBut-v2 Moral, NYCC, and MemeQA, indicating that deliberative reasoning does not consistently improve humor understanding.

7 Challenges and Future Directions

Multimodal humor understanding remains limited by a gap between literal scene recognition and socio-cultural, rhetorical, and value-laden interpretation. Future progress requires open-ended evaluation, broader cultural and temporal coverage, explicit narrative and intent reasoning, integrated capability training, and context-sensitive safety and provenance practices.

  • Evaluation: Current MCQ and binary benchmarks inflate discriminative accuracy through shortcut learning and poorly capture humor’s inherently open-ended interpretation.Rubric-based generative evaluation can assess incongruity detection, target identification, and cultural grounding while penalizing unsupported reasoning.
  • Knowledge and culture: MLLMs poorly reason about shared background knowledge, implicit norms, audience beliefs, and whose perspective or belief a humorous artifact expresses.Fixed knowledge cutoffs and predominantly English and Western training data create blind spots for ephemeral trends, non-Western symbols, and humor conventions.
  • Knowledge and culture: Future benchmarks should be multicultural and per-culture, annotate required background knowledge, and pair humor-aware retrieval with refreshed norms, discourse, timelines, and factual context.Existing multilingual and culture-specific datasets broaden coverage, while retrieval has improved hateful-meme detection and contextualized meme explanation but remains limited relative to global humor traditions.
  • Narrative reasoning: Sequential multi-panel humor exposes the widest model–human gap because systems must preserve entity identity, construct setup expectations, and localize narrative violations across panels.Panel-aware architectures, setup–punchline objectives, and diagnostic tests such as next-panel prediction and swapped-panel detection target these skills beyond aggregate accuracy.
  • Capability integration and safety: Recognition, interpretation, and generation are deeply intertwined, while incongruity, exaggeration, and irony make harmful humor difficult to detect with isolated pipelines.Joint capability training and self-consistency checks could test transfer and reveal mismatches between humor ratings and explanations.
  • Capability integration and safety: Humor safety requires rhetoric-aware moderation conditioned on whether content critiques or promotes harm, humor-specific red-teaming, and provenance records for origin, licensing, and consent.These practices move beyond keyword filtering and support auditability of datasets and models.

8 Conclusion

Multimodal visual humor remains challenging because its meaning extends beyond surface perception, requiring socially and rhetorically grounded reasoning for robust understanding and generation. The survey organizes creative understanding from recognition through interpretation to generation and identifies future directions for meaningful, reliable, and responsible engagement with human-created media.

  • Multimodal visual humor challenges AI systems because its meaning extends beyond surface perception.
  • Robust humor understanding and generation require socially and rhetorically grounded reasoning.
  • The survey uses a capability-centric task hierarchy spanning recognition, interpretation, and generation, while outlining meaningful, reliable, and responsible engagement with human-created media.

Limitations

The survey is limited by Western, internet-centric coverage that underrepresents other cultural traditions and non-mainstream media, as well as by observations that may not generalize to rapidly evolving multimodal models.

  • Coverage: Existing research and datasets focus predominantly on Western, internet-centric visual humor, leaving many cultural traditions and non-mainstream media underexplored.Examples include memes and cartoons.
  • Generalizability: Rapidly evolving multimodal models mean some survey observations may not fully generalize to future architectures or training paradigms.The limitation concerns changes in both model architectures and training approaches.

Ethical considerations

Ethical challenges in multimodal humor understanding arise because creative media can convey harmful or sensitive content implicitly, while dataset imbalance and unresolved ownership issues threaten fairness, robustness, and responsible use.

  • Ethical considerations: Implicit humor, irony, and symbolism can cause harmful-content misinterpretation, bias amplification, or over-censorship.These risks are distinct to interpreting creative expression, where meaning may not be explicit.
  • Ethical considerations: Dataset bias and cultural imbalance threaten fairness and robustness, particularly for marginalized communities.The passage identifies these concerns as additional ethical risks in creative-media datasets.
  • Ethical considerations: Creative datasets also raise unresolved copyright and ownership concerns.The passage flags these concerns especially as models increase their use of creative data.

A Evaluation Details · B Overview of Datasets and Benchmarks

The survey evaluates models zero-shot while preserving each benchmark’s original task and evaluation protocol, and catalogs datasets and benchmarks by the paper’s capability hierarchy. Results are intended for cross-benchmark comparison with protocol differences kept explicit.

  • A Evaluation Details: Models are evaluated in a zero-shot setting across benchmarks.The benchmark task definitions and evaluation procedures are retained.
  • A Evaluation Details: Each benchmark retains its reported task definition, prompt format, answer format, evaluation split, and scoring procedure.This preserves benchmark-specific evaluation conventions rather than imposing a single protocol.
  • A Evaluation Details: Except for GPT-4o, evaluated models use publicly available Hugging Face checkpoints.GPT-4o is evaluated through the fixed API version gpt-4o-2024-05-13.
  • A Evaluation Details: Each model is queried once per question, with do_sample=true used for open-source models.Other generation and evaluation settings follow the corresponding benchmark papers.
  • A Evaluation Details: Official evaluation splits are used when available; otherwise, the survey follows each benchmark paper’s split or sampling procedure.Prompts and evaluation protocols differ across benchmarks by design.
  • B Overview of Datasets and Benchmarks: The datasets and benchmarks are presented in a comprehensive tabular overview organized by the paper’s capability hierarchy.The overview covers the datasets and benchmarks studied in the survey.

C Task Specific Models Before MLLM Era

Before multimodal large language models, visual humor understanding relied on task-specific discriminative architectures that progressively improved multimodal fusion, structured representations, and knowledge injection. These gains remained task-dependent, with limited scalability and cross-domain generalization, motivating a shift toward multimodal foundation-model paradigms.

  • C Task Specific Models Before MLLM Era: Pre-MLLM systems were primarily task-specific discriminative architectures designed for sarcasm, humor, or metaphor detection.
  • C Task Specific Models Before MLLM Era: Early methods improved heterogeneous-feature fusion through separate encoders, concatenation, and hierarchical integration of text, images, and attributes.Schifanella et al. (2016) used separate encoders and concatenation, while Cai et al. (2019) modeled cross-modal interactions hierarchically.
  • C Task Specific Models Before MLLM Era: Later models addressed cross-modal incongruity and implicit intent using contrastive attention, humor knowledge, object-level representations, cross-modal graphs, and external commonsense knowledge.These mechanisms supported finer discrepancy detection, localized visual–textual conflict reasoning, richer contextual embeddings, and hierarchical congruity modeling.
  • C Task Specific Models Before MLLM Era: Despite progress in fusion, structured representations, and knowledge injection, non-MLLM gains remained task-dependent, with limited scalability and cross-domain generalization.These limitations motivated the later shift toward more general multimodal foundation-model paradigms.
Loading 2607.19011v1…