Source-linked AI summary

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

Xilin Jiang, Qiaolin Wang, Junkai Wu, Xiaomin He, Zhongweiyang Xu, Yinghao Ma, Minshuo Piao, Kaiyi Yang, Xiuwen Zheng, Riki Shimizu, Yicong Chen, Arsalan Firoozi, Gavin Mischler, Sukru Samet Dindar, Richard Antonello, Linyang He, Tsun-An Hsieh, Xulin Fan, Yulun Wu, Yuesheng Ma, Chaitanya Amballa, Weixiong Chen, Jiarui Hai, Ruisi Li, Vishal Choudhari, Cong Han, Yinghao Aaron Li, Adeen Flinker, Mounya Elhilali, Emmanouil Benetos, Mark Hasegawa-Johnson, Romit Roy Choudhury, Nima Mesgarani

arXiv:2601.17645v1cs.SDcs.CLcs.CVcs.MMeess.AS

TL;DR

Current multimodal models lack reliable evidence of understanding audio-visual meaning beyond text-like surface content, especially in cultural contexts. AVMeme Exam addresses this gap with a human-curated benchmark of over one thousand iconic clips, metadata, and questions spanning multiple levels of understanding, then evaluates MLLMs and humans. The results show persistent weaknesses on textless audio and contextual and cultural reasoning, with implications for human-aligned multimodal intelligence.

  • Problem

    Existing multimodal models must understand audio-visual signals and shared cultural meaning beyond what is said or shown on the surface.

  • Method

    AVMeme Exam manually collects over one thousand iconic audio-visual clips with human-annotated metadata and questions covering surface content, context, emotion, usage, and world knowledge.

  • Results

    Evaluation reveals consistent limitations in models’ ability to recognize, interpret, and culturally situate audio-visual clips, especially textless audio and contextual and cultural reasoning.

  • Takeaways & Limitations

    AVMeme Exam provides a benchmark for diagnosing contextual and cultural weaknesses in multimodal models and assessing human-aligned multimodal intelligence.

  • Takeaways & Limitations

    The benchmark’s cultural coverage reflects contributors who are highly educated researchers aged 22–35, and 30-second truncation may omit context essential for real-world understanding.

Abstract

from arXiv · show

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme Exam, a human-curated benchmark of over one thousand iconic Internet sounds and videos spanning speech, songs, music, and sound effects. Each meme is paired with a unique Q&A assessing levels of understanding from surface content to context and emotion to usage and world knowledge, along with metadata such as original year, transcript, summary, and sensitivity. We systematically evaluate state-of-the-art multimodal large language models (MLLMs) alongside human participants using this benchmark. Our results reveal a consistent limitation: current models perform poorly on textless music and sound effects, and struggle to think in context and in culture compared to surface content. These findings highlight a key gap in human-aligned multimodal intelligence and call for models that can perceive contextually and culturally beyond the surface of what they hear and see. Project page: avmemeexam.github.io/public

1 Introduction

AVMeme Exam addresses whether multimodal models can understand audio-visual memes beyond literal content, including context, emotion, usage, and cultural grounding. It evaluates this gap with a human-curated benchmark and finds persistent weaknesses in contextual and cultural understanding.

  • Audio-visual meaning depends on time-varying prosody, melody, pacing, and motion that language cannot fully describe.
  • Understanding what is said or shown is only a starting point; models must also infer intent, usage, recognizability, and cultural significance.
  • AVMeme Exam treats recognizable audio-visual clips as a testbed for literal content, context, emotion, usage, and cultural grounding.
  • The benchmark fills a gap left by prior audio-visual evaluations focused mainly on recognition, captioning, events, alignment, causality, or frame-level content.
  • Over one thousand iconic clips receive human-annotated metadata and human-written questions spanning surface understanding, contextual inference, emotion, humor, usage, and world knowledge.
  • Evaluation of state-of-the-art MLLMs and humans reveals consistent limitations in recognizing, interpreting, and culturally situating audio-visual clips.

2 AVMeme Exam

AVMeme Exam is a human-curated, audio-centric benchmark of 1,032 culturally diverse audiovisual memes, each paired with annotated metadata and a multiple-choice Q&A. Manual verification, text-cheat detection, and visual-cheat handling target genuine multimodal understanding across surface content, context, emotion, humor, and cultural knowledge.

  • Collection: The benchmark is human-collected and annotated by 27 researchers from multiple cultural and regional backgrounds, with sound as the primary medium.Its audio-centric scope includes speech, songs, music, and wordless sound effects, while visual information is complementary.
  • Collection: 1,032 audio-visual memes span more than ten languages and five sound categories, with metadata and a multiple-choice Q&A for each meme.Videos are sourced from YouTube (86.1%) and Bilibili (13.9%), and most exceed one million views.
  • Verification: Nine human verifiers review each entry’s transcript, summary, usage, sensitivity, language, video integrity, and Q&A, returning issues for correction.Question types are assigned by verifiers rather than contributors to support consistent labeling.
  • Cheat Detection: Text-only LLM checks identify text-cheat items, producing meme-full with 1,032 memes and meme-main with 846 after removing all such items.Contributors are prohibited from including names, authors, historical events, or other giveaway keywords, although some items guessed by all three LLMs are retained.
  • Cheat Detection: Visual-cheat clips are manually labeled, and clips whose on-screen text collapses the task into OCR or object detection are evaluated without visual input.Together with text-cheat detection, this is intended to evaluate genuine multimodal understanding without textual or visual shortcuts.
  • Question Design: Seven question types progress from audio and language analysis to contextual inference, emotion, humor and popularity, human conventions, and external knowledge.The first two target immediate clip information, while the middle types examine interpreted meaning, feeling, and humor; category boundaries remain subjective.
  • Question Design: The question-type boundaries are not absolute because interpretation depends on the listener’s experience, such as distinguishing contextual inference from world knowledge.An experienced musician may infer a genre or composer from a first listen, whereas an ordinary listener may need external information.

3 Main Results

AVMeme Exam evaluates MLLMs across surface, contextual, cultural, and multimodal understanding using over one thousand curated audio-visual memes. Models perform best on linguistic content but struggle with deeper cultural reasoning, textless sounds, lesser-known languages, and shortcut-free evaluation.

  • Evaluation setup: 19 state-of-the-art MLLMs were evaluated alongside human participants on AVMeme Exam.The benchmark includes audio-only and audio-visual models, while human evaluation used controlled answer formats aligned with model evaluation.
  • Overall performance: 76.6 audio-only and 80.0 audio–visual average accuracy were achieved by Gemini 3 Pro on meme-main, while Qwen3-Omni reached 55.4 and 57.4.GPT-4o Audio was strongest in audio-only evaluation at 50.1, and audio-visual models consistently outperformed audio-only counterparts.
  • Content versus context and culture: Language Analysis was easiest, while World Knowledge typically fell to 20–55% and deeper tasks showed 15–30% declines from Language Analysis.Contextual Inference, Humor & Popularity, Usage & Application, and World Knowledge were substantially harder than recognizing and parsing spoken language.
  • Textless sounds: Models performed best on speech, worse on music, and worst on sound effects, with leading audio models reaching only 35 to 45% on music and sound effects.Speech and songs reached around 60 to 65%, indicating that linguistic structure remains substantially easier to interpret than non-verbal audio.
  • Lesser-known languages: Leading models frequently dropped to 35–55% on Japanese, Korean, and Persian, while Gemini 3 Pro scored 56.1% in Persian on meme-main.English and Chinese were generally strongest, and visual input improved lesser-known languages and non-verbal sounds only marginally.
  • Evaluation validity: Removing easy questions lowered accuracies by 5–10%, while visual-cheatable conditions inflated accuracy by 40% or more on the affected subset.Providing meme names also increased accuracy by approximately 10%, motivating strict removal of text and visual shortcuts.

4 Conclusion

AVMeme Exam is a multimodal, multilingual, and multicultural benchmark for testing how MLLMs understand audio-visual meaning beyond surface content. Its evaluation reveals persistent weaknesses in textless audio, contextual and cultural reasoning, and alignment with human interpretations.

  • AVMeme Exam evaluates whether MLLMs understand how context, emotion, usage, and shared cultural knowledge construct meaning in audio-visual clips.
  • Models perform substantially worse on textless audio and struggle with contextual and cultural reasoning.
  • Models often fail to align with human interpretations, revealing a gap between current multimodal capabilities and human expectations of intent-aware, culturally grounded understanding.
  • The benchmark is intended to support broader cultural coverage and methods for advancing human-aligned multimodal intelligence.

Limitations

AVMeme Exam has cultural, temporal, technical, task-format, and interpretive limitations. Accordingly, it is a diagnostic and comparative reference benchmark rather than absolute ground truth for human meaning.

  • Cultural coverage reflects highly educated researchers aged 22–35 and does not fully represent global, intergenerational meme culture.
  • Annotations capture interpretations contemporary to the end of 2025 and cannot anticipate future cultural drift.
  • Strict model audio-video length limits require clips to be truncated to 30 seconds, potentially omitting context essential for real-world understanding.
  • Controlled multiple-choice questions on single clips do not cover real-world multi-turn dialogue, personalization, or open-ended scenarios.
  • Meme interpretation is inherently subjective, so majority understandings cannot fully represent equally valid alternative readings.
  • Current multimodal AIs remain weaker at audio-visual understanding than text and at contextual and cultural thinking than surface content.

Ethical Considerations

AVMeme Exam uses human curation and public online clips, with safeguards against sensitive content and participant research conducted under IRB oversight.

  • Contributors personally recognize and use the human-curated clips, which are drawn exclusively from publicly available online videos.
  • The dataset excludes private, paywalled, and confidential content.
  • Political materials and explicit depictions of sexual, violent, hateful, criminal, or drug-related content are prohibited.
  • Human evaluation was conducted under an Institutional Review Board protocol, with online sessions lasting approximately 30 minutes and $15 compensation.

B Human Evaluation Details

The human evaluation established a controlled reference aligned with model testing. Twenty participants evaluated native-language, culturally contextualized, or nonverbal clips individually without search or collaboration, using a staged familiarity and multiple-choice protocol.

  • 20 participants were recruited, including 10 native English speakers and 10 native Chinese speakers aged 18–35.
  • Participants were frequent online-video users who grew up in the U.S. or China and resided in the U.S. during the study.
  • Participants completed evaluations individually without web search, collaboration, or external assistance.
  • Each participant evaluated 37 or 38 clips, producing 750 total human-evaluated samples.
  • Participants evaluated clips in their native language and cultural context, along with clips without spoken language such as music or sound effects.
  • Participants first judged familiarity from the clip alone, then answered the corresponding multiple-choice question under the same format used for model evaluation.

C Supplementary Results

Figures 7 and 8 examine how model performance varies with visual text hints and the original year of meme clips.

  • Figure 7 measures model performance across different levels of text hint in the visual stream.
  • Figure 8 measures model performance against the original year of meme clips.

C.1 Study on On-screen Text

The study evaluates how on-screen textual hints affect multimodal model accuracy. Accuracy is highest when text directly reveals the solution and declines as visual hints weaken.

  • Study design: Nine human verifiers categorize clips by the strongest visual text into solution-revealing text, entity keywords, transcription, or no text.These categories represent a spectrum from the strongest to the mildest visual hints.
  • Results: Accuracy peaks when the visual stream directly reveals the solution and decreases monotonically as text hints weaken.
  • Results: The relationship between weaker text hints and lower accuracy is consistent across models, indicating that on-screen text provides a strong shortcut.
  • Results: For some models, clips without text perform slightly worse than clips with transcription.The passage attributes this pattern to some textless clips being harder to interpret.

C.2 Study on Meme’s Year

The study groups meme clips by original year to examine temporal variation in model accuracy. Performance is highest for clips from 1980–2000 and lower for both older and newer clips.

  • Study design: Original year is defined by the source video's upload year, or by the debut or premiere year for music, songs, and movies.
  • Study design: Clips are grouped into coarse temporal spans, with model accuracy aggregated within each span.
  • Results: Performance peaks on memes originating between 1980 and 2000.
  • Results: Accuracy drops for both older clips from before 1980 and more recent clips from after 2020.
  • Interpretation: The authors hypothesize that uneven Internet coverage may explain why middle-era memes are better represented and repeatedly circulated online.

C.3 Correct and Wrong Thinking Example

The examples contrast correct and incorrect multimodal reasoning on sound effects and music questions. They show that reasoning can connect an audio cue to its game context, but can also overweigh literal or imagined associations.

  • Overview: The supplementary section presents audio-visual examples of both effective and misguided reasoning alongside Gemini 3 Flash and Pro results.
  • Sound-effect example: A game-over sound-effect question asks where a listener would likely be after hearing the sound, with Hospital as the correct answer.
  • Sound-effect example: The model's analysis uses acoustic features such as a sharp impact and metallic ringing to identify the sound's typical use and cultural significance.
  • Sound-effect example: Gemini 3 Pro's reasoning identifies the sound as GTA 5's Wasted effect and connects it to respawning outside the nearest hospital.
  • Music example: For a music question, the benchmark describes an instrumental interlude used for absurd humor, irony, or mocksolemn moments and marks mundane tasks as the correct answer.
  • Music example: Gemini 3 Flash selects the correct mundane-task option, whereas Gemini 3 Pro's low- and high-thinking analyses choose incorrect options based on literal or mismatched musical associations.

D More AVMeme Samples

The samples show how AVMeme Exam pairs diverse audio-visual memes with metadata and questions testing usage, context, humor, language, and world knowledge.

  • D More AVMeme Samples: Each sample combines a transcription or summary with emotion, sensitivity, usage, and a multiple-choice question.The questions target usage, contextual inference, humor and popularity, language analysis, or world knowledge.
  • D More AVMeme Samples: Samples span speech, songs, music, and sound effects across languages, years, and cultural contexts.Examples include English, Chinese, and Japanese-language material, as well as music and operating-system sounds.
  • D More AVMeme Samples: Usage questions ask how memes function in recognizable social or cultural situations, such as cartoon chase scenes or flashy effects.Liszt’s Hungarian Rhapsody No.2 is associated with cartoon chase and slapstick scenes, while “Duang” describes an over-the-top sound or visual effect.
  • D More AVMeme Samples: Contextual-inference questions require interpreting speakers’ motives, emotional reactions, or humorous misidentifications beyond literal transcription.Examples include a girl misidentifying animals, a parody of a god complex, and a sand complaint used as awkward dialogue.
  • D More AVMeme Samples: World-knowledge and popularity questions test cultural associations, including the origin of Leekspin and the Windows error sound.Other examples ask about meme transformations and references from films, anime, and early Internet video culture.
Loading 2601.17645v1…