Source-linked AI summary

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo

arXiv:2608.23065v1cs.CVcs.AIcs.CLcs.IRcs.MM

TL;DR

Existing video-cultural benchmarks often test visible cues without separating symbolic naming, visual recognition, and temporal localization. CMB addresses this gap with a 306-concept Southeast Asian benchmark and a three-stage evaluation, finding low end-to-end success and ability- and modality-specific failures across six VLMs. The benchmark also shows that cultural performance depends on country-specific knowledge and provides a diagnostic framework for attributing failures.

  • Problem

    Existing video-cultural benchmarks often test visible cues and collapse naming, visual recognition, and temporal localization into a single score, limiting diagnosis of cultural understanding.

  • Method

    CMB evaluates 306 expert-curated concepts from seven Southeast Asian countries through naming, video-moment recognition, and free-form localization stages designed to isolate each ability.

  • Results

    Fewer than 30% of concepts are correct across all three stages for the two strongest closed-source models, while failure patterns vary by ability and modality.

  • Takeaways & Limitations

    CMB functions as a diagnostic harness, showing that abilities do not fully cascade and that cultural understanding requires country-specific knowledge.

  • Takeaways & Limitations

    CMB covers seven Southeast Asian countries and five categories, leaving other countries, traditions, and sub-national diversity under-represented.

Abstract

from arXiv · show

Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this into three abilities: naming what a concept symbolizes, visually recognizing it on video, and locating its sub-events in time. Existing video-cultural benchmarks tend to test what is seen, collapsing these three abilities into a single score that hides the bottleneck. We introduce the Cultural Moment Benchmark (CMB): 306 expert-curated concepts from seven countries in Southeast Asia across five categories. We evaluate each concept through three stages, one per ability. Given a description, Stage 1 (S1) selects from four candidate concept names, Stage 2 (S2) selects from four candidate video moments, and Stage 3 (S3) predicts the start and end times of the moment in a video. To keep each stage focused on a distinct ability, we use three design choices: semantic-similarity distractors (S1, S2), unlabeled video moments (S2), and free-form localization on a different example video (S3). Across six vision-language models, failure modes vary by ability and modality. i) Even the strongest closed-source models score below 30% when all three stages must be correct; ii) The three abilities do not fully cascade: naming a concept correctly helps half the models recognize it on video, but recognizing it has little effect on locating the sub-event in time; iii) Audio is complementary, redundant, or distracting depending on the concept, more often distracting in non-Latin-script countries; removing both audio and subtitles hurts Games and Music the most. Our 14-rater human study shows that even Expert raters score below chance on concepts from a neighboring country, indicating that CMB requires country-specific cultural knowledge. CMB acts as a diagnostic harness, attributing failures to a specific ability or modality.

1 Introduction

CMB separates cultural video understanding into naming, visual recognition, and temporal localization, then uses a staged evaluation to diagnose where models fail. Across six VLMs, end-to-end performance remains low and failure patterns vary by ability and modality.

  • A single accuracy number conflates naming, visual recognition, and temporal localization, so CMB scores each ability separately to expose model-specific bottlenecks.
  • CMB evaluates each concept through three stages: naming from four candidates, visual recognition among four video moments, and free-form localization of a sub-event.
  • Six VLMs show ability- and modality-specific failure modes, demonstrating why a single aggregate score is insufficient.
  • Fewer than 30% of concepts are correct across all three stages for Gemini 3.1 Pro and GPT-5.4, while four open-source models score in single digits and InternVL-14B is near zero.
  • CMB contributes a 3-stage × 3-mode framework with per-stage and conditional metrics, an end-to-end Joint score, and failure profiles for six VLMs.

2 Related Work

Prior multicultural benchmarks mostly evaluate cultural understanding from images or video through answer choices and do not isolate distinct cultural abilities. CMB extends this landscape with staged recognition and free-form temporal localization on a different video.

  • Most multicultural benchmarks use a single image and evaluate cultural understanding through textual or visual answer choices, with some also testing cultural reasoning.
  • CMB carries the visual-option design into video while adding naming from the same description and free-form temporal localization on a different video.
  • Recent video benchmarks cover diverse regions, languages, long-form reasoning, and specialized tasks, but use varied designs and modalities.
  • Existing video benchmarks do not decompose performance by ability, use video moments as cultural recognition options, or include free-form localization on a different example video.

3 Cultural Moment Benchmark

CMB combines expert-curated cultural concepts and videos with symbolic descriptions, staged questions, and layered human quality checks. Its construction keeps naming, visual recognition, and temporal localization tied to distinct evidence and contexts.

  • Benchmark construction: CMB uses human visual ground truth while LLMs scaffold textual artifacts across concept curation, stage construction, and quality validation.
  • Benchmark construction: Local Annotators validate concepts and videos, Cultural Annotators curate symbolic descriptions and temporal spans, and Reviewers run outsider and triage checks.
  • Concept curation: The benchmark organizes concepts by country, category, and concept across seven Southeast Asian countries and five categories, retaining candidates through local-annotator consensus.
  • Stage construction: S1 and S2 descriptions encode symbolic meaning rather than visual appearance, with LLM drafts selected and rewritten by Cultural Annotators.
  • Stage construction: S1 uses semantically similar distractors, including same-country concepts and cross-country analogues, to require fine-grained discrimination beyond broad regional priors.
  • Stage construction: S2 presents annotator-confirmed video moments matching the concept options, while S3 localizes a uniquely occurring sub-event in a separate example video.
  • Evaluation inputs: Models receive full video with native audio, subtitles, and on-screen text by default during S3 evaluation.
  • Quality validation: Quality control combines LLM audits, an Outsider Filter based on regional priors, and Reviewer triage of flagged S3 spans; validated concepts pass both checks.

4 Experiments

Experiments show that CMB exposes distinct bottlenecks across models, stages, modalities, and cultural settings rather than a single performance pattern. End-to-end success remains low, naming only partly supports recognition, recognition rarely supports localization, and modality effects depend on category and country.

  • 4.2 Overall Results: Below 30%: Gemini and GPT achieve this Joint score when naming, recognition, and localization must all be correct.The four open-source models score in single digits, with InternVL close to zero.
  • 4.2 Overall Results: Closed-source models consistently outperform open-source models across stages, while within-family rankings are less reliable.Every closed-source versus open-source comparison is significant on all three stages, but ordering within groups does not consistently survive resampling.
  • 4.3 Cascade Analysis: Correct naming raises S2 accuracy by about 18 points for Gemini and GPT, but the benefit shrinks or disappears for most open-source models.InternVL-14B (t) reverses the pattern, with conditional S2 below its unconditional rate.
  • 4.3 Cascade Analysis: Correct S2 recognition rarely improves S3 localization: five models show no measurable benefit, while InternVL’s two-point exception is practically negligible.Localization is evaluated on a fresh clip, and the exception occurs against a very low baseline.
  • 4.4 Modality Dependence Is Concept-Specific: Removing both audio and subtitles most harms Games and Music, where these modalities improve performance for slightly more than 30% of videos.Game commentary, callouts, score updates, and Music alignment provide temporal signals, whereas persistent Dance music offers limited localization cues.
  • 4.4 Modality Dependence Is Concept-Specific: Audio is distracting in 29% of videos from Cambodia, Myanmar, and Thailand versus 14% from Indonesia, Malaysia, the Philippines, and Vietnam.The distracting cases are dominated by continuously playing traditional solo instruments that lack reliable sub-event timing cues.
  • 4.5 Human Evaluation: Expert raters drop by 37 points on S1 and 23 on S2 for neighboring-country concepts, with neighboring-country S1 performance below the 25% chance level.On native concepts, experts outperform non-experts by 19 points on S1 and 7 on S2, but that ordering reverses for neighboring concepts.

5 Conclusion

Cultural video understanding comprises three abilities that current VLMs do not reliably compose: naming concepts, recognizing them visually, and temporally localizing sub-events. CMB separates these failures and identifies model-specific priorities.

  • Cultural video understanding comprises three abilities that current VLMs do not reliably compose.These are naming a concept, recognizing it on video, and locating its sub-events in time.
  • Correct naming helps only some models recognize concepts, while correct recognition rarely helps models localize sub-events temporally.
  • A 14-rater human study indicates that CMB measures country-specific cultural knowledge rather than broad regional cues.Model performance reaches Expert level for naming but only Non-expert level for visual recognition.
  • CMB functions as a diagnostic harness that points to model-specific priorities across cultural knowledge, visual recognition, and temporal localization.The stated priorities differ for InternVL, Qwen, and all models on temporal localization.

Limitations

CMB’s limitations concern cultural coverage, dataset scale, distributional balance, modality-ablation scope, static-bias diagnostics, and source-video reproducibility.

  • CMB covers seven Southeast Asian countries and five categories, leaving Brunei, Laos, Timor-Leste, Singapore, and traditions outside these categories out of scope.
  • 306 concepts and 624 source videos prioritize expert curation over scale, making CMB suitable for evaluation but not VLM training or fine-tuning.Further scaling would require semi-automated pipelines or expanded annotator capacity.
  • Concept counts and S3 span counts are uneven across categories and countries, while minority-language, regional, and diaspora diversity is under-represented.Per-category aggregate scores should therefore be read as averages over uneven concept and span counts.
  • The modality ablation reruns only three of six evaluated models, so its Complementary, Redundant, and Distracting labels may differ for the other models.
  • CMB lacks a single-frame baseline for directly isolating static cues from temporal motion in S2 accuracy.The low S2→S3 transfer provides indirect evidence, but a direct comparison remains future work.
  • CMB uses YouTube URLs and versioned releases, but a removed source video cannot be recovered for downstream users.

Ethical Considerations

CMB addresses cultural-integrity risks through local-expert review, constrained distractor construction, and a release format that avoids redistributing raw videos.

  • CMB releases YouTube identifiers, timestamps, spans, prompts, and metadata rather than raw video bytes, respecting platform terms and uploader copyright.
  • Local experts from all seven countries verify that questions and annotations reflect the cultural semantics of depicted artifacts and rituals.This human-in-the-loop process is intended to reduce misrepresentation and Western-centric interpretation.
  • Each concept receives same-country and different-country distractors, with country, ambiguity, repetition, and similarity constraints governing assignments.Two concepts from the same ambiguity group are never placed in one multiple-choice question.
  • The distractor algorithm fills all slots without country or ambiguity violations, and selected distractors have median Sentence-BERT similarity 0.70 versus 0.60 for non-selected candidates.
  • The released records encode concept, country, category, native-script name, description, stage options, correct indices, and derived modality labels.

A.5 Dataset Analysis

CMB’s dataset analysis describes its country/category composition, question and video durations, S3 span structure, and deliberately difficult semantic distractors.

  • CMB contains 306 concepts, with one S1 and one S2 question per concept, while Set B contains 631 S3 questions across 318 source videos.Some Set B videos contribute multiple sub-event spans.
  • Celebration has 91 concepts and Music 68, whereas Wedding has 11, making the category distribution strongly uneven.
  • S1 and S2 questions average 26.4 words, S3 questions average 23.5 words, and Set A S2 moments have a 20-second median duration.
  • The selected S1/S2 distractors are harder than non-selected candidates, with median Sentence-BERT cosine similarity of 0.70 versus 0.60.
  • The inventory spans seven Southeast Asian countries and five categories, including Celebration, Dance, Game, Music, and Wedding.

B.1 Human Evaluation Design

The human evaluation compares local and non-local cultural knowledge using structured rater assignments, overlap sampling, and matched S1/S2 judgments.

  • Fourteen qualified raters evaluate seven countries, with two Local and two Non-local raters assigned to each country.
  • The fixed +1/+2 rotation gives each country Non-local ratings from two different home countries, with Myanmar as the structural script-cluster exception.
  • Local and Non-local raters share a 20% overlap subset, while the remaining concepts are split into 40% solo assignments per rater.
  • Native and Non-native scores are means of per-concept Local and Non-local ratings computed on the same concept subsets as model results.
  • Per-rater S1 workload ranges from 42–64 multiple-choice questions, with approximately 732 total ratings collected.
  • S2 agreement is positive across all countries while S1 agreement fluctuates around zero, consistent with lower ambiguity from moment-grounded questions.

C.1 More on Implementation Details

The implementation evaluates end-to-end success through a Joint score while using concept-level cluster bootstrapping for comparable uncertainty estimates across models.

  • The Joint score averages the product of correct S1, correct S2, and each concept’s S3 IoU, assigning zero when no S3 prediction is produced.This makes end-to-end success require both naming and recognition before temporal localization contributes.
  • 138 concepts correct on both S1 and S2 with mean S3 IoU 0.32 yield Joint = 14.4%.
  • The main S1/S2/S3 experiments use a single Latin-common query form, making writing-system cluster a post hoc country property.
  • Each frame is resized to 224×224, while S1/S2 use four-way multiple-choice prompts and S3 requests a predicted time span in seconds.
  • Confidence intervals use 10,000-resample cluster bootstraps over concepts, the shared unit because questions within a concept share video and annotator.All six models use the same 306 concepts and 631 Stage-3 pairs, enabling direct cell comparisons.

D.2 Modality Ablation Details

Modality ablations show that audio and subtitles can help, remain redundant, or distract, with distraction differing by writing-system cluster and content category.

  • Three lowest-cost models were rerun on modality ablations, while GPT-5.4 and both InternVL3.5-14B configurations were excluded for budget reasons.The ablation results therefore cover only the three rerun models.
  • Audio’s Distracting-role rate differs by 15 points between writing-system clusters, while subtitles differ by 8.5 points in the same direction.The audio gap is reported as z=3.84 with p<.001; the subtitle gap is smaller.
  • Among 26 non-Latin Distracting concepts, 8 are music and 10 are traditional games, with solo-instrument music prominent in the subset.
  • The proposed explanation that audio is less reliably anchored for underrepresented languages cannot be verified because model training mixtures are undisclosed.
  • Redundant concepts remain the majority across all three ablations and all tested thresholds, ranging from 56.3% to 74.8%.Changing the threshold shifts mass between Redundant and active roles but does not affect the section’s claims.

D.6 Stage-3 result breakdowns

Stage-3 localization is sensitive to source-video duration and sub-event span length, with long-tailed IoU distributions and persistent model-specific bottlenecks.

  • Every model’s S3 mIoU falls sharply beyond videos of about ten minutes, while shorter-video mIoU rises with mean sub-event span length.Closed-source models degrade more gracefully than the others.
  • S3 per-clip IoU is long-tailed, with a heavy mass near zero for every model.
  • CMB compares S3 IoU@0.3 and mIoU across all six models, three modes, categories, and writing-system clusters.
  • CMB asks a culture-specific diagnostic question analogous to general temporal-understanding decomposition: whether models can localize cultural sub-events.
  • Temporal localization remains a persistent bottleneck across all evaluated models, motivating model-specific intervention priorities alongside regional adaptation.

E Benchmark Categorization

The categorization compares cultural video benchmarks by coverage, cascade structure, conditional metrics, temporal-localization format, and human-evaluation design.

  • SCB covers 138 Southeast Asian concepts with a two-stage image cascade, but its second stage is spatial segmentation rather than free-form temporal localization.
  • VideoNorms uses three parallel tasks rather than a headline stage cascade, and its evidence-quality metric is conditional only on correct labels.
  • GIMMICK includes image and video subsets spanning UNESCO cultural heritage, but its six tasks are evaluated independently and only in English.
  • ViMUL-Bench spans 14 languages and countries without Southeast Asian languages, uses parallel question formats, and relies on partial independent human validation.
  • MINERVA-Cultural partially covers Southeast Asia through Indonesia and Thailand, uses single-stage open-ended QA, and reports conditional cross-step success in IEI.
  • VideoVista-CulturalLingo lacks Southeast Asian focus and uses a single-stage MCQ design without cascade or conditional scoring.
Loading 2608.23065v1…