Source-linked AI summary

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

Bo Zeng, Linfeng Gao, Peiqin Lin, Yu Zhao, Mingyan Zeng, Yu Tong, Xintong Wang, Linlong Xu, Longyue Wang, Weihua Luo, Qinggang Zhang, Jinsong Su

arXiv:2609.03526v1cs.AI

TL;DR

CulturalMenuBench addresses whether strong food recognition reflects cultural understanding or visual matching by evaluating multimodal models on process-grounded culinary tasks. Across its benchmark, models show a substantial knowledge-application gap: near-ceiling standard MC performance drops sharply on Chinese regional attribution under the same format. The findings indicate that perception does not reliably activate cultural knowledge, motivating explicit connections among perception, procedure, and cultural context.

  • Problem

    Existing evidence does not establish whether multimodal models apply culinary cultural knowledge or merely match visual patterns, despite deployment in dietary recommendation and culinary tourism.

  • Method

    CulturalMenuBench evaluates multimodal models on 4,870 items pairing dish and cooking-process images with ingredients, procedural text, and cultural labels across 10 languages and 18 regions.

  • Results

    Models achieving 94–97% on standard four-way MC tasks reach at most 56% on sub-regional Chinese cuisine classification under the same interface, with diagnostics indicating visual shortcuts.

  • Takeaways & Limitations

    Near-perfect recognition can conceal an inability to apply cultural knowledge, indicating that scaling alone is insufficient without explicitly bridging perception and structured cultural knowledge.

  • Takeaways & Limitations

    Cuisine represents only one facet of cultural understanding, so the findings do not necessarily generalize to history, social customs, or arts.

Abstract

from arXiv · show

Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.

1 Introduction

CulturalMenuBench tests whether multimodal models can apply culinary knowledge beyond visual recognition by combining process evidence with fine-grained cultural labels. Evaluation reveals a large knowledge-application gap, with diagnostic analyses implicating visual shortcuts and procedural evidence supporting the benchmark’s process-grounded tasks.

  • Benchmark: CulturalMenuBench combines final-dish and step-by-step images, ingredients, procedural text, and regional labels across 10 tasks.It covers visual recognition, cross-modal matching, process-grounded reasoning, and cultural classification.
  • Benchmark: 4,870 items span 10 languages and 18 regions, including provincial labels for Chinese cuisines.The benchmark supports both cross-cultural and intra-cultural evaluation.
  • Results: 94–97% on standard four-way MC tasks falls to at most 56% for sub-regional Chinese cuisine classification under the same interface.The reported 40–50 point decline is presented as a knowledge-application gap rather than a response-format effect.
  • Results: Diagnostic errors are consistent with random guessing and accuracy tracks visually distinctive cuisines, indicating reliance on visual shortcuts rather than cultural reasoning.Sichuan reaches 78% accuracy versus 19% for Anhui in the cited comparison.
  • Results: Removing sequential context selectively degrades process tasks by 5.7–11.7pp while non-process tasks remain stable at ≤1.0pp.This double dissociation supports the benchmark’s requirement for procedural evidence.

2 CulturalMenuBench

CulturalMenuBench is constructed from curated, multimodally complete recipes and organized into recognition and process-grounded tasks spanning multiple cultural regions and languages. Its fixed criteria prioritize aligned procedural evidence and cultural coverage while accepting a small, nonuniform sample with limited per-cuisine resolution.

  • Benchmark construction: 4,870 evaluation items cover 10 languages and 18 regions, pairing 587 curated dishes with final images, step-by-step images, ingredient lists, and procedural text.The benchmark is built to support tracing technique sequences and intermediate states that distinguish regional traditions.
  • Menu collection pipeline: The collection pipeline uses URL deduplication, semantic aggregation, multimodal integrity checks, and final curation to retain complete, aligned recipe evidence.Automated filtering removes malformed metadata, followed by human inspection of multimodal presence, alignment, and consistency.
  • Curation boundary: The final selection is filtered coverage under documented criteria rather than uniform sampling, so thin cuisine classes limit the resolution of per-cuisine estimates.An additional quality audit removed 25 items with incorrect cultural attribution before release curation.
  • Dataset composition: The dataset contains 200 Chinese and 387 non-Chinese dishes, with fine-grained provincial labels for Chinese entries and country labels for non-Chinese entries.Chinese dishes represent 34.1% of the benchmark, reflecting richer source coverage and the intended study of inter-provincial diversity.
  • Task construction: The benchmark defines four recognition tasks and six process-grounded tasks, separating direct matching from reasoning over ordered procedures, ingredient alignment, and cross-modal evidence.Process-grounded tasks require models to track state changes across ordered visual frames and compose intermediate inferences.
  • Evaluation formats: Binary evaluation uses yes/no questions with a 50% random baseline, while multiple-choice evaluation uses four candidates with a 25% random baseline and plausible distractors.Tasks are independently evaluated, although they draw from the shared pool of 587 base dishes.

3 Experiments

CulturalMenuBench evaluation reveals that strong conventional recognition does not translate into reliable cultural attribution or process-grounded reasoning. Diagnostics indicate that models rely on visual shortcuts, while sequential-image ablations confirm that process tasks require procedural evidence.

  • 3.1 Setup and Human Baseline: 86.86% human accuracy with pairwise IAA of 0.946 indicates that the benchmark is demanding but has reliable labels.The best model reaches 81.4%, approximately six points below humans.
  • 3.3 Validating Process-Grounded Reasoning: 93.0% on an easier recognition task falls to 57.8% and 58.6% on harder relational and process-grounded tasks, a 35-point chasm.The pattern supports a distinct higher-order reasoning dimension beyond basic recognition.
  • 3.3 Validating Process-Grounded Reasoning: 5.7–11.7pp degradation on practice-image tasks versus ≤1.0pp on non-practice tasks shows that sequential context matters selectively for process reasoning.Overall accuracy decreases by 3.4–7.4 percentage points when sequential visual context is reduced.
  • 3.2 The Knowledge-Application Gap: 94–97% accuracy on standard four-way MC tasks contrasts with at most 56% on sub-regional Chinese cuisine classification under the same interface.The within-format comparison isolates a knowledge-application gap rather than a response-format effect.
  • 3.4 Diagnosing Cultural Reasoning Failure: 38–56% accuracy under the fixed four-way regional-attribution probe, versus up to 97% on standard tasks, shows that models cannot reliably infer cultural origins from images.Human majority-vote accuracy is 69.0%, while the best models trail humans by 12.9 points.
  • 3.4 Diagnosing Cultural Reasoning Failure: Cuisine accuracy follows visual distinctiveness, with Sichuan at 78.4% and Anhui at 18.8%, while confusion ratios remain near the 18.1% random baseline.Stronger models resolve more visually salient samples correctly rather than exhibiting superior cultural reasoning.
  • 3.4 Diagnosing Cultural Reasoning Failure: Dish-name-only classification improves performance by 7.4–17.5pp, indicating that cultural-geographic knowledge is present but inaccessible through visual input.The probe uses identical four-way options while removing images.

4 Related Work

Existing cultural benchmarks largely test factual recall, retrieval, or one-hop multimodal question answering. Food-centered benchmarks add regional or multilingual coverage but generally lack the procedural context and explicit knowledge-versus-application diagnosis provided by CulturalMenuBench.

  • Cultural Competence Benchmarks: Text-based cultural benchmarks probe factual recall and cultural reasoning, while multimodal suites assess cultural understanding through retrieval or one-hop QA.These settings provide limited evidence about process-grounded cultural attribution.
  • Food-centered Cultural Evaluation: FoodieQA covers 14 Chinese regional cuisines but lacks procedural context, whereas IndiFoodVQA adds multi-hop reasoning but is English-only.The comparison highlights complementary gaps in procedural grounding and language coverage.
  • Food-centered Cultural Evaluation: CulturalMenuBench addresses these gaps by coupling process-level metadata with sub-regional labels and a diagnostic protocol that disentangles knowledge possession from application.Its central distinction is whether models can apply cultural knowledge to multimodal evidence.

5 Conclusion

CulturalMenuBench evaluates process-grounded culinary reasoning across diverse languages and regions, revealing that near-ceiling standard performance can coexist with weak cultural attribution. The findings support explicitly bridging perception with structured cultural knowledge.

  • 5 Conclusion: CulturalMenuBench is a multimodal benchmark of 4,870 samples across 10 languages and 18 regions for process-grounded culinary reasoning and cultural knowledge.It evaluates a broad cultural and procedural scope rather than final-dish recognition alone.
  • 5 Conclusion: 94–97% standard MC performance falls to 38–56% on sub-region classification under the same four-way interface, with diagnostics indicating visual shortcuts.The result identifies a knowledge-application gap between recognizing dishes and attributing their cultural origins.
  • 5 Conclusion: The results suggest that scaling alone is insufficient and that future work should explicitly bridge perception and structured cultural knowledge.This conclusion follows from the observed failure to activate cultural knowledge through visual input.

Limitations

The benchmark has several scope and evaluation limitations, including food-specific coverage, single-platform sourcing, translation constraints, reused dish pools, and objective-only formats.

  • Food-specific scope: The findings concern cuisine as one facet of cultural understanding and may not generalize to history, social customs, or arts.The authors identify extending the diagnostic protocol to non-culinary cultural domains as future work.
  • Single-platform sourcing: All recipes originate from MeishiChina, so non-Chinese dishes reflect cross-cultural adaptations rather than natively authored recipes.The authors recommend adding cuisine-native platforms; a 50-dish Italian replication is reported as directional evidence.
  • Translation quality: 96.2% macro-average translation-audit pass rate does not eliminate risks from transliteration, culturally specific terminology, or shared LLM-judge blind spots.Targeted native-speaker validation supplemented the tribunal, while broader native review remains future work.
  • Data reuse across tasks: All 10 tasks reuse the same 587-dish pool, leaving potential memorization effects across task boundaries in joint analyses.Separate dish pools per task would strengthen independence guarantees.
  • Objective-only evaluation: Binary and multiple-choice evaluation excludes subjective interpretation and culturally contingent appropriateness relevant to authentic cross-cultural communication.The authors propose open-ended generation with human or LLM-based judging as a complement.

Ethical Considerations

The dataset uses publicly accessible user-contributed recipes without personal identifiers, but its cultural coverage and diagnostic evidence have important scope boundaries.

  • Data and evaluation: The released metadata contains no user names, account identifiers, or other personal information, and human evaluation reports only anonymized aggregate results.Recipes were collected solely for research evaluation, and human evaluation was voluntary and conducted by non-author colleagues.
  • Scope of evidence: The detailed confusion and saliency analyses are grounded in Chinese regional cuisines, the only subset with fine-grained sub-regional labels.The authors present broader multilingual and Italian-pilot evidence as suggestive, while noting that broader validation remains necessary.

B Native-Data Replication beyond Chinese Cuisine

A native-Italian pilot tests whether the recognition–region gap extends beyond Chinese-platform adaptations, finding the same directional pattern across four models.

  • Replication design: The pilot uses 50 native Italian dishes from Giallozafferano with regional labels from its “ricette regionali” taxonomy.Dish recognition and region attribution were rerun as four-way multiple-choice tasks with a 25% random baseline.
  • Results: The recognition–region gap appears across all four evaluated models in the native-Italian replication.The table caption identifies the replication as N = 50 and four-way MC.
  • Results: 16–38pp gap reproduces in direction for every model, with recognition at or above 92%.Because the pilot covers only 50 dishes, it provides directional rather than comprehensive Italian regional-cuisine evidence.
  • Diagnostic context: The diagnostic table evaluates cultural-reasoning errors across 11 model/configuration conditions and compares intra-family confusion against an 18.1% random baseline.The supplied table caption indicates that no condition shows meaningful above-chance evidence.

D Additional Analyses

Additional analyses show that multilingual gaps persist across languages, process tasks require sequential evidence, and translation quality varies systematically across fields and cuisines.

  • Cross-Lingual Consistency: 12–26pp MC-to-binary gap persists across all ten languages, despite higher performance in French and Italian than Korean and Thai.French reaches 83.4% and Italian 84.8%, while Korean and Thai each reach 73.2%.
  • Response Bias: Balanced four-way labels generally produce balanced model outputs, while binary responses show a 58–65% rejection bias despite balanced ground truth.GPT-5 Nano is an exception, preferring option A at 29.1% with p < 0.0001.
  • Disjoint-Subset and Bootstrap Analysis: Every disjoint subset and all 1,000 bootstrap samples retain a positive within-format gap across the six models evaluated on both layers.No subset gap falls below 0.30, and the proportion of bootstrap samples with a positive gap is 1.000 for every model.
  • Reasoning Chain Verification: Reasoning chains show visual observation in 83–100%, process reasoning in 87–93%, and ingredient–procedure cross-referencing in 50–93% of sampled cases.The analysis uses Claude 4.6 and Qwen3.6-Plus on 30 practice-image samples.
  • Translation Quality: 96.20% macro-average pass rate was achieved in the pre-curation translation audit using majority approval from three LLM judges.The audit covered 412 candidate non-Chinese dishes and 1,236 fields.
  • Translation Quality: Dish names have the lowest translation pass rate at 90.0%, versus 98.5% for procedures and 99.3% for ingredient lists.Indian and British dish names have lower pass rates of 80% and 83%, respectively.
  • Native-Speaker Validation: Native-speaker audits found no outright mistranslations but identified 2–3 under-specification cases per language across Hindi, British English, and Vietnamese.The audited items covered dish names, ingredients, and cooking steps.

J Error Analysis by Failure Type

Error analysis separates visual confusion, process reasoning failure, ingredient recognition failure, and language–cuisine effects across model capabilities.

  • Visual Confusion: 49.6% of GPT-5.1 errors are visual confusions, compared with 20.7% for GPT-5 Nano.These errors involve visually similar dishes; Indian cuisine contributes 24.7% of shared hard cases despite representing 11.8% of the data.
  • Process Reasoning Failure: 51.4% of GPT-5 Nano errors are process reasoning failures, compared with 22.6% for GPT-5.1.On practice_image→dish_image, Nano has a 46.0% error rate versus 1.0% for GPT-5.1.
  • Ingredient Recognition Failure: Ingredient recognition failure constitutes approximately 28% of errors for both models, while Nano has 6.3× more such errors in absolute count.The passage interprets this pattern as ingredient-identification difficulty scaling with overall capability.
  • Language and Cuisine Effects: Vietnamese has the highest error rate for both models, at 13.7% for GPT-5.1 and 36.8% for Nano.Vietnamese and Hindi account for 28.7% of GPT-5.1 errors despite representing 13.4% of the data, combining low-resource languages with visually homogeneous cuisines.
  • Language and Cuisine Effects: 74% of GPT-5.1’s 115 errors are shared with Nano, indicating that residual errors cluster on genuinely ambiguous items.The shared-error pattern complements the model-specific differences in visual and process failure modes.

K Experimental Setup Details

The evaluation covers multiple model configurations and controlled four-way prompts, with standardized inference settings and distractors designed to prevent surface-level elimination.

  • Model Configurations: Twelve models are evaluated in Layers 1–2, while Layer 3 uses 11 model/configuration conditions spanning 10 unique checkpoints.Model provenance is documented through corresponding system cards and technical reports.
  • Model Configurations: All OpenAI models use the GlobalStandard variant, while Kimi K2.6 serves Layer 3 and K2.5 serves Layers 1–2.Open-weight models use BF16 precision with a maximum sequence length of 8,192 tokens.
  • Inference Parameters: Layer 1–2 evaluation uses greedy decoding with temperature=0 and a 64-token output limit.Layer 3 uses GPT reasoning_effort=low, Claude Opus 4.6 extended thinking with a 1,024-token budget, and temperature=0 for the text-only probe.
  • Prompt Templates: Four-way sub-region tasks use image prompts naming four cuisine options, while text-only probes replace the image with a dish name.The prompts require only the answer letter for these standard conditions.
  • Prompt Templates: CoT prompts require models to identify the dish, analyze ingredients, seasonings, and techniques, then infer its regional cuisine.Dish names are cleaned to remove cuisine-revealing keywords, with 20 of 189 names requiring cleaning.
  • Prompt Templates: Text+Image prompts pair each dish image with its name to test whether textual anchoring bridges visual-to-cultural reasoning.All four-way tasks sample distractors from the same cultural family when possible, including at least one close regional alternative for sub-region classification.
  • Evaluation Protocol: Answer extraction selects the first or last [A-D] occurrence depending on task type, with empty or unparseable responses retried up to 3–5 times.Requests run concurrently at 4–10 parallel calls per model, depending on API rate limits.

L CoT and Text-Anchor Ablation

The ablation compares step-by-step reasoning and text anchoring against image-only classification, showing that dish names provide substantially more cultural signal than visual input alone.

  • CoT Prompting: CoT prompting asks models to identify dishes, analyze ingredients and techniques, and infer regional cuisine from the same 177 dish-image samples.Models use their strongest available reasoning configuration, including built-in or extended thinking where available.
  • Text+Image Combined: Text+Image prompting supplies the dish name alongside the image to test whether textual anchoring bridges the visual-to-cultural reasoning gap.The condition combines the name and image rather than replacing visual input.
  • Ablation Results: 10–14pp accuracy gains from providing dish names exceed the marginal improvement from CoT prompting in sub-region classification.The strongest text-anchored result, 67.8%, matches text-only performance rather than exceeding it.
  • Ablation Results: 7–18pp accuracy improvements occur when images are removed from the same 189-sample, four-way sub-region classification task.This pattern indicates that cultural-geographic knowledge exists but is not reliably activated through visual input.
  • Illustrative Tasks: Task 1 requires semantic disambiguation across textual, visual, and cultural cues, whereas Task 2 requires categorical equivalence despite plating variation.The examples illustrate multi-step reasoning integrating visual, procedural, and cultural knowledge.
Loading 2609.03526v1…