Source-linked AI summary
Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images
Yuechen Jiang, Enze Zhang, Md Mohsinul Kabir, Qianqian Xie, Stavroula Golfomitsou, Konstantinos Arvanitis, Sophia Ananiadou
TL;DR
Structured cultural metadata inference from images remains underexplored compared with perceptual captioning. The paper introduces Appear2Meaning, a cross-cultural benchmark with museum-grounded annotations and LLM-as-Judge evaluation, finding that VLMs capture partial cultural signals but struggle with exact, coherent inference. The study also identifies collection and geographic-proxy biases that constrain interpretation.
Problem
Existing heritage datasets and captioning systems provide limited evidence about inferring structured, culturally grounded metadata across object categories and cultural contexts.
Method
Appear2Meaning curates verified museum objects across four categories and four cultural regions, then evaluates VLM metadata predictions with semantic, attribute-level LLM-as-Judge comparisons.
Results
Models capture partial cultural signals, but exact metadata inference remains challenging, especially for culture, period, and origin, with performance varying across regions and metadata types.
Takeaways & Limitations
Current VLMs remain limited in producing coherent and fully grounded structured cultural metadata from visual input alone.
Takeaways & Limitations
Geographic regions proxy culture, while museum collections reflect institutional and curatorial biases that may be inherited or amplified by evaluated models.
Abstract
from arXiv · showhide
Recent advances in vision-language models (VLMs) have improved image captioning for cultural heritage. However, inferring structured cultural metadata (e.g., creator, origin, period) from visual input remains underexplored. We introduce a multi-category, cross-cultural benchmark for this task and evaluate VLMs using an LLM-as-Judge framework that measures semantic alignment with reference annotations. To assess cultural reasoning, we report exact-match, partial-match, and attribute-level accuracy across cultural regions. Results show that models capture fragmented signals and exhibit substantial performance variation across cultures and metadata types, leading to inconsistent and weakly grounded predictions. These findings highlight the limitations of current VLMs in structured cultural metadata inference beyond visual perception.
1 Introduction
Appear2Meaning addresses the limited study of structured cultural metadata inference from images, beyond perceptual captioning. The benchmark and evaluation reveal partial but inconsistent VLM performance across metadata fields and cultural regions.
- Existing heritage datasets largely emphasize visual, narrative, or emotional descriptions rather than structured cultural metadata such as period, origin, and attribution.
- Appear2Meaning introduces a multi-category, cross-cultural benchmark for structured metadata inference from image-only input.
- 9 SOTA VLMs achieve very low exact-match accuracy, while substantially higher partial-match rates show that models often capture some correct attributes without producing fully consistent predictions.
- Title and creator are easier for models than culture, period, and origin, while accuracy is higher in East Asia and lower in Europe and the Americas.
- The benchmark formalizes heritage understanding as structured prediction and evaluates attribute-level correctness using semantic alignment, classifier-based extraction, and human auditing.
2 Related Work
Prior cultural-heritage AI research remains fragmented across captioning, artifact-specific analysis, and collection-management applications. Most work emphasizes visual or emotional description, with limited structured metadata inference across diverse cultural contexts.
- General image captioning progressed from CNN-RNN systems to transformer-based and large-scale multimodal models that improve transferability and generation quality.
- Cultural-heritage captioning datasets often center on paintings, ceramics, Thangka paintings, or architectural heritage rather than broad cross-cultural coverage.
- Most cultural-heritage approaches emphasize visual descriptions or emotion-aware generation, leaving object-level structured cultural metadata inference limited.
- AI for cultural-heritage practice includes visual tagging, classification, and collection management, but machine-generated tags can diverge from expert-curated metadata.
3 Appear2Meaning Benchmark
Appear2Meaning formulates cultural heritage understanding as structured prediction from images, targeting culturally grounded attributes that are not directly observable. It combines museum-grounded curation with semantic, attribute-level evaluation across object categories and cultural regions.
- The benchmark predicts structured metadata from visual input and evaluates predictions against normalized and raw museum annotations using an LLM-as-Judge.
- The task targets culture, period, origin, and creator, whose labels are grounded in museum schemas and require cultural and historical knowledge.
- Models may generate intermediate captions, but these serve as auxiliary representations rather than the evaluation target.
- The dataset contains 750 objects from four categories and four cultural regions, with 50 artifacts sampled per culture–type combination from verified museum records.
- Curation combines rule-based metadata filtering with two-stage human verification to retain validated culture–type assignments.
- Evaluation reports exact match, partial match, outcome distributions, and attribute-level accuracy across cultural regions using labels of correct, partially correct, and incorrect.
4 Experiment
The experiment evaluates image-only structured cultural metadata inference across models, metrics, attributes, and cultural regions. Results show strong partial recovery but weak exact, coherent, and culturally grounded prediction, with recurring systematic errors.
- Overall Performance: All models produce low exact-match accuracy, while partial-match rates are substantially higher, indicating fragmented rather than fully consistent metadata predictions.Exact-match accuracy remains around 0.01–0.03; Qwen3-VL-Flash reaches a partial-match rate of 0.658.
- Overall Performance: Qwen3-VL-Flash achieves the highest partial-match rate at 0.658, followed by GPT-4.1-mini at 0.609 and Qwen-VL-Max at 0.560.
- Overall Performance: Title and creator are more accurate overall, whereas culture, period, and origin remain challenging across the evaluated models.Qwen3-VL-Flash scores 0.539 on title, while culture, period, and origin reach 0.367, 0.328, and 0.241, respectively.
- Per-Culture Analysis: East Asia yields the strongest regional performance, including a partial-match rate of 0.740 for Qwen3-VL-Flash and culture accuracy up to 0.793.In the Ancient Mediterranean region, partial match remains high but exact match is consistently low, with creator accuracy reaching 0.876 for Claude Haiku 4.5.
- Per-Culture Analysis: Open-weight Qwen models, particularly Qwen3-VL-Flash, lead partial-match performance across regions, while larger models provide slight exact-match gains and closed-source models favor title and creator.
- Error Analysis: Error analysis identifies cross-cultural misattribution, object-type confusion, period compression, and creator memorization without coherent multi-attribute integration.These structured errors explain why models often recover one or two plausible fields but rarely produce a consistent metadata profile.
5 Conclusion and Future Work
The paper concludes that current VLMs can capture partial cultural signals but struggle with exact metadata inference, especially for culture, period, and origin. Future work expands coverage and investigates knowledge-grounded methods for more representative and coherent evaluation.
- The benchmark evaluates image-only VLM inference for culture, period, origin, and creator using an LLM-as-Judge framework.
- Current models capture partial cultural signals, but exact metadata inference remains challenging, especially for culture, period, and origin.
- Future work targets finer-grained distinctions, broader object categories, larger balanced samples, external knowledge sources, and ontology-grounded connections to museum knowledge bases.
6 Ethical Considerations
The benchmark uses openly licensed museum-collection data, but those collections carry historical, institutional, and curatorial biases that models may inherit and amplify.
- Publicly available museum data may transmit historical, institutional, and curatorial biases into model training or evaluation.
- Performance disparities across cultural regions provide evidence that these collection-level biases may be inherited and amplified by models.
A Case Studies and Error Analysis
The error analysis examines model outputs across attributes and cultural contexts, using recurring patterns and representative experiment-log examples to characterize systematic discrepancies against reference metadata.
- Prediction outputs are analyzed across models, attributes, and cultural contexts to identify recurring error patterns.
- The analysis compares visually grounded, internally coherent descriptions with their alignment to normalized reference metadata.
Object ID: 1055_Butter Pat
Butter Pat illustrates recurring cross-cultural misattribution: models produced plausible style-based associations but shifted the American object toward European, Japanese, or Chinese contexts.
- Butter Pat is an American object dated 1885 and created by Union Porcelain Works.
- Multiple models repeatedly misclassified Butter Pat as European porcelain or as Japanese or Chinese objects.Representative predictions included French or European attribution, European 18th-century dating, Japanese Meiji attribution, and Chinese Qing attribution.
- The object lacks highly distinctive visual cues that uniquely identify one cultural context from the image alone.
- Learned associations with material, form, and decorative motifs can favor more frequently represented or visually dominant traditions.
Object ID: 1513_Celery vase
Celery vase demonstrates style-driven cross-cultural confusion: models associated visually familiar ceramic features with European traditions despite the object’s American provenance.
- Celery vase is an American object dated 1849–58 and made by the United States Pottery Company.
- Models variously attributed Celery vase to Dutch Delftware, English Wedgwood, British Staffordshire, or European modernism.
- Marbled surface patterns and vessel forms visually resembled ceramic traditions commonly associated with European production.
- Visual similarity across traditions and learned training-data associations can shift attribution toward better-documented traditions.
- Cultural metadata such as origin and creator often depends on contextual, historical, and institutional knowledge beyond the visual signal.
A.3 Case Study C: Partial Object Recognition without Cultural Attribution
The case studies show that models can recognize broad object forms while systematically missing cultural provenance, historical specificity, or canonical metadata. They also reveal over-specification and evaluation mismatches when plausible interpretations differ from fixed references.
- Partial Object Recognition without Cultural Attribution: Models broadly recognized an andiron as a fireplace-related metal artifact while shifting its cultural metadata toward European contexts.
- Partial Object Recognition without Cultural Attribution: Accurate functional recognition did not imply correct inference of cultural or historical context, which often depends on provenance information absent from visual features.
- Partial Object Recognition without Cultural Attribution: Models identified a classical female figure but missed the Muse Polyhymnia and sometimes inferred an incorrect funerary function.
- Partial Object Recognition without Cultural Attribution: Basin predictions matched Chinese ceramic culture but added unsupported dynastic periods, motifs, and workshop or export contexts.
- Partial Object Recognition without Cultural Attribution: Models generated detailed cultural narratives under strong stylistic cues, reflecting a tendency toward specificity beyond available evidence.
- Partial Object Recognition without Cultural Attribution: Field-by-field evaluation penalized coherent, art-historically grounded predictions that differed from canonical reference annotations.
- Partial Object Recognition without Cultural Attribution: Regional performance differences reflect interacting influences from training priors, dataset composition, visual signal quality, and evaluation constraints.