Source-linked AI summary
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries
Shudong Liu, Yiqiao Jin, Cheng Li, Derek F. Wong, Qingsong Wen, Lichao Sun, Haipeng Chen, Xing Xie, Jindong Wang
TL;DR
VLMs remain limited in cultural understanding because their training data are biased toward English and Western contexts, motivating broader multimodal evaluation and improvement. The paper constructs CultureVerse, evaluates 16 models, and fine-tunes CultureVLM; results show regional disparities and improved cultural perception with cross-cultural generalization without significant loss of general capabilities.
Problem
VLMs struggle with culturally specific understanding because training data underrepresent diverse cultures and are biased toward English and Western contexts.
Method
The paper constructs the CultureVerse multimodal benchmark and fine-tunes a series of CultureVLM models on it.
Results
Evaluations reveal stronger cultural understanding for the Americas than for Asia and Africa, while fine-tuning enhances cultural perception and preserves general capabilities.
Takeaways & Limitations
Culturally diverse multimodal training data can support improved cultural understanding and generalization across cultures, continents, concepts, and datasets.
Takeaways & Limitations
CultureVLM uses languages as proxies for cultural boundaries, fine-tuning covers only the current model set, and CultureVerse contains only multiple-choice questions.
Abstract
from arXiv · showhide
Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper, we construct CultureVerse, a large-scale multimodal benchmark covering 19, 682 cultural concepts, 188 countries/regions, 15 cultural concepts, and 3 question types, with the aim of characterizing and improving VLMs' multicultural understanding capabilities. Then, we propose CultureVLM, a series of VLMs fine-tuned on our dataset to achieve significant performance improvement in cultural understanding. Our evaluation of 16 models reveals significant disparities, with a stronger performance in Western concepts and weaker results in African and Asian contexts. Fine-tuning on our CultureVerse enhances cultural perception, demonstrating cross-cultural, cross-continent, and cross-dataset generalization without sacrificing performance on models' general VLM benchmarks. We further present insights on cultural generalization and forgetting. We hope that this work could lay the foundation for more equitable and culturally aware multimodal AI systems.
1 Introduction
The paper addresses cultural limitations in VLMs by introducing CultureVerse for broad benchmarking and CultureVLM for targeted improvement. Evaluations reveal regional disparities, while fine-tuning improves cultural understanding, generalization, and cultural coverage without significantly reducing general capabilities.
- Motivation: VLMs struggle with cultural understanding because training data underrepresent culturally specific content and overrepresent English-language, Western contexts.These limitations particularly affect understanding of culturally significant symbols and cultures in the global south.
- Findings: All evaluated VLMs show regional disparities, performing best on the Americas and weakest on Asia and Africa.Europe and Oceania fall between these regional extremes.
- Findings: Fine-tuning enhances cultural perception, narrows regional and category gaps, and does not significantly compromise general capabilities.The CultureVLM series is built through a flexible, cost-effective data process and dataset fine-tuning.
- Findings: Cultural understanding generally increases with model size, larger fine-tuning datasets yield larger improvements, and fine-tuning generalizes across cultures, concepts, continents, and datasets.The size relationship is not absolute: Llama 3.2-11B performs comparably to Qwen 2-72B.
- Contributions: CultureVerse contains 19,682 cultural concepts and 228,053 samples spanning 3 tasks, 188 countries or regions, and 15 cultural concepts.Its test set includes 11,085 recognized concepts and 31,382 corresponding samples.
- Contributions: The authors evaluate cultural concepts across 16 open-source and proprietary VLMs of varying scales.The benchmark is intended to characterize multicultural understanding across models and cultural contexts.
2 Related Work
Prior work documents cultural bias in language and vision-language models, but VLM cultural research remains limited by scarce multimodal data and incomplete representation. Existing datasets often lack scale, regional coverage, or cultural relevance, while culturally aware VLM development remains underdeveloped.
- Cultural Bias in LLMs and VLMs: Research on LLMs reports alignment with dominant U.S. or Western cultural perspectives and limited representation of other cultures.Reported gaps include proverb reasoning, translation, and responses to English prompts.
- Cultural Bias in LLMs and VLMs: VLM cultural-bias research remains preliminary, with many efforts relying on manual collection and datasets such as MaRVL and CVQA.These efforts seek broader language and cultural coverage through image hierarchies or multilingual visual question answering.
- Datasets and Models for Cultural Understanding: Existing multimodal cultural datasets often have limited size, insufficient regional or national representation, and weak cultural relevance.The related-work discussion identifies the lack of significant model-level efforts to develop culturally aware VLMs.
3 CultureVerse: A Scalable Benchmark for VLM Cultural Understanding
CultureVerse is a scalable multimodal benchmark designed to address limited and biased cultural coverage in VLM datasets. It combines automated collection with expert quality assurance across thousands of concepts, countries, and question types.
- Construction Pipeline: The dataset uses a scalable pipeline combining automated web crawling for diversity and scalability with expert human annotation for reliability.The pipeline includes cultural concept collection, question-answer generation, and quality assurance.
- Question Types: Three VQA tasks assess cultural understanding through image recognition, cultural knowledge, and culturally situated scene reasoning.Recognition identifies concepts in images, while scene reasoning evaluates contextually appropriate responses rather than factual recall alone.
- Quality Assurance: Human annotators check concept-image alignment, image quality, and whether generated questions and answers are clear, logical, and uniquely answerable.The process also removes redundancy and resolves ambiguities in questions and answer options.
- Quality Assurance: Over 98% of evaluation samples were correctly annotated by the automated process, while remaining erroneous or challenging samples were filtered out.Human annotations were used for evaluation data, whereas the larger training set used the automated annotation pipeline.
- Scalability: Compared with manually constructed datasets, CultureVerse expands coverage to tens of thousands of concepts across approximately two hundred countries or regions.Its construction process also supports adding images and synthesizing open-ended, multiple-choice, and reasoning questions at scale.
4 Analysis of CultureVerse
CultureVerse organizes evaluation across tasks, continents, and cultural-topic categories, while CultureVLM analyzes the effects of fine-tuning on cultural understanding. The section uses these analyses to examine performance patterns and fine-tuning outcomes.
- CultureVerse Distribution: CultureVerse distributes its data across three tasks, five continents, 188 countries, and 15 cultural topics.North and South America are combined into one regional category, and country-level collection naturally produces larger datasets in regions containing more countries.
- CultureVLM Analysis: Figure 4 presents results and analyses for CultureVLM models fine-tuned on CultureVerse.The figure focuses on the effects of dataset-based fine-tuning on cultural understanding.
5 Experiments with CultureVerse
Experiments evaluate CultureVerse across tasks, models, prompts, continents, and benchmarks. CultureVLM improves cultural perception and generalizes across regions while largely preserving general VLM capabilities.
- Experiment Setup: CultureVerse separates training and test images across all countries and regions to assess transferability without data leakage.The test set uses common cultural concepts with manual quality checks, while training includes all cultural concepts.
- Task Difficulty: Scene reasoning tends to outperform image recognition and cultural knowledge questions because it combines visual evidence with contextual background.Recognition depends heavily on diverse image data, while cultural knowledge benefits from richer text-based memory.
- Generalization and Robustness: CultureVLM performs best in-distribution while retaining strong out-of-domain generalization across continents and CultureVerse categories.The figure compares continental transfer and category-based training/evaluation, including CHT, HL, and NELR.
- Model-Level Variation: GPT-4o achieves the best results, while model size alone does not determine cultural understanding.LLaVA-1.5 and Qwen2-VL show similar performance despite size differences, and smaller models can perform strongly with comparable training data.
- Prompting: Stepwise reasoning does not improve cultural recognition and often impairs performance because explanations frequently contain hallucinations.The comparison contrasts direct identification with describing the image and analyzing options before answering.
- Training CultureVLM: Fine-tuning three open models produces substantial cultural-understanding gains, with performance reaching levels comparable to closed-source models.Reducing fine-tuning data lowers performance only minimally, indicating that smaller training subsets can still improve multicultural awareness.
- Case Study: Explanations added during training enhance CultureVLM’s cultural recognition and understanding in the reported case study.Figure 7 compares LLaVA and CultureVLM under fine-tuning and prompt variations.
- Catastrophic Forgetting: Fine-tuning has comparable performance on general VQA benchmarks, indicating that cultural adaptation preserves natural-language understanding and commonsense reasoning.The evaluation includes ScienceQA and TextVQA to test for loss of general VLM capabilities.
6 Conclusion and Limitation
The paper introduces CultureVerse and CultureVLM to measure and improve multicultural VLM understanding. It reports regional disparities, fine-tuning gains, and limitations concerning cultural representation, model coverage, and question format.
- Conclusion: CultureVerse characterizes multicultural VLM understanding, while CultureVLM improves cultural perception and cross-cultural generalization.The paper presents culturally diverse training data as important for improving VLMs.
- Conclusion: VLMs perform more strongly on Western cultural contexts and more weakly in underrepresented regions such as Africa and Asia.The reported disparities span regions and tasks.
- Limitations: Language is used as a proxy for cultural boundaries, although language alone does not capture culture’s full complexity.The authors state that the pipeline remains flexible for incorporating additional languages and cultures.
- Limitations: Fine-tuning experiments cover only the current model set, limiting insight into a wider range of models.The authors identify resource constraints as the reason for this scope boundary.
- Limitations: CultureVerse currently contains only multiple-choice questions, leaving open-ended assessment for future exploration.The authors identify open-ended questions as an additional assessment avenue.
A Details of the CultureVerse Dataset
CultureVerse defines a large collection of cultural concepts and organizes them into a broad evaluation dataset spanning many countries and regions.
- Dataset Scale: The dataset contains 11,085 evaluation samples and 19,682 training samples from 188 countries and regions.The evaluation set and training set are reported separately.
- Dataset Organization: Table 2 presents the dataset’s cultural concepts, overall categories, and descriptions.The table documents how concepts are organized and described.
B.1 Experiment Setup
The experiments evaluate a broad set of open-source and proprietary VLMs and fine-tune selected models with specified training settings and quality-controlled synthetic data.
- Evaluation Models: The evaluation includes open-source models such as LLaVA, LLaMA, Qwen2-VL, InternVL-2, Phi-3-Vision, MiniCPM, GLM-4V, and Pixtral.The setup also includes proprietary GPT-4o and Gemini-1.5-Pro.
- Training Configuration: Fine-tuning uses one training epoch on four A100 80GB GPUs with model-specific batch sizes and learning-rate settings.The reported learning-rate decay uses gamma 0.85, and the LLaMA-3.2 setup uses a 1 × 10^-5 learning rate with no weight decay.
- Data Quality: Training data are synthesized from concepts that passed either GPT-4o or human quality assurance.The authors report that this process improves data accuracy without large-scale human annotation.
B.2 Detailed Main Results
The paper reports zero-shot accuracy across three tasks, five continents, and three cultural categories, with detailed results provided in Table 3.
- Evaluation scope: Detailed results are organized by task, continent, and cultural category.
- Evaluation scope: Table 3 reports zero-shot accuracy for open-source and proprietary models across three tasks, five continents, and three cultural categories.The categories are Cultural Heritage and Traditions, History and Landmarks, and Natural Environment and Local Resources.
B.3 Detailed Fine-tuning Results
The paper presents fine-tuned-model results across tasks, continents, and categories, alongside temperature analyses and human validation of the dataset.
- Fine-tuning results: Detailed fine-tuned-model results are provided in Table 4, with temperature-specific results shown in Figure 9.
- Fine-tuning results: Table 4 reports fine-tuned-model performance across three tasks, five continents, and three cultural categories.The categories are Cultural Heritage and Traditions, History and Landmarks, and Natural Environment and Local Resources.
- Human validation: Ten expert annotators evaluated question-answer correctness, consistency, and relatedness, with each instance labeled by two experts and verified by another.Annotators also manually checked results using search engines, and operations followed local laws and regulations.
C.2 Accuracy of Human Annotation
Human annotation was used to validate CultureVerse, while the dataset construction prompts operationalize cultural concepts, categories, and question generation across varied examples.
- Annotation accuracy: Table 9 reports the accuracy of CultureVerse based on human annotations.
- Examples: CultureVerse examples span image recognition, cultural knowledge, and scene reasoning questions for concepts from Laos, China, and Italy.Examples include Pha That Luang, Peking Opera, and Florence Cathedral.
- Question generation: Question-generation prompts produce image-identification questions with four answer options, using country and concept information as inputs.Additional prompts request detailed introductions and cultural explanations for selected concepts.
- Prompted concept construction: The concept-judgment prompts require concepts to be specific, indivisible, culturally representative, and sufficiently distinctive within a country.Broad categories and widely distributed items are rejected, while specific items such as Peking Duck and Schweinshaxe are accepted.
- Prompted concept construction: The prompts extract cultural elements across categories such as food, plants, animals, landmarks, festivals, artifacts, figures, and clothing.They limit outputs to the most famous elements, with no more than 10 elements and NA for categories without famous elements.