Source-linked AI summary
NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
Akhila Yerukola, Fabrice Y Harel-Canada, Simran Khanuja, Abhinav Sukumar Rao, Ashima Suvarna, Nanyun Peng, Saadia Gabriel, Maarten Sap
TL;DR
AI systems have limited ability to reason about visually observable behaviors through local social norms, a gap left by text-based and artifact-focused cultural evaluation. The paper introduces human-validated contrastive and explanatory datasets for this task, finding low baseline pair accuracy but substantial relative finetuning gains that remain below 30%.
Problem
Visual norm understanding—reasoning about visually observable behaviors through local social norms—remains underexamined beyond text-only and artifact-focused cultural evaluation.
Method
The paper introduces NormViz-Bench, a human-validated contrastive benchmark, and NormViz-Train, a 64k-image synthetic dataset paired with explanations of behavior and cultural significance.
Results
Current VLMs perform poorly on contrastive visual norm understanding, while finetuning on NormViz-Train significantly improves Qwen3-VL group accuracy but leaves absolute performance below 30%.
Takeaways & Limitations
Visual norm understanding remains an open challenge for equitable, globally competent multimodal AI.
Takeaways & Limitations
The benchmark uses country boundaries as a practical cultural proxy, although broader evaluation through diverse cultural proxies is needed.
Abstract
from arXiv · showhide
AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior work on cultural understanding evaluates AI systems on text-only settings or on visual artifact recognition (e.g. foods, clothing). The ability to reason about visually observable behaviors through local social norms, which we call visual norm understanding, remains unexamined. We introduce NormViz-Bench, a high quality, human-validated benchmark of 3,268 contrastive image pairs (6,536 images) spanning 16 countries. Each pair varies only in the culturally relevant behavior (e.g., objects, attributes, spatial relations, and actions) that alters how each image is interpreted. Each image is labeled as conforming to, violating, or irrelevant to local social norms, and pair-level evaluation requires both images to be correctly classified, thereby preventing reliance on superficial visual shortcuts. Even the strongest VLMs, Gemini 3.0 Flash and Qwen2.5 VL 7B, succeed on only 26.6% and 21.6% of pairs, struggling most with identifying violating and culturally benign visual behaviors. Towards bridging this, we introduce NormViz-Train, a training dataset of 64k images paired with explanations. Though absolute performance remains low (<30%), finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance. Together, NormViz-Bench and NormViz-Train establish visual norm understanding as a challenging and consequential frontier for multimodal AI.
1 Introduction
Visual norm understanding is the ability to connect visually observable behaviors to local social and cultural norms, but prior AI evaluation largely emphasized text or artifact recognition. NormViz-Bench and NormViz-Train address this gap, while results show that current VLMs remain weak despite training gains.
- Motivation: Visual norm understanding connects observable visual behaviors with local social and cultural norms.A Japanese gift containing four koi fish illustrates how a visually small difference can change cultural interpretation.
- Research gap: Prior cultural-understanding evaluations largely used text-only inputs, while multimodal work emphasized foods, clothing, landmarks, and other artifacts.Situated behaviors such as gifting practices are difficult to isolate in web-scale image data.
- Contributions: NormViz-Bench evaluates 3,272 contrastive image pairs across 16 countries, requiring both images in each pair to be correctly classified.Contrastive pairs differ only in the culturally relevant behavior, reducing reliance on superficial visual shortcuts.
- Results: 25.3% and 23.0% group accuracy were achieved by Gemini 3.0 Flash and Qwen2.5 VL-7B, respectively, with failures concentrated on violating, irrelevant, and Global South behaviors.Providing descriptions alongside images or replacing images with descriptions improved performance, but the strongest models remained below 45.3% group accuracy.
- Contributions: NormViz-Train contains 64k images paired with explanations of depicted behavior and cultural significance.The dataset is introduced to improve visual norm understanding through culturally grounded training examples.
- Results: 86% and 36% relative improvements in group accuracy were obtained for Qwen3-VL 4B and 8B after finetuning, although absolute accuracy remained below 30%.The corresponding improvements were 14.9% to 27.7% and 19.6% to 26.7%.
2 Related Work
Cultural AI benchmarks have mainly assessed text-based knowledge or visual artifact recognition, leaving compositional reasoning about situated social behaviors underrepresented. Visual training data is especially difficult to collect because norms are hard to isolate, label, and index in naturally occurring images.
- Existing benchmarks: Text-based cultural benchmarks assess artifacts, facts, values, beliefs, and social norms through language.These evaluations generally provide cultural context explicitly rather than requiring visual inference from behavior.
- Existing benchmarks: Multimodal benchmarks have primarily recognized cultural artifacts such as food, clothing, and landmarks rather than compositional social behavior.Cultural norms require reasoning about how multiple visual elements form a situated behavior.
- Data bottleneck: Culturally grounded visual data is scarce because norms and social practices are difficult to isolate, label, and index in web data.Existing multimodal approaches therefore remain largely artifact-centric or rely on curated real images for cultural safety.
- Data bottleneck: No existing work had used culturally grounded textual knowledge to generate visual training data covering social behaviors and norms systematically absent from natural image corpora.This identifies the modality gap that NormViz-Train is designed to address.
NORMVIZ-BENCH Construction
NormViz-Bench converts culturally grounded norm descriptions into minimally differing image pairs through metadata filtering, prompt construction, multi-source curation, and human validation. The resulting benchmark spans 16 countries and combines synthetic, edited, and retrieved images, while retaining important distributional limitations.
- Pipeline: The benchmark pipeline transforms cultural norm descriptions into contrastive image pairs through five stages ending in human validation.The stages cover norm extraction, metadata and visual-detectability filtering, prompt construction, multi-source curation, and validation.
- Task design: Contrastive evaluation changes only culturally relevant behavior and requires both images to be classified correctly.This pairwise design is intended to expose reliance on superficial features.
- Task design: Each image is classified as Follow, Avoid, or Benign using country context, with group accuracy requiring correct classification of both paired images.The depicted behaviors include gestures, object attributes, spatial arrangements, and actions.
- Norm processing: Norms are sourced from three text benchmarks, restructured into norm–contrast pairs, and filtered for visual detectability before prompt generation.The pipeline excludes norms without concrete visual anchors and retains highly visual norms scoring at least 7/10.
- Image curation: The curation stage combines synthetic and real images to diversify visual style and capture both generated and naturally occurring manifestations of cultural behavior.Sources include text-to-image generation, inpainting, and country-specific image retrieval, followed by automatic quality filtering.
- Validation: 3,272 image pairs were retained from 8,456 candidate pairs after annotations from 212 annotators across 16 countries and rigorous quality filtering.The final benchmark contains 6,544 images spanning 16 countries.
- Benchmark characteristics: 65.2% of pairs use text-to-image generation, 27.6% use inpainting, and 7.2% use natural-image retrieval.Pair types are Follow–Benign (63.7%), Avoid–Benign (21.7%), and Follow–Avoid (14.6%).
- Benchmark characteristics: Country coverage is uneven, with Vietnam contributing 410 pairs, China 344, and the United States 92.The distribution reflects inherited seed-benchmark coverage and the differing visual detectability of nonverbal norms across contexts.
4 Benchmarking: Setup, Results, and Analysis
The benchmark evaluates 21 VLMs on culturally situated visual behaviors using pairwise group accuracy and single-image macro F1. Models perform poorly overall, with failures reflecting visual perception, vision-language alignment, and especially cultural normative reasoning.
- Setup and Metrics: 21 VLMs classify images into culturally meaningful, inappropriate, or benign categories using group accuracy and macro F1.Group accuracy evaluates whether both images in a contrastive pair are correctly classified, while macro F1 evaluates single-image classification.
- Main Results: 45.3% group accuracy is the best result after replacing images with descriptions, leaving cultural normative reasoning as the largest barrier.The oracle ablations improve performance, but even the best description-only setting remains below half of contrastive pairs.
- Main Results: 25.3–19.8% group accuracy and 55.8–44.2 macro F1 characterize the strongest models, showing a sharp gap between pairwise and single-image evaluation.The gap suggests that models may rely on shallow cultural associations rather than distinguishing subtle culturally relevant visual details.
- Three Barriers: +9.0 group accuracy and +7.3 F1 follow from adding image descriptions, while removing images improves averages further to 32.4% and 54.9.These results identify visual perception and vision-language misalignment as bottlenecks, while performance remains limited by cultural normative reasoning.
- Performance by Data Source: Retrieval-based images are hardest, inpainted images follow, and independently generated images are easiest across data sources.Retrieval images likely contain greater visual variability and noise, while inpainting makes minimal pair differences difficult to detect.
- Performance by Social Acceptability Labels: Follow behavior reaches 58–67 F1, whereas Avoid behavior reaches 14–54 F1, making offensive behavior the most difficult label to detect.For benign images, models more reliably identify counterparts of Avoid norms than counterparts of Follow norms.
5 Improving Cultural Visual Reasoning
The paper introduces a synthetic, explanation-rich training pipeline for visual norm understanding and evaluates how finetuning changes performance across models, labels, countries, and cultural dimensions. Finetuning improves results broadly, with especially large gains for smaller models and culturally underrepresented regions, but absolute accuracy remains below 30%.
- NORMVIZ-TRAIN: Norm decontamination removes candidate norms semantically matching benchmark norms before training-data construction.GPT-4o compares candidate norms with each country’s full set of test norms, discarding semantic matches regardless of wording.
- NORMVIZ-TRAIN: 64k training examples pair images with Follow, Avoid, or Benign labels and structured explanations covering visual grounding and cultural appropriateness.The dataset is built from culturally sourced norms, filtered to avoid overlap with the benchmark, and formatted as VQA instances.
- Finetuning results: 86% and 36% group-accuracy gains occur for Qwen3-VL 4B and 8B, respectively, under follow-heavy 64k training.The finetuned 4B model reaches 27.7%, compared with 26.7% for the finetuned 8B model and 25.3% for Gemini 3.0 Flash.
- Finetuning results: Gains are strongest for offensive-content detection, increasing by 145% for Qwen3-VL 4B and 82% for Qwen3-VL 8B.Improvements remain consistent across image sources and social-acceptability labels.
- Cross-country and semantic gains: The largest country-level gains occur in Southern Asia, South-eastern Asia, and Sub-Saharan Africa, reaching 15–25 points.European countries show more modest 8–11-point gains, while Argentina regresses by 11.5 points.
- Cross-country and semantic gains: Significant improvements concentrate in Presence/Absence, Action/Activity, Gestures/Body Language, Clothing/Accessories, and several norm domains.The strongest norm-domain gains are in Rituals, Offerings & Ancestral Customs, Public Behavior & Activities, and Dining Etiquette.
- Limitations and future directions: Absolute group accuracy remains below 30%, and the demonstrated finetuning results cover only the Qwen model family.The discussion suggests richer spatial supervision, contrastive objectives, new architectures, and retrieval as possible directions for further improvement.
6 Conclusion
The paper presents visual norm understanding as an open challenge for globally competent multimodal AI. It combines a human-validated benchmark with synthetic explanation-based training that improves Qwen3-VL performance but does not resolve the task.
- Benchmark and challenge: NORMVIZ-BENCH contains 3,272 human-validated image pairs across 16 countries for evaluating visual norm understanding.VLMs struggle most with offensive behaviors and non-Western cultures, with failures linked to cultural perception, vision-language misalignment, and normative reasoning.
- Training and conclusion: A 64k synthetic training dataset improves Qwen3-VL performance and exceeds Gemini 3.0 Flash in the reported comparison.The conclusion frames visual norm understanding as an open challenge for equitable, globally competent AI.
Ethics Statement
The paper discusses cultural representation, annotation safety, synthetic-data bias, and risks from documenting offensive behaviors. It frames these issues as constraints requiring broader evaluation, validation, and responsible use.
- Cultural representation: Country boundaries serve as a practical cultural proxy, but broader proxies are needed because norms vary within countries, regions, and social groups.The authors note that country-based evaluation cannot capture the full cultural landscape.
- Annotation safety: The annotation study mitigates exposure risks through informed consent, fair compensation, limited demographic collection, and Institutional Review Board supervision.Annotators were compensated at $12/hour.
- Synthetic-data bias: LLM and text-to-image generation may introduce Western-centric or stereotypical defaults into the dataset.The authors mitigate this risk through multi-stage validation by in-country annotators across 16 countries.
- Harm prevention and intended use: Documenting culturally offensive behaviors carries risks of misuse and misrepresentation, so the authors explicitly reject harmful applications.The stated intended use is to support AI systems that are less likely to cause cultural offense or misinterpretations.
B Annotation Framework Details
The annotation framework converts cultural descriptions into visually testable norm contrasts, then structures and filters them for semantic coverage, detectability, and human validation across countries.
- Annotation and validation: Annotations were collected from country- and ethnicity-matched Prolific participants who passed a qualification task.The study was covered by the organization’s institutional review board.
- Dataset construction: The benchmark spans 16 countries and combines text-to-image, inpainting, and natural-image-retrieval pairs.Human validation yields Follow-Benign, Avoid-Benign, and Follow-Avoid pair types.
- Norm sourcing: The source pool contains 7,463 norms from CulturalBench, NormAd, and SafeWorld, with multi-behavior descriptions split into individual norms.The source benchmarks cover 16 specified countries.
- Metadata and detectability: GPT-4o extracts key items, concrete objects, visual contrasts, and other metadata from norm–contrast pairs to assess visual detectability and social acceptability.Pair-based extraction is reported as better than using norms alone because it makes the contrast explicit.
- Metadata and detectability: Visual detectability is scored from 0–10 using clarity, adherence-versus-violation distinguishability, and required contextual information.Scores map to Low Visibility at 0–3, Moderately Visual at 4–7, and Highly Visual at 8–10.
- Prompt construction: For each Follow or Avoid norm, GPT-4o creates behavior descriptions, culturally neutral counterparts, and inpainting instructions across multiple prompt variations.The resulting contrastive framing preserves the norm’s meaning while changing the depicted behavior.
D.4 Stage 4: Multi-Source Image Generation
The framework generates contrastive images through text-to-image synthesis, inpainting, and country-specific real-image retrieval, followed by automated quality filtering.
- Image generation: Imagen 3.0 generates two images per prompt variation for Follow/Avoid behaviors and their benign counterparts.The generation covers both the socially relevant behavior and its neutral comparison.
- Image generation: Gemini 2.0 Flash applies inpainting instructions to generated Follow/Avoid images to create coherent benign counterparts.The edits are intended to preserve image coherence while changing targeted visual details.
- Real-image retrieval: DataComp retrieval searches country-specific subsets using CLIP ViT-bigG-14 embeddings and cosine similarity, returning up to 20 candidates per prompt.The subsets are assigned using country-code top-level domains in image URLs.
- Quality filtering: Quality filtering first retains images with VQAScore above 0.7, then applies a GPT-4o filter for image–prompt alignment.The threshold was selected after qualitative evaluation of images and associated scores.
E.1 Main Results
The main evaluation finds substantial variation across countries and model families, with weaker performance concentrated in several Global South and non-Western settings.
- Evaluation design: The evaluation reports 21 VLMs using group accuracy and oracle ablations, with full breakdowns by source and label in Tables 4–6.Group accuracy requires correct classification of both images in a contrastive pair.
- Model scaling: Model size correlates weakly positively with F1 (r = 0.29) and group accuracy (r = 0.26), while scaling varies substantially across families.Qwen 3 improves consistently from 4B to 8B, whereas other families show inconsistent or minimal changes.
- Country-level results: Kenya and India have 17 of 21 models underperforming their overall accuracy, followed by Indonesia with 16 of 21.China, Peru, the United States, and Vietnam each have 15 of 21 underperforming models.
- Training-data preparation: Norm decontamination removes semantic matches between CANDLE candidates and benchmark test norms before constructing NORMVIZ-TRAIN.GPT-4o compares each candidate against the full set of country-specific test norms.
- Training-data preparation: After visual-detectability filtering, the training pipeline retains 15,331 highly visual cultural descriptors tagged as Follow or Avoid.The retained descriptors are used for subsequent synthetic image generation.
F.2 Experimental Setup
The finetuning study varies label distributions and data scale for Qwen3-VL models, while noting that the experiments remain confined to the Qwen family.
- Experimental setup: Training uses balanced, follow-heavy, and avoid-heavy label distributions with 32k and 64k example scales for Qwen3-VL 4B and 8B.The datasets contain 32,144 and 64,288 selected examples.
- Scope: The experiments are limited to the Qwen model family, although Qwen2.5-VL and Qwen3-VL differ in architecture and vision-encoder training data.The authors report consistent gains across the evaluated variants.
- Results: Finetuning consistently helps across label configurations and data scales, with balanced and follow-heavy training performing competitively.Avoid-heavy training yields relatively lower gains, likely because the test set is skewed toward Follow.
- Results: Scaling from 32k to 64k produces modest gains in most configurations, with group-accuracy changes of +0.1–2.6 and F1 changes of +0.3–2.0.The exception is follow-heavy 64k, which drops by -0.5.
F.4 Probing for training-test leakage
A targeted experiment trained on the benchmark’s test-set norms produced only marginal improvements, suggesting that missing norm-level information alone does not explain current finetuning limits. Broader finetuning results nevertheless show gains across model sizes and countries, while absolute performance remains low.
- Leakage probe: 23.9 group accuracy and F1 52.0 were achieved by Qwen3-VL 4B after training directly on test-set norms, up from 23.2 and 48.5.The experiment used a balanced label distribution and 32k data size.
- Interpretation: The marginal gains suggest that lacking norm-level information is not the bottleneck for the current finetuning setup.The authors instead identify richer training signals, stronger cultural vision-language alignment, or better-suited architectures as possible requirements for meaningful gains.
- Country-level results: Qwen3-VL 4B shows significant improvement for culturally underrepresented non-Western countries and can surpass finetuned Qwen3-VL 8B and Gemini 3.0 Flash.Country-wise performance is reported using group accuracy and single-image F1 score.
- Granular analyses: Finetuned Qwen3-VL 4B improvements are evaluated across cultural norm domains and visual contrast types using group accuracy and single-image classification F1.The corresponding analyses are shown separately for Qwen3-VL 4B and 8B.