Source-linked AI summary

Visually Grounded Reasoning across Languages and Cultures

Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, Desmond Elliott

arXiv:2109.13238v2cs.CLcs.AIcs.CV

TL;DR

Existing vision-and-language resources largely inherit ImageNet’s English and Western cultural bias, raising questions about their multilingual and multicultural representativeness. The paper introduces a native-speaker-driven protocol and MaRVL across five languages, then finds substantial cross-lingual transfer difficulty compared with English benchmarks. It argues that MaRVL provides a more faithful test of model suitability beyond narrow linguistic and cultural domains.

  • Problem

    ImageNet-derived concepts and images may not adequately represent multiple languages and cultures because their coverage is rooted in English WordNet and biased data sources.

  • Method

    The paper uses native speakers to select concepts and images across five languages and elicits true-or-false statements about image pairs for MaRVL.

  • Results

    Cross-lingual and multilingual baselines on MaRVL sometimes perform just above chance and suffer considerably compared with English datasets.

  • Takeaways & Limitations

    MaRVL may offer a more faithful estimate of state-of-the-art models’ suitability in real-world applications outside narrow linguistic and cultural domains.

  • Takeaways & Limitations

    ImageNet’s concepts are often unfamiliar or absent across languages and can be overly specific relative to basic-level concepts.

Abstract

from arXiv · show

The design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to progress in computer vision, it is mostly derived from lexical databases and image queries in English, resulting in source material with a North American or Western European bias. Therefore, we devise a new protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. In particular, we let the selection of both concepts and images be entirely driven by native speakers, rather than scraping them automatically. Specifically, we focus on a typologically diverse set of languages, namely, Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. On top of the concepts and images obtained through this new protocol, we create a multilingual dataset for {M}ulticultur{a}l {R}easoning over {V}ision and {L}anguage (MaRVL) by eliciting statements from native speaker annotators about pairs of images. The task consists of discriminating whether each grounded statement is true or false. We establish a series of baselines using state-of-the-art models and find that their cross-lingual transfer performance lags dramatically behind supervised performance in English. These results invite us to reassess the robustness and accuracy of current state-of-the-art models beyond a narrow domain, but also open up new exciting challenges for the development of truly multilingual and multicultural systems.

1 Introduction

ImageNet-derived vision-and-language resources inherit English and Western cultural biases, motivating a native-speaker-driven multilingual and multicultural dataset. MaRVL evaluates grounded statements about image pairs and exposes substantial cross-lingual transfer difficulty.

  • ImageNet underpins major vision and multimodal datasets through a hierarchy of concepts derived from English WordNet.Its 1,000-concept ILSVRC subset also supports datasets such as NLVR2 and models such as ResNet.
  • ImageNet’s concepts and images may inadequately represent languages and cultures beyond North America and Europe.The paper links this concern to skewed data origins and content, and to cultural variation in salient concepts and prototypical members.
  • MaRVL selects concepts and images through native speakers across Indonesian, Swahili, Tamil, Turkish, and Mandarin Chinese.Annotators also compare image pairs with native-language descriptions centered on grounded true-or-false judgments.
  • The MaRVL task asks models to determine whether grounded descriptions of image pairs are true or false.It is designed to require information integration across modalities and deep linguistic understanding rather than superficial feature matching.
  • Cross-lingual transfer performance deteriorates considerably compared with the English NLVR2 benchmark.The paper attributes the challenge to the combined domain shift in concepts, images, and language variety.

2 Motivation

ImageNet-derived datasets have limited multilingual and multicultural coverage because their concepts, image retrieval, and cleanup procedures are not designed around global diversity. The section motivates a more comprehensive examination of these biases and their consequences for multimodal reasoning.

  • Dataset scope: ImageNet and related multimodal datasets rely on concepts and images whose ability to represent multiple languages and cultures is questioned.ImageNet’s 1,000-concept subset also underlies datasets such as NLVR2.
  • Concepts: Basic Level and Prototypes: Concept prototypes, basic-level categories, and category boundaries vary across cultures, environments, and individual experience.Different cultures may adopt different basic-level concepts, and perceptive, cognitive, environmental, and cultural factors constrain categorization.
  • Limitations of ImageNet: ImageNet’s original annotation was not intended to ensure universal, human-salient concepts, limiting its suitability for everyday reasoning across languages and cultures.These design choices may be especially important for multimodal systems intended to handle everyday scenarios globally.
  • Limitations of ImageNet: ImageNet concepts are culturally uneven: most synsets occur in 30 or fewer languages, few are universal, and coverage is concentrated in Eurasian languages and macro-areas.The analysis maps synsets to Wikipedia language availability and uses WALS to examine language-family and macro-area coverage.
  • Limitations of ImageNet: ImageNet favors finer-grained English WordNet synsets over the higher-level labels people use, with average depths of 10.61 versus 8.92.The comparison uses 447 concepts and corresponding human object labels; the mismatch may be aggravated in other cultures, as illustrated by KOTO.
  • Sources of Bias: Bias enters through random concept selection, English-centered search retrieval, and cleanup by annotators whose cultural and linguistic representativeness is unknown.The ILSVRC subset may overrepresent non-basic levels; search results do not match real-world demographic distributions, and disagreement may reflect cultural variation rather than annotation error.

3 MaRVL: Dataset Annotation

MaRVL uses a native-speaker-driven protocol to select culturally relevant concepts and images across five diverse languages, then constructs grounded True/False caption judgments from image pairs.

  • The protocol has five phases: selecting languages, universal concepts, language-specific concepts, images, and captions.
  • MaRVL covers Indonesian, Swahili, Tamil, Turkish, and Mandarin Chinese, spanning diverse language families, writing systems, geographies, and resource levels.
  • Concept selection: Native annotators provide culturally representative concepts from shared semantic fields, while selection avoids relying on translated WordNets and retains concepts supported by multiple votes.
  • Image selection: Native annotators select images that are commonly found or representative in their speaking population, alongside requirements such as multiple instances, interactions, activities, natural imagery, and CC licensing.
  • Caption annotation: Each annotation instance randomly pairs eight concept images into four pairs, with a caption written to be true for two pairs and false for two others.Each instance contributes four data points: two true pairs and two false pairs.
  • Validation: A separate validation stage hides the original labels, relabels every image-caption pair, flags language errors, and returns disagreements for revision.

4 Dataset Analysis

Dataset analyses show strong human agreement, culturally specific concept coverage, and image distributions that differ across languages and from English benchmarks, while recruitment and coverage remain constrained for low-resource languages.

  • Human validation: At least 0.887 Fleiss’ kappa was achieved across languages among caption writers and validators, indicating almost perfect inter-annotator agreement.Validator accuracy was mostly in the high 90%s, except for Swahili at 93.0%.
  • Concept statistics: MaRVL includes culturally specific concepts absent from English WordNet, including yağlı güreş, 四合院, and DOSA.
  • Image distribution: Chinese MaRVL images have distributions very different from English NLVR2 images, whose clusters include many dog species associated with ImageNet’s granularity.
  • Image distribution: Indonesian and Swahili MaRVL image distributions also vary, largely because their concept sets include different regional animal species.
  • Multilingual and multicultural statistics: MaRVL concepts occur across more languages, language families, and macro-areas than concepts in ImageNet and NLVR2.
  • Limitations: Recruiting qualified annotators for low-resource languages remains difficult, with only 2–4 caption writers per language and possible amplification of individual-annotator bias.

5 Baselines

The paper benchmarks multilingual and monolingual vision-and-language Transformer models using controlled pre-training, English NLVR2 fine-tuning, and zero-shot or translation-based cross-lingual transfer.

  • Multilingual models: M3P extends Unicoder-VL with multilingual input encoding, providing a multilingual multimodal BERT-like architecture.
  • Model architecture: The UNITER input concatenates language embeddings with visual features from a pre-trained object detector and an image-level feature.
  • Experimental setup: The authors pre-train models with the same data and hyperparameters as a controlled setup, then fine-tune them on NLVR2 for fair monolingual–multilingual comparison.
  • Baselines: Five monolingual models—UNITER, VL-BERT, VisualBERT, ViLBERT, and LXMERT—are fine-tuned on English NLVR2 and evaluated using translate-test cross-lingual transfer.

6 Results

MaRVL evaluations show substantial performance degradation under multilingual and multicultural transfer, with out-of-distribution concepts contributing most strongly to errors. Translation improves results but does not close the gap with English performance.

  • Evaluation setup: Accuracy and consistency are the two reported metrics for MaRVL and NLVR2, with translate test evaluating MaRVL after translation into English.The table caption notes that bolded best scores do not imply statistical significance.
  • Baseline comparisons: Differences among models using the same transfer method are not statistically significant, suggesting limited impact from neural architecture variation at equal pre-training data scale.The comparison covers the baseline evaluations reported for MaRVL.
  • Zero-shot vs. translate test: Zero-shot multilingual transfer drops 10–20 percentage points on MaRVL, remaining just above chance even for Mandarin Chinese.Translate-test baselines gain 4–15% across languages, with Turkish improving the most, but remain more than 10% below English NLVR2 performance.
  • Disentangling shifts in distribution: Out-of-distribution concepts cause the largest share of errors, producing an average accuracy drop of 10%.Manually translating MaRVL-ZH into English improves each model by only 1–2%, except mUNITER, indicating that machine translation is not the main source of difficulty.
  • Disentangling shifts in distribution: Both mUNITER and xUNITER lose 16% accuracy when English NLVR2 examples are manually translated into Mandarin Chinese despite remaining in-domain.This controlled comparison attributes the gap to cross-lingual transfer from English into Chinese.
  • Translate train: mUNITER scores 62.5/18.7 and xUNITER 61.8/16.7 when evaluating translated NLVR2 training data on MaRVL-ZH.These results are close to the models’ performance when MaRVL-ZH is machine translated into English.

7 Related Work

MaRVL extends grounded language reasoning to a multilingual and multicultural setting while following the binary classification formulation used by NLVR2. It builds on earlier multilingual image-caption datasets but targets reasoning over paired real-world images.

  • Grounded language reasoning: MaRVL is the first multilingual and multicultural dataset for grounded language reasoning.It follows NLVR2 by evaluating binary judgments of human-written descriptions grounded in pairs of real-world photographs.
  • Grounded language reasoning: NLVR2 extends NLVR from synthetic image pairs to real-world photographs while retaining compare-and-contrast reasoning over descriptions.The task evaluates whether grounded descriptions are true or false.
  • Multilingual multimodal datasets: Earlier multilingual multimodal datasets primarily added translated or newly written captions for existing image collections.Examples include Multi30k and caption extensions for MS-COCO images in German, French, Japanese, and Chinese.

8 Conclusions and Future Work

The paper concludes that existing visiolinguistic datasets underrepresent concepts and images salient across languages and cultures, motivating native-speaker-driven data construction. MaRVL provides a challenging benchmark whose results expose limitations of current transfer methods and suggest directions for future work.

  • Conclusions: Existing visiolinguistic datasets likely contain concepts and images that are neither salient nor prototypical outside English-speaking and Western cultural contexts.The conclusion attributes this assessment to the paper’s empirical and theoretical analyses.
  • Conclusions: MaRVL selects images and captions through native speakers and covers Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish.The dataset and annotation guidelines are publicly released for further expansion.
  • Conclusions: Baseline performance on MaRVL can remain just above chance because concepts, images, and languages are out of distribution relative to English datasets.The authors argue that this makes MaRVL a more faithful estimate of model suitability beyond a narrow linguistic and cultural domain.
  • Future work: Future work should improve out-of-distribution crosslingual transfer, study how visual stimuli affect language, and test additional tasks and contrastive multilingual models.The paper specifically mentions object recognition and multilingual visiolinguistic extensions based on contrastive learning.

Ethics Statement

The study gathers data from workers across under-represented languages and language families, using multiple crowdsourcing platforms and staged annotation procedures. Quality control includes training, filtering, image requirements, validation, and human accuracy checks.

  • Motivation: The data collection targets languages under-represented in vision-and-language datasets and spans different language families.The stated rationale is to support lower-resourced languages and broader linguistic diversity.
  • Crowdsourcing: Workers were recruited through Prolific and Proz, with Proz workers predominantly based in Indonesia, India, Turkey, China, Tanzania, Kenya, and Somalia.Prolific workers were predominantly based in North American and Western European countries; all workers were paid £15–£20/hour.
  • Image selection: Image collection excludes synthetic images, collages, watermarks, and low-resolution images, and requires Creative Commons licensing.Collected images undergo a second quality-check round to remove unqualified items.
  • Annotation procedures: Description writers receive training, submit samples, and revise work when it does not align with the guidelines.Description writers are recruited through Proz, while validators come from both Proz and Prolific.
  • Validation: Final-round validators prioritize logical correctness over grammaticality and fluency when relabelling examples.Validators are recruited from both crowdsourcing platforms.

E Additional Dataset Statistics

The dataset includes language-specific description-length and image-feature analyses, with descriptions generally longer than those in NLVR2 except in Turkish and Tamil.

  • Description lengths: Descriptions are longer than NLVR2 for Indonesian, Mandarin Chinese, and Swahili, but not for Turkish or Tamil.The comparison is based on the description-length distributions plotted in Figure 6.
  • Image features: Image-feature distributions are additionally visualized for Tamil and Turkish, alongside all MaRVL languages and NLVR2.These supplementary visualizations extend the main-text analysis to the remaining languages and provide a full cross-dataset graph.

G Baselines Details

The baselines extend UNITER with multilingual pre-training and are fine-tuned on English NLVR2 before zero-shot and cross-lingual evaluation on MaRVL. Analyses compare model behavior across semantic chapters and inspect the visual and data-design choices underlying the benchmark.

  • Dataset analysis: Image examples show that the same concept, basketball, can have drastically different visual representations across languages and cultures.Differences include players’ personal attributes and field backgrounds.
  • Architecture: mUNITER and xUNITER extend UNITER’s single-stream Transformer architecture, concatenating sub-word text inputs with visual features.The architecture uses a single stack of Transformer layers similar to BERT, with language and vision inputs combined.
  • Pre-training: Multilingual pre-training alternates text-only batches using masked language modelling with English multimodal batches using MLM, MRC-KL, and ITM objectives.The multimodal objective combines masked language modelling, masked region classification with KL-divergence, and image–text matching.
  • Dataset analysis: MaRVL and NLVR2 image-feature distributions are compared across languages, including supplementary distributions for Tamil and Turkish.The full multilingual comparison appears in Figure 8.
  • Fine-tuning and evaluation: The models are fine-tuned on English NLVR2, where they classify whether a description is valid for both images, then evaluated zero-shot on MaRVL.The best NLVR2 validation parameter sets are used for zero-shot and cross-lingual evaluation on MaRVL.
  • Performance by chapter: Per-chapter accuracy remains close to overall accuracy, with no chapter consistently easier or harder across languages for both models.Chapter-level fluctuations vary by language, model, and semantic field, including opposite behavior on “Motion” in Swahili.

I More Examples from MaRVL

Additional MaRVL examples illustrate grounded true-or-false statements across Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. The examples cover culturally situated concepts and comparisons between paired images.

  • Language coverage: The appendix provides two additional examples for each of the five MaRVL languages.The examples are organized as MaRVL-TA, MaRVL-ZH, MaRVL-SW, MaRVL-ID, and MaRVL-TR figures.
  • Tamil example: A Tamil example states that in one picture a finger shows the vote and assigns the grounded statement a true label.The example pairs the native-language statement with the concept INK.
  • Swahili example: A Swahili example compares one person blowing a flute in the left image with two people doing so in the right image.The statement explicitly grounds the comparison in the number of flute players across the paired images.
  • Turkish example: A Turkish example labels a statement about multiple people holding qanuns on their knees as true.Qanun is identified in the example as a popular instrument in Turkey.
Loading 2109.13238v2…