Source-linked AI summary

Assessing Gender Bias in Machine Translation -- A Case Study with Google Translate

Marcelo O. R. Prates, Pedro H. C. Avelar, Luis Lamb

arXiv:1809.02208v4cs.CYcs.CL

TL;DR

The paper asks whether statistical translation tools reproduce or amplify societal gender asymmetries. It translates occupational sentences from gender-neutral languages into English, compares pronoun outputs with Bureau of Labor Statistics data, and finds exaggerated male defaults, especially in STEM-related occupations.

  • Problem

    The paper addresses the limited ability to systematically study gender bias in AI by examining whether machine translation exposes such asymmetries.

  • Method

    The study translates job-position sentences from 12 gender-neutral languages into English with Google Translate and compares pronoun frequencies with U.S. Bureau of Labor Statistics participation data.

  • Results

    Google Translate shows prominent and exaggerated male defaults in stereotyped fields, and its female-pronoun frequencies do not correlate with women’s participation across occupations.

  • Takeaways & Limitations

    The findings support probing statistical translation outputs for stereotypical gender roles and considering debiasing techniques for translation systems.

  • Takeaways & Limitations

    The study does not test whether its findings extend to translations between gender-neutral and non-English gendered languages, and Korean pronoun omission cannot be framed clearly in its sentence templates.

Abstract

from arXiv · show

Recently there has been a growing concern about machine bias, where trained statistical models grow to reflect controversial societal asymmetries, such as gender or racial bias. A significant number of AI tools have recently been suggested to be harmfully biased towards some minority, with reports of racist criminal behavior predictors, Iphone X failing to differentiate between two Asian people and Google photos' mistakenly classifying black people as gorillas. Although a systematic study of such biases can be difficult, we believe that automated translation tools can be exploited through gender neutral languages to yield a window into the phenomenon of gender bias in AI. In this paper, we start with a comprehensive list of job positions from the U.S. Bureau of Labor Statistics (BLS) and used it to build sentences in constructions like "He/She is an Engineer" in 12 different gender neutral languages such as Hungarian, Chinese, Yoruba, and several others. We translate these sentences into English using the Google Translate API, and collect statistics about the frequency of female, male and gender-neutral pronouns in the translated output. We show that GT exhibits a strong tendency towards male defaults, in particular for fields linked to unbalanced gender distribution such as STEM jobs. We ran these statistics against BLS' data for the frequency of female participation in each job position, showing that GT fails to reproduce a real-world distribution of female workers. We provide experimental evidence that even if one does not expect in principle a 50:50 pronominal gender distribution, GT yields male defaults much more frequently than what would be expected from demographic data alone. We are hopeful that this work will ignite a debate about the need to augment current statistical translation tools with debiasing techniques which can already be found in the scientific literature.

1. Introduction

The paper situates machine translation within debates about statistical models and investigates whether Google Translate reproduces gender stereotypes in English outputs.

  • Statistical machine translation has faced criticism for approximating unanalyzed data rather than modeling underlying scientific relationships.Chomsky contrasts such systems with traditional scientific models, while Norvig defends statistical inference as compatible with scientific practice.
  • The debate over statistical models reflects broader uncertainty about their role in science and their potential societal risks.The paper presents the Chomsky–Norvig disagreement as unresolved while emphasizing concerns about increasingly sophisticated statistical tools.
  • Recent examples of algorithmic bias include image labeling that classified dark-skinned people as gorillas and alleged racial bias in criminal-risk prediction.These cases motivate examining gender asymmetries in machine-learning systems.
  • The study proposes quantitatively probing gender bias by translating sentences from gender-neutral languages into English with Google Translate.Its examples suggest stereotypical roles are often rendered with female pronouns for nurses and bakers but male pronouns for engineers and CEOs.

2. Motivation

The motivation section frames Google Translate as a large-scale system whose gendered outputs can be studied alongside existing evidence that algorithmic bias can be measured and reduced.

  • Figure 1 shows traditionally male-dominated occupations interpreted with male pronouns, while nurse, baker, and wedding organizer are interpreted with female pronouns.The figure illustrates how translating from a gender-neutral language into English can expose occupational gender associations.
  • Prior debiasing work reduced stereotypical word-embedding analogies from 19% to 6% without significant performance compromise.This result motivates investigating whether comparable debiasing could be applied to Google Translate.

3. Assumptions and Preliminaries

The paper assumes translation should not amplify occupational gender inequality beyond demographic patterns and argues that gender-neutral language can test this assumption.

  • The study assumes statistical translation tools should reflect at most the gender inequality present in society.This assumption links possible translation bias to the real-world data used for training.
  • Grammatical gender distinctions may influence perceptions of the world, with prior studies relating them to sexism and gender inequalities.These claims provide background for treating language and translation as relevant to gender bias.
  • Gender-neutral communication is presented as a means of preserving neutrality instead of defaulting to male or female variants.The paper treats this as a possible way to promote improved gender equality in languages where neutrality can be achieved.
  • Google Translate is defined as negatively gender-biased when male defaults overestimate the male-to-female employee distribution in an occupation.The paper therefore does not require translated pronouns to follow a 50:50 distribution.

4. Materials and Methods

The study measures gender bias by translating gender-neutral sentences into English across selected languages, occupations, and adjectives. It uses curated language and labor-data choices while excluding languages whose grammar or script posed methodological problems.

  • Study design: Gender bias is assessed by mapping sentences from gender-neutral languages into English with an automated translation tool.The paper uses a Hungarian nurse example that Google Translate renders with a female pronoun and an engineer example that receives a male pronoun.
  • Occupations: The dataset centers on occupations from the U.S. Bureau of Labor Statistics, filtered and grouped into broader categories for analysis.The resulting collection contains 1019 occupations from 22 distinct categories.
  • Adjectives: The study adds 21 human-applicable adjectives selected from the 1000 most frequent adjectives in the COCA corpus and manually curated.This provides complementary evidence beyond the professional-occupation setting.
  • Language selection: The language set includes Google Translate-supported languages lacking mandatory gender demarcation on simple phrases.Table 1 groups languages by family and records whether they enforce male/female demarcation, with omitted languages marked separately.
  • Language exceptions: Korean and Nepali were omitted because of technical or linguistic limitations in constructing reliable gender-neutral templates.Korean context could not make omitted gender pronouns sufficiently expressive, while Nepali grammar expertise was unavailable to the authors.

5. Distribution of translated gender pronouns per occupation category

Google Translate produces more male than female or gender-neutral pronouns overall, with male defaults especially prominent in STEM and Legal occupations. Healthcare, Production, and Education show more balanced or female and neutral outputs, while many sentences receive no extractable gender pronoun.

  • Google Translate translates occupation sentences with male pronouns more frequently than with female or gender-neutral pronouns overall.
  • STEM fields concentrate near zero translated female and gender-neutral pronouns, while male-pronoun counts extend toward higher values.
  • Healthcare has the most balanced translated-pronoun distribution in comparison with STEM, where male defaults are second to most prominent.
  • Occupation-category bar plots contrast male-default predominance in Legal and STEM with larger female and neutral proportions in Healthcare and Education.
  • STEM fields are second only to Legal in male-default prominence, followed by Arts & Entertainment and Corporate; Healthcare, Production, and Education lie at the opposite end.
  • One-sided t-tests rejected the null that male pronouns were not more frequent than female pronouns for every language and occupation category examined.

6. Distribution of translated gender pronouns per language

Language-level results vary, but the distributions generally suggest male defaults and substantial difficulty extracting gender pronouns from some languages. Basque is the only tested language with more gender-neutral than male pronouns.

  • Basque was the only tested language yielding more gender-neutral than male pronouns, with Bengali and Yoruba following it in this pattern.
  • Language-group bar plots were used to examine cultural variation and the difficulty of extracting gender pronouns from different languages.
  • Female pronouns reached as low as 0.196% for Japanese and 1.865% for Chinese in the language-level distributions.
  • Rows do not generally total 100% because many translated sentences lack an obtainable gender pronoun, particularly in Basque.

7. Distribution of translated gender pronouns for varied adjectives

Preliminary adjective experiments also indicate male defaults, while showing adjective-specific variation in whether translations use female, male, or neutral pronouns.

  • The adjective analysis used a filtered subset of frequently used English adjectives that fit the human-description templates.
  • Male defaults reappeared across adjectives, but Shy, Attractive, Happy, Kind, and Ashamed were predominantly translated with female pronouns.
  • Arrogant, Cruel, and Guilty were disproportionately translated with male pronouns, and Guilty was never translated with female or neutral pronouns.
  • Figure 12 associates Shy, Desirable, Sad, and Dumb with female pronouns, while Proud, Guilty, Cruel, and Brave are almost exclusively associated with male pronouns.

8. Comparison with women participation data across job positions

The study compares Google Translate’s female-pronoun frequencies with U.S. Bureau of Labor Statistics women-participation data across occupations. Translation outputs substantially underrepresent female participation, so male defaults cannot be understood as a straightforward reflection of workplace demographics.

  • The study tests whether Google Translate’s male defaults merely reflect low female participation in male-dominated occupations.The authors frame workplace demographics as a possible alternative explanation for the observed translation asymmetry.
  • Female pronouns accumulate in the first 12-quantile, whereas BLS female participation peaks in the fourth and remains significant in later quantiles.
  • The authors report no correlation between female-worker distribution and female-translation frequency across job positions.
  • 11.76% of translations use female pronouns versus 35.94% average female participation in BLS occupation data.Translation variance is approximately 0.028, compared with approximately 0.067 for the BLS data.
  • The one-sided t-test yielded p ≈6.210−94 against α = 0.005, leading the authors to reject equal-frequency and conclude that female translations underestimate female participation.
  • The prominence of male defaults therefore remains without a clear justification from workplace demographics.

9. Discussion

The discussion places the findings in the context of Google Translate’s historical single-output design and emerging efforts to reduce gender bias. It argues that debiasing is feasible, while noting that the newer feature does not address all shortcomings and has limited language coverage.

  • The experiments analyze Google Translate as it existed in August 2018, before later interface changes.
  • Historically, Google Translate supplied one translation even when feminine and masculine forms were possible, inadvertently reproducing biases in web-based training examples.
  • Google later introduced feminine and masculine translation forms, which the authors view as evidence that translation systems can be debiased without requiring balanced training data.
  • The authors describe the new feature as an important first step rather than a complete solution to the shortcomings identified in the paper.
  • Evidence from word-embedding research suggests that gender bias is a statistical phenomenon not restricted to one proprietary translation tool.
  • The discussion recommends engineering post-training debiasing solutions because unbiased training texts are probably scarce.

10. Conclusions

The paper concludes that Google Translate exhibits gender-biased male defaults that can be studied by translating occupational and descriptive sentences from gender-neutral languages. These defaults are exaggerated in stereotyped fields and do not match U.S. workplace demographics, motivating further debiasing work.

  • The authors use translations of professional sentences from gender-neutral languages into English to measure asymmetry in female and male pronouns.
  • Google Translate shows prominent and exaggerated male defaults in STEM occupations and other fields associated with gender stereotypes.
  • Female-pronoun proportions also vary substantially with adjectives: Shy and Desirable favor female pronouns, while Guilty and Cruel are almost exclusively male.
  • Hungarian produces a more balanced male–female distribution than Chinese, while Yoruba and Basque often produce gender-neutral pronouns.The authors note that Basque also has a high frequency of phrases from which no gender pronoun could be automatically extracted.
  • Female translation frequencies underestimate BLS female participation, showing that Google Translate does not reproduce the real-world distribution of women workers.
  • The paper calls for discussion of how AI engineers can minimize harmful effects of machine bias and points to existing debiasing algorithms as a possible resource.
Loading 1809.02208v4…