Source-linked AI summary

On Measuring Gender Bias in Translation of Gender-neutral Pronouns

Won Ik Cho, Ji Won Kim, Seok Min Kim, Nam Soo Kim

arXiv:1905.11684v1cs.CL

TL;DR

Gender bias in machine translation is difficult to evaluate across languages, especially when gender-neutral Korean pronouns may become gender-specific in English. The paper constructs a controlled Korean-English test corpus and introduces TGBI to measure preservation of gender neutrality. Google Translator scores highest overall among the evaluated candidates, while interpretation must distinguish corpus-volume bias from social bias.

  • Problem

    Gender-neutral Korean pronouns can be translated into gender-specific English forms, but gender bias in machine translation has not been thoroughly evaluated.

  • Method

    The paper constructs a 4,236-utterance Korean-English evaluation corpus and measures systems with the translation gender bias index (TGBI).

  • Results

    Google Translator scored highest overall among the three evaluated candidates.

  • Takeaways & Limitations

    Gender bias should be reduced when the translation context does not require gender specification.

  • Takeaways & Limitations

    A higher TGBI or pw does not necessarily establish that a system is unbiased because corpus-volume bias and social bias can both affect the result.

Abstract

from arXiv · show

Ethics regarding social bias has recently thrown striking issues in natural language processing. Especially for gender-related topics, the need for a system that reduces the model bias has grown in areas such as image captioning, content recommendation, and automated employment. However, detection and evaluation of gender bias in the machine translation systems are not yet thoroughly investigated, for the task being cross-lingual and challenging to define. In this paper, we propose a scheme for making up a test set that evaluates the gender bias in a machine translation system, with Korean, a language with gender-neutral pronouns. Three word/phrase sets are primarily constructed, each incorporating positive/negative expressions or occupations; all the terms are gender-independent or at least not biased to one side severely. Then, additional sentence lists are constructed concerning formality of the pronouns and politeness of the sentences. With the generated sentence set of size 4,236 in total, we evaluate gender bias in conventional machine translation systems utilizing the proposed measure, which is termed here as translation gender bias index (TGBI). The corpus and the code for evaluation is available on-line.

1 Introduction

The paper addresses gender bias when Korean gender-neutral pronouns are translated into gender-specific English pronouns. It proposes a test corpus and measure for evaluating whether KR-EN systems preserve gender neutrality.

  • Motivation: Korean gender-neutral pronouns can be translated into gender-specific English forms because corpus associations link neutral source words with target pronouns.This can reduce generality, insert social bias, and produce outputs that are incorrect with respect to gender neutrality.
  • Approach: The evaluation uses Korean templates containing sentiment or occupation terms and measures whether the pronoun is translated as female, male, or gender-neutral.The template varies the lexicon, pronoun formality, and sentence politeness.
  • Corpus: The resulting equity evaluation corpus contains 4,236 utterances built from 1,059 sentiment and occupation terms.The corpus includes 324 sentiment-related phrases and 735 occupation words.
  • Contributions: The paper contributes a KR-EN test corpus, a comparison measure for preserving gender neutrality, and an analysis of why neutral pronoun preservation matters.The corpus is accompanied by detailed construction guidelines.

2 Related Work

Prior work documents gender bias across NLP tasks, including machine translation, but Korean was omitted from a closely related multilingual translation study. This paper fills that gap with a KR-EN evaluation scheme and cautions that reduced male dependency may reflect another social bias.

  • Existing evidence: Gender bias has been observed in image captioning, language modeling, coreference resolution, machine translation, classification, and recommendation.Reported examples include gendered associations in captions and male-skewed applicant recommendations.
  • Machine translation: Prates et al. (2018) found strong male dependency in Google Translator across twelve languages using occupations and adjectives, but Korean was omitted.The paper identifies Korean as a missing cross-lingual case despite its relevance to gender-neutral pronoun translation.
  • Research gap: This work proposes the first concrete KR-EN scheme for evaluating gender bias involving sentiment words and occupations, together with an inter-system measure.It also warns that reduced male dependency may indicate involvement of another social bias rather than reduced overall bias.

3 Proposed Method

The proposed method creates an equity evaluation corpus and applies it to machine-translation models to evaluate gender bias.

  • Method: The EEC is created as a test corpus and then used to evaluate gender bias in machine-translation models.

3.1 Corpus generation

The corpus-generation scheme targets gender-neutral Korean pronoun translation with controlled lexical and stylistic variation. It combines neutral pronouns, sentiment and occupation lexicons, formality, politeness, and polarity criteria.

  • Template and lexicon design: The corpus uses templates of the form “S/he is [xx]” with sentiment and occupation terms to test pronoun translation without strongly gendered lexical cues.Occupation terms are treated as neutral, while sentiment terms are divided by positive and negative polarity.
  • Gender-neutral pronouns: Korean lacks grammatical gender, but speakers use gender-specific 그/그녀 and gender-neutral 걔 when a person’s gender is unidentified.The study also adopts 그사람 as a more formal variation of 걔.
  • Stylistic controls: Formality, politeness, and polarity define the test-set construction criteria, with the polite suffix 요 added or the sentence ending transformed when necessary.
  • Sentiment lexicon: Sentiment terms were filtered through native-speaker consensus and exclusions of categories judged prejudicial, including appearance and richness.The source sentiment dictionary supplied positive and negative items before filtering.
  • Occupation lexicon: The occupation list contains 735 official job terms, with gender-specific and hateful expressions removed to avoid sentiment polarity.The terms were collected from an official government employment website and checked for redundancy.

3.2 Measure

The measure scores how closely translations preserve gender neutrality, then averages those scores across sentence sets to form TGBI. Its interpretation must distinguish corpus-volume bias from social prejudice encoded in lexical choices.

  • Measure: For each sentence set, pw, pm, and pn denote the portions translated as female, male, and gender-neutral, respectively.The Korean template targets pronouns whose gender neutrality should be preserved in translation.
  • Measure: The per-set score ranges from 0 to 1, maximizing when pn is 1 and minimizing when either pw or pm is 1.This matches the goal of gender-neutral pronoun translation and treats balanced gendered guesses as preferable for a fixed pn.
  • Measure: The Korean template changes to a noun-phrase form with -ya instead of -hay when [xx] is a noun phrase.This preserves the appropriate grammatical form for occupation expressions.
  • Measure: TGBI is the nonweighted average of per-set scores across all sentence sets, yielding 1 when every prediction uses gender-neutral terms.Nonweighted averaging prevents sentence sets with smaller volumes from being overlooked.
  • A remark on interpretation: Interpretation separates VBias from corpus frequency and SBias from social prejudice projected through the evaluated lexicons.A high pw can reflect SBias rather than reduced bias, so higher scores for a sentence set do not by themselves establish an unbiased system.

4 Evaluation

Across three conventional translation systems, gender bias varied by content and style: occupation sentences were most biased, formal pronouns skewed male, and qualitative analysis revealed system-specific patterns.

  • Overall comparison: GT scored highest overall and KT lowest, while the proposed comparison examines both volume bias and social bias rather than relying on raw gender counts alone.The authors caution that a lower occupation score does not by itself establish greater bias because one system’s corpus bias can attenuate another bias component.
  • Content-related features: Positive sentiment words yielded relatively more female inferences than negative words for GT and NP, although absolute values still indicated male-oriented volume bias.This contradicted the initial expectation that negative sentiment would be more associated with female translations.
  • Style-related features: Formal, gender-neutral pronouns were significantly more male-biased than informal ones, whereas politeness produced no significant system difference.The authors relate the formality pattern to corpus context and possible real-world occupational gender ratios, while cautioning that this interpretation is controversial.
  • Content-related features: Occupation sentences were more biased than sentiment or style subsets across all systems, with GT reflecting occupational stereotypes and KT showing stronger male-volume dominance.NP produced unexpected female translations for researchers and engineers and male translations for cuisine-related occupations, which the authors associate with technical bias mitigation.
  • Comparison with other schemes: The evaluation corpus supports inter- and intra-system analysis across content and style subsets, while qualitative inspection remains necessary because quantitative scores do not reveal translation semantics.The scheme also retains gender-neutral outputs and is intended to show how bias is distributed rather than label any one gender-skewed system intrinsically biased.
  • Discussion: The methodology is most directly extensible to other source languages with gender-neutral pronouns and comparable formality and politeness distinctions, such as Japanese.The authors note that target-language behavior depends on whether the target has gender-neutral pronouns.

5 Conclusion

The paper introduces a 4,236-sentence Korean–English test corpus and a measure for evaluating gender-neutral pronoun preservation, then compares translation systems and identifies context-aware post-processing as future work.

  • 5 Conclusion: The corpus contains seven sentence subsets covering formality, politeness, sentiment polarity, and occupation, and evaluation averages PS across subsets.PS combines the portions translated into female, male, and gender-neutral terms.
  • 5 Conclusion: Google Translator scored highest overall among the three evaluated systems, while Kakao Translator scored lowest.Qualitative analysis suggested that Naver Papago may implement an algorithmic modification for occupation-related cases.
  • 5 Conclusion: The authors argue that gender bias should be reduced when gender specification is not required.
  • 5 Conclusion: Future work will explore context-aware post-processing and expanded contextual test sets to assign gender specificity or neutrality more appropriately.The authors also plan deeper analysis of model architectures and training datasets.

A Proof on the boundedness of the measure

The proof establishes that the proposed measure is bounded between zero and one and reaches its maximum under fully gender-neutral translation and its minimum under entirely female or male translation.

  • A Proof on the boundedness of the measure: The optimization analyzes W over a compact, convex domain using a Lagrangian and KKT conditions.Interior optimality leads to a contradiction, so optimal points lie on the boundary, which is decomposed into three independent segments.
  • A Proof on the boundedness of the measure: The boundary analysis separately considers cases constrained by x + y = 1, y + z = 1, and the associated nonnegative variables.
  • A Proof on the boundedness of the measure: 0 ≤ W0 ≤ 1 establishes boundedness of the proposed measure.
  • A Proof on the boundedness of the measure: W0 is maximized when pn = 1 and minimized when either pw or pm = 1.

B A brief demonstration on the utility of adopting multiple sentence subsets

The section explains why evaluating multiple sentence subsets is useful: deterministic systems can produce uneven subset scores, making simple arithmetic averaging misleading, while the proposed measure preserves variation across sociolinguistic subsets.

  • B A brief demonstration on the utility of adopting multiple sentence subsets: Conventional translation services provide a determined answer for each input sentence, which can hinder achieving a high score with the proposed measure and EEC.
  • B A brief demonstration on the utility of adopting multiple sentence subsets: Deterministic translation services can yield high scores for some subset pairs but lower scores for others, so arithmetic averaging may obscure bias.
  • B A brief demonstration on the utility of adopting multiple sentence subsets: The proposed measure uses pn and a square-root function to prevent averages across many subsets from converging to one specific value.
  • B A brief demonstration on the utility of adopting multiple sentence subsets: The authors therefore retain all sentence sets in the EEC to observe tendencies across varied sociolinguistic dimensions.
Loading 1905.11684v1…