Source-linked AI summary
Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models
Melina Morch, Daniel Braun
TL;DR
Anti-LGBTQ bias evaluation remains limited beyond English and often misses cultural and linguistic variation. This paper combines a German WinoQueer adaptation with community-sourced data, evaluates multilingual models, and tests fine-tuning mitigation. The models reproduce anti-queer stereotypes, culturally adapted data differs from translated data, and mitigation is small and inconsistent.
Problem
Anti-queer bias benchmarks are limited beyond English and often fail to capture culturally situated, intersectional, and linguistically varied understandings of gender and sexuality.
Method
The paper combines a German WinoQueer adaptation with a community survey dataset, evaluates German and English-German models, and tests fine-tuning on queer-oriented media content.
Results
Language models reproduce anti-queer stereotypes, with bias varying by identity and model; genderqueer identities show the strongest bias, and masked models exhibit higher bias than autoregressive models.
Takeaways & Limitations
Differences between translated and community-based data indicate that cultural adaptation is central to robust multilingual bias evaluation, while translation alone is insufficient.
Takeaways & Limitations
The bias score is threshold-free and relative, while fine-tuning data is not explicitly matched to measurable mitigation success for each identity group.
Abstract
from arXiv · showhide
While gender and racial biases in language models have been widely studied, anti-LGBTQ biases remain underexplored, particularly beyond English. Existing benchmarks often do not capture cultural and linguistic variation and rely on gender representations. This paper introduces a multilingual German-English benchmark dataset for the evaluation of anti-LGBTQ biases in language models. It combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer. The data is used to evaluate eight language models across sizes and architectures and explore mitigation through fine-tuning on community and progressive media content. Results show that language models reproduce anti-queer stereotypes, with variation across identities and models. Differences between the translated and community-based data highlight the importance of cultural adaptation for multilingual bias evaluation. Fine-tuning reduces bias on average, but not consistently across models and identities. Warning: This text contains examples of anti-queer hateful language and stereotypes.
1 Introduction
Anti-queer bias in language models is underexplored outside English, especially where gender and sexuality are culturally situated and linguistically complex. The paper addresses this gap with German adaptations, community-grounded data, a broader bias score, multilingual evaluation, and targeted fine-tuning.
- Research gap: Existing anti-queer benchmarks are mainly English-language and often reduce gender to binary categories, limiting coverage of fluid, intersectional, and culturally situated discrimination.German adds complexity because gender is encoded morphologically and syntactically.
- Research gap: Multilingual models may transfer, reinforce, or transform queer stereotypes across linguistic and cultural contexts.The paper therefore examines how queer identities are represented and marginalized in German and English-German models.
- Contributions: The paper introduces a German adaptation and extension of WinoQueer for evaluating anti-queer bias beyond English-language settings.This creates a benchmark suited to multilingual evaluation in a grammatically gendered language.
- Contributions: The paper adds a survey-derived dataset from more than 100 German-speaking queer individuals, grounding evaluation in lived experiences and community-specific language practices.The dataset captures culturally situated stereotypes and assumptions often absent from existing resources.
- Contributions: The paper systematically evaluates German and English-German multilingual models across queer identities and broader assumptions about LGBTQ communities.The dataset and accompanying code are publicly available on GitHub.
- Contributions: The study introduces a more nuanced and comprehensive score for assessing and quantifying language-model biases.It also explores domain-specific fine-tuning on community-oriented and progressive media content for multilingual English-German and German models.
2 Related Work
Prior bias research has expanded from static embeddings to contextual language-model benchmarks, but queer-inclusive evaluation remains limited outside English. This paper extends that work with German WinoQueer and a culturally grounded German dataset.
- Prior bias evaluation: Bias research spans gender, ethnicity, temporal, and social biases in word embeddings, while transformer models motivated direct evaluation of outputs and task performance.Contextualized representations cannot be analyzed as simply through geometric methods as static embeddings.
- Prior bias evaluation: Benchmarks such as CrowS-Pairs and StereoSet broadened evaluated bias categories and used sentence-level comparisons between stereotypical and anti-stereotypical variants.Their scoring relies on masked or comparative likelihoods.
- Queer-inclusive benchmarks: WinoQueer uses sentence pairs derived from survey responses of 295 LGBTQ individuals describing lived experiences with bias and stereotypes.Its evaluation masks shared tokens while holding modified tokens constant before calculating sentence likelihoods.
- Queer-inclusive benchmarks: WinoQueer’s bias score is the percentage of examples where the stereotypical sentence has higher likelihood than the less stereotypical sentence.Prior English results reported less anti-queer bias in masked language models than in autoregressive models.
- Multilingual gap: Queer-inclusive research has a substantial gap beyond English, including limited study of grammatical gender in languages such as German or Spanish.Prior multilingual work may include stereotypes across languages while still using binary gender definitions.
- Multilingual gap: Bergstrand and Gambäck’s Norwegian extension reported an average bias score of 68.27% across models, indicating greater likelihood of LGBTQIA+ stereotypes than anti-stereotypes.This result illustrates the scale of bias measurable in non-English benchmark adaptations.
- This paper’s contribution: This paper extends prior evaluation by translating WinoQueer into German and introducing a German culturally grounded dataset based on queer people’s lived experiences.The two datasets combine multilingual coverage with community-specific context.
3 Methodology
The study constructs complementary translated and community-gathered German datasets, evaluates eight language models with binary and soft bias metrics, and prepares queer-language corpora for mitigation fine-tuning.
- Dataset construction: The methodology combines a German translation of WinoQueer with a community-gathered dataset documenting lived experiences of bias in German-language contexts.The community dataset was created through a participatory, community-in-the-loop approach.
- Dataset construction: 45,540 WinoQueer sentence pairs were semi-automatically translated into German and reviewed through native-speaker inspection and rule-based scripts.English names were also replaced with German names, and 64% of translations received manual or automated postprocessing.
- Dataset construction: The community survey produced 430 reported bias experiences from 103 queer participants, which were converted into 387 CrowS-pairs after duplicate removal.Frequently reported experiences included claims that queerness is a phase, media-driven, dangerous for children, or tied to pedophilia.
- Dataset construction: The resulting community dataset contains a larger share of non-binary sentence pairs than trans, lesbian, and pansexual sentence pairs.Non-binary identities comprise 26.4%, trans identities 13.3%, lesbians 10.4%, and pansexual people 8.2% of the dataset.
- Models and evaluation: Eight models span German-only, bilingual, and multilingual settings across masked and autoregressive architectures.The evaluated systems include German BERT, multilingual BERT, GPT-2, OPT, BLOOM, XLM-RoBERTa, XLM-CLM-ENDE, and XLM-MLM-ENDE.
- Models and evaluation: Bias evaluation masks differing tokens for masked models or predicts shared tokens left to right for autoregressive models, then compares sentence scores.Binary scores identify whether the biased sentence is preferred, while soft scoring expresses preference intensity from 0% to 100%.
- Mitigation: Mitigation training uses 2,208 filtered posts from 11 queer-themed or queer-friendly German Mastodon instances and queer-focused articles from taz.The corpus preparation removes sexually explicit posts, automated server messages, non-relevant hashtags, and links.
4 Results
Across eight language models, anti-LGBTQ bias varied by identity, model, dataset, and scoring method. Transgender and non-binary identities showed especially high bias, while the community-based dataset produced higher overall bias than translated WinoQueer and fine-tuning reduced bias inconsistently.
- Translated WinoQueer: 49.3% bias and 49.6% soft score were the mean results across models on translated WinoQueer, indicating an almost balanced overall score.Scores above 50% count as biased.
- Translated WinoQueer: Transgender and non-binary identities had the highest average bias and substantial variation across models, ranging from 11.8% anti-trans bias in GPT-2 to 96.2% in BLOOM-560m.The general LGBTQ category had the lowest bias in GPT-2 at 8.4% and remained below 50% for all models except XLM-MLM.
- Cross-dataset comparison: Translated WinoQueer produced lower bias than the original evaluation, with only transgender and non-binary groups exceeding 60% mean bias in German compared with all identities in the original.The translated dataset did not reproduce the original evaluation’s highest mean bias against asexual people.
- Community-based dataset: 56.3% bias and 52.1% soft score were the corresponding means for the newly created community-based dataset, higher than the translated dataset.All models except XLM-MLM-ENDE and XLM-RoBERTa scored above 50% overall bias on the new dataset.
- Community-based dataset: Bias in the community-based dataset was highest for queer, pansexual, and demisexual individuals, while German BERT was notably biased across most identities.Non-binary bias was low in most models but very high in German BERT; XLM-MLM-ENDE and XLM-RoBERTa tended to score low.
- Mitigation: Fine-tuning reduced bias by -9.9% in bias score and -5.2% in soft score on the survey dataset, but mitigation was slight on translated data and reversed for some models.XLM-MLM-ENDE showed reinforced bias, especially against transgender and asexual identities; the smaller survey dataset also produced higher variance.
5 Conclusion
The evaluation finds strong anti-queer bias, especially against genderqueer identities, while mitigation effects are generally small and inconsistent. Differences between translated and community-based data underscore the importance of cultural adaptation and more integrated queer-linguistic evaluation.
- Bias against genderqueer identities, comprising non-binary and transgender people, is consistently strongest across models.
- The translated dataset shows no notable bias-score difference between monolingual and multilingual models, whereas German BERT is much more biased than its multilingual counterpart on the community dataset.A possible explanation is stronger internalization of the community dataset’s cultural stereotypes by the monolingual model.
- The translated benchmark and community survey produce different bias patterns, indicating that cultural adaptation is central to multilingual bias evaluation.The survey captures culturally grounded, intersectional expressions more representative of German-speaking contexts.
- Less than 10% ∆ improvement is observed on average after mitigation, despite reductions for most models and identities and amplification in some cases.
- Translation can weaken culturally embedded stereotypes because idiomatic usage and discourse-specific associations are often lost.
- The authors call for deeper integration of queer linguistics and evaluation designs beyond purely statistical identity representations.
Limitations
The paper identifies limitations in counterfactual wording, scoring interpretation, mitigation design, model scale, and dataset stability. Several boundaries concern whether the evaluation faithfully captures culturally situated and fine-grained queer identities.
- Counterfactual sentences using “straight”, “cisgender” and “heterosexual” may be less effective because queer-sensitive contexts make these terms unusually explicit.Cis- and heteronormative media may leave such norms undescribed.
- The threshold-free bias score counts any difference between factual and counterfactual sentences as bias or anti-bias, making results sensitive to the chosen threshold and dataset composition.The soft score instead averages bias across sentence pairs to represent preference strength.
- Cross-lingual fine-tuning remains unevaluated, limiting comparison with prior work that used 8.2 million sentences rather than 4.9 million words.
- The fine-tuning data was selected using identity keywords but was not explicitly matched to measurable mitigation success.Identity-specific training and evaluation could better pinpoint mitigation effects.
- The study mostly focuses on smaller language models because of comparability and resource restrictions, although large models are more important in practice.
- The survey and evaluation address sensitive discrimination and verbal-violence topics, with voluntary, anonymous participation and the option to skip or stop.
C.1 Motivation for Dataset Creation
The dataset was created to study and mitigate anti-LGBTQ bias in German language models using stereotypes grounded in LGBTQ community experiences. Its intended use is academic analysis of real-world linguistic discrimination, with acknowledged misuse risk.
- The dataset targets anti-LGBTQ bias in German language models using real-world stereotypes experienced by LGBTQ community members.
- The captured stereotypes could be misused to reproduce the same harmful stereotypes.
C.2 Dataset Composition
The dataset consists of crowd-sourced stereotype pairs designed for evaluation rather than model training. Each pair contrasts an LGBTQ-targeting stereotypical statement with a counterfactual statement about non-LGBTQ people, and the authors recommend soft scoring.
- Each instance is a Crowd-sourced Stereotype Pair consisting of two sentences.
- The dataset contains 387 sentence pairs, with targeted-group counts exceeding the number of pairs because statements may target multiple groups.
- Each pair contains one stereotypical or offensive attribution to LGBTQ people and one counterfactual attribution to non-LGBTQ people.
- No recommended data split is provided because the dataset is not designed for model training.
- The authors recommend the paper’s soft scoring method as the evaluation measure.
C.3 Data Collection Process
The dataset was collected through an online survey of unpaid volunteers and converted participants’ reported stereotypes into sentence pairs. It covers multiple LGBTQ groups but is not claimed to be complete or representative.
- Data collection used an online survey hosted on SoSci Survey.
- Participants were recruited through university email lists and queer social media as unpaid volunteers.
- The data was collected in 2025 from survey responses reported by participants.
- Authors created sentence pairs from stereotypes that participants reported experiencing.
- The supplied description does not specify the sampling strategy, representativeness, or missing-data rationale.
- The dataset spans multiple LGBTQ groups but does not claim completeness.
C.4 Dataset Distribution
The dataset is archived on GitHub, released with the paper’s publication, and distributed under CC-BY-4.0 without stated access or export restrictions.
- The dataset is archived in a GitHub repository.
- The dataset’s first distribution is tied to publication of the paper.
- The dataset is distributed under the CC-BY-4.0 license.
C.5 Dataset Maintenance
The dataset is maintained through its GitHub repository, with updates planned only for important mistakes and extensions suggested through GitHub forks. The documentation states that participants were informed, the dataset is non-identifiable, and it contains offensive content.
- Dataset Maintenance: The GitHub repository is the stated location for dataset support, maintenance, updates, and obsolescence communication.
- Dataset Maintenance: Updates are not planned unless important mistakes become clear.
- Dataset Maintenance: Users can extend the dataset by creating a fork on GitHub, with a repository for linking papers or systems that use it.
- Ethics and Data Protection: Participants were informed about the survey purpose, and the resulting dataset contains no information linked to participants.
- Ethics and Data Protection: The documentation states that participants consented during the survey and that no harm or unfair disadvantage was identified.
- Ethics and Data Protection: The dataset is described as containing no sensitive or confidential information but does contain inappropriate or offensive information.