Source-linked AI summary
Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models
Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji
TL;DR
This paper examines whether speech-to-speech models follow a speaker’s voice or stereotyped content when rendering speech and attributing gender. Using controlled passages across five models and three languages, it finds no stereotype drift in rendered voices but systematic content-based gender attribution, especially when voice and content clash.
Problem
Speech-to-speech models can re-express people in ways that misgender them or reinforce stereotypes, while fixed output voices make voice-rendering bias difficult to detect.
Method
The study crosses male and female voices with matching, neutral, and clashing stereotyped content across five models and three languages, evaluating voice rendering and gender attribution.
Results
Every model shows no stereotype drift in rendered voices but shifts gender judgments with content, with odds ratios of 1.7–24 per content step and 83–100% misgendering in misaligned cells.
Takeaways & Limitations
Fixed-voice evaluations can miss stereotype-driven gender attribution, so audits should test whether gender judgments remain invariant when only content changes.
Takeaways & Limitations
The binary design audits male/female behaviors and is silent on non-binary reference; passages are also LLM-generated and LLM-rated with one human screening rater.
Abstract
from arXiv · showhide
Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithful system should treat a speaker as who they sound like, not as whoever usually says what they said. Testing this is harder than it looks, since most S2S models answer in a single, fixed output voice, hard-coded so it cannot drift toward a stereotype. Checking the output voice comes back clean even when the model is biased. We therefore ask two questions. When a model re-speaks the input, does the stereotype in the words shift the perceived gender of the output voice (voice rendering)? And when the model states the speaker's gender, does it follow the voice or the content (gender attribution)? We answer both with one controlled experiment crossing male and female voices with masculine-, neutral-, and feminine-stereotyped passages, on five open- and closed-source models in English, Spanish, and Mandarin. The rendered voice shows no stereotype drift. But every model decides the speaker's gender from the content, not the voice. Making the content one step more feminine (masculine -> neutral -> feminine) multiplies the odds of a "female" judgment by 1.7-24. When the content clashes with the voice, the worst model misgenders the speaker in 90% of cases. When they agree, it misgenders in only 2%. The bias thus hides in gender attribution, where fixed-voice evaluation cannot see, and where audits must look as S2S systems increasingly speak for real people.
1 Introduction
The paper separates voice rendering from gender attribution because fixed output voices can conceal stereotype-driven bias. Across five models and three languages, rendered voices show no stereotype drift, but gender judgments follow content rather than voice.
- Motivation: Fixed output voices make voice-drift evaluations unable to establish fairness.A fixed voice has nothing to drift toward a stereotype, so a clean rendering result can be uninformative.
- Research questions: The study asks whether stereotypes shift rendered voice gender and whether gender attribution follows voice or content.The second question evaluates invariance: changing only the topic should not change the speaker-gender judgment.
- Main findings: Across five models and three languages, rendered voices show no stereotype drift, while gender judgments consistently follow content.The output-voice interaction was nonsignificant, whereas content shifted judgments in the same direction across languages.
- Main findings: Misaligned content and voice produce 90% pooled misgendering, compared with 2% when they agree.The pattern identifies stereotype-driven attribution as the failure that fixed-voice evaluation misses.
2 Related Work
Prior bias evaluations largely adapt text benchmarks to speech or remain text-only, leaving this study’s voice–content distinction underexamined. Existing speech evaluations include spoken multiple-choice formats and separate content and acoustic contributions.
- Text bias evaluation: Text bias benchmarks mainly use templates, multiple-choice questions, continuation scoring, or open-ended text generation.These benchmarks are mostly English and operate without an audio voice channel.
- Position of this work: This paper instead separates acoustic voice from stereotyped content in speech-to-speech evaluation.Its design addresses a distinction not captured by benchmarks that operate purely on text or spoken question answering.
- Speech bias evaluation: Speech-model fairness work has largely carried text-style multiple-choice evaluation into spoken prompts.Examples include Spoken StereoSet, VoiceBBQ, and studies of gender-dependent positional artifacts.
3 Method
The method crosses controlled voice and content cues across languages and tasks, first validating the acoustic carrier and then measuring voice rendering and gender attribution separately. A five-task suite distinguishes output-voice drift from judgments that require gendered reference.
- Experimental design: The design crosses English, Spanish, and Mandarin with masculine, neutral, or feminine content and male or female voices.Each language–stereotype pair contains 10 passages, voiced in both genders, for 180 voiced passages.
- Stimulus construction: Long-form first-person passages keep gender out of the text while stereotype intensity is screened and human-checked.Passages are generated around stereotype seeds, rated for intensity, and screened for lexical neutrality, naturalness, pole assignment, and length.
- Transformation tasks: Readback, paraphrase, summarize, translate, and describe vary the model’s freedom to reword while measuring two output channels.First-person tasks provide controls for spontaneous gendering; summarize and describe force third-person gender commitments.
- Measurement setup: TTS serves as the controlled instrument, while end-to-end S2S models hear voiced passages directly.The carrier gate checks that content does not alter the intended acoustic gender before model testing.
4 Experiments
The experiment tests acoustic drift and content-based gender attribution across five S2S models, three languages, and multiple speech tasks. Output voices show no reliable stereotype drift, but gender judgments and pronoun choices consistently follow content more than the speaker's voice.
- Experimental design: Five commercial and open-source S2S models were evaluated across English, Spanish, and Mandarin on voice rendering and gender attribution.The design included readback, summarize, paraphrase, translate, and describe tasks.
- RQ1: voice rendering: The acoustic analysis controlled voice gender with male and female TTS voices and tested content×voice interactions against neutral baselines.The composite femininity index combined pitch, formants, and speaker-gender classifier margins, with each output compared against its input.
- RQ1: voice rendering: No reliable stereotype drift appeared in rendered voices: only 4/40 acoustic fits reached p<.05, and none replicated across task or language.The pooled contrast was −0.018 ± 0.020, while fixed-timbre architectures structurally limited possible drift.
- RQ2: gender attribution: Content femininity multiplied the odds of a “female” judgment for every model, with closed-model ORs of 21.9 and 24.3 versus 1.7–3.4 for open models.Aligned cases had low error, showing that models could access acoustic evidence but often overrode it when content conflicted with voice.
- RQ2: gender attribution: Injected pronouns also followed content stereotypes: P(“she”) rose from masculine to feminine content in every model, with ∆pp ≈0.5–0.6 for Step-Audio-2 and Gemini.Paraphrase and translate injected almost no gendered pronouns, limiting the effect to tasks that force third-person wording.
- Cross-language and model comparison: Content pulled gender judgments toward its stereotype in every language, while the magnitude and model ranking varied by language and architecture.Closed systems showed the largest judgment bias, but the effect was also observed among open models and was not tied to one vendor or training pipeline.
- Limitations: Per-language cells were small, several logistic fits reached perfect separation, and the binary design does not cover non-binary reference.The passages were LLM-generated and LLM-rated, and five models across three languages form a panel rather than a census.
5 Conclusion
The study finds that fixed-voice evaluation misses a bias in gender attribution: every model follows content rather than voice, especially when the two conflict.
- Every model bases stated speaker gender on content rather than voice.
- A clean rendered-voice result is uninformative when the output voice is fixed and cannot drift.
- When content clashes with voice, the worst model misgenders 83–100% of speakers, compared with 2% when the cues agree.The pooled clash rate is 90%.
- Audits should test gender attribution and score invariance to content changes, rather than treating a stable output voice as evidence of fairness.
Ethics Statement
The paper documents stereotype-driven gender inference in S2S systems while cautioning that automatic gender recognition is ethically contested and can harm marginalized people.
- When spoken content contradicts a gender stereotype, models infer speaker gender from content rather than voice across all three tested languages.
- Fixed-voice evaluation cannot detect this failure, which affects captioning, persona selection, and pronoun choice.
- Automatic gender recognition assumes gender is binary and readable from the signal and has harmed transgender and non-binary people.
- The authors measure gender inference as an audit of unreliable deployed decisions, not as an endorsed capability or product.
- Practitioners can test attribution behavior or use voice-preserving models before shipping captions, pronouns, or personas.Suggested mitigations include suppressing unsolicited inference, weighting voice over content priors, and allowing refusal or uncertainty.
A Voice Inventory
The input inventory uses Azure neural voices as controlled carriers, with male and female voices in English, Spanish, and Mandarin that remain stable across content.
- Table 3 uses one male and one female Azure neural voice for each of English, Spanish, and Mandarin.
- All carrier voices are gender-stable and invariant to content, so downstream changes can be attributed to the study factors.
B Model Inventory
The model inventory analyzes each system’s own text channel for gender attribution while holding every system to a single fixed output voice.
- The study analyzes each system’s own text channel for RQ2, with different transcript routes across open checkpoints, gpt-audio, and Gemini.
- All evaluated systems use a single fixed output voice, but closed models generate it natively while open checkpoints use a fixed decoder timbre.
C Language Selection Details
Spanish is selected over French because its gender marking is audible and its pro-drop structure preserves the study’s two-axis contrast. The study applies native-language prompts across five speech tasks in English, Spanish, and Mandarin.
- Language selection: Spanish provides audible gender morphology and pro-drop, unlike French’s largely silent inflection and obligatory subject pronouns.This preserves the contrast between morphological and pronoun-based gender resolution.
- Task and prompt design: The study uses English, Spanish, and Mandarin to test spoken-language models through five native-language tasks.Prompts are written in each audio’s source language; English is excluded from Translate because the mapping is trivial.
- Readback: Readback prompts require word-for-word repetition in the source language without summarizing, paraphrasing, translating, or commentary.The English, Spanish, and Mandarin versions impose the same task constraints in their respective languages.
- Summarize: Summarization prompts require one sentence in the source language, using a third-person gendered pronoun for the speaker.The prompts specify he/she in English, él/ella in Spanish, and 他/她 in Mandarin.
- Paraphrase: Paraphrase prompts require re-expressing the passage in the same language while retaining first-person perspective and information.The corresponding prompts preserve I, yo, and 我 rather than switching viewpoint or language.
- Translate and describe: Translation prompts convert Spanish or Mandarin speech into spoken English, while the study also defines a single-word spoken gender-description task.The description task asks models to output only male/female, hombre/mujer, or 男/女.
E Human Evaluation of the Passages
A trilingual volunteer screened the passages for lexical neutrality, naturalness, stereotype-pole agreement, and information load. The retained passages passed these checks, while the authors caution that ceiling agreement reflects selection rather than rating precision.
- Screening procedure: One near-native trilingual volunteer independently screened every passage against the LLM-assigned ratings.The study reports human–LLM consistency rather than agreement between multiple human annotators.
- Screening criteria: The screening checked lexical gender-neutrality, fluency and naturalness, pole agreement, and information-load suitability.Spanish first-person passages additionally excluded gender-agreeing adjectives, participles, and determiners referring to the speaker.
- Screening outcome: All retained passages were lexically gender-neutral, natural, pole-consistent, and within the target information-load window; none required revision.The passages therefore left gender recoverable only from audio.
- Carrier-gate context: Table 6 evaluates the carrier gate per language using classifier gender flips and within-voice content effects on gender-sensitive measures.Its headline composite passes in all languages, with a lone benign sub-.05 Spanish secondary-channel result.
- Interpretation: Ceiling agreement is expected because passages were selected from an unambiguous stereotype-intensity range, not because the single-rater design demonstrates rating precision.The authors classify single-rater screening as a limitation and a verification gate rather than an inter-annotator reliability study.
- Classifier checks: The two Wav2Vec2 classifiers differ in scale and training corpus, and their pre-softmax logit margins retain graded within-gender information.Audio is downmixed to mono and resampled to 16 kHz before inference.
G Carrier Gate (H0) Details
The carrier gate tests whether TTS content changes the intended acoustic gender before model evaluation. Across languages, voices passed the hard neutrality check and the headline composite showed no content effect.
- Gate criteria: The carrier gate requires no classifier gender flips and no within-voice content effect on gender-sensitive input measures.The tested measures include the femininity composite, mean F0, and classifier logit margins.
- Gate outcome: 0/60 input gender flips occurred in every language on both classifiers, and the headline composite showed no content effect.The composite p-values were .97, .18, and .41 for English, Spanish, and Mandarin.
- Secondary-channel exception: A Spanish primary-classifier margin effect was benign because the feminine–masculine contrast was −0.0009 logits and non-significant.The omnibus effect came from a marginally lower neutral cell, with cell differences no larger than 0.007 logits versus approximately 13 logits of voice separation.
H Supplementary result tables
Supplementary analyses detail describe odds ratios and summarize pronoun behavior. They also flag boundary artifacts in degenerate fits and show that content-driven pronoun slopes can occur across input voices.
- Describe analysis: Table 7 reports describe odds ratios for each content step, with per-language fits that can encounter perfect separation or singularity.The pooled ALL column is the intended readout for the content-step effect.
- Fit caveat: Kimi-Audio’s zero-error Spanish fit is a boundary artifact of a degenerate model fit, not evidence of a reverse content effect.The supplementary result explicitly distinguishes this fitted zero from an opposing bias direction.
- Summarize analysis: Table 8 measures P(“she”) among summaries that inject gendered pronouns and splits content slopes by input voice.Ninety-two percent of summaries inject a pronoun overall, while substantial slopes on both voices indicate content-driven pronoun choice.