Source-linked AI summary

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

Beatrice Savoldi, Sara Papi, Wafa Aissa, Matteo Negri, Luisa Bentivogli

arXiv:2606.26968v1cs.CL

TL;DR

Speech-model safety and fairness remain sparsely assessed beyond English and under naturalistic conditions. RedVox benchmarks eight models across five languages using natural voices, finding vulnerabilities that worsen outside English and intensify with spoken inputs.

  • Problem

    Safety and fairness reporting for speech models is sparse and English-centric, leaving multilingual and naturalistic conditions insufficiently assessed.

  • Method

    The paper surveys speech-model reporting, introduces the natural-voice REDVOX benchmark across five languages, and evaluates eight state-of-the-art speech models.

  • Results

    Safety and fairness vulnerabilities persist under naturalistic conditions, systematically worsen in non-English languages, and intensify with spoken input; one model produces harmful or controversial responses in up to 44% of cases.

  • Takeaways & Limitations

    Naturalistic multilingual speech evaluation reveals persistent vulnerabilities and participant privacy challenges that remain underexplored in speech safety research.

  • Takeaways & Limitations

    The evaluation targets naturalistic harmful requests in simplified single-turn scenarios rather than deliberate jailbreaks or multi-turn interactions.

Abstract

from arXiv · show

Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and under naturalistic conditions remain understudied. We survey safety reporting practices across state-of-the-art speech model releases, finding that only 8% document any multilingual analysis. To address this gap, we introduce RedVox, a multilingual safety and fairness benchmark for audio and speech built on real voices, covering unsafe and unfair stereotypical requests across five languages (English, French, Italian, Spanish, and German). Evaluating eight state-of-the-art models, we find that vulnerabilities persist even under non-adversarial conditions, worsen in non-English languages, and are amplified when the request comes from a spoken input. Finally, by surveying the participants who contributed to RedVox, we document the unique personal and privacy challenges of collecting speech data with human participants, pointing to broader sociotechnical challenges in naturalistic speech safety research.

1 Introduction

Speech-capable models expand human-AI interaction but introduce heightened safety and fairness risks, especially beyond English and under naturalistic spoken conditions. RedVox addresses these gaps with multilingual evaluation of real voices and shows vulnerabilities across eight state-of-the-art speech models.

  • Motivation: Speech-capable models process spoken input natively, preserving paralinguistic features while reducing latency and supporting more accessible, hands-free interaction.These capabilities broaden voice-based human-AI interaction beyond traditional text interfaces.
  • Motivation: Audio-rich interaction expands the vulnerability surface, making safety and fairness urgent as speech systems become more widely deployed.Prior work has identified stereotypes, toxicity, value-alignment failures, and multimodal vulnerabilities.
  • Research gap: Existing speech safety and fairness analyses are overwhelmingly English-centric and rely on synthetic voices, leaving blind spots under naturalistic, multilingual conditions.These limitations raise questions about how AI benefits and risks are distributed across languages and communities.
  • Contributions: 8% of surveyed speech model releases report any evaluation beyond English.The review identifies sparse and English-centric safety-vulnerability documentation.
  • Contributions: RedVox is introduced as a speech and audio benchmark spanning multiple languages and real-voice conditions to evaluate unsafe and unfair behavior.The benchmark is designed to address the lack of multilingual and naturalistic speech safety and fairness analysis.
  • Findings: Eight state-of-the-art speech models exhibit safety and fairness vulnerabilities under naturalistic, non-challenging conditions, with risks worsening in non-English languages and spoken input acting as an additional stressor.The findings indicate that spoken voice elicits vulnerabilities beyond those triggered by text alone.

2 On Reporting Speech Vulnerabilities

A review of 38 open-ended, instruction-following speech models finds substantial gaps in safety reporting: only 11 document safety evaluation, and reported evaluations are overwhelmingly English-only. The review covers SpeechLLMs and OmniLLMs, using a broad safety notion that includes fairness, toxicity, and adversarial robustness.

  • Scope and methodology: The review includes SpeechLLMs and OmniLLMs capable of speech input in open-ended generation settings where risks are most consequential and user-facing.Task-specific models such as Whisper and Canary are excluded.
  • Scope and methodology: Safety is defined broadly to encompass fairness, toxicity, and adversarial robustness.For each model, the review annotates the languages and modalities evaluated and the evaluation methodology.
  • Reporting coverage: Only 11 of 38 examined speech models document safety evaluation.The review covers model cards, technical reports, and accompanying papers, including four proprietary models.
  • Language coverage: Reported evaluations are overwhelmingly English-only, with SeaLLMs-Audio and bilingual (en/ko) Raon Speech as exceptions.Three additional models address safety without specifying the evaluation languages.
  • Language coverage: Phi-4-Multimodal has the broadest reported coverage, evaluating speech inputs across 8 languages.The review selects the most recent model version within each family and annotates languages, modalities, and evaluation methodology.

BENCHMARK DETAILS

REDVOX evaluates safety and fairness through two request types that combine harmful content with speech, text, or distracting audio, using an increasing-severity response workflow. Its design addresses sparse, English-centric multilingual assessment in speech models.

  • Request Types: REDVOX includes two request types targeting safety and fairness: spoken harmful content with textual follow-up, and text-only harmful content paired with distracting audio.Request Type I is Speech; Request Type II is Audio.
  • Evaluation Workflow: The evaluation workflow assesses model responses on an increasing severity scale.
  • Motivation: REDVOX is motivated by sparse, English-centric speech-model safety reporting and limited systematic multilingual assessment.

3 The REDVOX Dataset

REDVOX is a multilingual, multimodal benchmark testing speech-model vulnerabilities across safety and fairness, using parallel adaptations of established textual resources and naturalistic participant-recorded inputs. Its public release contains 3,414 entries from 26 voices, while robustness analyses show that this subsample preserves full-dataset conclusions.

  • Benchmark scope: REDVOX separates safety requests about criminal, hazardous, or violent acts from fairness requests involving stereotypical generalizations about social and identity groups.The authors use these as high-level umbrella categories to distinguish stereotypical content, which is often absent from safety policies and frameworks.
  • Collection and release: 52 researchers from 7 European institutions collected 6,118 entries totaling almost 10 hours of audio and speech, but public release consent reduced REDVOX to 26 voices and 3,414 entries.Participants contributed voluntarily, with language-specific participation of 18 English, 10 German, 9 Italian, 8 French, and 8 Spanish individuals.
  • Data sources: 181 parallel SHADES stereotypes and 350 M-ALERT safety instances provide the benchmark’s multilingual textual foundation across English, German, Spanish, French, and Italian.The 350 M-ALERT instances are equally distributed across Criminal Planning, Substances, Sexual Content, Suicide & Self-Harm, and Guns & Illegal Weapons.
  • Multimodal construction: Each entry becomes a Speech request with harmful content vocalized plus follow-up text, or an Audio request with harmful text paired with silence or background noise.Audio conditions use silence, ambient noise, or babble noise, while the speech+text design reduces unrelated outputs caused by speech-request misunderstanding.
  • Robustness: ρ = 0.98, p < .01 for Spearman’s correlation on safe-response percentages, while Cramér’s V ≤ .09, indicate that the released subsample preserves model rankings and categorical evaluation patterns.These robustness checks compare the public subset against the entirety of collected data.

4 Experimental Setting

The evaluation measures both harmfulness and fairness-related behavior and whether models understand requests, using a two-step GPT-5.5 judge validated against manual annotations. It covers eight state-of-the-art systems supporting five languages.

  • Evaluation dimensions: The assessment evaluates responses along safety and fairness, plus relatedness: whether the model understood the input request.Relatedness distinguishes genuine safe behavior from hallucination or repetition of the input.
  • Response taxonomy: Responses are categorized as Safe, Safe by Accident, Controversial, or Unsafe based on harmfulness and relatedness.Safe-by-Accident responses are harmless because the model misunderstood the request, whereas Unsafe responses endorse, support, or comply with harmful requests.
  • LLM-as-a-judge: GPT-5.5 performs separate safety/fairness and relatedness labeling steps with language-specific few-shot exemplars, then recombines the labels into the taxonomy.The judge receives both the input request and model output, unlike existing multilingual moderation models that lack relatedness assessment.
  • Judge validation: 0.94 Macro F1 is achieved on relatedness, while safety and fairness scores reach 0.89 for binary labels and 0.79 for ternary labels, outperforming Qwen3Guard.The judge was validated against 250 manually annotated entries sampled across languages, models, and dataset variables.
  • Experimental systems: The experiments evaluate 8 state-of-the-art systems supporting five languages, including 5 freely available HuggingFace models and 3 proprietary models.The open models are Qwen2-Audio, Phi4Multimodal, Voxtral, Qwen3-Omni, and Gemma 4; proprietary systems include Gemini 3.1 Flash-Lite, Gemini 3.1 Pro-Preview, and GPT-realtime-2.

5 Results

RedVox reveals substantial safety and fairness risks, with unsafe behavior generally higher in non-English languages and in spoken-input settings. Proprietary models usually show the lowest unsafe response rates, while stereotypes frequently produce controversial rather than clearly refused responses.

  • Overall model behavior: Proprietary models systematically show the lowest unsafe response rates (≤3.1%), except Qwen3-Omni at 3.4%; Voxtral reaches roughly 1 in 4 fully harmful responses.Across nearly all models, controversial responses outnumber fully unsafe ones, indicating a substantial grey area.
  • Multilingual results: 5.1% of English requests are unsafe versus 10.0% in other languages (∆96%), with the gap systematic across models except Gemini-Pro3.1 Pro on Spanish.Open models drive much of the gap: Voxtral reaches 28% unsafe responses in Spanish and French, a +15% increase from English, versus 13.0% for non-English requests overall (∆26%).
  • Vulnerability types: Fairness requests involving stereotypes have the highest controversial-response rates, because models engage with nuanced, benign-sounding generalizations rather than issuing clear refusals.Qualitative inspection found models endorsing positive stereotypes such as “Germans are hard worker”.
  • Multimodality: Speech is the most vulnerable input setting, with controversial-or-unsafe response rates reaching 10–44%; open models broadly worsen from text to audio to speech.Even silence or noise paired with harmful requests increases unsafe rates, suggesting audio presence stresses non-proprietary models independently of semantic content.

6 Discussion

REDVOX shows that speech-model safety and fairness vulnerabilities persist in naturalistic settings, widen across non-English languages, and remain evident even in safety-aligned open-weight models. The discussion also identifies multilingual speech-data scarcity and distinctive psychological, ethical, and privacy challenges in collecting human voices.

  • Model vulnerabilities: REDVOX finds safety and fairness vulnerabilities in speech models even under naturalistic, non-adversarial conditions, including in open-weight models that explicitly report safety alignment.The discussion describes these vulnerabilities as a present concern, especially across many recent open-weight models.
  • Model vulnerabilities: Gaps widen in non-English languages, likely reflecting scarce multilingual resources for model alignment, while multimodal open-weight models handle textual requests better than spoken inputs.The discussion connects the multilingual gap to alignment-resource scarcity and echoes prior findings on limits of cross-modal safety transfer.
  • Data collection challenges: 52 participants voluntarily contributed to REDVOX, but only 50% consented to public data release, illustrating the resource bottleneck and challenges of collecting multilingual human speech.The discussion states that multilingual speech data for safety and fairness evaluation is scarce and that voice collection raises challenges without a direct analogue in text red teaming.
  • Data collection challenges: 56.4% of participants felt pronouncing harmful prompts made them more personally responsible, while participants also feared voice identification.Voice elicitation produced lower comfort than text even for generic release, with a stark drop for harmful recordings; the discussion calls for privacy-preserving collection and attention to speech labor ethics.

7 Related Work

Prior work finds that speech-capable models are more vulnerable than their language-model backbones, with multimodal inputs, audio perturbations, and delivery style affecting safety responses. However, existing evaluations often rely on unreleased datasets and focus solely on safety, motivating REDVOX’s distinctive scope.

  • Prior findings: SpeechLLMs are more vulnerable to attack than their LLM backbones, and multimodal inputs further weaken safety alignment.These findings are reported by Yang et al. (2025), Lu et al. (2025b), Peng et al. (2026), and Pan et al. (2025).
  • Comparison with related resources: Table 2 compares speech vulnerability resources across fairness, safety, modality, language coverage, and voice type.The table defines T as Text, S as Speech or Audio, and distinguishes natural from synthesized voice.
  • Prior findings: Optimized audio perturbations and delivery style condition model responses in speech vulnerability evaluations.Song et al. (2025) studies optimized audio perturbations, while Yu et al. (2026) and Cheng et al. (2026) examine delivery style.
  • Limitations of prior work: Many existing datasets are not publicly released, and some studies focus solely on safety rather than broader evaluation dimensions.The cited studies include Aloufi et al. (2026), Yang et al. (2025), Song et al. (2025), and Yu et al. (2026).

8 Conclusions

REDVOX is a multilingual speech safety and fairness benchmark built on naturalistic voices, used to evaluate eight state-of-the-art speech models across five languages. The results show persistent safety vulnerabilities under non-adversarial conditions, systematically worse performance in non-English languages, and consistent stress from audio and speech relative to text, while participant questionnaires reveal psychological and methodological challenges in collecting harmful speech data.

  • REDVOX is a multilingual speech safety and fairness benchmark built on naturalistic voices.
  • Eight state-of-the-art speech models were evaluated across five languages using REDVOX.
  • Safety vulnerabilities persist under non-adversarial conditions, are systematically worse in non-English languages, and are consistently stressed by audio and speech relative to text.
  • Participant questionnaires documented psychological and methodological challenges in collecting harmful speech data.These challenges represent an underexplored bottleneck for the field.

Limitations

REDVOX’s limitations concern language coverage, evaluation conditions, speech modalities, sample size, semantic scope, and privacy infrastructure. These constraints limit generalizability, statistical power, and the scenarios assessed.

  • Language coverage: REDVOX covers five high-resource Indo-European languages, so findings may not generalize to typologically distinct or underserved languages.Safety alignment tends to be weaker in underserved languages.
  • Evaluation conditions: The benchmark targets naturalistic harmful requests rather than deliberate jailbreaking, potentially missing higher harmful-output rates elicited by optimized attacks.This setting is intended to surface vulnerabilities in unoptimized interactions that ordinary users might naturally produce.
  • Speech modality: REDVOX excludes requests delivered entirely through speech, leaving one scenario unexplored because speech-only instructions may reduce comprehension and introduce unrelated responses.Participant availability constrained the choice of evaluation conditions.
  • Scope and privacy infrastructure: The study focuses on semantic content rather than paralinguistic speaker features, while unreleased full-data evaluation and stronger privacy guarantees would require sustained funding.A platform-based infrastructure could assess models without releasing raw voice data and enable use of the full REDVOX.

Ethical Considerations · A Model cards · B REDVOX Details

The paper describes safeguards for voluntary speech-data participation, controlled research use, and gated REDVOX access, while documenting how model safety and fairness reporting was systematically assessed. It also points readers to participant-distribution details, REDVOX examples, and participation guidelines.

  • Ethical Considerations: Participation was voluntary and non-compensated, with separate consent for participation and public data release and withdrawal permitted at any time.Participants could contribute without agreeing to public release of their data.
  • Ethical Considerations: Participants were informed about the activity’s nature, scope, objectives, and risks, with ongoing communication, breaks, and segmented sessions encouraged.Participants were also told they could skip requests they did not want to generate.
  • Ethical Considerations: REDVOX is intended solely for research evaluating speech-model safety and fairness, with speech files decoupled from direct participant identifiers.The dataset’s stated purpose is research use rather than general deployment.
  • Ethical Considerations: REDVOX uses gated access requiring acceptance of a customized license and binding terms governing the included data.The access model is presented as a mitigation for residual data-related risks.
  • A Model cards: The model-card review covered open-weight HuggingFace models and proprietary speech-input, instruction-following models.The review examined each model’s technical report and model card.
  • A Model cards: The documentation search systematically checked for “red teaming,” “toxicity,” “safety,” “jailbreak,” “responsible,” “security,” “fairness,” and “attack.”The results are reported in Table 12.
  • B REDVOX Details: REDVOX details include examples in Table 6 and a participant-pool breakdown in Table 3.Table 3 reports self-declared gender by language and native versus non-native English speakers; participants in other languages were native speakers.
  • B REDVOX Details: The activity participation guidelines are available in the project repository.The passage provides the repository link for these guidelines.

B.1 REDVOX Full and Released Sets … E AI Use Statement

The released REDVOX subset preserves the full set’s model rankings and largely preserves label distributions, while the paper details inference, computation, evaluation validation, complementary analyses, and AI-assisted preparation.

  • B.1 REDVOX Full and Released Sets: REDVOX’s released subset preserves relative model safety rankings, with label-distribution differences generally negligible across languages and dimensions.Spanish is the only language with significant deviations in both safety and relatedness; Cramér’s V remains ≤.09, indicating minor label imbalances rather than a systematic representational gap.
  • C.1 Inference Details: Inference uses HuggingFace transformers with model-specific versions and default decoding parameters, while proprietary models use standard Google and OpenAI APIs.Inference code and outputs will be released open source upon paper acceptance.
  • C.2 Computational Details: 165h, 155h, 83h, 66h, and 3h were required to process the released benchmark for Gemma, Qwen3-Omni, Voxtral, Qwen2-Audio, and Phi4-Multimodal.The corresponding full-benchmark times were 296h, 278h, 149h, 119h, and 5.5h; proprietary-model costs were also reported for both sets.
  • C.3.1 Evaluation Testbed: 250 manually sampled outputs were independently annotated by two native-speaker evaluators per language across relatedness, safety/fairness, and refusal.The sample was equally stratified across anonymized models, five languages, and two harmful request conditions.
  • D Complementary results: Complementary results report outcomes by model, language, input type, speaker nativeness, gender, and participant discomfort across vulnerability categories.These analyses include native-versus-non-native English speech requests and gender comparisons across speech inputs.
  • E AI Use Statement: The authors used coding tools to streamline visual-artifact generation and writing assistants to polish parts of the manuscript.This disclosure appears in the paper’s AI Use Statement.
Loading 2606.26968v1…