Source-linked AI summary
Language-Specific Gaps in AI Safety Training Datasets
Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru
TL;DR
Multilingual safety coverage claims may not reflect the rigor of evaluation for each individual language. This paper audits language slices across Hausa, Swahili, and French and finds recurring, resource-linked gaps whose patterns differ across audit dimensions.
Problem
Collection-level multilingual safety claims provide limited evidence about the rigor of evaluation for individual non-English languages.
Method
The paper audits language slices independently across 20 datasets and evaluates provenance, annotation, access, taxonomy coverage, reuse, and quantified quality.
Results
Gap patterns were neither uniformly monotonic nor random: provenance and access worsened steadily as resource level fell, while annotation, taxonomy, and reuse followed different patterns.
Takeaways & Limitations
The findings support auditing language slices independently rather than treating collection-level multilingual coverage as sufficient evidence of comparable safety evaluation.
Takeaways & Limitations
The full audit was conducted by a single annotator per slice without independent double-coding or measured inter-rater reliability.
Abstract
from arXiv · showhide
Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage claims frequently do not survive inspection at the level of an individual language. Auditing 21 resources across 25 language slices, of which 20 count as datasets under our counting rules, spanning three languages chosen to represent low- (Hausa), mid- (Swahili), and high-resource (French) tiers, we find that gaps in provenance, annotation reliability, access, harm-taxonomy coverage, and data reuse recur in patterns that partially, but not fully, track resource level. Using a controlled within-pipeline comparison, we show a Hausa-language slice falling below its own paper's translation-quality acceptance threshold while the same pipeline's Swahili output clears the same bar comfortably; this is evidence that these gaps are measurable and addressable, not inherent. We further show that self-harm and sexual-content categories have no native-language coverage in either African-language tier we studied, a total rather than gradated gap that a purely resource-level account does not predict. We connect these findings to a documented, persistent asymmetry in multilingual jailbreak robustness (single-turn attacks largely mitigated, multi-turn attacks still effective), arguing that this asymmetry is structurally consistent with where our audit finds training and evaluation data thinnest. We contribute a reusable slice-level audit methodology, a cross-tier empirical comparison, and concrete recommendations for dataset creators, model providers, and venues aiming to make ``multilingual coverage'' claims verifiable rather than merely stated. Dataset: https://huggingface.co/datasets/ChialukaOnuoha/safety-slice-audit
1 Introduction
The paper argues that collection-level multilingual safety claims often overstate the rigor of individual language slices. Through a cross-tier audit and controlled case study, it links recurring dataset gaps to persistent low-resource-language jailbreak vulnerabilities and proposes slice-level verification practices.
- Motivation and scope: The audit examines 21 resources across 25 language slices spanning Hausa, Swahili, and French, showing that aggregate multilingual coverage often fails to represent individual-language rigor.The study includes 20 datasets under its counting rules and compares low-, mid-, and high-resource tiers.
- Audit findings: Gaps cluster in recurring forms, including overstated provenance, missing inter-annotator agreement, and restrictive or poorly documented access conditions.The paper distinguishes data-quality problems from artifact-accessibility problems rather than treating collection-level statistics as representative.
- Audit findings: Most gap types and severities correlate with resource level, but the paper emphasizes that this relationship is incomplete and does not explain every measured dimension.The comparative audit yields a typology of recurring gap types across Hausa, Swahili, and French.
- Downstream safety implications: Approximately 83% of Swahili jailbreak attempts produced unsafe responses in a 2023-era model, while later multi-turn attacks in low-resource languages still achieved harmful-response rates of roughly 42% to 71%.The cited replication found single-turn translation-based attacks largely closed but multi-turn conversational jailbreaks remained highly effective.
- Contributions: The paper contributes a reusable slice-level audit methodology, a controlled within-pipeline quality case study, downstream-safety analysis, and recommendations for verifiable multilingual coverage claims.Its recommendations target dataset creators, model providers, and venues.
2 Related Work
Prior work provides frameworks for dataset documentation, resource-tier classification, multilingual harmful-content datasets, and uneven cross-language safety alignment. This paper’s gap is a systematic, slice-level audit linking evidentiary weaknesses to persistent multi-turn jailbreak vulnerability.
- Dataset documentation and transparency: Dataset documentation frameworks require evaluating provenance, collection, annotation, and intended use rather than relying on headline descriptions.The paper builds on Datasheets for Datasets and Data Statements for NLP; some audited datasets explicitly adopt data statements.
- Resource-level taxonomies: Resource-level taxonomies classify languages by available labeled and unlabeled data and document NLP research’s concentration in high-resource languages.Some audited datasets operationalize resource level using the language’s share of the CommonCrawl corpus.
- Multilingual and low-resource hate speech / offensive-content resources: African-language hate-speech and offensive-content research includes native-reviewed datasets and community-embedded collection designed to reduce annotator demographic mismatch.Examples include Hausa Offensive Content and XTREMESPEECH, which recruits local fact-checkers rather than external annotators.
- Multilingual jailbreak and safety-training-gap literature: Low-resource-language jailbreak susceptibility has been reported as roughly three times that of high-resource languages under translation-based attacks, with related effects shown against GPT-4.More recent work finds single-turn translation attacks substantially mitigated but multi-turn conversational attacks still effective; translation quality measured by BERTScore, METEOR, and BLEU explains residual cross-language variance in attack success.
- Government and multi-institution safety evaluation efforts: Joint government and multi-institution evaluations now include non-English safety domains, including a Kenya AISI-contributed Kiswahili component on fraud and sensitive-information leakage.The paper treats this resource as a boundary case because its harm taxonomy falls outside the hate-speech, harassment, self-harm, extremism, sexual-content, and misinformation framework used elsewhere.
- Gap in the literature: No prior work systematically audits language-slice evidentiary quality across multilingual AI safety datasets with a shared resource-tier-stratified methodology while linking gaps to residual multi-turn jailbreak vulnerability.This combined methodological and downstream-safety connection defines the literature gap addressed by the paper.
3 Scope and Definitions
The study focuses on content safety and jailbreak safety, using a six-category harm framework while excluding agentic/operational safety from its primary analysis. It compares Hausa, Swahili, and French across resource tiers and defines language-specific gaps as divergences between collection-level claims and slice-level evidence.
- Scope: The study covers content safety and adversarial/jailbreak safety, including hate, harassment, self-harm, extremism, sexual content, and misinformation.Agentic/operational safety is excluded from the primary analysis.
- Analytical framework: The six-category framework is adapted from the Trust and Safety Professional Association taxonomy and serves as the study’s primary analytical lens.Each dataset is assessed against all six categories, with resource tier used as the cross-cutting comparison.
- Resource tiers: The audit compares low-resource Hausa, mid-resource Swahili, and high-resource French.The languages were selected for non-trivial representation across multiple datasets, enabling within-language and cross-tier comparisons.
- Resource tiers: The low/mid/high stratification is contested because some CommonCrawl-share thresholds classify Swahili as low-resource, but the study retains it to reflect practical dataset availability.The authors note that Swahili’s audited outcomes are intermediate between Hausa’s and French’s on most measured gap dimensions.
- Gap definition: A gap is any divergence between a dataset’s collection-level coverage claim and the evidentiary quality available for a specific language slice.The audit examines provenance, annotation/agreement, access/licensing, taxonomy coverage, reuse/double-counting, and quantified quality gaps.
4 Audit Methodology
The audit treats each language slice—not the dataset collection—as the unit of analysis, applying a common 31-field schema and source-checking protocol. It makes multilingual coverage comparable through explicit taxonomy mapping, inclusion rules, severity criteria, and reproducible rationales.
- Unit of analysis: Each language within a dataset receives an independent audit row, so slice-level conclusions and counts are not conflated with dataset-level totals.A dataset covering Hausa and Swahili therefore yields two separately assessed rows.
- Audit instrument: 31 fields are applied identically to every language slice across seven field families, with each field carrying a rationale.The schema was iteratively extended from Datasheets for Datasets to include provenance-depth and usability fields.
- Verification protocol: Primary-source checking prioritizes the paper’s per-language tables for data-making claims and the live release for availability claims.The protocol caught language-assignment, provenance, release-state, and attribution errors that abstracts and secondary descriptions missed.
- Taxonomy mapping: Taxonomy mapping uses source-defined label meanings, records labels spanning multiple categories against each applicable category, and excludes labels outside the six-category framework.The mapping is documented judgment rather than direct measurement.
- Scoring criteria: Severity increases from native authorship toward unreviewed machine translation or synthetic generation, and from open access toward unreleased resources.Annotation severity likewise rises when slice-level agreement is unreported, only aggregate, or below conventional thresholds.
- Reliability and limitations: Assignments were made by one annotator per slice, with written source-based rationales and joint review of flagged rows to support reproducibility and contestability.Single-annotator assignment remains a stated limitation.
5 Corpus Under Study
The audit covers 25 language slices from 21 resources, resolving to 20 datasets after exclusions and counting rules. Its corpus is uneven in provenance and format: Swahili has more slices than French largely because of repeated evaluations, while African-language data remains concentrated in Twitter-derived, Latin-script resources.
- Corpus composition: 25 language slices from 21 resources resolve to 20 datasets; Ubisoft’s ToxBuster slice is excluded because its data is not public.One resource is a derivative re-test rather than a dataset, and 19 datasets fall within the defined safety scope.
- Corpus composition: Swahili has more audited slices than French, but its count is inflated by re-annotations and repeated evaluations of shared tweet and prompt pools.French slices are largely independent resources, so slice counts do not directly measure coverage quality.
- Provenance and harm domains: Native provenance is not comparable across tiers: Hausa native slices are mostly hate or offensive-speech resources, while Swahili native slices span hate, offensive speech, political misinformation, and incitement.AfriSenti supplies native Hausa and Swahili data without harm labels, and Hausa LSR encodes individually targeted physical violence outside the audit’s extremism category.
- Format: The corpus is dominated by single-turn prompt-only and labelled-item resources; multi-turn data comes from one synthetic pipeline for Hausa and Swahili and one French guard-training corpus.The paper relates this format distribution to literature reporting that single-turn translation-based jailbreaks are substantially mitigated while multi-turn conversational attacks remain effective.
- Domains: Every native Hausa and Swahili slice except one draws on Twitter or X, whereas French is the only tier represented by materially different domains such as Wikipedia talk pages and web-crawl text.One Hausa dataset supplements Twitter with Facebook material collected using dictionary guidance.
- Script and language use: Every audited slice uses Latin script, with no Ajami coverage; code-switching is explicitly modelled in four slices.The code-switching slices include one Hausa-English benchmark component and three Swahili-English resources.
6 Findings … 6.3 Provenance gaps
The audit finds multilingual safety-data gaps that vary by category and resource tier, with provenance and access worsening as resources decline. A controlled comparison shows measurable translation-quality inequality, while provenance differences across Hausa, Swahili, and French reveal distinct native-versus-translated coverage patterns.
- 6 Findings: Findings are organized by gap category, using resource tier as the cross-cutting comparison and testing whether gap severity tracks resource level.The controlled within-pipeline comparison comes first because it calibrates interpretation of later categories, while §6.9 addresses the paper’s central question.
- 6.1 Overview: gap incidence by tier: Two categories—provenance and access and licensing—worsen steadily as resource level falls, while other categories show nonmonotonic or concentrated patterns.Taxonomy coverage is worst in the low tier but levels between mid and high; annotation and agreement is worst at mid tier; reuse and double-counting concentrates by resource level.
- 6.2 Quantified quality gaps: a controlled within-pipeline comparison: The controlled pipeline holds models, translation, validation, and acceptance scoring constant, isolating translation and validation quality from dataset-origin and annotation confounds.The same expert-written English seed, models, machine-translation stage, native-validation protocol, and automatic scoring were used for Hausa and Swahili.
- 6.2 Quantified quality gaps: a controlled within-pipeline comparison: 66.37 is the Hausa policy-material score, below the authors’ acceptance threshold of 70, whereas Swahili scores 93.30 on policy material and 96.99 on transcripts.Hausa scores 93.31 on transcripts, showing that the deficit is localized to structured policy material rather than a general Hausa translation failure.
- 6.2 Quantified quality gaps: a controlled within-pipeline comparison: The quality deficit is especially probative because it is published by the authors, judged against their own threshold, and produced despite nominally equal validation effort.One native validator reviewed twenty sampled pairs for each language, yet the equal treatment yielded unequal reliability.
- 6.3 Provenance gaps: Provenance gaps are the most frequent category and stratify most clearly by resource level when crossed with harm domain rather than counted in aggregate.This framing shows why aggregate provenance counts can obscure language- and domain-specific differences.
- 6.3 Provenance gaps: Hausa has narrow native provenance, Swahili extends native coverage beyond hate and offensive speech, and French has the mildest but still incomplete provenance gap.Hausa’s native slices are concentrated in hate or offensive speech, Swahili spans political misinformation and partly incitement, and French combines native red-teaming and toxicity data with translated and synthetic breadth.
6.4 Annotation and agreement gaps · 6.5 Access and licensing gaps
Annotation agreement is sparsely reported and uneven, with strong Hausa results alongside the audit’s weakest agreement in Swahili. Access, release fidelity, and documentation problems concentrate especially in low-resource slices, creating measurable barriers to verification.
- 6.4 Annotation and agreement gaps: Only 6 of 25 slices report language-specific numeric inter-annotator agreement, including AfriHate Hausa κ = 0.75 and Swahili κ = 0.55.Other reported values include AfriSenti Hausa κ = 0.66, XTREMESPEECH Swahili κ = 0.13, Onyango et al. Swahili κ rising from 0.435 to 0.552, and PolyGuard French Krippendorff’s α ≈0.47.
- 6.4 Annotation and agreement gaps: Reported agreement does not track resource level straightforwardly: Hausa reaches κ = 0.75 and κ = 0.66, while XTREMESPEECH’s Kenya slice records κ = 0.13.The Kenya slice belongs to a methodologically careful resource that recruits local fact-checkers to address annotator-comprehension challenges.
- 6.4 Annotation and agreement gaps: The audit interprets low agreement primarily as evidence that taxonomy boundaries do not match distinctions shared by the annotator community.Agreement is described as highest in taxonomies designed by and for the relevant language community and lowest elsewhere.
- 6.4 Annotation and agreement gaps: Unreported agreement can reflect ethical limits on redundant exposure to abusive material, but those limits leave small-annotator-pool languages least verified.The paper treats this as locally defensible while identifying its broader distributional consequence as inequitable for least-resourced slices.
- 6.5 Access and licensing gaps: For Hausa, NaijaOffens was announced but never released, HOC is request-only without a stated license or reported size, and TukaBench releases 300 items against 986 described.The TukaBench discrepancy reflects two AfriJail components that have not been uploaded.
- 6.5 Access and licensing gaps: Paper-versus-release mismatch is a distinct failure mode occurring only for the low-resource language, because repositories can exist while artifacts still differ from paper descriptions.Detecting this mismatch requires opening releases and counting contents, which collection-level checks skip.
- 6.5 Access and licensing gaps: Documentation is also incomplete: fully public annotation guidelines are a minority, partial guidelines are common, and absent guidelines plus missing size information concentrate in Hausa resources.HOC’s missing size was verified across three paper versions and recorded as a documentation finding.
6.6 Taxonomy-coverage gaps · 6.7 Reuse and double-counting gaps
Taxonomy coverage does not follow a clean resource-level gradient: French and Swahili each cover four categories, Hausa two, while African-language slices lack native self-harm and sexual-content coverage. Reuse further compresses apparent corpus diversity, inflating Swahili volume and French methodological independence while creating opportunities to study taxonomy effects.
- 6.6 Taxonomy-coverage gaps: 4 categories are covered natively in French and Swahili, versus 2 in Hausa, so the predicted resource-level gradient is absent between high and mid tiers.Only the low tier is clearly behind.
- 6.6 Taxonomy-coverage gaps: No native self-harm or sexual-content coverage exists in either African language, and Hausa self-harm lacks coverage of any kind.The Hausa self-harm gap includes native, translated, transcreated, and synthetic resources.
- 6.6 Taxonomy-coverage gaps: Misinformation has native Swahili coverage but no native French coverage, showing that local research and civic priorities can reverse resource-level expectations.The Swahili source is an election-driven code-switched corpus.
- 6.6 Taxonomy-coverage gaps: Dominant taxonomies originate in English-language frameworks and are commonly inherited wholesale by translated resources, while local harm conceptualization remains a minority practice.The audit identifies jailbreak, hazard, safety, and toxicity taxonomies among the dominant imported frameworks.
- 6.7 Reuse and double-counting gaps: Swahili’s apparent corpus diversity is compressed because AfriHate re-annotates three earlier corpora, making slice counts materially exceed underlying data diversity.The cited predecessors include AfriSenti’s negative class, PolitiKweli, and Hate_Speech_Kenya.
- 6.7 Reuse and double-counting gaps: French benchmarks also share English red-teaming prompts and toxicity seeds, so apparent cross-benchmark corroboration partly reflects translated common source material.This reuse inflates apparent methodological independence rather than data volume.
- 6.7 Reuse and double-counting gaps: Summing Swahili items across the corpus overstates the underlying pool, and evaluating a model on two reused resources is not independent evaluation twice.The consequences follow directly from the documented re-annotation chains.
- 6.7 Reuse and double-counting gaps: Re-annotation chains provide a natural experiment: identical items labeled under different taxonomies can reveal how much harm judgments depend on taxonomy rather than text.The paper treats this both as an opportunity and as a hazard.
6.8 The boundary case: safety coverage outside the framework · 6.9 Does gap severity track resource level?
A native-reviewed Kiswahili safety resource falls outside the paper’s content-safety framework, illustrating why collection-level coverage counts can mislead. Across six gap categories, resource-level gradients are mixed: some worsen toward Hausa, while others invert or follow research-community mechanisms.
- 6.8 The boundary case: safety coverage outside the framework: A native-reviewed Kiswahili resource covers agentic harms, fraud, and sensitive-information leakage through multi-turn tool-use trajectories, but none of its categories map to the framework.It is therefore recorded as out-of-framework and excluded from Table 6.
- 6.8 The boundary case: safety coverage outside the framework: Collection-level recognition of Kiswahili safety evaluation can falsely imply that the paper’s content-safety and jailbreak gaps were assessed.The two facts are compatible because the resource’s coverage leaves those documented gaps untouched.
- 6.9 Does gap severity track resource level?: Gap severity tracks resource level cleanly for two categories, leaves low clearly worst but high and mid level for a third, and varies or inverts across the remaining categories.The paper therefore gives a qualified rather than unqualified answer to the central question.
- 6.9 Does gap severity track resource level?: Native provenance, access, and licensing worsen monotonically from French through Swahili to Hausa, with serious use-preventing access failures occurring only for Hausa.Taxonomy coverage is four categories for French, four for Swahili, and two for Hausa, leaving the low tier behind while high and mid are level.
- 6.9 Does gap severity track resource level?: Agreement quality is worst at the mid tier, while native misinformation coverage exists for Swahili but not French.These inversions reflect research-community priorities and capacity rather than language web-crawl share.
- 6.9 Does gap severity track resource level?: Reuse and double-counting follows research-community concentration around repeatedly re-annotated foundational corpora rather than resource level.In this corpus, that concentration occurs for Swahili but could occur at any tier.
- 6.9 Does gap severity track resource level?: For Hausa, narrow harm coverage compounds with unreleased, unlicensed, undocumented, and translated material that falls below its own pipeline’s quality bar.Collection-level coverage claims conceal these combined deficits.
7 Discussion
The audit argues that multilingual safety-data gaps are shaped by research-community structure and harm priorities, not resource level alone. Its central methodological warning is that narrow coverage, weak verification, restricted access, and marginal translation quality can co-occur within the same language slice.
- Resource level and community structure: Gap severity tracks resource level cleanly for only two of six categories, while other categories show tier inversions, parity, or stronger links to research-community structure.The paper presents this as more precise than claiming that low-resource languages simply have less safety data.
- Slice-level risk: For Hausa, the narrowest harm-coverage slices are also most likely to be unreleased, unlicensed, undocumented in size, and below their own translation pipeline’s quality bar.The paper identifies this co-occurrence—not any single dataset-card field—as the transferable downstream risk.
- Jailbreak robustness and data gaps: Single-turn translation-based jailbreaks have been substantially mitigated, but multi-turn attacks remain open, corresponding structurally to thin multilingual multi-turn safety resources rather than proving causation.The corpus contained only one in-scope multi-turn resource for either African language, and it was synthetic and machine-translated; the authors lack providers’ training and evaluation data.
- Bounds on the resource-level thesis: Self-harm and sexual content have zero native coverage in both Hausa and Swahili, while Hausa self-harm has no coverage of any kind.This total gap does not distinguish the low and mid tiers and is especially concerning because local idiom, euphemism, and indirection may carry safety signals.
- Bounds on the resource-level thesis: PolygloToxicityPrompts shows that a French, high-resource, natively collected resource can lack human harm annotation, so scale and native provenance do not guarantee verification.The paper cautions against treating “high-resource” as synonymous with “fully verified.”
8 Recommendations
The recommendations require slice-level transparency, precise provenance and reuse reporting, and verification before making language-specific safety claims. They also prioritize native-authored, multi-turn, agreement-verified data collection and venue requirements that enforce these practices.
- Dataset creators: Report each language’s size, provenance, and format in a standard per-language table rather than requiring auditors to reconstruct slices from scattered tables.This recommendation is intended as the default for multilingual collections.
- Dataset creators: Distinguish provenance precisely and audit claims against methodology, while making source-to-subset reuse mappings prominent to prevent double-counting and clarify re-annotation.The audit found two cases where secondary descriptions implied native authorship contradicted by methodology sections.
- Venues and funders: Require per-language reporting tables in submission checklists for dataset papers claiming multilingual coverage, specifically enforcing the slice-versus-collection distinction.This recommendation is framed as analogous to existing reproducibility and data-statement checklists.
- Venues and funders: Fund native-authored, multi-turn, agreement-verified collection because synthetic-and-translate pipelines can fall below their own quality thresholds for lower-resource languages on structurally complex material.Such pipelines may still help fill volume temporarily but are not substitutes for native collection at the frontier of coverage.
Limitations
The audit is limited by single-annotator judgments, a non-exhaustive corpus, contested tier framing, time-bound resource states, and unvalidated taxonomy mappings. Five of six gap categories rely on ordinal judgments rather than independently validated measurements.
- Annotation and measurement: Single-annotator assignments were not independently double-coded across the full corpus, and the audit did not measure inter-rater reliability.Flagged rows were reviewed jointly by the author team, while anchoring criteria were documented for contestability and reproduction.
- Corpus scope: The low/mid/high framing covers only three languages, and Swahili’s tier assignment is contested depending on the operationalization of resource level.Results specifically dependent on Swahili’s tier position should therefore be interpreted cautiously.
- Corpus scope: The corpus was built through citation-following and targeted search rather than a systematic, preregistered review, so it should be treated as a substantial structured sample rather than a census.The authors consider it likely that safety-relevant Hausa- or Swahili-language resources were missed because of visibility bias.
- Annotation and measurement: Five of six gap categories use ordinal judgments rather than independently validated measurements; only quantified quality gaps rely on a metric independent of author judgment.The provenance, agreement, access, taxonomy-coverage, and reuse categories were assigned using observable anchoring criteria, but the four-level scale has no claimed known validity.
- Temporal and taxonomy validity: The findings are time-bound because audited resources and releases can change, while the harm-taxonomy mapping was not independently validated against native-speaker judgment.The authors expect a re-audit, even a year later, to produce a different picture for at least some access and release gaps.
Ethics Statement
The paper audits published resources without creating harmful content or recruiting human subjects, frames its findings as a systemic critique, and releases a detailed audit instrument to support verification. It cautions that documenting data gaps should motivate reinvestment rather than disinvestment.
- Research conduct: The study created no new harmful content, relying on already-published examples with original authors’ content warnings.It did not introduce new offensive examples beyond source materials.
- Research conduct: The paper presents its findings as a systemic critique rather than an assessment of individual researchers.It attributes some gaps, including missing inter-annotator agreement, to funding, annotator availability, and ethical constraints.
- Risks and interpretation: The authors warn that documenting thinner, less-verified low-resource-language datasets could be misused to justify disinvestment instead of reinvestment.They state that every documented gap concerns available data, not whether closing it is worthwhile.
- Research conduct: No human subjects were recruited, and the study conducted no new annotation, translation, or user studies.All reviewed data came from published papers and public or previously accessed repositories.
- Reproducibility: The authors release all 25 language slices and completed audit instruments scored against a thirty-one-field schema and six gap categories.Per-field rationales enable readers to verify specific severity judgments against the underlying evidence and reasoning.