Source-linked AI summary
AV-SafetyBench: A Safety Benchmark for Text-to-Audio-Video Generation
Suah Choi, Tae-Young Lee, Gyeong-Moon Park
TL;DR
T2AV models can convey unsafe meaning through video, audio, or their interaction, while existing benchmarks largely evaluate modalities separately. AV-SafetyBench addresses this gap with a four-axis, 13-category taxonomy, 5,200 manually reviewed prompts, and three-view risk attribution, finding substantial unsafe rates and many risks missed by video-only evaluation.
Problem
T2AV safety evaluation lacks benchmarks designed to capture unsafe meaning conveyed through audio or audiovisual interaction.
Method
AV-SafetyBench combines 5,200 manually reviewed prompts with Full-AV, Video-Only, and Audio-Only judgments that attribute unsafe outputs to four risk sources.
Results
For four of five models, Audio-Only and AV-Joint cases comprise 41.6–48.3% of assigned Full-AV unsafe outputs, while Full-AV Unsafe Rates range from 25.1% to 49.4%.
Takeaways & Limitations
T2AV safety evaluation should assess visual and audio tracks together rather than relying on isolated-modality evaluation.
Takeaways & Limitations
Full-AV Unsafe Rate reflects both safeguards and whether a model can realize the requested content, so it should be interpreted with category-level results.
Abstract
from arXiv · showhide
Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a single text prompt. This capability poses new challenges for safety evaluation, as unsafe content may be conveyed through the audio track or arise only when the visual and audio tracks are interpreted jointly. Existing safety benchmarks largely focus on either generated video or generated audio in isolation and are therefore not designed to capture these risks. To close this gap, we introduce AV-SafetyBench, the first safety benchmark developed specifically for T2AV generation. AV-SafetyBench comprises a four-axis, 13-category taxonomy and 5,200 manually reviewed prompts that specify visual scenes, speech, and non-speech audio. Our evaluation protocol assesses each output under three views: Full-AV, Video-Only, and Audio-Only. It then uses the Video-Only and Audio-Only judgments to assign Full-AV unsafe outputs to one of four risk sources: Video-Only, Audio-Only, AV-Both, or AV-Joint. We evaluate five open-source T2AV models and validate the automated Full-AV judgments against human annotations. Across the five models, Full-AV Unsafe Rates range from 25.1% to 49.4%. Beyond these aggregate rates, risk-source analysis reveals that, for four of the five models, Audio-Only and AV-Joint cases-unsafe outputs missed by video-only evaluation-account for 41.6-48.3% of Full-AV unsafe outputs for which a risk source could be assigned. In the Cross-Modal Harm Emergence category, AV-Joint accounts for 87.5% of unsafe outputs withan assigned risk source. Together, these findings demonstrate the value of AV-SafetyBench for evaluating T2AV safety across the visual and audio modalities and their interaction.
1 Introduction
T2AV models create safety risks through visual content, audio content, and interactions between the two, while existing evaluations largely examine only one modality. AV-SafetyBench addresses this gap with a multi-view benchmark and finds substantial unsafe-output rates and risks missed by video-only evaluation.
- Motivation: T2AV outputs can convey unsafe meaning through the visual track, audio track, or their interaction.These pathways include unsafe speech or sound paired with benign visuals and harms that emerge only when both tracks are interpreted together.
- Motivation: Existing video and audio safety benchmarks evaluate modalities separately and are not designed for joint audiovisual risks.T2VSafetyBench omits audio, while TTA-Bench evaluates audio in isolation.
- Findings: 25.1%–49.4%: Full-AV Unsafe Rates across five open-source T2AV models.At the lower end, one prompt in four yields an unsafe output.
- Benchmark: AV-SafetyBench pairs 5,200 prompts across a four-axis, 13-category taxonomy with three-view risk-source attribution.The protocol evaluates Full-AV, Video-Only, and Audio-Only views, then attributes unsafe outputs to Video-Only, Audio-Only, AV-Both, or AV-Joint sources.
- Findings: 41.6–48.3%: Audio-Only and AV-Joint cases comprise this share of assigned Full-AV unsafe outputs for four of five models.In Cross-Modal Harm Emergence, AV-Joint accounts for 87.5% of attributed unsafe outputs.
2 Related Work
Prior safety benchmarks generally assess generated images, videos, or audio in isolation. AV-SafetyBench extends this work by evaluating complete T2AV outputs and identifying whether unsafe meaning comes from either track or their interaction.
- T2AV generation: T2AV models jointly generate video, speech, sound effects, and ambience from one prompt, expanding safety evaluation beyond silent video.This multimodal generation setting motivates evaluation of audiovisual meaning rather than visual content alone.
- Existing evaluation: Existing image, video, and audio benchmarks largely treat each output modality separately.The cited prior work covers safety evaluation across individual media types but not native T2AV outputs.
- AV-SafetyBench: AV-SafetyBench evaluates complete audiovisual outputs while using isolated-view judgments to characterize the source of unsafe meaning.This design distinguishes visual, audio, and interaction-based risk sources.
3 AV-SafetyBench
AV-SafetyBench constructs a policy- and literature-grounded taxonomy and a manually reviewed T2AV prompt set, then evaluates outputs through Full-AV, Video-Only, and Audio-Only views. Its protocol assigns risk sources by comparing isolated-view judgments with the Full-AV judgment.
- Risk Taxonomy: The taxonomy contains 13 categories organized under four axes: Depicted Content, Identity and Rights, Social and Deceptive, and T2AV Compositional Risks.The taxonomy is derived from major provider usage policies and prior safety-evaluation and prompt-collection work.
- Prompt Construction: Prompts specify a visual scene together with speech, non-speech audio, or both, allowing risks to arise in either track or through interaction.Speech specification varies by category, from 6.0% for Graphic Injury to 100% for several social and deceptive categories.
- Prompt Construction: 176 manually written seed prompts are pilot-validated for realization of their intended visual and audio cues before revision.Pilot generations use LTX-2.3, which supports joint video, speech, and non-speech audio generation.
- Prompt Construction: Each category contains 400 accepted prompts after inspection for category fit, audiovisual completeness, clarity, and duplication, totaling 5,200 prompts.Expanded candidates are removed or revised until every category reaches the target count.
- Evaluation Protocol: Each clip is judged under Full-AV, Video-Only, and Audio-Only views using the corresponding visual, audio, transcript, prompt, and category-rubric inputs.The automated judge is Gemini 3.5 Flash; ASR transcripts are generated with Whisper large-v3.
- Evaluation Protocol: Risk-source assignment labels Full-AV unsafe outputs as Video-Only, Audio-Only, AV-Both, or AV-Joint based on isolated-view judgments.AV-Joint applies when both isolated views are judged safe, and assignment requires parseable labels from both isolated views.
4 Experiments
Across five open-source T2AV models, Full-AV unsafe rates vary substantially, while risk sources differ by category and modality. Audio-related and audiovisual interaction risks expose blind spots in video-only evaluation, although observed rates also reflect whether models realize the requested content.
- Main results: 25.1% to 49.4%: automated macro-averaged Full-AV Unsafe Rates across the five evaluated T2AV models.Human annotations label slightly more outputs unsafe overall, with 43.4% versus the judge’s 41.2%; mean category-level Cohen’s κ is 0.762.
- Depicted Content Risks: daVinci records the highest FUR for Explicit Sexual Content, Suggestive Sexualization, and Graphic Injury, while Physical Harm ranges from 38.3% to 51.7%.Explicit Sexual Content and Graphic Injury reverse ordering between Ovi-1.1 and LTX-2.3, showing category-specific model variation.
- Identity and Rights Risks: Likeness Misuse ranges from 18.3% to 28.3%, whereas IP Misuse reaches 88.3% for Ovi-1.1 and 86.7% for LTX-2.3 versus 41.7% for daVinci.The IP Misuse values represent the widest spread on the Identity and Rights Risks axis.
- Social and Deceptive Risks: JavisDiT++ has near-zero FURs for Targeted Abuse, Fabricated Communication, and Deceptive Solicitation, while the other models show substantially higher rates in these speech-centered categories.JavisDiT++ often fails to produce intelligible speech central to the intended risk; Illegal Activities ranges from 13.3% to 25.0%.
- T2AV Compositional Risks: Cross-Modal Harm Emergence reaches 28.3% for LTX-2.3 and Ovi-1.1, while Temporal Harm Emergence remains at or below 6.7% and NAVA records no unsafe outputs there.Identity-Claim Attribution reaches 61.7% for LTX-2.3 and 53.3% for NAVA, while JavisDiT++ records none.
- Risk-Source Analysis: 41.6–48.3% of assigned-risk unsafe outputs are missed by video-only evaluation for four models, while AV-Joint constitutes 87.5% of Cross-Modal Harm Emergence cases.Audio-Only and AV-Joint cases are the missed outputs; AV-Joint cases receive human Full-AV agreement of 91.4%.
5 Conclusion
AV-SafetyBench evaluates T2AV safety across visual and audio tracks and their interaction. Across five open-source models, video-only evaluation missed unsafe outputs carried by audio or emerging through audiovisual interaction.
- AV-SafetyBench comprises 5,200 manually reviewed prompts across a 13-category taxonomy and a three-view protocol.
- Five open-source model experiments show that video-only evaluation misses unsafe outputs carried by audio or arising through audiovisual interaction.
- The missed share varies widely across models, so measured safety profiles reflect both safeguards and the risks models can realize.
Ethical Considerations
The annotation study involved consenting adult volunteers who were warned about potentially offensive or disturbing material. Participants retained control over their participation and were compensated.
- Annotators were adults who participated voluntarily and gave written informed consent.
- Participants were warned in advance that the material might be offensive or disturbing.
- Annotators could skip items or withdraw at any time, and all were compensated.
Supplementary Material for AV-SafetyBench
The supplementary material documents the benchmark’s judging rubrics and prompt-construction process. Validated seed prompts were expanded by an LLM and manually reviewed to produce the released dataset.
- A.1 Category Definitions and Judging Rubrics: Table A1 reproduces the 13 category rubrics given verbatim to every judge.
- A.2 Prompt Construction: Prompt validation generates one LTX-2.3 clip per seed, revising prompts until their specified cues convey the intended category risk.
- A.2 Prompt Construction: Validated category seeds are expanded by an LLM using the corresponding rubric and then manually reviewed to remove duplicates and off-category prompts.
- A.2 Prompt Construction: The expansion process yields 400 prompts per category, totaling 5,200 prompts.
A.3 Dataset Statistics
Dataset statistics characterize where speech and non-speech audio enter the prompts and document the evaluation and annotation setup. Audio is especially load-bearing in speech-centered risk categories.
- Speech-Careing Prompts: Speech is specified at or near 100% for categories 7–9 and 13, whose risks are carried by speech, and least often for depicted-content categories.
- Dataset Scope: All prompts specify adults only, with source prompts involving minors excluded and expanded candidates checked during manual review.
- Generation Settings: Each evaluated model generates one native audio-video clip per prompt under its default configuration, without prompt rewriting or output selection.
- Annotation and Judging: The judging materials include the category rubric, generation prompt, generated clip, ASR transcript, and three-way verdict.
- Annotation and Judging: Full-AV judging considers video and audio together, while Video-Only excludes audio and Audio-Only excludes visual input.
B.4 Protocol and Subset Validation
The appendix validates both the three-view judging protocol and the 780-prompt tinyset used for evaluation. Separate-call judging creates inconsistencies, while the reported single-pass results are conservative and tinyset metrics remain representative.
- Protocol validation: 471 monotonicity violations occur under separate-call judging, versus none under the single-pass protocol.The violations are cases where an isolated view is unsafe while the containing Full-AV view is safe.
- Protocol validation: The single-pass protocol reports a lower AV-Joint share and video-only miss rate than separate-call judging.The reported main-text protocol therefore biases against, rather than toward, the claim that video-only evaluation misses unsafe outputs.
- Subset validation: 10,000 stratified subsamples show every tinyset metric, including risk-source shares, within the central 95% sampling band.The subsamples use the identical stratified design drawn from the 5,200-prompt full set.
C.1 Judging Reliability
The automated Full-AV judge agrees substantially with human annotations, but judge choice materially changes reported unsafe rates. Verdict coverage is nearly complete, and unsafe realizations generally match their intended categories.
- Judging reliability: 91.4% of outputs judged unsafe by the automated judge are also unsafe under human annotations.The judge reports a lower unsafe rate than annotators overall and has lower recall than precision.
- Judge dependence: 70.5–89.4% pooled unsafe rates from open-source detectors exceed the 41.2% Gemini reference.Judge selection changes the reported rate by more than the evaluated models’ 24-point FUR gap.
- Coverage: Unclear verdict rates remain below 1% in every view, with no judging call failing to parse.Risk-source assignment coverage is at least 98.9% for every model.
- Intended-risk realization: 96.5% overall intended-category fit shows that unsafe generations generally realize the risk specified by their prompts.Fit falls to 83.4% for Explicit Sexual Content and 72.7% for Temporal Harm.
C.2 Risk-Source Validity
Risk-source assignments are supported by human agreement and category concentration, especially for AV-Joint cases. However, judge choice affects absolute shares, and low unsafe rates can reflect missing speech capability rather than safety intervention.
- Risk-source validity: Agreement stays within three points across all four risk sources, with 89.8% human confirmation for video-only misses.AV-Joint outputs are confirmed unsafe as often as Video-Only outputs.
- AV-Joint concentration: 87.5% of Cross-Modal Harm Emergence unsafe outputs with assigned sources are AV-Joint, compared with 1.4% outside that category.The category contributes 49 of 56 AV-Joint cases, while eight categories contain none.
- Judge robustness: AV-Joint concentration in Cross-Modal Harm Emergence persists under five of six judges, despite macro-averaged FURs spanning 41.2–89.4%.Absolute shares vary with judge strictness, but the category concentration remains.
- Judge robustness: Nemotron without thinking is an exception: AV-Joint is 0.0% inside and 0.06% outside Cross-Modal Harm Emergence.Enabling thinking restores a 61.5% within-category share, linking the exception to over-flagging.
- Capability confound: JavisDiT++ has near-zero FUR in speech-carried categories alongside 4.2% mean utterance recall.Its low rate coincides with incomplete speech realization rather than a measurable safety intervention.
C.4 Benchmark and Model Comparison
The benchmark comparison uses transferred video-only prompts, commercial case studies, and qualitative samples to examine whether audiovisual risk patterns generalize across prompt sets and model settings. Native audiovisual prompting changes risk-source composition toward audio-related and joint risks.
- Benchmark comparison: 699 non-adversarial T2VSafetyBench prompts are transferred to the five T2AV models under the same judging protocol.Prompts targeting text-to-image tokenizer behavior are excluded from the comparison.
- Commercial comparison: Commercial case studies test audio-carried and cross-modal failure modes but are too small to rank models.Only outputs that passed provider-side filtering are included.
- Risk-source comparison: Native audiovisual prompting shifts Full-AV unsafe-output composition from Video-Only toward Audio-Only and AV-Joint sources.Figure A5 compares transferred T2VSafetyBench and AV-SafetyBench prompts generated and judged under matched conditions.
- Generation failures: Figures A6–A7 illustrate safe outputs caused by generation failures in Likeness Misuse and Temporal Harm cases.The intended likeness or harmful temporal turn is not realized, so the outputs are judged safe.
- Qualitative coverage: Figures A8–A14 provide three-frame, waveform, and transcript samples spanning all 13 risk categories.Sexual and graphic content is moderately masked where applicable.
E Limitations
The evidence is bounded by unvalidated isolated-view risk-source judgments, judge-dependent rates, limited uncertainty reporting, and temporal sampling that can miss short-lived events. The claim about missed unsafe outputs applies to the benchmark’s evaluation views, not deployed moderation systems.
- Human validation covers Full-AV judgments, but not the isolated-view verdicts used to assign risk sources.AV-Joint outputs were confirmed unsafe within three points of the other sources, but the isolated-view labels themselves were not human-validated.
- Per-cell rates lack confidence intervals, although stratified resampling quantifies subset uncertainty and per-cell counts expose small denominators.
- 61.2% to 87.5%: the AV-Joint share within Cross-Modal Harm Emergence varies across judges, although its concentration reproduces.The main-text figure reports the upper end of this judge-dependent range.
- One frame per second can miss staged progressions between samples, potentially contributing to the low reported Temporal Harm Emergence rates.This limitation concerns the judge’s temporal sampling, alongside generation failures shown in Figure A7.
- The claim that video-only evaluation misses unsafe outputs concerns this three-view protocol, not deployed moderation models.Whether practical video and audio moderation tools fail on Audio-Only and AV-Joint outputs remains unanswered by the benchmark.