Source-linked AI summary
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
Kaichao Jiang, Changtao Miao, Baiqi Wu, Zhiyuan Lu, Kang Yang, Peiwei Zhao, Junchi Chen, Yunfeng Diao, He Liu, Qi Chu, Tao Gong, Nenghai Yu
TL;DR
Existing safety benchmarks do not systematically evaluate risks whose evidence is distributed across multimodal and temporal conditioning. MULTI2AV-SAFETY fills this gap with a factorized benchmark covering all 11 non-singleton T/I/A/V configurations and 11,024 attacks, revealing Composed Harm and Diluted Harm as complementary weaknesses in current safety guards.
Problem
Existing safety benchmarks remain largely prompt-centric or tied to fixed conditioning interfaces, making compositional risks across modalities and time difficult to study systematically.
Method
MULTI2AV-SAFETY systematically evaluates multimodal-conditioned audio-video generation across conditioning configurations, attack mechanisms, harm-evidence structures, and harm categories.
Results
The evaluation reveals Composed Harm from jointly benign inputs and Diluted Harm when explicit harmful cues are mixed with benign multimodal context, exposing systematic guard weaknesses.
Takeaways & Limitations
Safe audio-video generation requires compositional risk perception: correctly integrating safety implications across conditioning inputs, not merely observing them.
Takeaways & Limitations
Image- and video-side adversarial perturbations are excluded because available methods did not transfer reliably and reproducibly to the target generator.
Abstract
from arXiv · showhide
Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no longer reside in any single input, but instead emerge from how otherwise benign or weakly harmful conditions interact across modalities and time. Existing safety benchmarks, however, remain largely prompt-centric or tied to fixed conditioning interfaces, leaving such compositional risks difficult to study systematically. To bridge this gap, we introduce Multi2AV-Safety, the first safety benchmark, to the best of our knowledge, to cover all 11 non-singleton T/I/A/V conditioning configurations for audio-video generation, comprising 11,024 attack instances. Evaluation on Multi2AV-Safety reveals systematic weaknesses in representative multimodal safety guards across attack mechanisms and harm-evidence structures. Our evaluation reveals two complementary failure modes: harmful semantics can emerge from the combination of individually benign inputs, while explicit harmful cues can become harder to detect when mixed with benign multimodal context. Together, these results identify \emph{compositional risk perception} as a central capability gap in safeguarding multimodal-conditioned audio-video generation: current safety guards fail to reliably integrate safety evidence across modalities and time, even when all conditioning inputs are observable. The dataset will be publicly released in October 2026.
1 INTRODUCTION
Multimodal conditioning distributes safety-relevant evidence across inputs and time, creating Composed Harm and Diluted Harm that existing benchmarks and guards do not reliably capture. MULTI2AV-SAFETY addresses this gap with systematic coverage and evidence-aware evaluation.
- Multimodal conditioning creates Composed Harm from jointly benign inputs and Diluted Harm when explicit cues are surrounded by benign context.These risks arise across modalities or time rather than necessarily within a single prompt.
- Most existing benchmarks examine only one or two input modalities, limiting systematic analysis of distributed or composition-dependent safety evidence.This also complicates attribution when safety performance degrades.
- MULTI2AV-SAFETY covers all 11 non-singleton text, image, audio, and video conditioning combinations for multimodal-conditioned audio-video generation.The benchmark contains 11,024 attack instances spanning two- to four-modal conditioning, four attack mechanisms, and five harm categories.
- Its factorized design makes attack mechanisms and harm-evidence structures explicit, enabling mechanism-stratified and evidence-aware attribution of safety failures.Paired input- and output-side evaluation assesses whether successful attacks are intercepted.
- Evaluation identifies a shared weakness across Composed Harm and Diluted Harm: guards struggle to recognize risks arising from multimodal composition.This weakness persists despite access to the conditioning modalities.
2 RELATED WORK
Prior media-generation safety benchmarks provide broad harmful-content and attack coverage but generally evaluate fixed conditioning interfaces. Recent compositional-risk studies broaden the problem, while MULTI2AV-SAFETY organizes these risks in a common conditioning space.
- Existing image- and video-generation benchmarks primarily study safety under predefined conditioning settings rather than interactions among conditioning inputs.Examples include I2P, T2ISafety, SafeSora, and T2VSafetyBench.
- Recent work identifies cross-modal and temporally fragmented compositional risks through benchmarks including SafeGen-Bench, Multimodal Pragmatic Jailbreak, and SceneSplit.These phenomena are studied under different conditioning and attack settings.
- Table 1 distinguishes “Multi” harm, requiring multiple independently harmful input carriers, from “Joint” harm, unavailable from any input in isolation.These labels organize how harmful evidence is distributed across conditioning inputs.
- MULTI2AV-SAFETY places prior compositional phenomena in a common conditioning space while keeping attack mechanisms and harmful-evidence structures explicit.The paper summarizes this factorized comparison in Table 1.
- Multimodal guardrails have progressed from image-text and video moderation toward generation-specific and reasoning-based multimodal detection.The related systems include Llama Guard 3 Vision, SafeWatch, ConceptGuard, GuardReasoner-VL, GuardTrace-VL, and GuardReasoner-Omni.
3 MULTI2AV-SAFETY
MULTI2AV-SAFETY factorizes multimodal audio-video safety evaluation across conditioning composition, attack mechanism, and harm-evidence structure. It covers all 11 non-singleton T/I/A/V configurations and constructs matched scenarios that localize or distribute harmful evidence across modalities and time.
- Factorized benchmark design: Each instance factors a target scenario, active modality set, realized modality inputs, attack mechanism, harm-evidence structure, and harm category.The active set contains 2 to 4 modalities, while the design distinguishes how harm is introduced from where or when decisive evidence appears.
- Factorized benchmark design: The benchmark covers all 11 non-singleton T/I/A/V configurations across two- to four-modal conditioning.It uses six two-modal, four three-modal, and one four-modal configuration.
- Factorized benchmark design: Harm evidence is classified as Composed Harm, Diluted Harm, or Full Harm according to how many active modalities are harmful in isolation.Composed Harm has k = 0, Diluted Harm has 0 < k < |S|, and Full Harm has k = |S|.
- Scenario and condition construction: The construction uses harmful–benign matched pairs and aligned multimodal conditions to control whether safety evidence is individual, joint, or temporally distributed.Selected bimodal scenarios are extended with aligned conditions for matched three- and four-modal comparisons.
- Multimodal attack construction: Four attack mechanisms—Direct, Jailbreak, Adversarial, and Temporal—define distinct ways of delivering harmful semantics without assuming an ordering of difficulty.Temporal attacks distribute decisive semantics across ordered segments, whereas joint jailbreaks require cross-modal integration at the same time.
- Multimodal attack construction: Image- and video-side adversarial perturbations are excluded because available methods did not transfer reliably and reproducibly to the target generator.This is described as the only systematic exception to the intended modality-carrier coverage.
- Coverage and quality control: 11,024 attack instances comprise 8,849 bimodal, 1,500 trimodal, and 675 four-modal cases.The distribution includes 6,475 Direct, 2,849 Jailbreak, 975 Adversarial, and 725 Temporal instances.
4 BENCHMARK EVALUATION
The benchmark evaluates target realization and residual risk across multimodal audio-video conditioning, then traces guard failures to attack mechanisms, modality access, and harmful-evidence structure. Results show broad vulnerability across configurations and persistent failures even with full modality access, including compositional and diluted harm.
- Evaluation protocol: Three domain experts review each generated output for target-aligned category harm, while conditioning inputs and attack metadata remain hidden.Majority vote assigns Yi=1 only when the output is both category-harmful and target-aligned.
- Evaluation protocol: ASR measures target realization before moderation, whereas RHR measures harmful attacks that survive the guard; lower ASR and RHR are better, while higher recall is better.
- Attack success and residual risk: ASR remains high across all 11 configurations at 83.7–99.1%, while complete-modality recall varies sharply within the same conditioning order.Configuration averages combine attack mechanism and evidence structure, limiting their ability to explain guard failures.
- Attack success and residual risk: 92.6%, 87.4%, and 89.3% are the ASR values for two-, three-, and four-modal inputs, while RHR rises with order for both evaluated full-access guards.For Qwen3-Omni, RHR rises from 32.8% to 40.4%; for GR-Omni-3B, it rises from 22.1% to 32.9% between two and four modalities.
- Tracing guardrail failures: Direct attacks are easiest to intercept, whereas Jailbreak and Adversarial attacks are harder in several settings, but substantial misses remain within individual mechanisms.The aggregate gap is therefore not simply an artifact of attack mixture.
- Tracing guardrail failures: 52.6% and 61.2% are the four-modal recall values for full-access Qwen3-Omni and GR-Omni-3B, respectively, showing that visible evidence is not reliably integrated.T/I-only guards are access-limited diagnostics because they cannot inspect audio/video-only evidence.
- Compositional risk perception: At k=0, Composed Harm reaches 85.7% bimodal ASR while both full-access guards recall fewer than half of these attacks.The unsafe meaning exists only through composition of inputs or temporal fragments that are benign in isolation.
- Compositional risk perception: For k ≥1, recall is lowest at k=1 and rises as harmful evidence spans more carriers across every conditioning order for both guards.A single explicit harmful signal can be easier to miss when surrounded by benign context, defining Diluted Harm.
5 ETHICAL CONSIDERATIONS
MULTI2AV-SAFETY contains safety-critical multimodal content and therefore applies research-use safeguards to construction, annotation, and release.
- The benchmark includes violence, sexual content, hate, illegal activities, and politically sensitive material.
- Data construction and annotation are restricted to research purposes, with minimized harmful-content exposure and no personally identifiable information collection.
- Annotators are informed about the task and may opt out of examples they consider inappropriate.
- Planned release safeguards include a research-oriented license, usage warnings, and restricted access to high-risk media or attack artifacts where necessary.
- The benchmark is intended solely to evaluate and improve multimodal generative-system safety.
6 CONCLUSION
MULTI2AV-SAFETY evaluates safety at the multimodal conditioning-set level, separating attack mechanism, modality access, and harm-evidence structure. Its conclusion identifies complementary failures in integrating distributed or context-obscured safety evidence and frames compositional risk perception as the capability being tested.
- 6 CONCLUSION: MULTI2AV-SAFETY evaluates safety at the multimodal conditioning-set level, where harmful evidence may be distributed across inputs or emerge through their interaction.
- 6 CONCLUSION: The evaluation separates attack mechanism, modality access, and harm-evidence structure when analyzing current safety-model failures.
- 6 CONCLUSION: Table 7 reports Within-Direct recall (%) by harmful-carrier cardinality, highlighting the single-carrier setting and marking the better model within each order and k.
- 6 CONCLUSION: Composed Harm arises when individually benign inputs jointly realize unsafe semantics, whereas Diluted Harm makes explicit harmful cues harder to detect in benign multimodal context.
- 6 CONCLUSION: Compositional risk perception denotes correctly integrating joint safety implications across modalities and time, beyond merely observing all conditioning inputs.