Source-linked AI summary

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning

Tianle Chen, Deepti Ghadiyaram

arXiv:2604.03995v1cs.CVcs.SD

TL;DR

Audio-visual MLLMs may process semantically similar cues inconsistently across modalities, creating a robustness gap relevant to safety-sensitive use. The paper introduces Multi-Modal Typography to test speech, visual, and text attacks across tasks and benchmarks, finding that coordinated audio-visual attacks are substantially stronger than single-modality attacks. It also shows that misleading speech can affect visually grounded reasoning and weaken harmful-content detection.

  • Problem

    The study addresses limited understanding of whether semantically similar typographic perturbations are treated consistently across audio, visual, and text modalities in audio-visual MLLMs.

  • Method

    Multi-Modal Typography injects misleading TTS-generated speech into videos while keeping the visual stream unchanged, and evaluates coordinated audio, visual, and text perturbations across MLLMs, tasks, and benchmarks.

  • Results

    83.43% ASR for aligned audio-visual attacks versus 34.93% for audio-only attacks on audio questions shows substantially stronger targeted steering on MMA-Bench.

  • Takeaways & Limitations

    Misleading speech can degrade visually grounded reasoning and weaken content moderation even when visual evidence remains unchanged.

Abstract

from arXiv · show

As audio-visual multi-modal large language models (MLLMs) are increasingly deployed in safety-critical applications, understanding their vulnerabilities is crucial. To this end, we introduce Multi-Modal Typography, a systematic study examining how typographic attacks across multiple modalities adversely influence MLLMs. While prior work focuses narrowly on unimodal attacks, we expose the cross-modal fragility of MLLMs. We analyze the interactions between audio, visual, and text perturbations and reveal that coordinated multi-modal attack creates a significantly more potent threat than single-modality attacks (attack success rate = $83.43\%$ vs $34.93\%$).Our findings across multiple frontier MLLMs, tasks, and common-sense reasoning and content moderation benchmarks establishes multi-modal typography as a critical and underexplored attack strategy in multi-modal reasoning. Code and data will be publicly available.

1 Introduction

This study examines whether semantically similar perturbations are processed consistently across text, spoken-audio, and visual-text modalities in audio-visual MLLMs. It introduces coordinated typographic attacks and finds that combining modalities produces stronger steering and disruption than single-modality attacks.

  • Prior typographic attacks show that overlaid text and logos can override highly relevant visual content, exposing vision-language models’ sensitivity to textual cues.
  • The study asks whether semantically similar perturbations receive different treatment across text prompts, spoken audio, and on-screen visual text.
  • Multi-Modal Typography injects misleading TTS-generated speech while keeping the visual stream unchanged, enabling controlled modality conflicts and testing of prediction steering.
  • 64.03% ASR on WorldSense shows that spoken typography can reliably steer Qwen2.5-Omni-7B toward injected targets.
  • 83.13% ASR on visual questions and 83.43% on audio questions show that aligned audio-visual attacks substantially strengthen failures on Qwen2.5-Omni-7B.
  • A ∼13% decrease in detection capability shows that safe injected speech can hijack content moderation for visually harmful videos.

2 Related Work

Prior work establishes visual typography, malicious audio, and conflicting modality streams as separate robustness concerns. This study extends those lines toward speech-based attacks in multimodal reasoning, including visually grounded tasks.

  • Visual Typography and Prompt Injection: Visual prompt-injection studies show that overlaid text and logos can hijack classification, question answering, and generation despite weak relevance to the underlying image.
  • Audio Injection and Speech-Centric Robustness: Audio robustness research identifies malicious or misleading audio as a viable attack surface, but primarily studies audio-only or speech-centric contexts.
  • Audio Injection and Speech-Centric Robustness: This work extends audio-centric robustness studies to multimodal reasoning, where misleading speech can affect both audio-grounded and visually grounded tasks.
  • Multi-modal Robustness Under Conflicting Streams: Prior multimodal robustness studies find that missing, noisy, or inconsistent streams can expose uneven modality reliance and brittle behavior under cross-modal disagreement.

3 Our Approach

The approach focuses on speech-based attacks as a direct semantic channel in audio-visual reasoning. It synthesizes speech into original video audio and evaluates both disruption of correct reasoning and targeted steering.

  • Audio typography injects synthesized speech carrying a semantic content sequence into a video’s original audio track.
  • The method focuses on speech rather than generic audio perturbations because spoken content directly conveys semantics and resembles natural narration or conversation.
  • Attack Success Rate measures the fraction of examples redirected to the injected target label, while Ground-Truth Accuracy measures prediction accuracy under clean and attacked inputs.

4 Experiments

Experiments evaluate audio typography across multiple audio-visual MLLMs, datasets, modalities, and attack configurations. Spoken and coordinated cross-modal injections degrade performance and redirect predictions toward injected targets, with aligned audio–visual attacks producing the strongest failures.

  • Experimental Setup: Multiple audio-visual MLLMs are evaluated on MMA-Bench, Music-AVQA, WorldSense, and safety benchmarks.The model set includes Qwen2.5-Omni-7B, Qwen3-Omni-30B, PandaGPT, ChatBridge, Gemini-2.5-Flash-Lite, and Gemini-3.1-Flash-Lite-preview.
  • Experimental Setup: Audio typography injects synthesized speech targeting class c* into a video's audio stream while preserving the visual stream.The attack compares predictions after misleading spoken semantic perturbations with clean-input accuracy and targeted attack success.
  • Audio Typography: 64.03% ASR on WorldSense for Qwen2.5-Omni-7B demonstrates targeted redirection toward injected labels rather than merely arbitrary prediction errors.Across MMA-Bench, the same model's audio-only ASR is 34.93% on audio questions and 24.27% on visual questions.
  • Audio Typography: 12.85% accuracy drop on Qwen2.5-Omni-7B visual questions shows spoken perturbations affect visually grounded reasoning despite untouched frames.Music-AVQA visual-only queries show a corresponding 10.76% drop.
  • Per-modality Attacks: Attack strength depends on delivery modality and model: text is strongest for Qwen2.5-Omni-7B, whereas visual attacks are strongest for Gemini-3.1-Flash-Lite-preview on MMA-Bench.Spoken injection remains effective, but its relative strength varies across models and question types.
  • Aligned and Conflicting Audio–Visual Typography: 83.43% ASR on audio questions under aligned audio–visual injection exceeds audio-only 34.93% and visual-only 45.19% for Qwen2.5-Omni-7B.On visual questions, aligned injection reaches 83.13% ASR versus 24.27% audio-only and 50.34% visual-only; Gemini shows the same weaker qualitative pattern.

5 Analysis of Attack Effectiveness

Audio typography strength depends on controllable parameters and involves an effectiveness–stealth trade-off. Richer semantic cues and stronger aligned or repeated injections produce more targeted disruption, while stealth varies across attack families.

  • Audio Typography Parameters: 34.72% ASR for audio questions and 29.78% for visual questions occur at volume multiplier 8.0, rising from 15.59% and 12.04%.Higher volume strengthens attacks on both question types.
  • Audio Typography Parameters: 19.60% injected-target rate for visual questions and 23.77% for audio questions occur when speech starts at 80%, versus 15.28% and 18.67% at 0%.Later placement generally produces stronger attacks, possibly because it is closer to the model’s final decision.
  • Audio Typography Parameters: 33.85% injected-target rate for audio questions and 23.80% for visual questions follow four repetitions, compared with 22.53% and 19.29% after one presentation.High repetition frequency strengthens attacks across both question types.
  • Audio Typography Parameters: 22.07% ASR is highest for female voices on audio questions, while visual-question ASR reaches 18.21%; voice identity has a comparatively modest effect.Male and neutral voices produce lower ASR than female voices under fixed injected semantics.
  • Effectiveness–Stealth Trade-Off: Volume produces the strongest attacks but the largest stealth cost, whereas repetition offers a more favorable effectiveness–stealth balance.Temporal position causes moderate degradation while leaving relative RMS nearly unchanged.
  • Semantic Richness: 64.03% ASR for Qwen2.5-Omni-7B occurs with strong target cues, compared with 23.16% for weak cues; random noise and speech have little effect.Stronger cues recite the target option’s semantic content rather than merely naming it.

6 Safety Application: Harmful-Content Detection

The study evaluates spoken benign cues as attacks on harmful-content detection. Stronger spoken manipulation lowers harmful-content detection, including when harmful evidence remains visually present.

  • Harmful-Content Detection: 8.04% harmful rate for Qwen2.5-Omni-7B under the prompt-style attack follows 20.41% under the keyword attack and 26.16% on original inputs.Lower harmful rate indicates less successful identification of harmful videos under attack.
  • Harmful-Content Detection: Stronger spoken manipulation increasingly weakens harmful-content detection on MetaHarm and high-risk generated content from I2P.The harmful evidence remains present in the video despite the spoken injection.
  • Safety Implications: Benign spoken injection reduces harmful-content detection and increases unsafe-to-safe errors on I2P and MetaHarm.These evaluations frame misclassification of unsafe video as safe as a safety-sensitive risk.

7 Discussion and Future Work

The study identifies audio typography as a semantic robustness gap in audio-visual MLLMs because spoken content integrates naturally into video audio. It proposes further work on realistic interference, mechanisms, defenses, and perceptual stealth.

  • Discussion: Audio typography is a highly effective semantic attack because it naturally integrates into a video’s audio.The authors connect this robustness gap to risks in safety-sensitive content moderation.
  • Future Work: Future work should test overlapping speakers and background narration as realistic interference vulnerabilities.These settings extend evaluation beyond the studied injected-speech conditions.
  • Future Work: Future work should develop modality-aware consistency checks and train models with semantically perturbed data.The paper also calls for mechanistic interpretation of competing modality cues.
  • Future Work: Human perceptual evaluations are needed to quantify the real-world threat posed by perceptually stealthy attacks.The authors identify perceptual stealth effectiveness as an open research direction.

A Ethics Statement

The paper frames spoken semantic injection as a safety and evaluation problem for audio-visual MLLMs, while acknowledging dual-use risks and controlled-study limitations. It argues that revealing this vulnerability can support safer multimodal systems and stronger evaluation.

  • Spoken semantic cues can steer model predictions despite unchanged visual evidence, defining audio typography as an underexplored robustness and safety failure mode.
  • The study's potential benefits include better robustness benchmarks, modality-aware consistency checks, grounding objectives, and training procedures against misleading semantic cues.
  • The experiments indicate that spoken perturbations affect audio-grounded questions, visually grounded reasoning, and harmful-content detection.
  • The authors caution that the findings could be misused to manipulate moderation, retrieval, recommendation, or decision-support systems, while not presenting a covert-abuse recipe.
  • The study uses synthesized speech, short explicit attacks, and controlled evaluation to reduce impersonation concerns and isolate spoken-content effects.
  • The controlled cues do not cover overlapping conversation, natural narration, or speaker-specific deception, so the work is not a complete estimate of real-world abuse prevalence.
  • The authors position disclosure as supporting safer multimodal systems because speech is a native, natural component of video.

Appendix Contents

The appendix provides implementation details, dataset-specific templates, ablations, stealth analyses, and qualitative case studies for the audio-typography experiments. It covers both methodological settings and additional results across reasoning and safety benchmarks.

  • The appendix includes the audio-typography generation pipeline and default settings, including speech synthesis, insertion, repetition, and related parameters.
  • Dataset-specific spoken injection templates adapt audio typography to class-label, multiple-choice, and safety benchmarks.
  • Additional analyses examine stealth metrics, qualitative examples, and the effectiveness–stealth trade-off under different attack settings.
  • WorldSense and related benchmarks receive ablations on target-directed speech content, semantic richness, and safety-related spoken manipulation.
  • Qualitative case studies cover clean controls, attack failures, successful targeted attacks, and safety-related spoken-injection examples.
  • Figure 4 shows short wrong-answer statements for class-label tasks, incorrect option content for WorldSense, and benign spoken cues for MetaHarm.
  • Table 5 summarizes the default standalone audio-typography setup and controlled variations in gain, position, repetition, and voice.

A.1 Dataset-Specific Spoken Injection Templates

The spoken injection remains short and answer-oriented across tasks, while its wording is adapted to each benchmark's target answer space. This preserves a consistent attack format while keeping the cue semantically targeted.

  • Short spoken cues are adapted to each dataset's answer space so the attack remains comparable across tasks.

A.2 Audio Typography Generation Pipeline and Default Settings

Audio typography inserts a short misleading spoken phrase into the original audio while leaving the visual stream unchanged. Its unified pipeline constructs the target phrase, synthesizes speech, inserts and mixes it, then applies fixed defaults and controlled parameter studies.

  • Audio typography injects a short misleading spoken phrase into the original audio track while leaving the visual stream unchanged.
  • The generation pipeline has three stages: target phrase construction, text-to-speech synthesis, and temporal insertion with waveform mixing.
  • The default setup uses simple, semantically targeted speech as a controlled symbolic cue rather than long-form adversarial narration.
  • Injected speech is repeated until it spans the original audio duration, avoiding a fixed repetition count and improving comparability across videos.
  • Templates name wrong semantic labels for class-label tasks, include incorrect option content for WorldSense, and use benign safety language for MetaHarm.

A.3 Extended Stealth Metrics and Additional Analysis

The extended analysis evaluates attack effectiveness against multiple stealth measures and finds that attack-family rankings remain consistent across metrics, benchmarks, and models. Stronger settings selectively redirect predictions toward injected targets, while acoustic prominence and repetition are more robust than temporal placement or voice identity.

  • Effectiveness–stealth trade-offs: Gain produces the largest reduction in average accuracy but also the greatest movement along every stealth axis, especially RMS and speech-recognition shift.Repetition offers a more favorable effectiveness–stealth curve, while temporal placement changes attack strength with little RMS change and voice identity has only modest effects.
  • Effectiveness–stealth trade-offs: RMS has the strongest monotonic relation with average accuracy (ρs = −0.62), followed by CLAP variance (ρs = −0.53) and spectral entropy (ρs = −0.48).Spectral flatness is weaker (ρs = −0.22), while speech-recognition shift is moderately monotonic (ρs = −0.32) and captures lexical recoverability rather than generic acoustic change.
  • Targeted prediction effects: Stronger attacks selectively reallocate probability mass from the ground-truth class toward the injected target rather than dispersing errors randomly.The qualitative evidence extends this targeted redirection to standard tasks and safety-related settings, including harmful-content misclassification under spoken semantic injection.
  • Parameter sensitivity: 67.81% ASR and 19.31% label accuracy result from increasing gain from 0.5× to 16× for Qwen2.5-Omni-7B on WorldSense.Repetition similarly raises ASR from 44.04% at ×1 to 61.67% at ×50 while accuracy falls from 33.69% to 22.14%.
  • Parameter sensitivity: 59.34%–62.30% ASR across Qwen2.5-Omni-7B voices and 45.91%–47.47% across Gemini voices shows that voice identity is a secondary factor.The passage attributes the modest variation to injected semantics being present across speaker styles.
  • Parameter sensitivity: 61.85%–61.97% ASR across temporal placements for Qwen2.5-Omni-7B on WorldSense indicates that insertion timing has little effect in this setting.The authors qualify this against MMA-Bench, where later placement was mildly beneficial, and describe temporal placement as dataset-dependent.
Loading 2604.03995v1…