Source-linked AI summary
CommanderSong: A Systematic Approach for Practical Adversarial Voice Recognition
Xuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long, Xiaokang Liu, Kai Chen, Shengzhi Zhang, Heqing Huang, Xiaofeng Wang, Carl A. Gunter
TL;DR
The paper asks whether practical, stealthy adversarial attacks can control modern ASR systems beyond noisy or proximity-dependent attacks. It automatically embeds commands into songs using ASR-guided synthesis, achieving high attack success and human inconspicuousness while proposing defenses. The evaluation also identifies an environmental-noise scope boundary.
Problem
Existing ASR attacks were limited by noise-like commands, ultrasonic hardware exploitation, proximity requirements, or insufficient evidence for stealthy attacks against modern DNN-based systems.
Method
CommanderSong uses Kaldi acoustic and language models, gradient-based synthesis, pdf-id sequence matching, and noise modeling to embed commands into songs for direct or over-the-air recognition.
Results
CommanderSong achieved 100% WTA and 96% WAA success against Kaldi across over 200 generated songs, and none of over 200 participants detected the commands.
Takeaways & Limitations
The attack can operate over the air, transfer to a black-box commercial ASR system, and spread through YouTube and radio, while audio turbulence and squeezing provide proposed defenses.
Takeaways & Limitations
The evaluation does not model invariance to varying background noise across environments such as grocery stores, restaurants, and offices.
Abstract
from arXiv · showhide
The popularity of ASR (automatic speech recognition) systems, like Google Voice, Cortana, brings in security concerns, as demonstrated by recent attacks. The impacts of such threats, however, are less clear, since they are either less stealthy (producing noise-like voice commands) or requiring the physical presence of an attack device (using ultrasound). In this paper, we demonstrate that not only are more practical and surreptitious attacks feasible but they can even be automatically constructed. Specifically, we find that the voice commands can be stealthily embedded into songs, which, when played, can effectively control the target system through ASR without being noticed. For this purpose, we developed novel techniques that address a key technical challenge: integrating the commands into a song in a way that can be effectively recognized by ASR through the air, in the presence of background noise, while not being detected by a human listener. Our research shows that this can be done automatically against real world ASR applications. We also demonstrate that such CommanderSongs can be spread through Internet (e.g., YouTube) and radio, potentially affecting millions of ASR users. We further present a new mitigation technique that controls this threat.
1 Introduction
The paper addresses the practical security risks of DNN-based ASR by automatically embedding commands into songs that remain inconspicuous to listeners but can control real-world systems. It also evaluates remote delivery and proposes defenses against this attack.
- Motivation: DNN-based ASR systems create security concerns because adversarial perturbations can affect systems that execute voice-controlled operations.Prior adversarial-speech attacks were limited by noise-like commands, ultrasonic delivery, hardware exploitation, or proximity requirements.
- Design goals: The attack targets three practical requirements: effectiveness amid speaker and environmental noise, stealthiness to ordinary users, and remote delivery at scale.The paper states that all three challenges were addressable in its research.
- Approach: CommanderSong automatically embeds voice commands into randomly selected songs, preserving normal human perception while enabling ASR-based control.The method uses Kaldi and gradient descent to minimize perturbations; a noise model supports over-the-air robustness.
- Evaluation: None of over 200 Mechanical Turk participants identified the embedded commands, while CommanderSong also transferred to iFLYTEK and could spread through YouTube.The authors further developed defense solutions and reported their effectiveness.
2 Background
The background introduces ASR architecture and explains why adversarial attacks are harder in speech than in images. Existing speech attacks had practical limitations involving recognition models, hardware, stealth, and delivery.
- Speech recognition: ASR systems extract acoustic features and decode them using pretrained acoustic and language models.Kaldi and other open-source toolkits support this general architecture.
- Speech recognition: Feature extraction uses short-time analysis, with MFCC among the most frequently used acoustic-feature algorithms.Other listed approaches include LPC and PLP.
- Speech recognition: Deep neural network models were adopted because GMMs are limited in describing nonlinear data manifolds.The passage places DNN-HMM systems within the evolution of statistical speech recognition.
- Adversarial attacks: Adversarial attacks can modify inputs or training data, with black-box and white-box settings distinguished by the attacker’s knowledge of system algorithms and parameters.The background contrasts these settings primarily in image-recognition research.
- Adversarial attacks: Speech attacks had not matched image-attack progress because speech is a complex continuous time-domain signal with more features than static images.Prior hidden-voice-command work generated obfuscated commands but did not resolve the paper’s practical goals.
3 Overview
The overview frames practical adversarial ASR as a set of open questions about physical robustness, stealth, and scalable delivery. It proposes songs as carriers and develops model-guided perturbations for direct and over-the-air attacks.
- Research questions: The paper asks whether adversarial samples can fool DNN-based ASR in complicated physical environments and whether they can remain stealthy and remotely deliverable.These questions address speaker noise, background noise, human perception, and large-scale delivery.
- Carrier selection: Songs are chosen as command carriers because people commonly listen to music through stations, online libraries, and YouTube.Random songs are used regardless of the desired command to avoid lyrics that directly reveal it.
- Attack design: Generating a CommanderSong trades off preserving the original song’s fidelity against making the desired command recognizable by ASR.The pure voice command serves as a reference during revision.
- Attack design: The approach separately processes the carrier song and command through ASR, intercepts acoustic-model outputs and command phoneme information, then crafts an adversarial sample.The crafted output is inverted through the acoustic model and feature extraction to produce audio.
- Attack settings: WTA attacks feed generated WAV samples directly to ASR APIs, whereas WAA attacks play them through speakers to interact with devices over the air.WAA incorporates a generic noise model to address loudspeaker and environmental interference.
4 Attack Approach
The attack analyzes ASR intermediate representations to embed a command into a song while minimizing audible modification, then models playback noise for over-the-air operation.
- Attack design: The approach addresses imperceptibility and physical practicality by minimizing song perturbations and modeling speaker, receiver, and background noise.The authors use pdf-id sequence matching and gradient descent for the first challenge, then add noise modeling for playback attacks.
- Kaldi Platform: Kaldi decoding exposes acoustic-model outputs, transition-ids, phonemes, and decoded words used to connect audio signals with ASR representations.The paper explains that phoneme identifiers are sequences of transition-ids, and transition-ids map to pdf-ids.
- Gradient Descent to Craft Audio: For each song frame, the method selects the pdf-id with the highest DNN probability and forms a sequence representing the song.The DNN output matrix contains pdf-id probabilities for each frame; the resulting sequence is denoted m.
- Gradient Descent to Craft Audio: The command produces a target pdf-id sequence, and the attack minimizes the L1 distance between song and command sequences through a perturbation δ(t).The modified audio is constrained by |δ(t)| ≤ l to limit deviation from the original song.
- Gradient Descent to Craft Audio: Gradient descent iteratively revises x(t) into x′(t) = x(t) + δ(t) to minimize the objective function.The optimization repeats until the objective value becomes stable, yielding a local minimum.
- Practical Attack over the Air: The over-the-air sample adds a captured noise model to the perturbation, while random noise improves robustness across speakers and receivers.The resulting adversarial audio is x′(t) = x(t) + µ(t), and the evaluation reports robustness to different speakers and receivers.
5 Evaluation
The evaluation tests CommanderSong against machine recognition, human perception, cross-platform transfer, remote delivery, and generation efficiency. Results show strong recognition performance with limited human detection, though transferability varies by command and platform.
- Experiment Setup: More than 200 adversarial songs were generated from 26 songs and 12 commands for evaluation.The commands included common actions such as turning on GPS and making a credit-card payment.
- Machine Recognition: 100% WTA success was achieved against Kaldi, with every command recognized correctly.The success rate counts correctly decoded words, treating a one-character difference as incorrect.
- Machine Recognition: Up to 96% WAA success was achieved over the air using a JBL speaker, although perturbations had SNRs below 2 dB.The pseudo-device used an iPhone 6S microphone and Kaldi decoder, with the speaker 1.5 meters away in a meeting room.
- Human Comprehension: Soft music was the best carrier for hiding commands, with as few as 15% of participants noticing abnormality.No participants recognized any word of the injected command; in WAA samples, fewer than 1% believed there were words beyond the lyrics.
- Transferability: Transfer to iFLYTEK succeeded for some commands but failed or weakened for others, while direct transfer to DeepSpeech was unsuccessful.“Open the door” and “good night” transferred well, whereas “airplane mode” achieved 66% on iFLYREC and 0% on iFLYTEK Input.
- Remote Delivery: CommanderSongs could be delivered through YouTube and simulated radio broadcasts, with iFLYTEK Input decoding commands at distances up to 0.5 meter.The radio experiment used FM 103.4 MHz and tested multiple smartphones.
- Efficiency: Most samples were generated in under two hours, while simple commands could be completed within half an hour.Commands containing words such as GPS or airplane, and some rock songs, required more generation time.
6 Understanding the Attacks
The analysis examines how songs support stealthy adversarial commands, how noise affects robustness and similarity, and how machine and human recognition diverge. It finds that song content and perturbation amplitude jointly shape attack effectiveness and detectability.
- Songs conceal injected commands from listeners and can be distributed through YouTube, radio, and television.
- Soft music is the best carrier because its components can align with phonemes or smaller units in the target command.
- Increasing noise lowers correlation with the original song because stronger perturbations are needed for robustness against interference.
- At SNR = 4 dB, CommanderSong reaches an 88% success rate while retaining 90% correlation with the original song.
- WTA can be recognized by Kaldi while remaining unnoticed by humans, whereas WAA becomes human-recognizable as noise increases.
7 Defense
The paper evaluates audio turbulence and audio squeezing as defenses for CommanderSong. Audio turbulence mainly defeats WTA, while audio squeezing reduces both WTA and WAA success.
- Audio turbulence: Audio turbulence adds noise before ASR processing and checks whether the perturbed input is decoded as different words.
- Audio turbulence: At SNR = 15 dB, WTA almost always fails while the clean command remains recognizable.
- Audio turbulence: WAA remains highly successful against audio turbulence because its random-noise construction is robust to turbulence noise.
- Audio squeezing: Audio squeezing detects CommanderSong when the original and downsampled inputs produce different ASR transcriptions.
- Audio squeezing: At 1/M = 0.7, WTA and WAA success rates are 0% and 8%, respectively, while the clean command succeeds at 91%.
8 Related Work
Related work covers attacks on speech-controlled systems, adversarial machine learning in physical settings, and defenses against adversarial examples. The paper distinguishes CommanderSong from prior methods by targeting DNN-based ASR with practical over-the-air attacks.
- Attack on ASR system: Prior speech attacks exploited analog sensors, electromagnetic interference, permission bypasses, voice impersonation, or microphone hardware vulnerabilities.
- Attack on ASR system: Hidden voice command attacked GMM-based ASR, while CommanderSong targets DNN-based ASR systems.
- Attack on ASR system: DeepSpeech adversarial examples required directly uploading an adversarial WAV file to the speech recognizer.
- Adversarial research on machine learning: Physical-world image attacks demonstrated robust perturbations and classifier failures, whereas this study targets speech recognition.
- Defense of Adversarial on machine learning: Existing adversarial defenses include adversarial training and defensive distillation.
9 Conclusion
The paper presents CommanderSong as a practical attack that embeds voice commands into songs for over-the-air ASR control. It reports transfer to iFLYTEK, broad distribution channels, and two defenses.
- CommanderSong injects voice commands into songs that ASR systems can execute over the air without users noticing.
- The attack transfers to iFLYTEK and can affect popular applications including WeChat, Sina Weibo, and JD.
- CommanderSongs can spread through YouTube and radio, while audio turbulence and audio squeezing defend against the attack.
Appendix
The appendix reports song familiarity and Spotify streaming counts for WTA survey samples, and provides user-scale information for sample iFLYTEK applications.
- Song survey samples: Table 7 reports individual-song human-comprehension survey results for WTA samples alongside Spotify streaming counts.The song Selling Brick in Street is absent from Spotify, so no streaming count is provided.
- Song survey samples: MTurk workers’ average familiarity with the songs was lower than expected.
- iFLYTEK applications: Table 8 lists sample applications using iFLYTEK as voice input, including Google Play downloads and total user amounts.
- iFLYTEK applications: Each listed iFLYTEK application has over 0.2 billion users worldwide.
- iFLYTEK applications: Google Play downloads may not correspond to total users because Google Services are inaccessible in China and Apple App Store information was not collected.