Source-linked AI summary

Light Commands: Laser-Based Audio Injection Attacks on Voice-Controllable Systems

Takeshi Sugawara, Benjamin Cyr, Sara Rampazzi, Daniel Genkin, Kevin Fu

arXiv:2006.11946v1cs.CR

TL;DR

Voice-controllable systems often lack robust authentication, raising the question of whether attackers can inject commands remotely and stealthily. LightCommands modulates laser light to exploit microphones’ unintended light response and convert the light into audio. The authors demonstrate command injection across voice assistants at distances up to 110 meters, including across buildings, with consequences for connected locks, garages, purchases, and vehicles.

  • Problem

    Voice-controllable systems often lack proper user authentication, while existing stealthy command-injection techniques are limited to proximity ranges and weakened by physical barriers.

  • Method

    LightCommands amplitude-modulates laser light with an audio signal so a target microphone converts the light into injected audio and voice commands.

  • Results

    The authors demonstrate successful light-based command injection on popular Alexa, Siri, Portal, and Google Assistant systems at distances up to 110 meters, including through closed glass windows and across buildings.

  • Takeaways & Limitations

    Light-injected commands can compromise connected locks, garage doors, e-commerce accounts, and vehicles when voice systems execute commands without adequate authentication.

  • Takeaways & Limitations

    LightCommands assumes line-of-sight access and does not properly penetrate opaque obstacles; precise aiming and higher laser power can also make attacks on mobile devices challenging.

Abstract

from arXiv · show

We propose a new class of signal injection attacks on microphones by physically converting light to sound. We show how an attacker can inject arbitrary audio signals to a target microphone by aiming an amplitude-modulated light at the microphone's aperture. We then proceed to show how this effect leads to a remote voice-command injection attack on voice-controllable systems. Examining various products that use Amazon's Alexa, Apple's Siri, Facebook's Portal, and Google Assistant, we show how to use light to obtain control over these devices at distances up to 110 meters and from two separate buildings. Next, we show that user authentication on these devices is often lacking, allowing the attacker to use light-injected voice commands to unlock the target's smartlock-protected front doors, open garage doors, shop on e-commerce websites at the target's expense, or even unlock and start various vehicles connected to the target's Google account (e.g., Tesla and Ford). Finally, we conclude with possible software and hardware defenses against our attacks.

1 Introduction

Voice-controllable systems offer convenient hands-free control but often lack robust user authentication, leaving long-distance command injection an open security question. LightCommands introduces light-based audio injection and demonstrates remote control of voice assistants and connected devices.

  • Motivation: Voice-controllable systems increasingly control services and IoT devices through natural-language commands without physical interaction.The paper frames Alexa, Siri, Portal, and Google Assistant as widely deployed interfaces for connected products.
  • Security Gap: Prior work identifies inadequate user authentication as a major limitation of voice-only interaction, enabling commands from malicious sources.Stealthier attacks also aim to prevent owners from hearing or recognizing injected commands.
  • Security Gap: Existing injection techniques generally rely on proximity, with a reported open-space range of about 25 ft (7.62 m), reduced further by physical barriers.The paper asks whether commands can instead be injected remotely and stealthily under realistic conditions.
  • LightCommands: LightCommands exploits microphones’ unintended response to modulated light to inject audio and remotely control voice-controllable systems.The authors characterize popular Alexa, Siri, Portal, and Google Assistant devices across distances and laser powers.
  • LightCommands: 110 meters is the demonstrated maximum command-injection distance, including attacks across buildings and through closed glass windows.A telephoto lens focuses the laser for long-range operation; the reported distance was the maximum safely available to the researchers.
  • Implications: Light-injected commands can unlock doors, open garages, make e-commerce purchases, and locate, unlock, or start connected Tesla and Ford vehicles.The paper also presents a cheap setup using commercially available laser pointers and drivers, with infrared lasers and volume features used to reduce discovery risk.
  • Mitigations: The paper discusses software and hardware countermeasures and reports responsible disclosure to affected vendors and relevant organizations.The disclosed vendors include Google, Amazon, Apple, Facebook, August, Ford, Tesla, and Analog Devices.

2 Background

The paper situates LightCommands among attacks on voice systems, sensors, and computing hardware, then reviews voice-controllable systems, MEMS microphones, and laser sources. Prior acoustic attacks face substantial range limits, while the proposed approach relies on focused, amplitude-modulated laser light and has safety concerns.

  • Voice-Controllable Systems: A voice-controllable system primarily uses natural-language voice commands and may execute them immediately without further interaction.Its typical pipeline includes voice capture, speech recognition, and command execution.
  • Attacks on Voice Systems: Earlier attacks injected audible, camouflaged, or inaudible commands into voice-controllable systems, while skill-squatting attacks exploit recognition errors.Ultrasonic approaches use microphone nonlinearities or word modulation to hide commands from human listeners.
  • Attacks on Voice Systems: 2 cm to 175 cm is the reported range for two ultrasonic attacks, while a later 61-speaker approach reached about 25 ft (7.62 m) in open space.Higher transmitting power can reintroduce audible leakage, and windows and air absorption further attenuate ultrasonic signals.
  • Related Signal-Injection Attacks: Prior work also used acoustic or light-based injection against IMUs, ultrasonic sensors, LiDARs, infusion pumps, cameras, and semiconductor memory.These studies demonstrate denial of service, spoofing, precise sensor control, illusory objects, and light-induced faults.
  • MEMS Microphones: MEMS microphones use a diaphragm and back plate as a variable capacitor, with an ASIC converting capacitive changes into an output voltage.Their small footprints and low prices make MEMS microphones popular in mobile and embedded devices.
  • Laser Sources: The paper focuses on laser-emitting diodes because their beam remains narrow over distance and their intensity can encode analog signals through amplitude modulation.The discussed low-power Class 3R systems emit less than 5 mW at visible wavelengths, while measured commercial products sometimes reached 1 W despite 5 mW labels.

3 Threat Model

The threat model considers remote, stealthy command injection against voice-controllable devices without physical access or owner interaction, assuming line of sight to the device and microphones.

  • Attackers are assumed to lack physical access and cannot make the owner perform useful interactions such as pressing buttons or unlocking the screen.
  • The attacker must have remote line of sight to the target device and its microphones, potentially through closed glass windows.
  • The model targets voice-controllable devices whose microphones may be visible despite physical obstructions protecting the system from nearby attackers.
  • Experiments were empirically verified on other available devices of the same model without instance-specific calibration.

4 Injecting Sound via Laser Light

The paper demonstrates that amplitude-modulated laser light can inject audio into MEMS microphones and investigates the electrical and mechanical transduction mechanisms underlying the effect.

  • 4.1 Signal Injection Feasibility: A laser driver amplitude-modulates encoded audio onto diode current, producing corresponding intensity changes that the microphone records as sound.The setup used a laser diode, driver, audio amplifier, microphone breakout board, and oscilloscope.
  • 4.1 Signal Injection Feasibility: A 1 kHz injected sine wave appeared clearly at the microphone output with matching frequency and no noticeable distortion.
  • 4.2 Characterizing Laser Audio Injection: Above each diode’s threshold current, optical power increased linearly with current when IDC − Ipp/2 > Ith.
  • 4.2 Characterizing Laser Audio Injection: Increasing the driving AC current Ipp linearly increased the sound volume received by the microphone for both blue and red laser diodes.
  • 4.2 Characterizing Laser Audio Injection: 20 Hz–20 kHz frequency responses for both lasers covered the entire audible band, implying injection of arbitrary audio signals.
  • 4.3 Mechanical or Electrical Transduction?: Less than 0.1 mW of laser power saturated the microphone when a focused spot targeted its ASIC, while opaque epoxy eliminated ASIC-targeted signals but not diaphragm-targeted signals.
  • 4.3 Mechanical or Electrical Transduction?: The results support both photoelectric transduction in the ASIC and light-induced mechanical vibration in the MEMS diaphragm.For attacks through the acoustic port, the authors hypothesize that both components are illuminated and contribute jointly.

5 Attacking Voice-Controllable Systems

The paper evaluates laser-based command injection across voice-controllable devices, varying laser power, distance, hardware, and authentication settings. All tested devices were susceptible, with sensitivity and reliability depending on device, power, distance, and recognition backend.

  • Device selection: 17 voice-controllable and third-party devices were benchmarked across Alexa, Siri, Portal, Google Assistant, and built-in speech-recognition systems.Different device generations were included to examine hardware variation, and the EcoBee thermostat represented a third-party device.
  • Evaluation criteria: The evaluation defined injection success as recognizing all four commands in three consecutive attempts, while individual command feasibility required recognizing every command word.Commands requiring attached IoT devices were assessed for command recognition rather than end-to-end actuation in this section.
  • Minimum power requirements: All tested devices were susceptible to laser-based command injection, including microphones covered by fabric or foam.The result indicates that acoustic-port coverings did not prevent the tested devices from responding to injected light-based audio.
  • Backend and hardware effects: Portal Mini required 6× more minimum power for Alexa than for “Hey Portal,” and Facebook’s backend failed to recognize “laser” in the final command.Because both experiments used the same setup and microphone, the authors attribute the difference to algorithmic differences between the recognition backends.
  • Attack range: 5 mW enabled command injection over tens of meters for especially sensitive Google Home and Eco Plus speakers, while most devices required 60 mW.The 5 mW tests used a 110 m hallway and the 60 mW tests used a 50 m hallway; a plus sign marked ranges truncated by hallway length.
  • Attack success probability: Google Home Mini injection was nearly always successful at 20 m, but no successful injections were observed at 27 m.Across 40 consecutive command injections, performance dropped sharply with distance, suggesting factors beyond command phonemes influence success probability.
  • Speaker authentication: Smart speakers generally lacked speaker authentication: previously unheard voices could execute security-critical commands such as unlocking doors, although voice purchasing was blocked for unrecognized voices.The paper distinguishes speaker recognition for personalization from speaker authentication for access control.
  • Speaker authentication: Phone and tablet authentication could be bypassed with recordings or synthesized owner-like voices, and the authors found matching artificial voices for all four tested tablets.The authors conclude that voice recognition alone does not provide sufficient entropy as a countermeasure to command injection.

6 Exploring Various Attack Scenarios

The paper tests LightCommands under realistic attack conditions, examines authentication weaknesses in connected devices, and evaluates ways to reduce acoustic and optical detectability.

  • 6.1 A Low-Power Cross-Building Attack: The cross-building setup used a telescope and telephoto lens to locate visible microphone ports and aim the laser at the target device.The experiment was conducted at night, while illuminated long-range attacks were reported elsewhere in the paper.
  • 6.1 A Low-Power Cross-Building Attack: 5 mW laser power successfully injected a voice command into an upright Google Home through a closed double-pane window from another building.The beam struck top-facing microphones at a 21.8-degree angle despite wind-induced wobbling.
  • 6.2 Attacking Authentication: PIN authentication can remain weak because Google Assistant delegates it to vendors, while August permits 1- to 6-digit PINs without attempt limits or delays.The paper reports that exhaustive enumeration of a 4-digit PIN space required about 36 hours, with successful unlocking when the correct PIN was reached.
  • 6.2 Attacking Authentication: Authentication bypasses expanded the attack surface: some garage-door commands required no authentication, and connected Tesla and Ford vehicles exposed different control limitations.Tesla voice commands controlled several vehicle functions without a PIN but could not start the car without key proximity; Ford allowed remote doors and engine start, but shifting out of Park stopped the engine.
  • 6.4 Exploring Stealthy Attacks: Attackers reduced detectability by lowering device response volume or using features such as Alexa whisper mode, and by aiming infrared lasers invisible to the human eye.Infrared injection succeeded at about 30 centimeters, but range experiments were omitted because prolonged exposure to invisible laser beams could damage eyes.
  • 6.5 Avoiding the Need for Precise Aiming: Precise aiming at microphone ports remained a limitation, although the authors used geared tripod heads and telescopes to support targeting.The telescope also helped determine the assistant type and microphone location from the device’s appearance.

7 Countermeasures and Limitations

The paper discusses software and hardware defenses for detecting or blocking light-based injection, while identifying line of sight, aiming, and device mobility as important limitations.

  • 7 Countermeasures: Randomized questions or additional authentication can mitigate command injection when attackers cannot eavesdrop on device responses.The paper also notes that manufacturers can compare signals across multiple microphones to detect illumination of only one microphone.
  • 7 Countermeasures: Sensor-rich devices may use intrusion detection because LightCommands differ from normal audible commands in their sensor signatures.The authors leave further exploration of this direction to future work.
  • 7 Countermeasures: Light-blocking barriers, diffracting films, silicon plates, movable shutters, or opaque microphone covers can reduce light reaching the microphone diaphragm.These designs aim to block straight light beams while allowing sound waves to detour around the barrier.
  • 7.3 Limitations: LightCommands assumes line-of-sight access and generally cannot penetrate opaque obstacles that sound can traverse.The authors note that fabric-covered microphones may resist attacks when the cover is sufficiently thick.
  • 7.3 Limitations: The attack also requires careful aiming, and mobile devices are especially challenging because they move and may require quicker targeting and higher laser power.A telescope can partially mitigate targeting difficulty by identifying microphone locations from device appearance.

8 Conclusions and Future Work

LightCommands uses modulated light to inject commands into voice-controllable systems from large distances, including through clear glass. The attack exposes further compromises of connected hardware and suggests broader sensor-injection applications.

  • LightCommands converts modulated light into audio within a microphone, enabling command injection into voice-controllable systems from more than 100 meters and through clear glass.The demonstrated systems used Siri, Portal, Google Assistant, and Alexa.
  • Security deficiencies in voice-controllable systems enabled additional compromises of connected third-party hardware, including locks and cars.
  • The same light-based physical principle may support acoustic injection attacks against other sensors, while laser heating may inject false sensor signals.
Loading 2006.11946v1…