Source-linked AI summary
Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
Wenpeng Xing, Minghao Li, Mohan Li, Meng Han
TL;DR
Embodied AI systems face vulnerabilities across sensor–actuator–algorithm interactions, while training and data limitations increase susceptibility to adversarial manipulation. This survey categorizes these vulnerabilities across exogenous, endogenous, and inter-dimensional dimensions and proposes strategies including world grounding, multimodal integration, and robust control.
Problem
Interactions between sensors, actuators, and algorithms expose embodied AI systems to a broad spectrum of vulnerabilities, while training and data limitations exacerbate vulnerability to adversarial manipulation and misuse.
Method
The survey categorizes embodied AI vulnerabilities into exogenous, endogenous, and inter-dimensional dimensions and proposes targeted safety and reliability strategies.
Results
The survey provides a framework spanning these vulnerability dimensions and recommends world grounding, multimodal integration, and robust control and adaptation mechanisms.
Takeaways & Limitations
Improving embodied AI safety requires combining stronger world grounding, multimodal integration, and robust control and adaptation mechanisms.
Takeaways & Limitations
Pre-training on uncurated internet data and limited mitigation from safety-dataset fine-tuning constrain protection against adversarial manipulation and misuse.
Abstract
from arXiv · showhide
Embodied AI systems, including robots and autonomous vehicles, are increasingly integrated into real-world applications, where they encounter a range of vulnerabilities stemming from both environmental and system-level factors. These vulnerabilities manifest through sensor spoofing, adversarial attacks, and failures in task and motion planning, posing significant challenges to robustness and safety. Despite the growing body of research, existing reviews rarely focus specifically on the unique safety and security challenges of embodied AI systems. Most prior work either addresses general AI vulnerabilities or focuses on isolated aspects, lacking a dedicated and unified framework tailored to embodied AI. This survey fills this critical gap by: (1) categorizing vulnerabilities specific to embodied AI into exogenous (e.g., physical attacks, cybersecurity threats) and endogenous (e.g., sensor failures, software flaws) origins; (2) systematically analyzing adversarial attack paradigms unique to embodied AI, with a focus on their impact on perception, decision-making, and embodied interaction; (3) investigating attack vectors targeting large vision-language models (LVLMs) and large language models (LLMs) within embodied systems, such as jailbreak attacks and instruction misinterpretation; (4) evaluating robustness challenges in algorithms for embodied perception, decision-making, and task planning; and (5) proposing targeted strategies to enhance the safety and reliability of embodied AI systems. By integrating these dimensions, we provide a comprehensive framework for understanding the interplay between vulnerabilities and safety in embodied AI.
1 INTRODUCTION
Embodied AI combines autonomy, physical interaction, and cognition to perform complex real-world tasks, but its interconnected sensors, actuators, and algorithms create broad security and safety vulnerabilities. The survey organizes these vulnerabilities, attack strategies, failure modes, evaluation datasets, and mitigation directions into a unified framework.
- Embodied AI systems integrate perception, decision-making, and actuation for autonomous operation in dynamic real-world environments.
- Interactions among sensors, actuators, and algorithms expose embodied systems to environmental complexity, sensor jamming and spoofing, and system failures.
- These vulnerabilities can cause decision-making errors, unauthorized actions, physical risks, and damage to system reputation, motivating robust security measures.
- The survey classifies vulnerabilities as exogenous, endogenous, and inter-dimensional, including physical attacks, sensor validation failures, and ethical challenges in interactive agents.
- It analyzes cybersecurity, sensor spoofing, and adversarial attacks, including attacks on LLMs and LVLMs through logits, prompts, and cross-modality interactions.
- The survey also examines embodied-AI failure modes and evaluation datasets before proposing future research directions for safety and reliability.
2 VULNERABILITY CATEGORY
The survey divides embodied-AI risks into exogenous, endogenous, and inter-dimensional vulnerabilities, then details environmental, physical, adversarial, cybersecurity, and human-interaction threats. These risks can produce unsafe decisions, unauthorized control, physical damage, and operational disruption.
- Exogenous vulnerabilities arise from external environments or malicious actors, while endogenous vulnerabilities originate in hardware, software, or design flaws.
- Inter-dimensional vulnerabilities occur when external and internal factors interact to exacerbate system fragility.
- The exogenous taxonomy covers dynamic environmental factors, physical attacks, adversarial attacks, cybersecurity threats, and human-interaction or safety-protocol failures.
- Dynamic Environmental Factors: Environmental changes and adversarial perturbations can mislead perception systems, and insufficient neural-network verification complicates safety-critical deployment and regulatory compliance.
- Adversarial Attacks: Adversarial examples can deceive perception and decision-making in autonomous vehicles and robotic surgery, while robustness against real-world attacks remains an open problem.
- Cybersecurity Threats: Cybersecurity weaknesses such as default passwords or unencrypted communications can enable unauthorized access, operational disruption, DDoS attacks, and GPS-spoofing hijacks.
- Human Interaction and Safety Protocol Failures: Compromised safety protocols in collaborative robots can cause severe human consequences, including a worker death after a malfunctioning cobot failed at a Volkswagen plant.
2.2 Endogenous Vulnerability
Endogenous vulnerabilities arise inside embodied systems through hardware failures, software flaws, design defects, and failures in instruction handling or generalization. These internal weaknesses can produce unsafe actions, operational disruption, and harm in real-world deployments.
- Endogenous risks originate within embodied systems and include hardware failures, software bugs, and design flaws that may have severe consequences without mitigation.
- Sensor Failures: Sensor malfunctions or improper input validation can cause incorrect environmental assessments, including missed obstacles or misjudged distances in autonomous vehicles.
- Hardware and Mechanical Failures: Mechanical wear and failures in industrial and medical robots can disrupt operations, create safety hazards, or require emergency conversion from robotic to open surgery.
- Software and Design Flaws: Software bugs, inadequate edge-case handling, insufficient testing, and flawed control algorithms can produce unsafe driving or disrupt industrial-robot operations.
- Interacting Risks: External attacks can exploit internal software weaknesses, while temperature or humidity can accelerate hardware degradation and mechanical failure.
- Instruction Misinterpretation: Instruction misinterpretation in driving and navigation tasks can cause dangerous behavior, incorrect manipulation, environmental damage, or human harm.
- Interactive Agents and Unseen Environments: Interactive agents may provide misleading information, while poor generalization to unseen environments can cause navigation failures, delays, or hazardous situations.
3 ATTACK TO VULNERABILITIES
The survey examines how embodied systems and multimodal large models expand the attack surface across perception, reasoning, language, and action. It structures this analysis through vulnerability assessment, threat modeling, and taxonomies of cybersecurity, sensor-spoofing, and adversarial attacks.
- LVLMs improve perception, reasoning, and interaction across vision-language tasks, but multimodal inputs expand their attack surface.
- Attackers can exploit visual inputs, textual prompts, or interactions between modalities to create sophisticated attack vectors against existing defenses.
- Vulnerability Analysis: The survey analyzes expanded attack surfaces, adversarial vulnerabilities, and risks arising when systems transition from text to physical actions.
- Threat Model and Attack Taxonomy: Its threat model describes attacker capabilities and potential targets before presenting taxonomies of exogenous, endogenous, and inter-dimensional vulnerability-centric attacks.
- Attack Methodologies: Specific attack categories include cybersecurity threats, sensor spoofing attacks, and adversarial attacks targeting embodied systems.
3.1 Vulnerability Analysis
LLM and LVLM vulnerabilities arise from limitations in training, data, neural-network behavior, and multimodal processing. In embodied systems, these weaknesses extend to action planning and structured outputs that downstream controllers execute.
- Training and Data Limitations: LLM next-word training diverges from helpful, truthful, and harmless response goals, while safety fine-tuning offers limited mitigation.Uncurated pre-training data can introduce bias and toxic content, increasing vulnerability to manipulation and misuse.
- Adversarial Vulnerabilities: DNNs are vulnerable to adversarial examples because small input perturbations can cause significant prediction shifts.Incomplete training coverage, steep gradients near decision boundaries, and sensitivity to high-frequency components further expose model–human perception mismatches.
- Expanded Attack Surface: LVLM multimodality allows adversarial signals in one modality to affect others, expanding risks for embodied and autonomous-driving systems.Manipulated visuals can disrupt textual reasoning, and reported attacks pose significant autonomous-driving risks.
- Embodied-System Attack Paradigms: Traditional jailbreak attacks may not fully apply to embodied LLMs because these systems plan and execute physical actions.Structured JSON or YAML outputs can be manipulated by attackers, causing downstream controllers to produce unsafe or unintended behavior.
3.2 Threat Model
The threat model characterizes attacks by attacker access and targeted system components. It distinguishes white-, gray-, and black-box capabilities and highlights perception, control, and communication targets.
- Attacker Capabilities: White-box attackers access system architecture, parameters, and APIs, enabling highly targeted attacks such as FGSM, PGD, APGD, and CW.This setting is common in open-source embodied AI systems or simulators.
- Attacker Capabilities: Gray-box attackers use partial access through high-level APIs or external interfaces and often exploit sensor data or user commands.They lack control over lower-level components.
- Attacker Capabilities: Black-box attackers lack internal system knowledge and interact through input queries, yet can still exploit external inputs.The threat persists despite proprietary protections in commercial systems.
- Attack Targets: Perception attacks on cameras, LiDAR, or GPS can cause faulty environmental interpretation and decision-making.Sensor spoofing may make autonomous vehicles misjudge distances or miss obstacles; GPS spoofing has also hijacked a civilian drone.
- Attack Targets: Control attacks can produce unsafe robot movements, while communication attacks can disrupt critical updates or commands through jamming or MitM.The section presents these components as distinct attack targets in embodied AI systems.
3.3 Attack Taxonomy
The attack taxonomy organizes threats by attacker capabilities, targeted components, and vulnerability origins. It spans adversarial and sensor manipulation, command and jailbreak attacks, infrastructure compromise, and coordinated threats.
- Input Manipulation Attacks: Adversarial attacks manipulate inputs to produce incorrect predictions, whereas sensor spoofing directly manipulates physical environments or sensor signals.Spoofed LiDAR, GPS, or camera data can cause unintended behavior, collisions, navigation failures, or malicious voice-command effects.
- Input Manipulation Attacks: Command injection targets voice or text interfaces to trigger unintended actions such as unlocking doors, disabling security, or dangerous maneuvers.Jailbreak attacks instead bypass LLM or LVLM safety mechanisms through prompt manipulation or adversarial prompting.
- System and Infrastructure Attacks: System and infrastructure attacks include API manipulation, denial-of-service, and MitM attacks that disrupt navigation, responsiveness, data integrity, or control commands.Strong encryption and intrusion detection are described as countermeasures for MitM attacks.
- Vulnerability-Centric Attacks: Endogenous attacks exploit model, firmware, side-channel, zero-day, and data vulnerabilities, while exogenous attacks include supply-chain compromise.Firmware attacks can persist below higher-level defenses, and data poisoning alters what models learn.
- Sophisticated and Coordinated Attacks: Inter-dimensional and coordinated threats include advanced persistent threats and ransomware that stealthily compromise systems or disable critical operations.Ransomware can lock control software, create cascading failures, and render autonomous vehicles or industrial robots inoperable.
3.4 Cybersecurity Threat
Cybersecurity threats target communications, sensor integrity, firmware, side channels, and operational availability in connected embodied AI systems. The survey pairs these risks with defenses including encryption, intrusion detection, sensor fusion, anomaly detection, secure boot, and recovery mechanisms.
- Communication and Infrastructure Threats: MitM and related communication attacks can compromise system integrity and confidentiality by intercepting or modifying data across connected components.Weak encryption broadens the attack surface, while TLS, IPsec, and intrusion detection are proposed defenses.
- Sensor-to-Model Attacks: Sensor-to-model attacks can corrupt real-time perception and control, causing unsafe rerouting, mission failures, collisions, production defects, or human harm.Sensor fusion, anomaly detection, and cryptographic protection of sensor communication are described as resilience measures.
- Sensor-to-Model Attacks: Spoofed LiDAR or camera data can create nonexistent obstacles, prompting unnecessary evasive maneuvers or collisions in autonomous driving.Manipulated sensor data can similarly cause industrial robots to misinterpret environments and create safety hazards.
- Firmware Attacks: Firmware compromise can provide persistent control over sensors, actuators, or communication protocols and bypass higher-level security mechanisms.A malicious drone firmware update may alter flight-control parameters and cause erratic or unsafe behavior; secure boot and runtime verification are proposed defenses.
- Side-Channel Attacks: Side-channel attacks can infer sensitive information from power, electromagnetic, or timing characteristics during cryptographic operations.The survey motivates cryptographic algorithms resistant to such leakage to protect confidentiality and critical operations.
- Ransomware Attacks: Ransomware can immobilize autonomous vehicles or halt industrial workflows by encrypting critical files or locking control software.Backup, recovery, and proactive detection mechanisms are proposed to reduce operational and safety-critical impacts.
3.5 Sensor Spoofing Attacks
Sensor spoofing attacks exploit vulnerabilities across tactile, vision, and other sensing pipelines to corrupt embodied agents’ perception and trigger unsafe behavior. The survey describes attack mechanisms and their effects on recognition, navigation, tracking, and control.
- Sensor-based attacks target data acquisition, processing, and interpretation, allowing attackers to bypass safeguards and compromise system integrity.The survey frames spoofing as a security concern spanning the sensor data pipeline rather than a single sensor type.
- Attack on Tactile Sensors: Tactile sensors face wear, manipulated readings, tampering, force or pressure attacks, and capacitive cross-talk that can produce erroneous responses and compromise decision-making.These weaknesses affect both sensing accuracy and downstream robotic actions.
- Attack on Vision Sensors: Vision attacks use perturbations, patches, projected light, or invisible-spectrum illumination to induce object and scene misinterpretation.The described techniques include sticker-pasting, light projection, and infrared manipulation of camera sensors.
- Attack on Vision Sensors: Perception manipulation can disrupt object recognition and lane detection, causing hazardous control decisions, unintended path deviations, and impaired real-time navigation.Examples include misclassifying stop signs, fabricating or erasing obstacles, altering lane markings, and disrupting object tracking.
- IMU and Other Sensor Spoofing Attacks: Attacks on velocity, navigation, or vehicle-control inputs can destabilize drones, mislead autonomous vehicles, cause red-light violations, and induce off-road or oncoming-traffic deviations.The reported consequences extend from flight-control instability to dangerous autonomous-driving behavior.
- Attacks on Proximity and Other Sensors: Spoofing can create false obstacles, hide real hazards, distort object locations, and produce inaccurate readings that lead to abrupt stops, unexpected lane changes, or navigation errors.The survey also reports sensor-specific effects on obstacle detection and spatial interpretation.
3.6 Adversarial Attack to LLMs and LVLMs
The survey organizes LLM/LVLM attacks into white-box and black-box paradigms, covering logits manipulation, fine-tuning, adversarial prompts, cross-modality attacks, transferability, evaluation, and mitigation. Reported findings show vulnerabilities in visual components, safety alignment, watermarking, and attack transfer across models.
- Attack taxonomy: LLM and LVLM attacks are categorized into white-box methods, including logits-based and fine-tuning attacks, and black-box methods centered on adversarial prompt generation.The survey also covers cross-modality attacks, attack transferability, evaluation strategies, and safety mitigation techniques.
- White-box attacks: VT-Attack targets encoded visual tokens in LVLMs, causing visual misinterpretations that produce incorrect or harmful outputs.The finding exposes vulnerabilities in LVLM visual components.
- White-box attacks: Fine-tuning attacks can steer models toward attacker objectives but may compromise generated-text naturalness and coherence.Even predominantly benign fine-tuning datasets can degrade model safety, while maliciously constructed datasets increase jailbreak vulnerability.
- Black-box attacks: Black-box prompt attacks include suffix embedding, prompt rewriting, template completion, multilingual prompting, heuristic substitution, and retrieval-augmented jailbreaks.Examples include AutoDAN suffixes that evade perplexity filters and low-resource-language prompts that bypass safety mechanisms.
- Black-box attacks: GCG iteratively optimizes discrete adversarial suffixes and transfers effectively to black-box models such as ChatGPT, Bard, and Claude.The method uses top-k gradient-based replacements, random token sampling, and best-replacement updates.
- Transferability and evaluation: Attack transferability varies across model families: it is limited across different LVLMs but improves among similarly trained models or closely related checkpoints.Broader targeting of highly similar LVLMs enhances transferability, while LoRA backdoors can persist across multiple adaptation modules.
- Safety mitigation: Text-domain forgetting significantly reduces LVLM attack success rates, whereas multimodal forgetting adds no reported advantage and requires substantially more computation.This supports prioritizing text-domain safety optimization in cross-modal safety alignment.
4 CHALLENGES AND FAILURE MODES
The survey describes how LLMs and LVLMs support embodied perception, reasoning, and action while introducing failure modes that threaten performance, safety, and reliability. It emphasizes shared weaknesses involving adversarial manipulation, distribution shifts, ambiguity, uncertainty, bias, and limited explainability.
- Embodied capabilities: LVLM- and LLM-based systems support embodied agents through visual-language reasoning, closed-loop action refinement, and object- or state-centric interaction.Examples include GPT-4V-based planning, ViLA’s visual feedback adaptation, and MultiPLY’s action and state tokens.
- Embodied capabilities: MultiPLY transitions between abstract reasoning and detailed multimodal observations, demonstrating versatility across diverse interactive scenarios.
- Failure modes: Despite these capabilities, LVLM failure modes can compromise performance, safety, and reliability in dynamic and complex environments.The survey presents addressing these limitations as necessary for realizing embodied AI’s potential.
- Shared limitations: LLMs and LVLMs face adversarial manipulation, including harmful prompts and small visual or textual perturbations that can cause substantial errors or unsafe behavior.The models also share risks from biased training data and limited explainability in safety-critical applications.
- Shared limitations: LLMs and LVLMs struggle with distribution shifts because reliance on training data limits adaptation to unfamiliar, dynamic, or unpredictable scenarios.For LLMs this can cause command misinterpretation or poor decisions; LVLMs similarly struggle when real-world inputs differ from training data.
- Shared limitations: Ambiguous or incomplete inputs can produce overconfident outputs because these models cannot effectively express or quantify uncertainty.The resulting errors are especially risky in environments where ambiguity is common or unavoidable.
5 DATASET TAXONOMY FOR LLMS AND LVLMS EVALUATION
The survey organizes LLM and LVLM evaluation datasets by their purposes, spanning general multimodal capabilities, adversarial robustness, safety, and alignment. These categories support assessment of model performance, resilience, ethical compliance, and instruction alignment.
- Dataset taxonomy: Evaluation datasets are categorized into General, Red Team, Robustness Evaluation, and Alignment Datasets, each serving a distinct assessment purpose.The taxonomy covers core capabilities, harmful-content testing, resilience to adversarial or ambiguous inputs, and model alignment.
- General Datasets: General datasets assess image classification, captioning, visual question answering, visual understanding, reasoning, and language generation.Examples include ImageNet, COCO Captions, RefCOCO, and VQA V2.
- Adversarial and robustness datasets: Adversarial datasets stress-test models against harmful content, adversarial visual instructions, and image-text mismatches.Red Team datasets target overtly harmful content, while robustness datasets include adversarial samples and sensitivity tests.
- Adversarial and robustness datasets: Many state-of-the-art models show significant performance degradation when evaluated on adversarial samples.The cited examples include diverse adversarial visual instructions and adversarial text-image inputs.
- Alignment Datasets: Alignment datasets support RLHF, preference modeling, and instruction tuning to balance helpfulness, harmlessness, safety, and usability.SPA-VL provides safety preference data, while VLFeedback contains over 82,000 multimodal instructions and AI-generated rationales.
6 CONCLUSION
The conclusion frames embodied AI as vulnerable across external, internal, and intersecting factors, with attacks targeting sensors, perception, and language models. It proposes comprehensive evaluation and safety strategies spanning multimodal integration, world grounding, robust control, and verification.
- Conclusion: Embodied AI systems face vulnerabilities because interactions among sensors, actuators, and algorithms expose them to security threats and system failures.The risks include dynamic environments, sensor jamming and spoofing, and failures that compromise safety and reliability.
- Vulnerability taxonomy: The survey categorizes vulnerabilities into exogenous, endogenous, and inter-dimensional dimensions.Exogenous vulnerabilities arise externally, endogenous vulnerabilities involve internal failures, and inter-dimensional vulnerabilities occur at their intersection.
- Attack paradigms: The survey examines attack vectors targeting sensor systems, perception, LVLMs, and LLMs through visual, textual, and multimodal perturbations.These attacks can deceive models and disrupt embodied-system operations.
- Evaluation and mitigation: The survey highlights comprehensive benchmarking and evaluation frameworks as necessary for assessing robustness.It also proposes targeted strategies to enhance embodied AI safety and reliability.
- Evaluation and mitigation: Proposed safety strategies include voxel-based and neural-radiance-field world grounding, multimodal integration, robust control, adaptation, formal verification, Sim2Real testing, and redundant safety mechanisms.Together, these strategies address representation, integration, control, adaptation, and safety-critical verification.
- Conclusion: The survey provides a comprehensive framework for understanding the interplay between vulnerabilities and safety in embodied AI systems.Its scope integrates vulnerability categories, attack vectors, robustness evaluation, and mitigation strategies.