Source-linked AI summary
Reinforcement Learning for IoT Security: A Comprehensive Survey
Aashma Uprety, Danda B. Rawat
TL;DR
IoT security is difficult because interconnected systems face diverse attacks, vulnerabilities, and privacy risks. This survey synthesizes RL and DRL countermeasures across IoT and CPS settings, including smart grids and transportation systems, and identifies research challenges. It concludes by organizing current attacks, countermeasures, and open directions for RL-based IoT security.
Problem
IoT systems face numerous attacks, security flaws, vulnerabilities, and privacy risks as connected-device usage expands.
Method
The paper surveys RL and DRL security solutions for IoT attacks and CPS systems, including smart grids and smart transportation systems.
Results
The survey summarizes RL-based countermeasures for attacks including jamming and spoofing, CPS security applications, and associated research challenges.
Takeaways & Limitations
The paper provides an overview of IoT security attacks, RL countermeasures, CPS applications, and open research directions.
Abstract
from arXiv · showhide
The number of connected smart devices has been increasing exponentially for different Internet-of-Things (IoT) applications. Security has been a long run challenge in the IoT systems which has many attack vectors, security flaws and vulnerabilities. Securing billions of B connected devices in IoT is a must task to realize the full potential of IoT applications. Recently, researchers have proposed many security solutions for IoT. Machine learning has been proposed as one of the emerging solutions for IoT security and Reinforcement learning is gaining more popularity for securing IoT systems. Reinforcement learning, unlike other machine learning techniques, can learn the environment by having minimum information about the parameters to be learned. It solves the optimization problem by interacting with the environment adapting the parameters on the fly. In this paper, we present an comprehensive survey of different types of cyber-attacks against different IoT systems and then we present reinforcement learning and deep reinforcement learning based security solutions to combat those different types of attacks in different IoT systems. Furthermore, we present the Reinforcement learning for securing CPS systems (i.e., IoT with feedback and control) such as smart grid and smart transportation system. The recent important attacks and countermeasures using reinforcement learning B in IoT are also summarized in the form of tables. With this paper, readers can have a more thorough understanding of IoT security attacks and countermeasures using Reinforcement Learning, as well as research trends in this area.
I. INTRODUCTION
IoT connects physical and digital systems through networked devices, but its expanding scale and dynamic communication create substantial security, privacy, and trust challenges. The paper surveys reinforcement-learning research addressing these challenges and explains why RL is relevant to IoT security.
- IoT uses sensors and actuators to exchange data between the physical and digital worlds and automate tasks.
- IoT security must address attacks, privacy, device trust, vulnerabilities, and runtime communication in increasingly connected systems.
- Reinforcement learning lets an agent interact with an environment to maximize numerical reward, although convergence to an optimal policy can be time-consuming.
- The paper reviews reinforcement learning research for securing IoT devices and compares RL with other machine-learning techniques.
- The survey covers RL-based protection against IoT threats, reinforcement learning fundamentals, CPS security, and open research challenges.
A. Reinforcement Learning
Reinforcement learning models sequential interaction between an agent and environment, using rewards to improve decisions. Deep reinforcement learning extends this process with neural networks that approximate values or policies for complex tasks.
- At each time step, an RL agent observes a state, selects an action, receives a reward, and transitions to a new state.
- The Bellman equation decomposes a state value into an immediate reward plus a discounted value for the next state.
- Deep reinforcement learning combines deep learning with RL to approximate value functions or policy gradients in complex environments.
- Deep networks can help RL agents optimize policies, while interaction with the environment generates training data for DRL.
- Unlike supervised learning, RL learns through environmental interaction rather than examples with known answers.
D. Why Reinforcement Learning in IoT
IoT security requires adaptive responses because devices operate in dynamic, complex environments and conventional learning methods depend on attack datasets. RL is presented as suitable because it can learn through interaction without prior datasets.
- IoT devices operate in highly dynamic and complex networked environments.
- Supervised and unsupervised methods support intrusion, malware, CPS-attack, and privacy tasks but cannot provide dynamic security responses.
- Collecting datasets for some IoT environments is extremely difficult, leaving no data with which to train conventional models.
- RL can learn without prior datasets by interacting with the environment and generating data during that process.
E. Reinforcement Learning for Securing IoT Against Adversarial Learning environment
The paper presents RL as a security approach for adversarial IoT environments because it incorporates environmental behavior into learning. This is relevant to IoT systems producing diverse, bursty, or continuous data streams.
- RL incorporates the environment’s behavior into the learning process concurrently for IoT security against adversarial environments.
- This feature is positioned for IoT settings where diverse devices produce large volumes of bursty or continuous data.
III. THREATS AND RL BASED SOLUTIONS IN IOT SECURITY
IoT systems span perception, network, and application layers and face diverse attacks that threaten availability, privacy, and safety. The paper surveys these threats and reinforcement-learning-based countermeasures.
- IoT security threats include DoS, eavesdropping, malware, virus injection, privacy leakage, and network, software, and physical attacks.
- DoS attacks flood networks or otherwise disrupt access, threatening human life, financial interests, and legitimate users’ network resources.The 2016 Mirai DDoS attack affected around 65,000 IoT devices within its first 20 hours.
- The IoT architecture comprises perception, network, and application layers that respectively sense physical objects, process and transmit information, and realize applications.
2) DoS Attack in IoT layers: •
DoS attacks affect all three IoT layers, while reinforcement-learning methods regulate traffic and learn network-control policies to mitigate flooding attacks.
- DoS attack types: Perception-layer DoS includes jamming, kill-command, and desynchronizing attacks; network-layer DoS includes ICMP, amplification, and reflection flooding.
- DoS attack types: Application-layer DoS includes path-based DoS and reprogramming attacks.
- RL-based countermeasures: Multiagent router throttling uses reinforcement-learning agents on routers to rate-limit traffic toward victim servers, but initial designs had scalability problems.
- RL-based countermeasures: Coordinated Team Learning combines hierarchical communication, task decomposition, and team rewards to improve scalability and adaptability with up to 100 agents.
- RL-based countermeasures: A DDPG-based software-defined IoT approach continuously controls traffic using rewards based on benign and attack traffic, mitigating TCP SYN, UDP, and ICMP flooding.
B. Jamming Attack
Jamming disrupts wireless IoT communication by interfering with signals and is especially serious for constrained devices. RL and DRL methods learn power, mobility, or spectrum-selection policies to avoid jammers.
- Attack characteristics: Jamming interrupts wireless information transmission by injecting interfering signals and reducing the receiver’s signal-to-noise ratio.
- DRL countermeasures: A CNN-DQN power-control scheme selects transmit power according to transmission status and jammer strength, and is implemented on USRPs.
- DRL countermeasures: A DQN scheme lets a secondary user leave heavily jammed areas or reconnect elsewhere while using spread spectrum and mobility without interfering with primary users.
- DRL countermeasures: For receiver mobility, DQN chooses whether to stay or leave; the method achieved faster convergence and higher SINR than Q-learning.
- DRL countermeasures: RCNN-based DRL addresses the infinite-state limitation of discrete-SINR approaches and produces an optimal anti-jamming strategy.
- Q-learning countermeasures: Q-learning enables WACR systems to learn sweeping-jammer patterns and switch sub-bands, while the reviewed approaches assume fixed jammers rather than adaptive cognitive jammers.
C. Spoofing Attack
Spoofing attacks impersonate trusted devices to access information or spread malware, creating serious risks in interconnected IoT systems. RL-based authentication adapts detection thresholds and supports attack-intention analysis.
- Attack characteristics: Spoofing impersonates another network device to gain trust, access legitimate information, or spread malware.
- Attack characteristics: In connected UAV systems, a spoofing attacker can join as a trusted node, sense critical battlefield information, and transmit false information.
- RL-based countermeasures: RL-based active authentication uses received signal strength and hypothesis testing, with Q-learning selecting the test threshold in dynamic environments.
- RL-based countermeasures: Dyna-Q produced a lower error rate than Q-learning, while both algorithms improved spoofing detection over a fixed-threshold approach.
- Proactive detection: Reachability analysis and inverse RL are combined to predict an attacker’s intention and identify compromised sensors in autonomous vehicles.
IV. REINFORCEMENT LEARNING IN CYBER PHYSICAL SYSTEMS
The survey examines reinforcement-learning defenses and attacks in smart grids, including sequential attacks, false-data injection, and adaptive attacker–defender interactions.
- Security in Smart Grid: Smart grids combine power infrastructure with information systems but face cyber, physical, blended, tampering, and eavesdropping attacks.
- Security in Smart Grid: Sequential attacks coordinate interdictions over time to change in-service lines into outages and can cause cascading blackouts.
- Security in Smart Grid: Q-learning was used to identify vulnerable smart-grid components and optimize stealthy false-data-injection strategies against automatic voltage control.
- Security in Smart Grid: A model-free RL defender detects low-magnitude deviations and responds online without requiring a known attack model.
- Security in Smart Grid: In attacker–defender games, multistage attacks caused greater outages, while defensive strategies reduced successful attacks and average generation loss.
B. Security in Smart Transportation System
The survey reviews reinforcement-learning security methods for smart transportation systems, covering UAV, VANET, V2V, and integrated UAV–vehicle networks under jamming and other attacks.
- Security in Smart Transportation System: Smart transportation systems connect vehicles, infrastructure, and pedestrians, creating security requirements for privacy, authorization, integrity, storage, and management.
- Security in Smart Transportation System: DQN selects UAV power allocations against smart attackers across multiple frequency channels using observed network state, SINR, and received-signal utility.
- Security in Smart Transportation System: Hotbooting-PHC accelerated UAV relay decisions by initializing Q-values and action probabilities from previously generated experimental data.
- Security in Smart Transportation System: The proposed UAV relay strategy decreased OBU-data BER and increased VANET utility relative to Q-learning, helping reroute information when an RSU is severely jammed.
- Security in Smart Transportation System: For hybrid attackers in integrated UAV–CAV networks, RL handled power control while UCB1 addressed channel selection as a multiarmed-bandit problem.
V. RESEARCH TRENDS AND OPEN RESEARCH CHALLENGES
The survey identifies IoT scale and dynamism as central challenges and highlights continuous action–state spaces as an open direction for reinforcement-learning security.
- RESEARCH TRENDS AND OPEN RESEARCH CHALLENGES: IoT generates massive transaction volumes in highly dynamic environments, making security approaches increasingly challenging.
- RESEARCH TRENDS AND OPEN RESEARCH CHALLENGES: Most existing work addresses finite action and state spaces, whereas real IoT environments require handling continuous spaces.
- RESEARCH TRENDS AND OPEN RESEARCH CHALLENGES: Discretizing action–state spaces is expensive and unsuitable for extremely nonlinear IoT problems.
- RESEARCH TRENDS AND OPEN RESEARCH CHALLENGES: Actor–critic and hierarchical deep reinforcement learning are proposed to address continuous actions and reduce dimensionality and scaling problems.
B. Learn with partially observable environment
The survey identifies partial observability, multi-agent coordination, and adversarial environments as unresolved challenges for reinforcement learning in IoT security.
- Learn with partially observable environment: IoT agents often observe only part of the environment because sensors and communication links have limited capacity.
- Learn with partially observable environment: DRL has been applied to partially observable environments, but the survey describes this as limited to small-scale IoT settings.
- Learn with partially observable environment: Recurrent neural networks combined with RL are suggested for learning policies under partial observability.
- Joint reward from multiple agents: Multi-agent RL remains difficult when agents perform different tasks because coordinating shared rewards and control is complex.
- Robustness against Adversarial RL: Few studies address adversarial environments, leaving robustness against uncertain environmental changes and adversarial attacks as an open challenge.