Source-linked AI summary
Application of Deep Reinforcement Learning for Intrusion Detection in Internet of Things: A Systematic Review
Saeid Jamshidi, Amin Nikanjam, Kawser Wazed Nafi, Foutse Khomh, Rasoul Rasta
TL;DR
IoT networks are difficult to secure because their threats and operating conditions evolve, while traditional IDS may not adapt effectively. This systematic review examines DRL-based IDS research from the past decade, synthesizing algorithms, applications, datasets, and research gaps. It finds that DRL improves adaptive threat detection and identifies dataset, reproducibility, scalability, and integration challenges for future work.
Problem
Traditional IDS struggle to adapt to IoT networks' dynamic conditions and evolving threat patterns, motivating examination of DRL-based alternatives.
Method
The paper conducts a systematic review of DRL-based IDS in IoT, analyzing research questions, algorithms, applications, datasets, and selected studies from the past ten years.
Results
The review reports that DRL-based IDS improve IoT security by adapting to operational environments for more accurate threat detection and fewer false positives.
Takeaways & Limitations
Future DRL-based IDS research should pursue federated learning, policy learning, high-level threat intelligence, and stronger integration with emerging IoT technologies.
Takeaways & Limitations
The field remains constrained by outdated datasets, reproducibility issues, class imbalance, and insufficient representation of heterogeneous, evolving IoT threats.
Abstract
from arXiv · showhide
The Internet of Things (IoT) has significantly expanded the digital landscape, interconnecting an unprecedented array of devices, from home appliances to industrial equipment. This growth enhances functionality, e.g., automation, remote monitoring, and control, and introduces substantial security challenges, especially in defending these devices against cyber threats. Intrusion Detection Systems (IDS) are crucial for securing IoT; however, traditional IDS often struggle to adapt to IoT networks' dynamic and evolving nature and threat patterns. A potential solution is using Deep Reinforcement Learning (DRL) to enhance IDS adaptability, enabling them to learn from and react to their operational environment dynamically. This systematic review examines the application of DRL to enhance IDS in IoT settings, covering research from the past ten years. This review underscores the state-of-the-art DRL techniques employed to improve adaptive threat detection and real-time security across IoT domains by analyzing various studies. Our findings demonstrate that DRL significantly enhances IDS capabilities by enabling systems to learn and adapt from their operational environment. This adaptability allows IDS to improve threat detection accuracy and minimize false positives, making it more effective in identifying genuine threats while reducing unnecessary alerts. Additionally, this systematic review identifies critical research gaps and future research directions, emphasizing the necessity for more diverse datasets, enhanced reproducibility, and improved integration with emerging IoT technologies. This review aims to foster the development of dynamic and adaptive IDS solutions essential for protecting IoT networks against sophisticated cyber threats.
1. Introduction
IoT security has become increasingly important as interconnected devices and systems expand, while DRL-based IDS offers an adaptive approach to detecting evolving threats. This review systematically analyzes DRL applications, algorithms, datasets, and research gaps in IoT security.
- IoT connects diverse devices and infrastructures, generating substantial data and creating a need to secure interconnected systems.Examples include home appliances, traffic lights, and industrial equipment.
- DRL-based IDS research is reviewed to assess state-of-the-art techniques for enhancing IoT security.
- The review evaluates DRL algorithms and their efficiency across IoT intrusion-detection applications.
- Datasets and benchmarks are analyzed for their relevance and representation of real-world IoT scenarios.
- The review identifies research gaps and challenges to guide future investigations in DRL-based IoT intrusion detection.
2. Background: IoT Security
IoT security challenges span perception, network, and application layers, where diverse attacks target devices, communications, and services. DRL supports adaptive intrusion detection by learning from environmental feedback and responding to evolving attack patterns.
- IoT Security Architecture: IoT systems contain perception, network, and application layers, each with distinct security threats requiring intrusion detection.
- Perception Layer: Resource-constrained perception-layer devices are vulnerable to physical tampering and require real-time anomaly detection and tamper-resistant hardware.
- Network Layer: Network-layer threats include man-in-the-middle attacks, DDoS attacks, and attacks on communication protocols, traffic flows, and routing.
- Application Layer: Application-layer threats include malware, ransomware, software vulnerabilities, and breaches involving sensitive data.
- IDS Technologies: Signature-based IDS detect known threats, whereas anomaly- and behavior-based IDS address unusual or unknown activity with differing false-positive characteristics.
- Role of DRL: DRL learns through environmental interaction and reward feedback, avoiding labeled-data requirements while adapting to emerging threats and reducing false positives relative to unsupervised methods.
3. Methodology
The review uses a structured, multi-stage methodology to examine DRL-based IDS in IoT over 2014–2024. It combines broad automated and manual searching with explicit eligibility criteria and analyzes the selected literature's research coverage and trends.
- Research scope and questions: The review examines DRL models for IoT-based IDS, identifying state-of-the-art approaches, datasets, limitations, and future research opportunities across the last ten years.Its research questions address current algorithms and applications, experimental datasets and real-world applicability, and remaining gaps in DRL-based IDS for IoT.
- Search strategy: The search combined automated database queries with manual searches of search engines and reference lists to compile relevant studies.The review used Web of Science, Compendex, and Inspec, followed by Google Scholar searches covering publications from 2014 to 2024.
- Search strategy: Search terms combined reinforcement-learning concepts with IoT and intrusion-detection terms, including acronyms and related keyword variations.The strategy used terms such as “Reinforcement Learning,” “Deep Reinforcement Learning,” RL, DRL, “Internet of Things,” IoT, “Intrusion Detection System,” and IDS.
- Selection criteria: Two independent reviewers assessed titles, abstracts, and, when necessary, full texts using predefined inclusion and exclusion criteria.Eligible studies were original peer-reviewed journal papers addressing DRL techniques for IDS in IoT and published between 2014 and 2024; conference papers and several other publication types were excluded.
4. Categorization of state-of-the-art papers
The reviewed studies apply diverse DRL algorithms to IoT intrusion detection across smart homes, smart grids, transportation, industrial, and other environments. These approaches dynamically adapt detection, routing, feature selection, and defense decisions to changing threats and network conditions.
- The review categorizes state-of-the-art IoT IDS studies by the types of DRL algorithms they apply.The categorization covers diverse approaches illustrated in Figure 7.
- DQN-based defenses dynamically adjust IP blocking to reduce false positives and avoid service disruption from incorrect threat identification.One reported system achieved 96% accuracy in a single-layered setting.
- MalBoT-DRL uses damped traffic statistics and an attention-based reward mechanism to address model drift, achieving 99.80% early and 99.40% late detection accuracy.The results came from trace-driven experiments on two representative datasets.
- DQL-based ReLFA detects and mitigates link flooding attacks in real time by combining Rényi entropy with dynamic routing adjustment.Simulation results report faster rerouting and mitigation than LFADefender and Woodpecker.
- SAC adapts an unsupervised classifier at IoT edge gateways, responding dynamically to changing network conditions while enabling collaborative detection among edge devices.The approach is designed to expedite detection and optimize resource utilization.
- Reviewed applications include Q-learning for adaptive DDoS defense in vehicular IoT, non-stationary MAB for smart-home anomalies, DDPG for smart-grid feature selection, and DDPG-based IDS for green IoT.Reported examples span V2I/V2V networks, smart homes, smart grids, and green IoT systems.
4.8. Inverse Reinforcement Learning (IRL)
This section covers reinforcement-learning approaches for adaptive IoT intrusion detection and defense across wireless, smart-city, IoT, industrial, and sensor-network settings. The studies emphasize adaptation to unknown attacks, dynamic traffic, energy constraints, and changing network conditions.
- Inverse Reinforcement Learning (IRL): IRL-based strategies detect intelligent backoff attacks under partial observability by identifying deviations from normal behavior and generalizing to unseen attack strategies.The reported goal is to improve wireless-network resilience beyond known attack patterns.
- Inverse Reinforcement Learning (IRL): One-shot learning enhanced with DRL provides a dynamic IDS for smart-city multi-access edge computing environments facing variable network behavior and emergent threats.The approach is reported to mitigate zero-day attacks and other emerging threats.
- Inverse Reinforcement Learning (IRL): A semi-supervised DDQN method combines autoencoder reconstruction, neural classification, and K-Means clustering to detect known and unknown IoT network anomalies.It achieved 83% binary-classification accuracy and improved F1-Score by approximately 6% over traditional DDQN.
- Inverse Reinforcement Learning (IRL): A cooperative trust mechanism with PG-DRL detects vampire nodes from energy-consumption deviations while dynamically selecting energy-efficient next-hop routes.The trust mechanism reached 100% detection accuracy in controlled scenarios.
- Inverse Reinforcement Learning (IRL): Other reviewed approaches include SARSA-based industrial-control IDS, GOA with Gain Actor-Critic and SVM, and MADDPG for cooperative SDN-IoT routing and DDoS defense.The GOA-GAC-SVM model reported accuracies of 99.71% on NSL-KDD, 99.11% on AWID, and 99.61% on CIC-IDS 2017.
4.16. Dueling Double Deep Q-Network (D3QN)
The Security Defense Strategy Algorithm uses D3QN to learn adaptive defense behaviors in adversarial IoT security scenarios involving one or multiple attackers.
- Dueling Double Deep Q-Network (D3QN): SDSA uses D3QN to optimize the allocation and coordination of defense resources through simulated adversarial scenarios.The method adaptively learns strategic behaviors for multiple defenders and attackers.
- Dueling Double Deep Q-Network (D3QN): SDSA improved defense effectiveness by approximately 87% and 85% over MADDPG and OptGradFP, respectively, with a single attacker.In multiple-attacker scenarios, it outperformed MADDPG and OptGradFP by 65% and 60%, respectively.
4.17. Neural Fitted Q-Iteration (NFQ)
The reviewed DRL-based IDS approaches apply adaptive learning across diverse IoT security settings, reporting strong detection performance alongside domain-specific limitations. NFQ with LSTM addresses high-dimensional and temporal DDoS detection in cloud–fog IoT environments.
- NFQ-based IDS: NFQ combined with LSTM dynamically analyzes malicious traffic across cloud, fog, and SDN-IoT environments.The design addresses high-dimensional state spaces and temporal dependencies while classifying and filtering infected packets.
- NFQ-based IDS: 98.34% classification accuracy, 98.88% precision, and 98.76% recall were achieved by the RDL-LSTM-2.2 variant.The scheme also reduced packet latency by 10.42% and computational complexity by 12.96% compared with existing approaches.
- Applications and performance: DRL-based IDS studies cover DDoS mitigation, routing, feature optimization, federated learning, smart cities, UAVs, wireless networks, and industrial systems.The comparative analysis organizes techniques by advantages, limitations, and use cases.
- Limitations: Limitations include restricted datasets, narrow application domains, high computational or resource overhead, limited scalability, and limited real-world evaluation.Examples include evaluation only on CICDDoS2019, smart-city-only generalizability, and high training overhead.
- Applications and performance: Reported results include 99.36% accuracy for ITS traffic learning, 99.98% DDoS detection in green IoT, and 98.8% accuracy for zero-day adaptation.Other studies report performance above 95% or accuracy near 99% in specialized IoT settings.
4.18. Distribution of Models in Literatures
DQN is the most frequently used DRL model in IoT intrusion-detection research, while other models serve more specialized applications. The distribution reflects differences in model capabilities and IoT security requirements.
- Model distribution: DQN accounts for 25.7% of reported DRL-model usage in IoT intrusion detection.Its prominence is associated with handling large state and action spaces through experience replay and off-policy learning.
- Model distribution: MDRL represents 11.4% of usage, indicating application in tailored IoT security settings.The passage describes MDRL as adaptable to specialized applications.
4.19. The existing DRL-based IDS
DRL-based IDS applications span multiple IoT domains and address varied attack types, including attacks represented across the reviewed datasets. The review organizes these applications by context and covered threats.
- IoT contexts: DRL-based IDS applications include smart homes, industrial IoT, healthcare IoT, and other specialized contexts.These applications use adaptive detection and response to address domain-specific security challenges.
- Attack coverage: The reviewed datasets cover attack categories including brute force, DDoS, botnets, reconnaissance, MITM, exploits, worms, and web attacks.The attack inventory also includes flooding, injection, and related network threats.
5. The datasets are used for DRL-based IDS
The review examines datasets ranging from general network-attack benchmarks to specialized industrial, home, wireless, energy, medical, and sector-specific IoT data. These datasets support training and evaluation across varied attack scenarios and operational settings.
- General-purpose datasets: General datasets such as ISCXIDS2012, UNSW-NB15, and CICIDS2017 represent DDoS, brute-force, and infiltration attacks.They are used to simulate realistic network environments for IDS training and evaluation.
- Specialized datasets: WUSTL-IIoT, N-BaIoT, and Hogzilla provide specialized industrial, home-IoT, malware, botnet, and advanced-threat scenarios.WUSTL-IIoT includes industrial traffic and APT scenarios, while Hogzilla represents polymorphic and metamorphic malware behaviors.
- Protocol and wireless datasets: IEC 60870-5-104 targets energy-sector control communications, while AWID focuses on wireless IoT flooding, injection, and impersonation attacks.These datasets represent protocol-specific and wireless-network security conditions.
- Large-scale attack datasets: BoT-IoT and CIC-DDoS2019 provide large-volume botnet and DDoS data for coordinated IoT attack detection.Their scenarios involve disruption caused by multiple compromised devices.
- Sector-specific datasets: TON-IoT, MedBIoT, and SCADA-based power-system datasets address industrial, medical, and critical-infrastructure security requirements.These datasets include threats such as malware propagation and command injection in specialized environments.
6. Discussion
The review shows that integrating DRL into IoT IDS enables dynamic adaptation to evolving threats through varied states, actions, and reward mechanisms.
- DRL enables IoT IDS to adapt dynamically and respond to new and evolving threats.This adaptability is important for complex and continuously expanding IoT networks.
- Reviewed studies define states ranging from simple network-traffic patterns to complex multidimensional data vectors.
- DRL-based IDS can select responses across a broad spectrum according to network conditions and threats.
7. The future research opportunities in DRL-based IDS
Future work should address dataset limitations, reproducibility, resource constraints, underexplored IoT domains, and incomplete integration with emerging technologies. The review also identifies directions including federated learning, transfer learning, policy learning, SDN integration, and real-time evaluation.
- Datasets: Current IoT datasets often lack the diversity and dynamism needed to represent evolving threats and real-world network complexity.This limitation can restrict the generalization of DRL-based IDS.
- Datasets: Class imbalance and inconsistent labeling can bias training toward benign traffic or produce inaccurate threat assessments.Oversampling, undersampling, synthetic generation, and improved labeling are identified as mitigation approaches.
- Reproducibility: Reproducibility remains limited because rapid and effective experimental evaluation is difficult, especially in real-world IoT settings.
- Application domains: Research is concentrated in Industrial IoT, healthcare, and smart cities, while Agricultural IoT, environmental monitoring, and retail and supply-chain IoT remain underexplored.
- Deployment constraints: DRL-based IDS must reduce CPU, computational, and energy demands while preserving threat-detection accuracy and scalability.
- Emerging directions: Future directions include federated learning, transfer learning, policy learning, high-level threat intelligence, SDN integration, and real-time testing.These directions target privacy, adaptation to unseen attacks, network-management integration, and practical evaluation.
8. Conclusion
The review finds substantial progress in DRL-based IDS for IoT from 2014 to 2024, while identifying persistent challenges involving datasets, reproducibility, scalability, and integration with emerging methods.
- DRL-based IDS research in IoT made significant progress between 2014 and 2024.
- DRL methodologies are described as adaptable and effective for improving IoT security against evolving cyber threats.
- The review identifies outdated datasets, reproducibility issues, and insufficiently scalable solutions as continuing challenges.
- Future research should examine federated learning, policy learning, and high-level threat intelligence integration.
Appendix
The appendix provides a table of abbreviations used in the research.
- The appendix contains a table of abbreviations used in this research.
- The abbreviation table serves as an appendix reference for terminology used throughout the research.
- The appendix is dedicated to abbreviation support rather than substantive findings or methods.