Source-linked AI summary
Edge AI: A Taxonomy, Systematic Review and Future Directions
Sukhpal Singh Gill, Muhammed Golec, Jianmin Hu, Minxian Xu, Junhui Du, Huaming Wu, Guneet Kaur Walia, Subramaniam Subramanian Murugesan, Babar Ali, Mohit Kumar, Kejiang Ye, Prabal Verma, Surendra Kumar, Felix Cuadrado, Steve Uhlig
TL;DR
Edge AI research lacks a consolidated view of its architectures, applications, resource requirements, and unresolved challenges. This paper conducts a systematic review and develops a taxonomy of Edge AI research, finding that much of the literature remains simulation-based and that resource, heterogeneity, security, and scalability issues require further study.
Problem
The paper addresses the need to systematically organize research on Edge AI across infrastructure, resource management, machine learning, and applications.
Method
The authors conduct a Kitchenham-guided systematic literature review and propose a taxonomy for classifying Edge AI systems across major technical domains.
Results
Most reviewed articles use simulation-based implementations, indicating that edge-enabled AI remains insufficiently tested in real-world IoT environments.
Takeaways & Limitations
The taxonomy and review provide a structured basis for examining Edge AI infrastructure, resource management, ML models, applications, and future research.
Takeaways & Limitations
The reviewed literature is dominated by simulation-based implementations rather than validation in real-world IoT environments.
Abstract
from arXiv · showhide
Edge Artificial Intelligence (AI) incorporates a network of interconnected systems and devices that receive, cache, process, and analyze data in close communication with the location where the data is captured with AI technology. Recent advancements in AI efficiency, the widespread use of Internet of Things (IoT) devices, and the emergence of edge computing have unlocked the enormous scope of Edge AI. Edge AI aims to optimize data processing efficiency and velocity while ensuring data confidentiality and integrity. Despite being a relatively new field of research from 2014 to the present, it has shown significant and rapid development over the last five years. This article presents a systematic literature review for Edge AI to discuss the existing research, recent advancements, and future research directions. We created a collaborative edge AI learning system for cloud and edge computing analysis, including an in-depth study of the architectures that facilitate this mechanism. The taxonomy for Edge AI facilitates the classification and configuration of Edge AI systems while examining its potential influence across many fields through compassing infrastructure, cloud computing, fog computing, services, use cases, ML and deep learning, and resource management. This study highlights the significance of Edge AI in processing real-time data at the edge of the network. Additionally, it emphasizes the research challenges encountered by Edge AI systems, including constraints on resources, vulnerabilities to security threats, and problems with scalability. Finally, this study highlights the potential future research directions that aim to address the current limitations of Edge AI by providing innovative solutions.
I. INTRODUCTION
Edge AI combines advances in AI, IoT, and edge computing to process data near its source, improving latency, bandwidth use, and real-time responsiveness. The review situates this approach across computing paradigms, applications, and research directions.
- Computing Context: Edge computing moves storage and processing closer to data sources than centralized cloud computing, reducing latency and bandwidth usage.This proximity supports real-time processing and applications including smart cities, autonomous vehicles, and industrial automation.
- Edge AI Concept: Edge AI integrates AI capabilities with edge devices to enable distributed intelligence and real-time data processing at the source.The approach addresses cloud-based IoT limitations involving privacy, connectivity, latency, and network congestion.
- Motivation: The integration is motivated by growing connected-device and data volumes, alongside cloud-centric limitations in latency, bandwidth, privacy, and responsiveness.Localized processing reduces reliance on distant cloud infrastructure and supports applications requiring instantaneous analysis.
- Objectives and Benefits: Edge AI offers bandwidth efficiency, reduced operational costs, lower network congestion, enhanced privacy and security, and scalability for data-intensive applications.These benefits follow from processing and distributing resources closer to data sources.
- Computing Paradigms: Cloud, fog, and edge computing differ in the placement of computing resources, with edge computing processing data closest to its source.Fog computing occupies an intermediate layer between cloud and edge, while cloud computing relies on remote centralized resources.
B. Integration of AI with Edge Technology
Integrating AI with edge technology distributes models and data processing across heterogeneous edge devices, enabling low-latency, privacy-aware applications. The section surveys applications and identifies resource, maintenance, energy, security, and scalability challenges.
- Architecture: Edge AI distributes AI algorithms to nodes such as mobile devices and IoT systems instead of processing data exclusively on cloud servers.This architecture enables fast computation near the data source.
- Benefits: Edge AI reduces latency for delay-sensitive healthcare and autonomous-vehicle applications by processing data in real time near its source.Cloud-based processing can introduce delays and unnecessary bandwidth use because data must travel to centralized servers.
- Benefits: Local processing and storage can improve security and privacy by reducing transmission and centralized storage of sensitive data.The cited examples include biometric and health data, whose communication and storage expose larger attack surfaces in cloud systems.
- Challenges: Heterogeneous edge devices create challenges involving limited processing capacity, energy consumption, maintenance, updates, task scheduling, and data synchronization.The section identifies specialized chips, containerization, orchestration, microservices, and load balancing as possible responses.
- Applications: Edge AI supports healthcare, smart parking, smart homes, computer vision, cybersecurity, and transportation through distributed sensing and real-time analysis.Applications include wearable-device diagnosis, energy optimization, biometric authentication, anomaly detection, and traffic-light management.
III. RELATED STUDIES AND SURVEYS
The related work surveys Edge AI applications across smart cities, manufacturing, vehicles, industrial automation, and healthcare, alongside optimization, resource management, privacy, and security. The review compares these studies with prior surveys and emphasizes Edge AI systems operating under constrained resources.
- Application domains: Related studies cover smart cities, smart manufacturing, autonomous vehicles, IoV, industrial automation, and healthcare monitoring.They examine how edge computing and AI support applications across these domains.
- Smart cities: Smart-city research emphasizes computation offloading and federation across edge, cloud, and fog computing.An intelligent offloading method is described as preserving privacy, improving edge utility, and increasing offloading efficiency.
- Intelligent manufacturing: AI-enabled manufacturing research addresses predictive maintenance, autonomous decisions, scheduling, and closed-loop cloud-manufacturing operations.The reviewed work includes AI-Mfg-Ops and Deep Q-network approaches for intelligent factory processes.
- Vehicles and IoT: Vehicle and IoT research focuses on real-time control, task optimization, security, communication, privacy, and multi-agent reinforcement learning.The studies include autonomous driving and heterogeneous IoT task-execution scenarios.
- Automation and healthcare: Industrial automation and healthcare studies deploy AI at the edge while addressing resource constraints, privacy, security, and decision-making.Examples include cloud-edge industrial control, secure AI microservices, federated learning in medical IoT, and genetic-based encryption.
- Survey comparison: The survey compares prior work across Edge AI advancement, constrained-environment optimization, training-data scarcity, and AI/ML-based resource management.Its comparison covers applications, optimization methods, and resource utilization, with a focus on constrained edge devices.
1) Related Surveys:
Prior surveys largely provide overviews or visions, leaving limited comprehensive synthesis and taxonomy for Edge AI. This paper responds with a systematic review, taxonomy, and analysis of applications, challenges, and future research directions.
- Related Surveys: Existing Edge AI reviews mostly offer overviews or visions rather than a comprehensive systematic review with a detailed taxonomy.The authors identify this as a gap in the existing survey literature.
- Our Contributions: The paper introduces Edge AI through its history, challenges, and prospects.This contribution frames the field before presenting the review and taxonomy.
- Our Contributions: A systematic review examines Edge AI research across many applications and identifies current trends and possible future directions.The review is conducted using a structured methodology based on Kitchenham et al.'s guidelines.
- Our Contributions: The proposed taxonomy classifies and arranges Edge AI systems while examining applications across disciplines.The broader review considers infrastructure, cloud and fog computing, services, use cases, learning methods, and resource management.
- Our Contributions: The paper emphasizes real-time edge processing and identifies resource limitations, security risks, and scaling issues as major challenges.These challenges motivate the paper's discussion of Edge AI research needs.
- Our Contributions: The paper proposes future research directions intended to address current Edge AI shortcomings.The stated directions seek innovative solutions and opportunities for further investigation.
- Review Methodology: The review methodology uses predefined research questions, database searches, keyword strategies, selection criteria, and quality assessment procedures.The process is presented as a way to reduce bias in study selection and analysis.
D. Inclusion and Exclusion Criterion
The review defines inclusion and exclusion criteria for selecting Edge AI studies, then screens candidate papers through title/abstract, full-text, and final-selection phases.
- Eligibility criteria: Studies had to address AI-based methodologies in edge computing and satisfy the review’s predetermined inclusion and exclusion criteria.The criteria included English conference, journal, or book-chapter articles available in full text and relevant to Edge AI infrastructure or related topics.
- Screening process: The screening process comprised title-and-abstract screening, full-text screening, and final selection using the criteria outlined in Fig. 6.Papers judged irrelevant by title and abstract were excluded, followed by full-text assessment and final elimination of studies failing the specified criteria.
- Research scope: The review questions covered Edge AI architectures, resource management, infrastructure choice, and reliable intelligent operation in dynamic or ambiguous conditions.The architecture question distinguishes monolithic and microservices designs, while resource questions address provisioning, allocation, deployment, and workload scheduling.
- Review protocol: The methodology followed established systematic-review standards to evaluate eligible studies and identify research gaps and future trends.The authors describe systematic review as a structured approach for analyzing positive and negative aspects of recent studies.
F. Extraction and Synthesis
The review extracted and synthesized evidence from selected studies, then organized Edge AI research across infrastructures, architectures, methods, use cases, and resource-management themes.
- Extraction and synthesis: Data extraction examined all 78 chosen studies and recorded relevant information in structured reports for later synthesis.The extraction stage created a mechanism for compiling data items from primary research studies.
- Taxonomy construction: The taxonomy classified 60 papers under 11 subheadings covering infrastructure, application architecture, IoT use cases, methods, resource management, model sizing, heterogeneity, security, scheduling, migration, and scaling.The authors developed the taxonomy by examining available papers, obtaining research summaries, and classifying solutions under the necessary headings.
- Infrastructure: Cloud computing pools virtualized resources in remote data centers, enabling flexible access to CPUs, GPUs, and memory for AI training and prediction.Cloud resources can be scaled according to business needs, improving resource utilization and reducing model-training and prediction costs.
- Infrastructure: Fog computing places distributed resources closer to users than cloud computing, supporting image, video, and natural-language processing with shorter response delays.The approach extends computing capacity from the network center toward the network edge.
- Infrastructure: Edge computing moves resources and intelligent services near objects or data sources, improving service quality and information security compared with cloud-centered processing.Edge AI applies this local-processing premise at device level, allowing some devices to operate without an Internet connection and make independent decisions.
B. Application Architecture
The review contrasts monolithic and microservice architectures and surveys Edge AI methods used for optimization, deployment, and resource allocation across IoT settings.
- Monolithic architecture: Monolithic architecture encapsulates functionality in one application, simplifying development, deployment, debugging, and monitoring while improving component-interaction performance.It may suit simple or initially developed AI systems with small business scale and infrequent requirement changes.
- Monolithic architecture: As application scale and demand increase, monolithic architecture becomes less scalable, adaptable, and maintainable because extending modules requires recompiling and redeploying the whole application.The resulting complexity can reduce resource efficiency and increase maintenance costs and error risk.
- Microservice architecture: Microservices decompose large applications into independent service units, enabling more flexible and efficient iterative upgrades than monolithic designs.The architecture still faces challenges involving service governance, network transmission efficiency, service expansion, and version iteration.
- IoT use cases: IoT use cases are categorized as static or mobile, with Edge AI supporting real-time analysis for sensors, cameras, wearables, and in-car telemetry.Wearable devices can analyze physiological data locally and provide health feedback without uploading it to the cloud.
- Methods: Heuristic, meta-heuristic, machine-learning, and deep-reinforcement-learning methods target accuracy, performance, model configuration, and resource allocation in Edge AI.Heuristics support rapid satisfactory solutions for resource scheduling, while meta-heuristics address model optimization and concurrent task allocation.
3) Machine Learning:
The section reviews machine learning and deep reinforcement learning at the edge alongside resource management, workload prediction, model-size reduction, and deployment constraints.
- Machine learning: Machine learning at the edge processes data near its source to reduce transmission delays, accelerate responses, and support real-time applications.Applications include autonomous vehicles, intelligent manufacturing, security, anomaly detection, product-quality monitoring, and patient-health monitoring.
- Deep reinforcement learning: Deep reinforcement learning combines deep learning with reinforcement learning to map perceptual inputs directly to actions through neural networks.The review identifies autonomous decision-making in edge-device applications, including in-car systems, as a relevant use case.
- Resource management: Growing edge-device numbers and application complexity make efficient resource management an urgent problem despite edge computing’s reductions in transmission delay and bandwidth demand.The review discusses provisioning, allocation, placement, prediction, and offloading as resource-management concerns.
- Resource management: Static provisioning suits stable workloads, whereas dynamic strategies use heuristic or machine-learning prediction for fluctuating resource demand.Dynamic approaches can anticipate changing requirements before allocating resources.
- Workload prediction: RNNs are considered promising for workload prediction because their inputs include information from a period of past time.The review contrasts this approach with earlier linear temporal-prediction methods such as ARIMA and Holt-Winters.
- Model sizing: Full-size deep-learning models are often impractical on edge devices because computing power, storage, and energy are limited, motivating model-size reduction.The review identifies pruning, quantization, MobileNet, and ShuffleNet as strategies or models for reducing deployment demands.
G. Heterogeneity
Edge AI environments are heterogeneous across computation, hardware, and platforms, creating distinct performance, deployment, security, and resource-management requirements. The review contrasts these differences and highlights challenges in supporting secure, efficient operation across edge devices.
- Edge-device heterogeneity comprises computational, hardware, and platform diversity.
- Computational Heterogeneity: AI applications favor parallel vector computation and GPU acceleration, whereas web services rely more on general-purpose or I/O-intensive processing.
- Computational Heterogeneity: Slow disk read/write operations can bottleneck microservices because much of their execution time is spent accessing databases.
- Hardware Heterogeneity: Edge devices differ in processor architectures, including ARM and AMD instruction sets, which can produce performance differences across software deployments.
- Platform Heterogeneity: Edge platforms include commercial systems such as AWS IoT Greengrass, Azure IoT Edge, and Cloud IoT Edge, alongside Kubernetes extensions including KubeEdge and OpenYurt.
- Security and Resource Constraints: Security must address platform, host, and network threats, while edge DDoS defenses remain constrained by limited aggregated traffic visibility and resource elasticity.
2) Task:
Edge task management covers scheduling, service deployment, migration, and validation across constrained distributed environments. The review emphasizes matching resource decisions and migration strategies to application state, network conditions, and operational objectives.
- Task Scheduling: Edge task scheduling addresses scheduling time and resource allocation, with machine learning and Markov Decision Processes used to model temporal resource-management decisions.
- Service Scheduling: Services are deployed and scheduled according to application requirements, resource availability, and network conditions to balance loads and reduce latency.
- Container Migration: Container migration supports recovery and load balancing, but stateful containers require data migration that considers data loss and integrity.
- Container Migration: Stateless containers migrate more simply because they retain no runtime state and can be recreated on another node.
- Migration Scope: Inter-cluster migration must account for cross-cluster latency and bandwidth, whereas intra-cluster migration redistributes workloads to address overloaded or degraded nodes.
- Testing and Validation: Container-migration evaluation combines simulations for early validation with real-world testbeds for assessing performance, reliability, and safety under operational conditions.
K. Container Scaling
The review frames edge AI scaling as a choice among proactive or reactive decisions and horizontal, vertical, or hybrid techniques. It also compares cloud, fog, and edge roles and contrasts monolithic with microservice architectures for deployment flexibility and scalability.
- Scaling Decisions: Proactive scaling predicts future container loads from historical data, whereas reactive scaling adjusts replicas and resources from real-time monitoring.
- Scaling Techniques: Horizontal scaling changes the number of container instances and is suited to stateless services and workloads with many concurrent requests.
- Scaling Techniques: Vertical scaling changes CPU, memory, storage, or other resources for one container instance and is suitable for stateful services.
- Scaling Techniques: Hybrid scaling combines horizontal and vertical expansion to adapt resource allocation to changing or difficult-to-predict application demand.
- Computational Paradigms: Cloud supports large-scale training and analysis, fog emphasizes lower-latency processing, and edge handles real-time inference near terminal devices while reducing bandwidth dependence.
- Architectures: Monolithic architectures favor simple deployment and performance optimization, while microservices favor flexible scaling and maintainability in dynamically adjusted edge AI applications.
C. IoT Use Cases
The review organizes edge AI use cases and methods by mobility, problem characteristics, resource demands, and deployment choices. It links algorithm selection and model deployment to application requirements, while structuring resource management as a sequence from provisioning to workload prediction.
- IoT Use Cases: IoT use cases are classified as static or dynamic according to user mobility, with edge AI examined across both categories.
- AI Methods: Heuristics, meta-heuristics, machine learning, and deep reinforcement learning are compared by scenario, complexity, data scale, training cost, real-time requirements, resource consumption, and generalization.
- AI Methods: Heuristics and meta-heuristics generally suit simple-to-medium problems with lower data and resource requirements, whereas ML and DRL target complex nonlinear problems with higher demands.
- Resource Management: Resource management proceeds from provisioning to allocation, application placement, workload distribution, and prediction-based adjustment.
- Model Deployment: Reduced-model and full-model deployment methods are compared by model size, inference speed, accuracy, training and deployment cost, and application scenario.
- Deployment Heterogeneity: Edge AI deployment involves computational, hardware, and platform heterogeneity, requiring placement and scheduling across diverse devices.
H. Security
The survey compares Edge AI security, scheduling, migration, and scaling approaches while analyzing publication trends and implementation evidence. It identifies recent growth, predominantly simulation-based evaluation, and future directions including energy optimization and next-generation networks.
- Security and Resource Management: Edge AI security is compared across platform, host, and network security categories.The survey also distinguishes container, task, pod, and service scheduling, comparing their emphasis, measures, and tools.
- Migration and Scaling: Edge AI container migrations span stateful or stateless, intra- or inter-cluster, cloud/fog/edge, and virtual or real-world testbed approaches.Container scaling is classified by active or passive response and horizontal, vertical, or hybrid expansion.
- Publication Analysis: A major chunk of referred articles is from 2023, indicating that the survey includes recent Edge AI research.The taxonomy draws on articles published from 2015 to 2024.
- Publication Analysis: The review collected 1253 articles and filtered them using redundancy and inclusion and exclusion criteria.The resulting literature was categorized as review, systematic literature review, simulation-based implementation, real/testbed-based implementation, or book.
- Implementation Evidence: Simulation-based implementations form the major portion of reviewed articles, showing that Edge AI remains insufficiently tested in real-life IoT use cases.The survey identifies practical validation in real-world IoT environments as a future research direction.
- Future Research: Future research directions include optimizing energy use, strengthening security, and integrating Edge AI with next-generation networks such as 6G.These directions are presented as ways to enhance Edge AI capabilities and applications.
B. Energy
Energy-focused Edge AI research combines adaptive algorithms, efficient scheduling, low-power hardware, and communication technologies to manage energy-constrained AIoT systems. The review emphasizes local processing for responsive, resilient energy applications and identifies manufacturing and smart-city opportunities.
- Energy Optimization: Energy optimization in AIoT requires dynamic energy distribution, task scheduling, workload management, and low-power hardware.RL methods such as DQN and MARL adapt energy use, while GA and PSO support task allocation in constrained environments.
- Communication and Infrastructure: LPWAN protocols such as LoRaWAN and NB-IoT support long-range AIoT communication with minimal power consumption.Their combination with edge AI enables real-time data exchange while conserving energy.
- Energy Applications: Edge AI reduces latency and bandwidth usage while enabling real-time decisions and resilience to network disruptions.Lightweight models such as MobileNet or TinyML can detect anomalies or optimize energy usage without substantially draining computational resources.
- Energy Applications: Federated learning and deep reinforcement learning are identified as approaches for optimizing energy distribution while preserving privacy and reducing communication overhead.Energy-aware federated updates can limit participation by devices with constrained energy resources.
- Manufacturing: Manufacturing research targets real-time predictive maintenance, quality control, and fault diagnosis to improve efficiency, reduce waste, and optimize resources.Edge devices can analyze machinery sensor data for early fault detection and immediate corrective action.
- Manufacturing and Smart Cities: Edge ML supports continuous manufacturing monitoring, early equipment-failure detection, and immediate corrective actions.The review also describes real-time edge processing for smart-city sensing, traffic adjustment, and urban mobility.
E. Smart Transport
Smart transport research applies deep reinforcement learning and mobile edge computing to balance traffic conditions, computation, energy, and communication latency. Future directions emphasize scalable learning, efficient hardware, cooperative communication, and 5G integration, while retaining trade-offs among model complexity and edge resources.
- DRL for Smart Transport: DQN-based mobile edge computing models learn traffic-state-to-action policies for local transport decisions such as traffic-light adjustment and vehicle rerouting.DQN maps states to expected rewards for candidate actions in smart transportation systems.
- Resource Trade-offs: Smart transport optimization must balance computational load, energy consumption, and communication latency as a multi-objective problem.Multi-agent DRL treats vehicles or edge devices as independent agents that optimize local performance while contributing to system objectives.
- Scalable Learning: Hierarchical reinforcement learning can decompose city-scale traffic control into smaller intersection-level subproblems, reducing computational overhead and improving scalability.This approach is proposed for further research rather than reported as an established result.
- Hardware and Communication: Energy-efficient accelerators such as Google’s Edge TPU and Nvidia’s Jetson are proposed for running DRL models locally with minimal power consumption.VANET communication can support real-time cooperative decisions among vehicles and roadside units.
- Next-Generation Networks: 5G integration and network slicing could provide ultra-low-latency communication and dedicated virtual resources for time-sensitive traffic tasks.Examples include emergency-vehicle routing and accident detection.
- Limitations and Future Work: Smart transport still requires research on trade-offs among DRL model complexity, edge computational resources, and energy consumption.The review also calls for improved communication protocols and network technologies.
H. Hardware
Edge AI hardware must balance computational performance, power efficiency, size, cost, and security as workloads become more complex and diverse. The review also highlights multi-prototype federated learning as a response to heterogeneous edge data, devices, and networks.
- Hardware constraints: Edge AI hardware must balance computational performance, power efficiency, size, and cost under the physical constraints of edge nodes.Specialized platforms include Nvidia’s Jetson TX2 and Google’s Edge TPU.
- Hardware directions: Future hardware directions include workload-specific accelerators, domain-specific architectures, heterogeneous computing, and hardware–software co-design.These designs target workloads such as natural language processing, computer vision, and reinforcement learning.
- Security and efficiency: Edge AI hardware evolution is expected to combine specialized accelerators, energy-efficient architectures, and secure computation technologies.Trusted execution environments are described as protecting sensitive data and model integrity on edge devices.
- Heterogeneous federated learning: Multi-prototype federated learning represents heterogeneous client data distributions with multiple weighted prototypes rather than one global prototype.Local prototypes capture client-specific characteristics, while weighted aggregation reflects the importance of corresponding data clusters.
- Heterogeneous federated learning: Multi-prototype aggregation improves accuracy and convergence rates in heterogeneous environments and can reduce the communication rounds needed for model optimization.The approach is presented as more representative of non-IID local data than traditional single-prototype methods.
- Security directions: Future edge AI security research combines quantum-safe cryptography, AI-driven threat detection, automation, and blockchain-supported communications.The stated goal is resilient protection against classical and quantum-based cyber threats.
K. Privacy
The review identifies privacy-preserving federated learning for health prediction and secure, scalable edge systems as important future directions. It also connects Edge AI’s development with 6G capabilities, while emphasizing infrastructure, resource management, and ML model sizing in the reviewed literature.
- Privacy-preserving health applications: Future health-prediction research should strengthen privacy preservation in federated learning while supporting real-time monitoring through wearable devices.Suggested techniques include differential privacy, homomorphic encryption, and scalable adaptive federated learning.
- 6G and beyond: 6G’s ultra-low latency and high bandwidth are expected to support edge AI deployment for real-time applications such as autonomous vehicles and smart cities.The review also identifies dynamic resource allocation, energy efficiency, secure edge computing, and blockchain as related priorities.
- Review scope and directions: The systematic review examines Edge AI across infrastructure, resource management, and ML model sizing using 78 selected studies.The studies were categorized into multiple domains to evaluate cloud, fog, and edge computing choices and their effects on application efficacy and resource usage.
- Future research directions: Future research directions include hybrid Edge–Fog–Cloud infrastructure, optimization of resource allocation and latency, stronger security and privacy methods, and real-world applications.Examples include smart cities and IoT-based healthcare systems.