Source-linked AI summary
A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
Naveed Khan, Besan Al Sbeihi, Maryam Alshehhi, Nasir Saeed
TL;DR
Wireless foundation-model research is fragmented across architectures, learning paradigms, and applications, lacking a unified survey. This paper organizes the field and reviews its methods, applications, evaluation, and remaining challenges, highlighting a transition toward unified, reusable wireless intelligence.
Problem
Wireless foundation-model research remains fragmented, lacking a comprehensive survey integrating architectures, pre-training, adaptation, datasets, evaluation, deployment, and applications.
Method
The survey develops a taxonomy by architecture, pre-training strategy, and deployment layer, then reviews learning methods, datasets, benchmarks, evaluation, and applications.
Results
The survey synthesizes WFMs as unified, reusable models supporting multiple wireless functions across physical-layer processing, signal intelligence, and network optimization.
Takeaways & Limitations
WFMs provide a framework for transferable wireless intelligence, while scalable deployment still depends on diverse standardized data, robustness, domain knowledge, and reliable evaluation.
Takeaways & Limitations
Large-scale deployment remains constrained by limited standardized datasets, distribution shifts, insufficient physical consistency, inefficient edge deployment, and unreliable evaluation protocols.
Abstract
from arXiv · showhide
Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless networks by enabling scalable, transferable, and data-efficient intelligence across diverse communication tasks. Unlike conventional deep learning models that are trained for individual applications, wireless foundation models (WFMs) learn generalized representations from large-scale heterogeneous wireless data and can be efficiently adapted to communication, sensing, localization, and network optimization tasks with minimal task-specific supervision. Despite rapid progress, current research remains fragmented across architectures, training paradigms, and application domains, with no unified survey dedicated to the design, learning, and deployment of WFMs. This survey presents a comprehensive and unified review of wireless foundation models. We first establish the fundamental concepts of WFMs and introduce a taxonomy that organizes the field according to model architectures, pre-training paradigms, and applications. We then review representative architectures, self-supervised pre-training strategies, parameter-efficient adaptation methods, datasets, benchmarks, and evaluation methodologies, highlighting their roles in enabling transferable wireless intelligence. Furthermore, we examine emerging applications spanning physical-layer signal processing, network intelligence, and cross-layer optimization, and discuss the key challenges of data availability, generalization, interpretability, efficient edge deployment, and standardization. Finally, we outline future research directions toward scalable, trustworthy, and general-purpose wireless intelligence for AI-native 6G networks. This survey provides a comprehensive reference for researchers and practitioners developing next-generation intelligent wireless systems.
I. INTRODUCTION · A. Motivation and Contributions
Wireless communications are moving toward data-driven intelligence for AI-native 6G, but task-specific learning and fragmented research motivate a unified wireless foundation model paradigm. This survey establishes its framework, taxonomy, reviews, applications, challenges, and roadmap.
- I. INTRODUCTION: Wireless systems are shifting from model-driven methods toward data-driven intelligence to satisfy AI-native 6G requirements for performance, scalability, and adaptability.Traditional methods supported channel estimation, signal detection, beamforming, resource allocation, and network management.
- I. INTRODUCTION: Deep learning has expanded from physical-layer tasks to localization, modulation recognition, intelligent receivers, and network-level resource optimization, but most approaches remain task-specific.Early successes included channel estimation, signal detection, and beamforming, while separate models are generally required for different tasks.
- I. INTRODUCTION: Foundation models address these limitations by pre-training on large-scale heterogeneous data with self-supervised or unsupervised objectives and adapting transferable representations to multiple applications.This paradigm has reshaped natural language processing and computer vision.
- I. INTRODUCTION: WFMs enable scalable, general-purpose wireless intelligence by adapting a shared pre-trained backbone across communication, sensing, localization, and network optimization tasks.Adaptation mechanisms include fine-tuning, parameter-efficient adaptation, prompt tuning, and in-context learning; examples include WirelessGPT and Large Wireless Models.
- A. Motivation and Contributions: Research on WFMs remains fragmented across communities and application domains, while existing surveys do not provide a unified treatment integrating architecture, self-supervised learning, and related aspects.The survey addresses the resulting lack of comprehensive coverage.
- A. Motivation and Contributions: The survey unifies WFM principles, architectures, learning paradigms, adaptation, deployment, evaluation, and applications while comparing representative models and strategies and identifying trends and future directions.It aims toward scalable, trustworthy, general-purpose wireless intelligence.
- A. Motivation and Contributions: Its contributions include a conceptual framework, a taxonomy by architectures, learning paradigms, and applications, reviews of models, datasets, benchmarks, metrics, and applications, and a roadmap addressing data scarcity, generalization, interpretability, deployment, adaptation, and standardization.Reviewed architectures include Transformer-based, physics-informed, graph-based, and multimodal models; applications span physical-layer processing, sensing, localization, resource management, and network intelligence.
B. Organization … B. Deep Learning for Wireless Communications
The survey traces wireless intelligence from model-based methods through task-specific deep learning to wireless foundation models, while outlining the survey’s organization around architectures, training, adaptation, applications, and evaluation. It emphasizes that reusable wireless representations are needed to overcome the scalability, adaptability, and generalization limits of task-specific learning in AI-native 6G networks.
- B. Organization: The survey reviews model-based and learning-based wireless communications, then covers WFM concepts, taxonomy, architectures, pre-training, adaptation, datasets, benchmarks, and applications.Its taxonomy is organized by architecture, training strategy, and protocol layer.
- II. EVOLUTION TOWARD WIRELESS FOUNDATION MODELS: WFMs emerge as the latest stage in wireless artificial intelligence because conventional model-based signal processing and task-specific deep learning face increasing AI-native 6G complexity.This evolution motivates large-scale transferable representation learning.
- II. EVOLUTION TOWARD WIRELESS FOUNDATION MODELS: The transition toward WFMs follows a progression from model-based wireless communications to deep learning, inspired by foundation-model advances in NLP and CV.The intended wireless systems use a shared pre-trained backbone to support multiple communication tasks.
- A. Model-Based Wireless Communications: Traditional wireless systems use analytical models, estimation theory, optimization, and statistical signal processing for channel estimation, detection, synchronization, beamforming, equalization, and resource allocation.Classical LS, MMSE, ML, and convex optimization methods provide interpretable solutions with theoretical guarantees.
- A. Model-Based Wireless Communications: Model-driven methods depend on accurate propagation and network descriptions, including statistical channel models, reliable CSI, and tractable optimization formulations.Under these assumptions, their performance is predictable and can often be analyzed rigorously.
- A. Model-Based Wireless Communications: AI-native 6G technologies such as massive MIMO, millimeter-wave and terahertz communications, ISAC, dense heterogeneous deployments, and dynamic propagation make accurate mathematical modeling and optimal signal processing more difficult.These challenges motivate data-driven approaches that extract knowledge directly from wireless measurements.
- B. Deep Learning for Wireless Communications: Deep learning learns complex nonlinear relationships from wireless observations and supports physical-layer tasks including channel estimation, detection, beamforming, localization, modulation recognition, and resource allocation.However, these systems are predominantly task-specific and often degrade under changes in propagation, antennas, mobility, hardware, or other distribution shifts.
- B. Deep Learning for Wireless Communications: Task-specific deep learning requires costly labeled data and substantial computation, memory, and training time, while its black-box nature challenges latency-sensitive and resource-constrained devices.These limitations motivate reusable wireless representations that transfer knowledge across heterogeneous tasks and deployment scenarios.
III. DEFINITIONS AND TAXONOMY OF WIRELESS FOUNDATION MODELS · A. Definition of Wireless Foundation Models · B. Taxonomy by Architecture
Wireless foundation models use large-scale heterogeneous wireless data to learn transferable representations that can be efficiently adapted across diverse downstream applications. Their taxonomy organizes models by architecture, pre-training strategy, and deployment/application, with architecture families offering complementary design philosophies.
- III. DEFINITIONS AND TAXONOMY OF WIRELESS FOUNDATION MODELS: The taxonomy classifies WFMs by model architecture, pre-training strategy, and deployment/application, covering physical-layer, MAC-layer, network-layer, and cross-layer optimization tasks.These dimensions collectively describe the model lifecycle from representation learning to downstream adaptation.
- A. Definition of Wireless Foundation Models: WFMs replace independently trained task-specific models with a unified pre-trained backbone that reuses transferable knowledge across multiple wireless tasks with limited supervision.This distinguishes WFMs from conventional machine learning and deep learning models optimized for a single task or deployment scenario.
- A. Definition of Wireless Foundation Models: WFM learning follows two stages: large-scale pre-training on heterogeneous wireless measurements, followed by downstream adaptation using lightweight or task-specific techniques.Pre-training can use masked reconstruction, contrastive learning, or generative modeling, while adaptation can use fine-tuning, parameter-efficient methods, prompt tuning, or prediction heads.
- B. Taxonomy by Architecture: Architectural backbones influence representation capability, scalability, computational efficiency, and adaptation performance, and current WFMs comprise four broad families.The four families are Transformer-based, CNN-based, physics-informed, and graph-based models.
- B. Taxonomy by Architecture: Transformer-based, CNN-based, physics-informed, and graph-based architectures provide complementary design philosophies for generalized wireless representation learning rather than competing solutions.The taxonomy emphasizes architectural diversity across communication scenarios.
- B. Taxonomy by Architecture: CNN-based models exploit local spatial and temporal correlations to provide lightweight feature extraction for wireless tasks requiring computational efficiency and low-latency inference.Their lower representation capacity than Transformers is offset by suitability for edge deployment and resource-constrained devices.
- B. Taxonomy by Architecture: Physics-informed architectures embed analytical models, optimization algorithms, or signal-processing principles into trainable networks, while graph-based models represent wireless entities and their relationships explicitly.Physics-informed methods can improve interpretability, robustness, sample efficiency, and generalization; graph neural networks capture topology-dependent spatial interactions for network tasks.
C. Taxonomy by Learning Paradigm · D. Taxonomy by Deployment
Wireless foundation models are primarily organized by self-supervised and multimodal learning paradigms that acquire reusable representations from heterogeneous wireless data. They are also classified by deployment layer, spanning physical signal processing, MAC resource management, network intelligence, and cross-layer coordination.
- C. Taxonomy by Learning Paradigm: Self-supervised pre-training is the dominant WFM learning paradigm, exploiting abundant unlabeled wireless measurements instead of manually annotated, task-specific datasets.It derives supervisory signals from input data through designed pretext tasks.
- C. Taxonomy by Learning Paradigm: WirelessGPT, LWM, WavesFM, IQFM, and CSI2Vec illustrate how self-supervised learning produces transferable or reusable representations across heterogeneous observations, signals, and tasks.These models reduce dependence on task-specific labels while enabling representation reuse.
- C. Taxonomy by Learning Paradigm: Contrastive learning separates unrelated samples while bringing related wireless observations together, whereas masked modeling reconstructs hidden inputs to capture spatial, temporal, and frequency-domain signal structure.Masked reconstruction can learn propagation and signal structure from correlated wireless data.
- C. Taxonomy by Learning Paradigm: Multimodal pre-training jointly represents wireless and contextual information, integrating modalities such as CSI, IQ samples, CIRs, radar, LiDAR, RGB images, GPS, traffic statistics, and environmental context.Multimodal WFM, MuSE-FM, and CSI-CLIP demonstrate shared-space integration and cross-modal contrastive alignment.
- D. Taxonomy by Deployment: Deployment taxonomy distinguishes protocol-stack layers because they use different data sources, optimization objectives, latency requirements, and decision timescales, while a shared backbone can span the wireless stack.The taxonomy extends from the physical layer toward higher-layer intelligence.
- D. Taxonomy by Deployment: At the physical layer, WFMs process CSI, IQ samples, CIRs, and RF signals for channel estimation, CSI feedback, MIMO detection, beam prediction, classification, localization, and integrated sensing.Large-scale pre-training benefits these tasks because of strong spatial, temporal, and frequency-domain correlations.
- D. Taxonomy by Deployment: At the MAC layer, WFMs support radio resource management, including spectrum access, scheduling, beam management, interference coordination, power allocation, and link adaptation.LLM-based approaches additionally address contextual resource allocation, task adaptation, and MAC-layer coordination.
- D. Taxonomy by Deployment: At the network layer, WFMs use traffic, mobility, topology, telemetry, service, and QoS information for prediction, routing, anomaly detection, slicing, and intent-driven networking; cross-layer models jointly represent coupled states and objectives.Higher-layer decisions influence radio allocation, energy management, and reliability, motivating cross-layer WFM design.
IV. ARCHITECTURES FOR WIRELESS FOUNDATION MODELS · A. Transformer-Based Architectures · B. Physics-Informed and Model-Driven Architectures
Wireless foundation model architectures prioritize transferable representations, with Transformers serving as the dominant scalable backbone and physics-informed or model-driven designs embedding communication principles into learning. Together, these approaches support adaptation across heterogeneous wireless tasks while addressing computational scalability, domain structure, data demands, and generalization challenges.
- IV. ARCHITECTURES FOR WIRELESS FOUNDATION MODELS: Architecture choice largely determines wireless foundation models’ representation learning capability, computational efficiency, scalability, and adaptation performance.
- A. Transformer-Based Architectures: Transformer-based architectures dominate because self-attention captures spatial, temporal, and frequency-domain correlations while scaling to large self-supervised wireless datasets.
- A. Transformer-Based Architectures: Transformers learn reusable representations from CSI, IQ samples, CIRs, spectrograms, and RF signals, then support downstream adaptation through fine-tuning, LoRA, adapters, or prompt tuning.These lightweight methods reduce labeled-data requirements and retraining costs.
- A. Transformer-Based Architectures: WirelessGPT, LWM, and WavesFM exemplify Transformer-based models for transferable representation learning across communication, sensing, localization, channel modeling, and communication intelligence.WavesFM combines a Vision Transformer backbone with masked pre-training and LoRA-based adaptation.
- A. Transformer-Based Architectures: Global self-attention has quadratic computational and memory complexity, limiting scalability for long CSI sequences, high-dimensional channel observations, and large wireless resource grids.Standard Transformers also lack inherent encoding of sparse multipath propagation, delay-Doppler structure, channel reciprocity, and antenna geometry.
- B. Physics-Informed and Model-Driven Architectures: Physics-informed and model-driven architectures address data-driven models’ dataset demands and unseen-environment generalization by incorporating channel propagation, estimation, detection, antenna processing, and optimization principles.
- B. Physics-Informed and Model-Driven Architectures: Deep unfolding converts iterative communication algorithms into trainable neural layers, including unfolded AMP, PGD, and WMMSE, while preserving conventional optimization structure.
- B. Physics-Informed and Model-Driven Architectures: Recent wireless foundation models incorporate physical priors into large-scale self-supervised pre-training, using channel reciprocity, sparse multipath, delay-Doppler structure, antenna geometry, and optimization constraints.This guides representation learning toward physically meaningful solutions rather than generic statistical representations.
C. Multimodal Architectures · D. Scalability Considerations
Multimodal wireless foundation models integrate heterogeneous wireless and contextual information to support communication, sensing, localization, and network intelligence. Their practical scalability depends on parameter-efficient adaptation, transferable training, and computation-aware architectures that balance representation capability with deployment constraints.
- C. Multimodal Architectures: Multimodal architectures jointly learn representations from communication signals and contextual information to support communication, sensing, localization, and network intelligence.They extend Transformer-based foundation models beyond single-wireless-modality processing.
- C. Multimodal Architectures: Recent multimodal wireless foundation models incorporate radar, LiDAR, cameras, GPS, network telemetry, and environmental measurements alongside wireless data.WirelessGPT instead learns generalized channel representations through large-scale self-supervised pre-training for transfer across communication and sensing tasks.
- C. Multimodal Architectures: Multimodal models use dedicated modality encoders and fusion mechanisms, including early fusion, late fusion, cross-attention, and multimodal Transformer encoders.These mechanisms align heterogeneous feature spaces while preserving modality-specific characteristics before learning shared representations.
- C. Multimodal Architectures: Multimodal architectures face modality alignment, synchronization, missing observations, higher computational complexity, and limited large-scale multimodal datasets.Scalable multimodal pre-training, efficient cross-modal representation learning, and standardized benchmark datasets are identified as needed responses.
- D. Scalability Considerations: Scalability is required for WFMs to operate across diverse communication tasks, heterogeneous wireless environments, channel configurations, frequency bands, antenna settings, and resource-constrained platforms.Model size and memory consumption are central considerations for practical deployment.
- D. Scalability Considerations: Larger backbone models improve representation learning and transferability but increase storage requirements, communication overhead, and fine-tuning costs.Training scalability is also constrained by large heterogeneous datasets, substantial pre-training resources, and variation across environments, hardware, frequency bands, and standards.
- D. Scalability Considerations: Shared pre-trained backbones combined with parameter-efficient adaptation allow multiple downstream tasks to reuse a common model without full retraining.LoRA updates only low-rank parameter components and substantially reduces trainable parameters during adaptation, including in WavesFM.
- D. Scalability Considerations: ComHymba, AirFM, lightweight encoders, windowed attention, and patch-based processing reduce inference or attention costs for long wireless sequences and high-dimensional inputs.Architectural choices balance representation capability, computational efficiency, data availability, latency, and deployment requirements; Transformers favor pre-training and transfer, while CNN-based and physics-informed models suit constrained applications.
V. PRE-TRAINING STRATEGIES · A. Self-Supervised Learning · B. Data Generation and Simulation-Based Pre-Training
Wireless foundation models primarily use self-supervised learning to extract transferable representations from abundant unlabeled wireless measurements, while simulation and hybrid datasets address the cost and scarcity of real-world data. These strategies combine masked and contrastive objectives with synthetic channel generation, but must address simulation-to-reality gaps.
- A. Self-Supervised Learning: Self-supervised learning dominates wireless foundation-model pre-training because it learns representations from large volumes of unlabeled wireless data.It avoids dependence on manually annotated datasets.
- A. Self-Supervised Learning: Raw CSI, IQ samples, CIR, and spectrograms provide readily available inputs for self-supervised pretext tasks, whereas accurate labels are expensive and time-consuming.The objective is to capture spatial, temporal, and frequency-domain characteristics of wireless data.
- A. Self-Supervised Learning: Masked signal modeling hides and reconstructs CSI elements, IQ samples, OFDM resource grids, or time-frequency patches to learn intrinsic wireless signal structure without explicit supervision.WavesFM and scalable masked channel models use reconstruction-based pre-training.
- A. Self-Supervised Learning: Modern wireless foundation models combine masked reconstruction with contrastive learning to capture signal structure while improving representation discrimination and transferability.This combination is described as the prevailing pre-training strategy.
- A. Self-Supervised Learning: Signal-domain augmentation creates multiple views of heterogeneous wireless observations, which are optimized with complementary reconstruction and contrastive objectives.The resulting shared representations capture both structural and discriminative signal characteristics.
- B. Data Generation and Simulation-Based Pre-Training: Large-scale, diverse pre-training datasets are critical, but real-world measurement collection is costly, time-consuming, and constrained by deployment, hardware, and privacy considerations.Labeled datasets for channel estimation, beam prediction, and localization require extensive campaigns across propagation conditions.
- B. Data Generation and Simulation-Based Pre-Training: Simulation-based pre-training uses stochastic channel models and ray-tracing to generate channel realizations under diverse propagation conditions for wireless foundation-model training.Ray-tracing represents reflections, diffraction, scattering, blockage, and other geometry-dependent phenomena.
- B. Data Generation and Simulation-Based Pre-Training: Hybrid pre-training combines synthetic and measured data to broaden scenario coverage and improve transfer, but synthetic simplifications create a persistent Sim2Real gap.Higher-fidelity simulators, domain adaptation, and hybrid datasets are proposed responses to hardware non-idealities, environmental dynamics, and unexpected interference.
C. Transfer Learning and Fine-Tuning … A. Channel Estimation and Equalization
Wireless foundation models support scalable wireless intelligence by transferring generalized representations to multiple tasks with limited labeled data, progressing from full fine-tuning toward parameter-efficient, prompt-based, and in-context adaptation. In channel estimation and equalization, self-supervised pre-training enables shared representations that improve robustness across diverse environments and hardware conditions.
- C. Transfer Learning and Fine-Tuning: Transfer learning adapts pre-trained wireless foundation models to multiple communication tasks using significantly less labeled data than training from scratch.A shared backbone can support channel estimation, signal detection, beam prediction, and localization through task-specific adaptation.
- C. Transfer Learning and Fine-Tuning: Full fine-tuning updates all model parameters and generally provides the highest task-specific performance, but requires substantial computational resources, memory, and storage.Maintaining separate models for each downstream application becomes impractical when many communication and sensing tasks must run simultaneously.
- C. Transfer Learning and Fine-Tuning: Parameter-efficient fine-tuning reduces adaptation costs by optimizing only a small subset of model parameters instead of retraining the full backbone.This approach addresses the resource demands associated with separate downstream models.
- D. Prompting and In-Context Learning: Prompt tuning guides a pre-trained model with learnable prompt embeddings or task instructions while preserving knowledge acquired during pre-training.It further reduces computational cost without modifying the model parameters.
- D. Prompting and In-Context Learning: In-context learning performs new wireless tasks from a small number of demonstration examples without updating model parameters.This enables rapid adaptation to unseen channel conditions, network configurations, and communication objectives with little or no retraining.
- D. Prompting and In-Context Learning: Parameter-efficient fine-tuning, prompt-based adaptation, and in-context learning form a progression toward lightweight, flexible mechanisms for scalable multi-task wireless intelligence.These strategies are expected to support heterogeneous AI-native 6G systems.
- A. Channel Estimation and Equalization: Channel estimation and equalization are mature wireless foundation model applications because they directly affect communication reliability, throughput, and spectral efficiency.Conventional models often degrade under unseen propagation environments, mobility conditions, or hardware configurations because they target specific channel models, antenna configurations, or SNR ranges.
- A. Channel Estimation and Equalization: Self-supervised pre-training learns generalized spatial, temporal, and frequency-domain channel representations that transfer to estimation, prediction, CSI feedback, and equalization.WirelessGPT, LWM, and WavesFM exemplify reusable representations, heterogeneous-environment robustness, and masked signal modeling for physical-layer tasks.
B. Signal Detection and Modulation Recognition · C. Beamforming and Resource Allocation · D. Localization and Integrated Sensing
Wireless foundation models support transferable intelligence across signal understanding, communication optimization, localization, and integrated sensing by learning generalized representations from heterogeneous wireless data. Across these domains, they enable adaptation, multitask operation, and unified processing for evolving AI-native 6G environments.
- B. Signal Detection and Modulation Recognition: Signal detection recovers transmitted symbols in MIMO and OFDM systems, while modulation recognition supports adaptive communication, spectrum monitoring, cognitive radio, and network management.
- B. Signal Detection and Modulation Recognition: Wireless foundation models learn generalized signal representations from large collections of unlabeled IQ samples, spectrograms, and channel measurements.
- B. Signal Detection and Modulation Recognition: Self-supervised pre-training captures signal characteristics robust to noise, fading, interference, and hardware impairments, enabling adaptation with limited task-specific supervision.The representations support signal detection, modulation classification, wireless technology recognition, and RF fingerprinting.
- B. Signal Detection and Modulation Recognition: Rapid adaptation to new modulation formats, communication protocols, and spectrum conditions supports cognitive radio, spectrum sensing, RF environment awareness, and intelligent network monitoring.
- C. Beamforming and Resource Allocation: Beamforming and resource allocation influence network capacity, spectral efficiency, energy efficiency, and QoS, but conventional methods require accurate CSI and iterative optimization for non-convex problems.
- C. Beamforming and Resource Allocation: MuSE-FM jointly exploits wireless measurements and environmental context for multitask communication optimization, while other models share representations across traffic forecasting, mobility prediction, and resource allocation.
- C. Beamforming and Resource Allocation: Large AI models introduce intent-driven wireless optimization by using natural-language reasoning for resource management and autonomous network control.This provides a direction for intelligent radio access networks beyond independently solving individual optimization problems.
- D. Localization and Integrated Sensing: Wireless localization and integrated sensing address 6G requirements for high-precision positioning, environmental perception, and reliable communication through ISAC.These capabilities require joint processing of heterogeneous measurements, including CSI, CIR, radar echoes, and IQ signals, under dynamic propagation environments.
E. Semantic and Goal-Oriented Communications · VII. DATASETS, BENCHMARKS, AND EVALUATION METRICS · A. Public Wireless Datasets
Wireless foundation models are advancing toward semantic, multimodal, and unified intelligence across communication, sensing, localization, and network optimization. Their development relies on Transformer backbones, self-supervised pre-training, parameter-efficient adaptation, and increasingly diverse datasets, although current benchmarks remain limited in scale and diversity.
- E. Semantic and Goal-Oriented Communications: Semantic and goal-oriented communications prioritize task-relevant meaning over faithful bit transmission, shifting objectives beyond BER, throughput, and spectral efficiency.Semantic communications preserve transmitted meaning, while goal-oriented communications focus on information relevant to a target task.
- E. Semantic and Goal-Oriented Communications: Large foundation models support semantic-aware wireless systems through user-intent understanding, semantic reasoning, autonomous network management, and multimodal sensing.Multimodal wireless foundation models combine wireless measurements with vision, radar, and environmental information.
- E. Semantic and Goal-Oriented Communications: Wireless foundation models are evolving from specialized physical-layer solutions into unified frameworks for communication, sensing, localization, and network optimization.Existing models still differ in architectures, pre-training strategies, adaptation mechanisms, protocol layers, and target applications.
- E. Semantic and Goal-Oriented Communications: Transformer-based architectures dominate because they model long-range spatial and temporal dependencies and support scalable multi-task learning across heterogeneous wireless modalities.Convolutional and physics-informed architectures remain valuable for specific applications.
- E. Semantic and Goal-Oriented Communications: Self-supervised learning predominates in pre-training, using masked modeling, contrastive learning, and multimodal representation learning to exploit unlabeled measurements and improve generalization.These methods reduce dependence on costly annotated datasets across diverse deployment scenarios.
- E. Semantic and Goal-Oriented Communications: Parameter-efficient adaptation, particularly LoRA, is replacing full-parameter fine-tuning by reducing computational and memory overhead across downstream communication and sensing tasks.Lightweight adaptation enables one pre-trained backbone to support multiple tasks.
- E. Semantic and Goal-Oriented Communications: The application scope now includes beamforming, resource allocation, localization, ISAC, semantic communications, and cross-layer network intelligence beyond channel estimation, prediction, and signal detection.This expansion reflects a transition from task-specific deep learning toward unified wireless intelligence.
- A. Public Wireless Datasets: Public wireless datasets underpin pre-training, benchmarking, and downstream adaptation, but existing benchmarks remain smaller, less diverse, and narrower than datasets used for natural-language and computer-vision foundation models.Current benchmarks often target specific tasks, frequency bands, or propagation environments, motivating large-scale multimodal datasets.
B. Performance Metrics · C. Generalization and Robustness Evaluation · VIII. OPEN CHALLENGES AND RESEARCH DIRECTIONS
Wireless foundation models require evaluation beyond task-specific communication metrics, covering downstream performance, generalization, robustness, efficiency, and adaptability across heterogeneous environments. Standardized benchmarks remain limited, motivating multi-domain, OOD, adversarial, continual-learning, and computational-efficiency evaluation suites.
- B. Performance Metrics: Wireless foundation model evaluation must jointly assess multi-application performance, generalization, computational efficiency, and adaptability across heterogeneous wireless environments.This broader evaluation scope differs from conventional single-task wireless algorithms.
- B. Performance Metrics: Higher-layer applications use spectral efficiency, energy efficiency, throughput, latency, and localization error, while ISAC additionally considers detection probability, false alarm rate, and sensing accuracy.Localization is commonly quantified using RMSE or mean positioning error.
- B. Performance Metrics: Cross-domain evaluation, few-shot adaptation, transfer learning performance, and OOD robustness are essential criteria for transferable wireless representations.These criteria reflect transfer across propagation environments, carrier frequencies, antenna configurations, mobility conditions, and communication tasks.
- B. Performance Metrics: Computational efficiency is critical because wireless foundation models must operate across cloud servers, edge platforms, and resource-constrained environments.Evaluation therefore extends beyond application performance to deployment feasibility.
- C. Generalization and Robustness Evaluation: Evaluating only BER or NMSE is insufficient; models must generalize across unseen propagation environments, network configurations, hardware platforms, and communication standards.Generalization evaluation compares pretrained models with deployment conditions differing from pre-training, including changes in SNR, carrier frequency, mobility, and hardware impairments.
- C. Generalization and Robustness Evaluation: Robustness evaluation examines resilience to noise, interference, adversarial attacks, and hardware non-idealities using tests such as FGSM and realistic channel and hardware impairments.Relevant conditions include channel estimation errors, synchronization offsets, phase noise, nonlinear distortion, quantization effects, and imperfect channel state information.
- VIII. OPEN CHALLENGES AND RESEARCH DIRECTIONS: Standardized evaluation protocols remain limited, making direct comparison difficult because studies use different datasets, channel models, and experimental settings.Future benchmark suites should integrate multi-domain datasets, standardized OOD evaluation, adversarial robustness testing, continual learning scenarios, and computational efficiency metrics.
A. Data Scarcity and Domain Shift · B. Interpretability and Physical Consistency · C. Robustness and Security
The survey identifies data scarcity, nonstationary wireless environments, limited physical interpretability, and robustness, security, and privacy vulnerabilities as central barriers to trustworthy wireless foundation models. It calls for multimodal datasets, proactive domain generalization, physics-aware evaluation, and integrated safeguards for AI-native 6G deployment.
- A. Data Scarcity and Domain Shift: Wireless foundation models are constrained by the lack of large-scale, diverse, representative, and standardized wireless datasets comparable to those available in language and vision.Existing public datasets are typically collected for specific communication tasks.
- A. Data Scarcity and Domain Shift: Wireless environments are nonstationary because mobility, environmental dynamics, network reconfiguration, hardware impairments, and spectrum utilization continuously change data distributions and violate IID assumptions.Most current pre-training strategies rely on independent and identically distributed data.
- A. Data Scarcity and Domain Shift: Wireless domain variation reflects propagation mechanisms, antenna configurations, frequencies, and hardware characteristics, so increasing dataset size alone cannot ensure robust generalization.Transfer learning, domain adaptation, and parameter-efficient fine-tuning improve adaptation efficiency but require target-domain data after deployment.
- A. Data Scarcity and Domain Shift: Multimodal datasets combining CSI, IQ samples, radar, LiDAR, environmental maps, mobility, traffic, and network telemetry are identified as prerequisites for general-purpose wireless intelligence.These datasets should jointly capture communication, sensing, localization, and networking information.
- B. Interpretability and Physical Consistency: Large Transformer-based architectures improve representation capability but reduce interpretability, requiring wireless foundation models to satisfy both statistical learning objectives and communication-theoretic constraints.This differs from conventional communication algorithms whose behavior can be explained through estimation theory, optimization, or information theory.
- B. Interpretability and Physical Consistency: Masked reconstruction and contrastive learning improve downstream performance but do not explicitly ensure that latent representations preserve meaningful channel characteristics or propagation behavior.The discrepancy between statistical optimization and physical consistency grows as model size and complexity increase.
- B. Interpretability and Physical Consistency: Future benchmarks should treat interpretability and physical consistency as evaluation criteria alongside BER and NMSE, quantifying physical plausibility, constraint satisfaction, uncertainty calibration, and decision reliability.Wireless-specific XAI methods should connect latent features to measurable physical phenomena, while standardized interpretability benchmarks support trustworthy models.
- C. Robustness and Security: Shared-backbone vulnerabilities can propagate across communication, sensing, localization, and network-management applications, making robustness and security fundamental design requirements.Robustness must cover adversarial perturbations, realistic channel and hardware impairments, and privacy threats involving sensitive wireless measurements and user data.
D. Energy Efficiency and Edge Deployment · E. Standardization and Practical Deployment · IX. FUTURE ROADMAP TOWARD AI-NATIVE 6G SYSTEMS
Wireless foundation models face major barriers to energy-efficient edge deployment, interoperability, standardization, and lifecycle management. The roadmap toward AI-native 6G progresses from foundational datasets and benchmarks to general-purpose, multimodal, and eventually autonomous shared wireless intelligence.
- D. Energy Efficiency and Edge Deployment: Large computational and memory requirements hinder practical edge deployment of wireless foundation models, whose millions or billions of parameters demand substantial pre-training and inference resources.Cloud infrastructures can support such models, but many future wireless applications require more resource-efficient deployment.
- D. Energy Efficiency and Edge Deployment: Lightweight adaptation and compression methods—including pruning, quantization, distillation, low-rank approximation, and PEFT—reduce computational and memory demands.LoRA and adapter modules update only a small parameter subset while allowing downstream tasks to share a pre-trained backbone.
- D. Energy Efficiency and Edge Deployment: Energy efficiency should become an explicit model-lifecycle objective, jointly balancing communication accuracy with computational efficiency through hardware-aware search, dynamic scaling, early exits, and energy-aware scheduling.These approaches can adapt computational complexity to available hardware resources and application requirements.
- E. Standardization and Practical Deployment: Operational deployment requires standardized interfaces, interoperable architectures, unified evaluation methodologies, and compatibility with existing wireless communication standards.Current studies remain concentrated on algorithm development and simulation-based validation.
- E. Standardization and Practical Deployment: Wireless foundation models need standardized datasets and benchmark suites because differing channel models, platforms, antenna configurations, carrier frequencies, and protocols impede direct comparison.Wireless communications lack a universally accepted counterpart to ImageNet for foundation model development.
- E. Standardization and Practical Deployment: Interoperability must span heterogeneous devices, radio access technologies, cloud-edge collaboration, distributed intelligence, hardware platforms, and communication standards while maintaining consistent performance.Standardized model interfaces are required to support this cross-platform operation.
- E. Standardization and Practical Deployment: Deployment frameworks must support secure distribution, version control, continual adaptation, validation, rollback, and legacy compatibility as propagation environments, spectrum bands, services, and topologies evolve.Standardization should advance alongside innovation through collaboration among academia, industry, 3GPP, ITU, ETSI, and the O-RAN Alliance.
- IX. FUTURE ROADMAP TOWARD AI-NATIVE 6G SYSTEMS: The roadmap advances from 2026-2028 foundational datasets and self-supervised pre-training, through 2028-2031 general-purpose multimodal models, toward 2031-2035 autonomous wireless intelligence engines.This evolution requires trustworthy AI, continual learning, privacy-preserving training, energy-efficient edge intelligence, interoperable frameworks, and standardized evaluation.
X. CONCLUSION
WFMs unify transferable wireless intelligence across communication, sensing, localization, and network intelligence through scalable pre-training and reusable models. Their large-scale deployment depends on addressing data, robustness, domain knowledge, deployment, and evaluation challenges while advancing toward multimodal, physics-aware, continually adaptive systems.
- Conclusion: WFMs replace isolated task-specific models with transferable representations supporting communication, sensing, localization, and network intelligence across heterogeneous environments.This paradigm also reduces dependence on large labeled datasets.
- Conclusion: The survey unifies architectures, learning paradigms, deployment, datasets, pre-training, adaptation, applications, and evaluation into one WFM landscape.It emphasizes scalable pre-training, self-supervised learning, parameter-efficient adaptation, multimodal representation learning, and physics-informed methods as technological foundations.
- Conclusion: WFMs are transitioning from narrowly optimized algorithms toward unified, reusable models that support multiple wireless functions within a single learning framework.This shift enables a common intelligence platform rather than collections of independently optimized algorithms.
- Conclusion: Large-scale WFM deployment requires diverse standardized datasets, robustness to distribution shifts, communication-domain knowledge integration, efficient cloud-edge deployment, and reliable evaluation protocols and benchmarks.The passage identifies these as fundamental remaining challenges before deployment at scale.
- Conclusion: Future WFMs are expected to become multimodal, physics-aware, and continually adaptive, reasoning across communication, sensing, localization, and network management.The roadmap progresses from large-scale data collection and self-supervised pre-training toward multimodal intelligence, cloud-edge collaboration, autonomous management, and AI-native 6G systems.