Source-linked AI summary
6G Native AI and Channel Foundation Models
Shugong Xu, Jun Jiang, Yuan Gao
TL;DR
6G native AI lacks a common definition and requires adaptable, scenario-generalizable intelligence beyond fragmented task-specific models. This paper positions channel foundation models as reusable channel-representation learners, with preliminary CSI-CLIP evidence showing improved positioning and beam prediction under limited task-specific supervision.
Problem
6G native AI lacks a common definition and must support adaptable intelligence across diverse tasks, propagation conditions, and operating scenarios.
Method
The paper positions channel foundation models as pretrained channel-centric representations learned from heterogeneous data for adaptation across downstream wireless tasks.
Results
21.57% average relative positioning improvement over the non-pretrained ViT baseline accompanied CSI-CLIP’s gains across all six beam-prediction scenarios.
Takeaways & Limitations
CFMs are presented as a possible technical option for 6G native AI through reusable channel representations adapted to multiple downstream wireless tasks.
Takeaways & Limitations
Wireless channel labels are expensive to obtain and difficult to maintain, while generalization can degrade across propagation environments, frequency bands, and antenna configurations.
Abstract
from arXiv · showhide
The integration of artificial intelligence (AI) and wireless communications is widely regarded as a core objective of sixth-generation (6G) systems. However, both the meaning of native AI and the type of AI capability that should be embedded into future wireless systems remain open to interpretation. This paper discusses 6G native AI from a system-design perspective and argues that native AI should be co-designed, optimized, and deployed as an intrinsic component of the wireless system rather than as a removable post-deployment add-on. From this perspective, conventional task-specific supervised models are difficult to use as the main technical basis of native AI because they depend heavily on labeled data, generalize poorly across propagation conditions, and require fragmented designs for different channel-related tasks. Motivated by these limitations, we position channel foundation models (CFMs) as a channel-centric foundation-model paradigm for 6G native AI. We define the scope of CFMs, clarify their differences from task-specific wireless AI models and large language models, and summarize three pretraining families: generative, discriminative, and hybrid pretraining. We further discuss how CFMs may support physical-layer processing, radio access network intelligence, and integrated sensing and communications. Preliminary CSI-CLIP-based results are included as bounded evidence that CFM-style pretraining can improve positioning and beam prediction when task-specific labels are limited.
I. INTRODUCTION
The section argues that 6G native AI should be designed as an intrinsic wireless-system capability requiring adaptable, generalizable, and scalable intelligence. It positions channel foundation models (CFMs) as a channel-centric path based on reusable representations and pretraining for diverse downstream wireless tasks.
- Channel foundation models: CFMs learn reusable channel representations from large-scale heterogeneous channel data and adapt them to downstream wireless tasks with limited task-specific supervision.Their design focus is reusable channel representation rather than merely increasing neural-network size.
- Native AI requirements: 6G native AI is framed as a system-level capability requiring task adaptability, scenario generalization, and deployment-aware scalability.These requirements arise from heterogeneous operating environments and diverse tasks including channel estimation, CSI feedback, beam management, positioning, and ISAC.
- Native AI requirements: Conventional task-specific wireless AI is insufficient as the main basis of native AI, motivating a shift toward foundation-model-style pretraining.The paper contrasts single-task supervised learning with multi-task learning and foundation-model approaches.
- Channel foundation models: The paper defines CFMs as channel-specialized foundation models and distinguishes them from task-specific supervised models, LLMs, and generic pretrained models.This establishes the channel-centric scope of the proposed foundation-model paradigm.
- CFM applications and evidence: CFM pretraining is organized into generative, discriminative, and hybrid strategies, with potential roles in physical-layer processing, RAN intelligence, and ISAC.Preliminary CSI-CLIP-based positioning and beam-prediction results provide bounded evidence under limited downstream supervision.
II. NATIVE AI REQUIREMENTS IN 6G · A. Native AI as a Built-in Capability
Native AI in 6G should be designed as an inherent wireless-system capability rather than added after deployment. Its requirements include cross-task adaptability, robustness to changing scenarios, and a balance between model capacity and deployment cost.
- A. Native AI as a Built-in Capability: Native AI must be included in the wireless system’s initial design space, not added as an external post-deployment function.This design space includes data collection, model updating, signaling overhead, inference latency, deployment location, and protocol interactions.
- A. Native AI as a Built-in Capability: Together, native-AI evaluation must consider system-level integration and practical deployment trade-offs alongside model performance.The requirements connect model evaluation to data, signaling, latency, placement, protocol interactions, and cross-scenario operation.
- A. Native AI as a Built-in Capability: Task adaptability requires reusing channel knowledge across communication, sensing, localization, and control rather than training independent models from scratch.These tasks may share channel state, multipath structure, mobility patterns, and blockage conditions despite differing labels and objectives.
- A. Native AI as a Built-in Capability: Scenario generalization requires useful representations across diverse 6G propagation conditions, including urban, indoor, industrial, aerial, high-speed, and satellite-terrestrial settings.When propagation distributions shift, the model should maintain useful representations or support efficient adaptation with limited new data.
- A. Native AI as a Built-in Capability: Native AI must account for storage, latency, energy, and signaling constraints because it is also a deployment problem, not only an offline training problem.Very large models may be difficult to deploy at the edge or in latency-sensitive radio functions.
- A. Native AI as a Built-in Capability: Very small task-specific models may be efficient but fail to generalize, whereas very large models may exceed edge and latency constraints.The paper frames native-AI design as a balance between representation capacity and deployment cost.
B. Limitations of Task-Specific Wireless AI
Task-specific wireless AI models can work in fixed settings but face four major limitations: dependence on labeled data, weak cross-scenario generalization, task fragmentation, and limited online adaptation. These limitations increase deployment and maintenance burdens and conflict with native AI as a coherent 6G system capability.
- Labeled-data dependence: Single-task supervised models depend strongly on the quality and quantity of labeled samples, which limits their reliability outside fixed settings.Examples include channel estimation, beam prediction, positioning, and channel feedback.
- Labeled-data dependence: Wireless channel labels are expensive and difficult to maintain because they require controlled measurements, simulations, annotations, positioning devices, or extensive calibration amid changing propagation conditions.Interference, blockage, mobility, and user density can vary over time, undermining dataset stability and representativeness.
- Weak generalization: Models trained for one propagation environment, frequency band, or antenna configuration may overfit and degrade sharply when transferred to another scenario.This weakness is especially problematic when terrestrial, aerial, maritime, and satellite domains coexist or when channel conditions exceed the training distribution.
- Task fragmentation: Conventional designs require dedicated architectures, feature pipelines, and training procedures for new wireless tasks, creating storage, maintenance, and integration burdens.The resulting proliferation of specialized models is inconsistent with native AI as a coherent system capability.
- Limited online adaptation: Many task-specific models adapt slowly to mobility, obstruction, hardware variation, and environmental dynamics, threatening low-latency and reliable operation.The limitation matters for applications including autonomous driving, industrial control, and remote healthcare.
III. FROM WIRELESS AI MODELS TO CFMS · A. Evolution of Wireless AI Paradigms
Wireless AI has evolved from isolated supervised models, through multi-task learning, to foundation-model-style pretraining with downstream adaptation. This progression shifts emphasis toward reusable channel representations learned from abundant unlabeled data rather than task-specific labels.
- A. Evolution of Wireless AI Paradigms: Single-task supervised learning follows a “one task, one dataset, one model” paradigm for specific wireless functions.It trains neural networks end to end for individual tasks.
- A. Evolution of Wireless AI Paradigms: Single-task models can perform strongly under matched data distributions but are label-hungry and limited in cross-task transfer.Their strengths therefore depend on alignment between training and deployment data.
- A. Evolution of Wireless AI Paradigms: Multi-task learning shares part of a model across related tasks, exploiting common structures and reducing duplicated training effort.Its benefits depend on compatible input-output structures or strong statistical relationships among tasks.
- A. Evolution of Wireless AI Paradigms: When objectives differ substantially, such as CSI reconstruction, beam selection, and user-location estimation, one multi-task architecture may require substantial task-specific design.These tasks do not necessarily provide sufficiently compatible input-output structures or statistical relationships.
- A. Evolution of Wireless AI Paradigms: Foundation-model-style pretraining first uses large-scale data, often with self-supervised objectives, before downstream adaptation.Adaptation can use lightweight finetuning or task-specific heads.
- A. Evolution of Wireless AI Paradigms: Wireless AI paradigms comprise three stages: single-task supervised learning, multi-task learning, and foundation-model-style pretraining followed by downstream adaptation.Foundation-model-style pretraining learns reusable representations before downstream adaptation.
- A. Evolution of Wireless AI Paradigms: Wireless foundation-model pretraining is attractive because unlabeled channel data are available from simulation, measurement, or network operation, whereas task labels remain expensive.The intended outcome is channel representations that are not tied to a single downstream task.
B. Definition and Scope of CFMs
Channel foundation models (CFMs) are wireless-channel-centered foundation models pretrained on heterogeneous channel data to learn reusable representations of propagation characteristics. Their scope spans channel-related inputs and multiple downstream tasks, distinguishing them from task-specific supervised models and language-focused LLMs.
- Definition: CFMs are specialized foundation models whose central object is the wireless channel, not language, images, or generic network logs.They are pretrained offline on large-scale heterogeneous channel data, including CSI, CIR, baseband I/Q signals, or other channel measurements.
- Input and Representation: CFM inputs are defined around channel observations, including complex CSI, amplitude-phase features, angle-delay features, antenna-subcarrier matrices, I/Q streams, and spectrograms.Their pretrained output is typically a latent representation encoding channel structure such as multipath behavior, spatial-frequency correlation, delay-domain sparsity, or blockage conditions.
- Task Scope: CFM tasks are channel-centered and include channel estimation, extrapolation, CSI compression and feedback, precoding support, scenario classification, positioning, beam management, and sensing.These tasks are assumed to depend on shared channel physics and benefit from a common channel representation.
- Relation to Other Models: CFMs differ from task-specific supervised models because their defining feature is large-scale pretraining for transferable channel representations rather than merely larger model backbones.They also differ from LLMs, which are pretrained mainly on language or general knowledge and do not directly learn channel physics without channel-specific adaptation.
C. Comparison with Other AI Paradigms
The section qualitatively compares task-specific supervised models, large language models (LLMs), and channel foundation models (CFMs) as candidates for native AI. It emphasizes trade-offs in deployment latency, adaptability, generalization, reasoning, parameter scale, and channel-centric pretraining rather than establishing a universal ranking.
- Design comparison: Table I presents a qualitative design comparison of task-specific supervised models, LLMs, and CFMs rather than a universal ranking.The comparison is framed from the perspective of native AI.
- Task-specific supervised models: Task-specific models support low-latency deployment but have limited task adaptability and scenario generalization.These models are attractive when low latency is a primary deployment consideration.
- Large language models: LLMs offer language-interface and general-reasoning capabilities, but their large parameter scale and non-channel-centric pretraining reduce their suitability for native wireless AI.The passage contrasts general-purpose capabilities with channel-specific requirements.
D. Core Properties of CFMs
CFMs are designed to generalize across heterogeneous wireless scenarios, adapt to multiple downstream channel tasks, and scale through increased model capacity and broader pretraining data. These properties reduce repeated learning and redesign requirements while supporting richer channel representations.
- Cross-scenario and cross-configuration generalization: CFMs use heterogeneous pretraining across environments, frequency bands, antenna configurations, user distributions, and mobility patterns to improve cross-scenario and cross-configuration generalization.This diversity helps models learn general channel statistics and reduces the need to relearn basic propagation features for each deployment.
- Downstream-task adaptability: CFMs provide reusable pretrained channel representations to support channel estimation, beam prediction, positioning, channel identification, and sensing through task-specific heads.The conceptual framework describes pretraining on heterogeneous wireless channel data followed by lightweight task-specific adaptation.
- Downstream-task adaptability: CFM adaptation reduces the labeled data and model redesign required for each downstream task, without eliminating the need for task-specific data.This property enables one pretrained channel representation to serve multiple channel-related applications.
- Scalability: CFM scalability has model and data dimensions: capacity can capture finer channel dynamics, while broader channel data expands pretraining coverage.Model scaling targets nonlinear multipath interactions, interference, and antenna-domain correlations; representation quality may improve with capacity and data under suitable objectives.
IV. PRETRAINING STRATEGIES FOR CFMS
CFM pretraining is defined by objectives that preserve channel structure and support downstream adaptation, with three families: generative, discriminative, and hybrid. Evaluation should prioritize transfer under limited labels, domain shifts, and new wireless tasks rather than pretraining loss alone.
- IV. PRETRAINING STRATEGIES FOR CFMS: Pretraining distinguishes CFMs from larger task-specific wireless AI models by aligning objectives with channel structure and expected downstream adaptation.The objective should preserve the relevant channel structure and support adaptation after pretraining.
- IV. PRETRAINING STRATEGIES FOR CFMS: Generative pretraining reconstructs or predicts missing channel observations, exploiting unlabeled data and correlations across antennas, subcarriers, time, and delay.Masked CSI, CIR, and radio representations are example targets; WiFo and WirelessGPT are representative reconstruction-oriented models.
- IV. PRETRAINING STRATEGIES FOR CFMS: Discriminative pretraining shapes representations by distinguishing related from unrelated channel samples, with CSI-CLIP aligning frequency-domain CSI and delay-domain CIR.Its effectiveness depends on physically meaningful positive and negative pairs rather than arbitrary data augmentations.
- IV. PRETRAINING STRATEGIES FOR CFMS: Hybrid pretraining combines reconstruction with contrastive alignment to preserve local channel structure while improving representation separation.The practical challenge is balancing losses so neither objective dominates and reduces downstream transferability.
- IV. PRETRAINING STRATEGIES FOR CFMS: CFM evaluation should target downstream transfer, especially adaptation under limited labels, domain shifts, or new wireless tasks, while meeting deployment constraints.Pretraining loss alone is not the evaluation target.
V. CFMS FOR 6G NATIVE AI · A. Physical-Layer Intelligence · B. Radio Access Network Intelligence
Channel foundation models can provide reusable, channel-aware representations for physical-layer and RAN intelligence, especially when labeled data are limited or channel conditions change. Their deployment should complement model-based signal processing and consolidate related RAN functions while addressing practical robustness and overhead constraints.
- A. Physical-Layer Intelligence: CFMs can support channel estimation, feedback, extrapolation, and precoding-related physical-layer tasks through reusable representations learned from heterogeneous channel data.These tasks otherwise often require separate task-specific training data and may degrade under channel-distribution changes.
- A. Physical-Layer Intelligence: Limited labeled data in a target scenario can be addressed by lightweight finetuning of a pretrained CFM representation.The passage presents finetuning as an adaptation mechanism for target scenarios with limited labels.
- A. Physical-Layer Intelligence: CFMs are intended to complement rather than replace model-based signal processing by combining learned representations with wireless priors.Relevant priors include channel sparsity, antenna geometry, delay-Doppler structure, and pilot design.
- A. Physical-Layer Intelligence: Physics-aware input representations and pretraining objectives can help avoid black-box models that work only under narrow simulated conditions.The passage identifies wireless priors as guidance for representation and objective design.
- B. Radio Access Network Intelligence: In RANs, CFMs may provide channel-aware representations for resource scheduling, interference management, mobility control, and beam management.Dense-urban beam selection is affected by blockage, user movement, cell-sector geometry, and inter-cell interference.
- B. Radio Access Network Intelligence: A shared channel representation across beam management, positioning assistance, and channel condition classification can reduce duplicated training and simplify deployment.This addresses the concern that RAN intelligence could otherwise become a collection of unrelated models.
- B. Radio Access Network Intelligence: Practical CFM deployment requires evaluation of latency, signaling overhead, model update frequency, and robustness to out-of-distribution channels.These requirements remain important even when related RAN functions share a common representation.
C. Integrated Sensing and Communications · D. Preliminary Evidence from CSI-CLIP
The section presents CFMs as unified channel representations for integrated sensing and communications, then offers bounded CSI-CLIP evidence of improved downstream transfer under limited labels. It emphasizes that broader validation is required before CFMs can be considered deployable native-AI components.
- C. Integrated Sensing and Communications: ISAC requires native AI to support communication and sensing over the same propagation environment despite their different objectives.Communication targets reliable data transmission, whereas sensing targets localization, tracking, imaging, or environment understanding.
- C. Integrated Sensing and Communications: CFMs can learn unified channel features supporting both communication and sensing, including multipath structure, delay spread, angular information, and blockage patterns.These features can support beam management and positioning within a shared representation.
- C. Integrated Sensing and Communications: CFM-based ISAC may reduce the gap between communication-oriented and sensing-oriented processing in spatially coupled wireless applications.The passage identifies intelligent transportation, indoor localization, and environment-aware networks as relevant settings.
- D. Preliminary Evidence from CSI-CLIP: CSI-CLIP experiments use a Vision Transformer encoder pretrained on more than 700,000 DeepMIMO samples spanning 35 representative wireless scenarios.The scenarios cover indoor and outdoor environments and sub-6 GHz, millimeter-wave, and terahertz frequencies, with a non-pretrained ViT baseline.
- D. Preliminary Evidence from CSI-CLIP: 21.57% average relative improvement in positioning error was achieved over the non-pretrained ViT baseline across nine city scenarios.CSI-CLIP reduced positioning error in every reported city scenario.
- D. Preliminary Evidence from CSI-CLIP: 1.75 to 2.78 percentage points of beam-prediction accuracy gains were obtained across all six scenarios.Together with the positioning results, this supports bounded evidence that channel pretraining can improve transfer for related downstream tasks with fewer task-specific labels.
- D. Preliminary Evidence from CSI-CLIP: The experiments do not settle the CFM research problem, requiring evaluation across broader datasets, cross-device and crosstime splits, hardware impairments, overhead, latency, compression, and online adaptation.These factors are necessary before CFMs can be treated as deployable native-AI components.
VI. CONCLUSION
The paper frames 6G native AI as an intrinsic system capability that must support task adaptability, scenario generalization, and deployment-aware scalability. It contrasts this goal with limitations of conventional task-specific supervised models and reports preliminary CSI-CLIP evidence under limited downstream supervision.
- Native AI should be embedded into 6G systems as an intrinsic capability rather than treated as an external addition.
- A viable native-AI design should provide task adaptability, scenario generalization, and deployment-aware scalability.
- Conventional task-specific supervised models are difficult to use as native AI’s main basis because they rely heavily on labeled data and generalize poorly across propagation conditions.
- The paper includes preliminary CSI-CLIP evidence evaluated under limited downstream supervision.