Source-linked AI summary
Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks
Qifei Wang, Zhen Gao, Li Qiao, Ziwei Wan, De Mi, Dapeng Li, Ying Sun
TL;DR
6G robotic vehicles need bandwidth-efficient visual communication that remains reliable under uplink NOMA interference and preserves perceptual image quality. KDG-SemNOMA combines channel-adaptive ConvNeXt DeepJSCC, OMA-teacher knowledge distillation, and channel-conditional GAN refinement; on FFHQ-256, it improves pixel-level fidelity and visual realism over state-of-the-art baselines.
Problem
RV visual data burdens bandwidth and energy resources, while uplink semantic NOMA systems face interference, varying channel conditions, and blurry pixel-wise reconstructions.
Method
KDG-SemNOMA integrates a ConvNeXt transceiver with CSI-conditioned AF features, two-stage OMA-to-NOMA knowledge distillation, and cGAN refinement conditioned on initial reconstructions and channel states.
Results
KDG-SemNOMA significantly improves pixel-level fidelity and visual realism on FFHQ-256; KD-SemNOMA additionally gains 0.3 ∼0.5 dB over the ConvNeXt-based SemNOMA framework.
Takeaways & Limitations
The framework offers a channel-adaptive approach for robust multi-user semantic communication in 6G robotic-vehicle networks.
Abstract
from arXiv · showhide
To achieve sustainable intelligent mobility, 6G-empowered robotic vehicles (RVs) require high-fidelity visual perception under stringent bandwidth and energy constraints. Semantic communication offers a spectral-efficient solution but suffers from severe interference in uplink non-orthogonal multiple access (NOMA) RV networks. To address this, we propose a knowledge distillation-driven and generative models-enhanced NOMA framework for robust and green RV communications, named KDG-SemNOMA. First, we develop a ConvNeXt-based deep joint source-channel coding (DeepJSCC) architecture with an enhanced attention feature (AF) module for dynamic channel adaptation. Second, to mitigate interference without inference overhead, an orthogonal transmission teacher model guides the NOMA student model via a two-stage knowledge distillation strategy. Finally, to address the over-smoothing artifacts of pixel-wise optimization, we introduce a channel-conditional GAN (cGAN). By explicitly taking the Stage-I initial reconstruction and channel states as conditional inputs, this module refines coarse outputs into high-fidelity images with realistic textures. Experiments on FFHQ-256 demonstrate that KDG-SemNOMA significantly outperforms state-of-the-art methods in both pixel-level accuracy and perceptual fidelity.
I. INTRODUCTION
KDG-SemNOMA addresses interference, channel variation, and blurry reconstructions in uplink semantic NOMA for robotic vehicles by combining channel adaptation, knowledge distillation, and cGAN refinement.
- Semantic communication reduces the bandwidth and energy burden of transmitting RV visual data by sending semantic features instead of raw bits.
- Existing DeepJSCC-based RV systems face insufficient NOMA interference cancellation, limited adaptation to varying CSI, and blurry pixel-wise reconstructions.
- KDG-SemNOMA introduces a ConvNeXt semantic transceiver with an enhanced AF module that uses SNR and Rayleigh fading parameters to modulate features dynamically.
- A two-stage KD strategy uses an interference-free OMA Teacher to guide the NOMA Student in extracting clean semantic features from superimposed signals without extra inference overhead.
- A channel-conditional GAN refines Stage-I reconstructions using spatial channel states, adaptively restoring high-frequency textures according to channel degradation.
II. SYSTEM MODEL AND NETWORK ARCHITECTURE
The system models uplink image transmission from multiple robotic-vehicle users to one base station over shared NOMA resources, using user embeddings, semantic encoding, channel distortion, and decoding.
- A. Semantic NOMA Transmission Model: N robotic-vehicle UEs transmit images to a single BS in an uplink multi-user semantic communication system.
- A. Semantic NOMA Transmission Model: Each image is concatenated with a user-specific identification vector before semantic encoding, helping the BS distinguish users in the NOMA superposition.
- A. Semantic NOMA Transmission Model: The encoder maps each combined image and identifier to a complex-valued semantic symbol sequence, with bandwidth compression ratio ρ = k/m.
- A. Semantic NOMA Transmission Model: Users transmit simultaneously over the same time-frequency resources, and the BS receives their channel-distorted signal superposition.
- A. Semantic NOMA Transmission Model: The model considers AWGN and Rayleigh fading channels, with SNR defined as γ = 10 log10(Pavg/σ2).
B. ConvNeXt-based Architecture with Enhanced AF-Module
The ConvNeXt-based transceiver uses an enhanced attention feature module to condition semantic feature modulation on instantaneous channel states. This adaptation emphasizes channel-robust features while suppressing noise-sensitive components, and the NOMA student is guided by an interference-free OMA teacher without added inference overhead.
- The ConvNeXt backbone integrates an enhanced AF module that conditions feature modulation on SNR, fading amplitude, and phase.The module projects channel statistics into an embedding, fuses them with global context, and generates a channel-wise attention mask.
- The AF mechanism adaptively emphasizes features robust to current channel conditions while suppressing noise-sensitive components.
- An interference-free OMA teacher guides the NOMA student, with only the student deployed during inference and no additional computational overhead.The teacher provides semantic supervision while the student operates under shared-resource interference.
A. Teacher Model Training
The teacher model is trained over interference-free orthogonal transmission using MAE reconstruction supervision. Its outputs and intermediate information provide the reference for initializing and training the NOMA student.
- The teacher network transmits each UE’s semantic features over an interference-free orthogonal channel.
- The teacher model is optimized with MAE reconstruction loss using the UE input image as the reconstruction target.
- The NOMA student is initialized from the pretrained teacher and trained under the NOMA channel model with three complementary losses.
1) Feature Affinity Distillation:
The distillation framework transfers structural and high-level semantic information from the teacher to the NOMA student. Feature affinity aligns spatial relationships, while CrossKD tests student features through the teacher’s decoder head.
- 1) Feature Affinity Distillation:: Feature affinity distillation minimizes differences between teacher and student spatial affinity matrices at selected decoder layers.The affinity matrices represent structural semantic relationships in the feature maps.
- 1) Feature Affinity Distillation:: The FA loss is normalized by the number of elements in each UE’s affinity matrix.
- 2) Cross-Head Prediction Distillation:: CrossKD feeds intermediate student decoder features into the teacher’s decoder head to produce a cross prediction matched to the teacher’s clean reconstruction.
- 2) Cross-Head Prediction Distillation:: CrossKD encourages compatibility with the teacher’s decision boundaries without directly forcing student and teacher outputs to coincide.
- 2) Cross-Head Prediction Distillation:: The proposed framework includes a separate Stage-II image-refinement architecture based on a channel-conditional GAN.
3) Reconstruction Loss and Total Objective:
The student objective combines reconstruction and two distillation losses through weighted summation. Stage-II then refines the Stage-I coarse reconstruction using channel-conditioned processing to restore high-frequency details.
- 3) Reconstruction Loss and Total Objective:: The student objective combines MAE, feature affinity, and CrossKD losses in a weighted sum.
- 3) Reconstruction Loss and Total Objective:: λ1, λ2, and λ3 balance reconstruction fidelity against distillation strength in the total objective.
- 3) Reconstruction Loss and Total Objective:: Stage-II refines the coarse Stage-I KD-SemNOMA reconstruction by restoring high-frequency textural details conditioned on instantaneous channel quality.
- 3) Reconstruction Loss and Total Objective:: The Stage-I reconstruction preserves fundamental semantic layout but lacks fine-grained details, while channel statistics encode SNR and fading parameters.The CSI vector uses one dimension for AWGN and three dimensions for Rayleigh fading scenarios.
- 3) Reconstruction Loss and Total Objective:: A linear projection and spatial expansion convert the low-dimensional CSI vector into a feature map matching the image resolution.
A. Conditional Generator and Discriminator
The Stage-II cGAN refines the Stage-I reconstruction using both image estimates and channel information. Its discriminator evaluates image authenticity conditioned on channel state, enabling distortion-specific restoration.
- The generator concatenates the initial estimate with a channel feature map to synthesize the refined image.
- The discriminator evaluates real and generated images conditioned on the channel state Mcsi.
- Channel conditioning guides adaptive restoration for different distortions, including heavy fading and mild noise.
B. Loss Functions
The generator uses a composite objective combining pixel consistency, perceptual similarity, and adversarial realism. The discriminator is trained with a Wasserstein-distance objective through alternating updates with the generator.
- The composite generator loss is designed to preserve pixel-wise accuracy while improving perceptual quality.
- LMAE enforces low-frequency consistency, while LLPIPS minimizes discrepancies in VGG feature space.
- The adversarial term LG_adv = −E[D([x̂, Mcsi])] represents the Wasserstein adversarial loss.
- The discriminator maximizes separation between real and synthesized distributions using an alternating optimization objective.
A. Simulation Setup
The evaluation uses FFHQ images downsampled to 256 × 256 and simulates two-user transmission over AWGN and Rayleigh fading channels. Comparisons include conventional, semantic NOMA, and orthogonal semantic baselines, using pixel-fidelity metrics.
- FFHQ images are downsampled to 256 × 256 and transmitted over AWGN and Rayleigh fading channels with SNR uniformly sampled from [0, 20] dB.
- The baselines are BPG+LDPC+QAM+SIC, DeepJSCC-NOMA, and SemOMA under equivalent and double transmission overheads.
- Figure 4 reports PSNR performance for two-user FFHQ-256 transmission at compression ratio ρ = 1/48.
- PSNR evaluates pixel-level reconstruction fidelity, while LPIPS and FID assess perceptual quality.
1) Performance analysis of KD-SemNOMA:
KDG-SemNOMA improves reconstruction fidelity and perceptual quality across the evaluated channel conditions. Its channel-conditional Stage-II refinement reduces perceptual metrics, while the complete framework yields gains over the reported baselines.
- Performance analysis of KD-SemNOMA: The ConvNeXt-based SemNOMA framework outperforms ResNet-based DeepJSCC-NOMA by 0.3 ∼0.4 dB in PSNR.
- Performance analysis of KD-SemNOMA: KD optimization adds 0.3 ∼0.5 dB over SemNOMA and exceeds same-overhead SemOMA by approximately 0.3 ∼0.5 dB.
- Effectiveness of Stage-II GAN Refinement: The two-stage KDG-SemNOMA framework significantly reduces LPIPS and FID across the entire SNR range under AWGN and Rayleigh fading.
- Impact of Channel Conditioning: The channel-conditional model consistently outperforms the variant without channel input, especially at low SNR and under complex Rayleigh fading.
- Conclusion: Simulation results demonstrate gains over state-of-the-art baselines in both pixel-level fidelity and visual realism.