Source-linked AI summary

Robust Decentralized Multi-Satellite Massive MIMO Transmission via Knowledge Distillation

Wenjing Cao, Zheng Lin, Yafei Wang, Wenjin Wang, Ye Wang, Rui Ding, Symeon Chatzinotas, Björn Ottersten

arXiv:2608.29620v1eess.SP

TL;DR

Cooperative multi-satellite massive MIMO must operate despite incomplete and imperfect sCSI, which limits decentralized precoding relative to centralized schemes. The paper distills clean global precoding knowledge from a centralized teacher into lightweight students and adds angle-phase calibration. Simulations show improved decentralized sum-rate performance and robustness across angle-error distributions and satellite-user scales.

  • Problem

    Decentralized satellites lack the clean global sCSI and extensive inter-satellite information needed to acquire cooperative precoding knowledge.

  • Method

    A centralized teacher trained on clean global sCSI distills cooperative precoding knowledge into lightweight students using partial noisy inputs and hybrid angle-phase correction.

  • Results

    The proposed framework significantly improves decentralized sum-rate performance and remains robust across angle-error distributions and satellite-user scales.

  • Takeaways & Limitations

    Hybrid distillation transfers interpretable steering-vector and precoder knowledge while supporting robust decentralized onboard inference.

Abstract

from arXiv · show

This paper investigates robust decentralized transmission for cooperative multi-satellite massive multiple-input multiple-output (MIMO) systems under imperfect statistical channel state information (sCSI). In the considered scenario, each satellite has complete access to its local information but receives partial information from other satellites due to limited inter-satellite links (ISLs), with only imperfect sCSI available. To address these challenges, we propose a knowledge distillation (KD) framework that transfers cooperative precoding knowledge from a centralized teacher neural network (NN) to lightweight decentralized student NNs. Specifically, a global-clean teacher, aggregating information from all satellites and accessing accurate sCSI during offline training, transfers its cooperative precoding knowledge to partial-noisy students, relying on complete local information, limited information exchanged by other satellites, and error-corrupted sCSI for local precoding. The teacher NN combines patch-wise self-attention with dual-axis attention to learn inter-user interference and inter-satellite coordination, whereas each student NN adopts a compact per-satellite architecture for efficient onboard inference. The teacher learns a high-quality weighted minimum mean square error precoding policy from global-clean inputs, which is then distilled into the students operating on partial-noisy inputs. To mitigate the resulting teacher-student performance gap, we develop a hybrid KD mechanism with explicit angle- and phase-error calibration. Simulation results demonstrate that the proposed framework significantly enhances the decentralized sum-rate performance and remains robust under diverse configurations.

I. INTRODUCTION

Cooperative multi-satellite massive MIMO addresses single-satellite limitations by jointly using distributed antennas, but centralized and decentralized precoding face complementary practical constraints. The paper proposes knowledge distillation from a clean global teacher to lightweight students operating with partial, noisy information.

  • Motivation: Cooperative multi-satellite massive MIMO aggregates distributed antenna resources to improve beamforming, spatial multiplexing, and inter-user interference suppression.Large-scale satellite arrays provide beamforming gains against propagation loss, while coordination supplies additional spatial degrees of freedom.
  • Motivation: Centralized precoding relies on global sCSI and jointly optimized precoders, whereas decentralized schemes use noisy partial sCSI and limited ISL information.The centralized approach incurs signaling and computational overhead, while decentralized inference suffers from restricted cooperative knowledge.
  • Proposed framework: The proposed KD framework trains a large centralized teacher on clean global information to guide lightweight decentralized student NNs.The teacher learns a high-performance precoding policy, while students infer local precoders with lower onboard complexity.
  • Proposed framework: The hybrid KD mechanism predicts angle and phase residuals, reconstructs corrected steering vectors, and transfers steering-vector-level and precoder-level knowledge.This calibration targets the teacher-student gap caused by partial and angle-corrupted student observations.
  • System model: The system comprises LEO satellites and mobile user terminals sharing time-frequency resources in non-coherent cooperative multi-stream transmission.Non-coherent transmission avoids the stringent synchronization required by coherent cooperation while supporting spatial multiplexing.

B. Statistical Characteristics and Angle Error Model

The system uses slowly varying statistical channel information because rapid fading is difficult to track in mobile LEO links. Angle perturbations model practical sCSI errors arising from estimation, attitude, calibration, and feedback imperfections.

  • Statistical characteristics: sCSI is used instead of instantaneous CSI because LEO mobility and propagation and feedback delays hinder timely fast-fading tracking.The available statistics are mainly geometric LoS information and large-scale channel parameters.
  • Statistical characteristics: The channel is modeled as Rician fading with normalized LoS and Gaussian NLoS components.The LoS and NLoS components have unit expected norm and normalized covariance trace, respectively.
  • Statistical characteristics: Satellite-side sCSI includes angles, average channel power, the Rician factor, and normalized NLoS covariance information.Angles are obtained from ephemeris, attitude, and UT position, while other parameters can be estimated from repeated downlink pilots.
  • Angle error model: Clean global sCSI is difficult to realize in decentralized transmission, so angle perturbations represent errors from estimation, attitude jitter, calibration uncertainty, and feedback delay.The model defines true angles and error vectors, with angle components assumed independently and identically distributed.
  • Angle error model: The angle-error bound ε specifies the maximum absolute deviation of each angle component, with larger ε indicating less accurate sCSI.Perturbed angles are then used to recompute student-side steering vectors and covariance terms.

C. Closed-Form Structure for Multi-Satellite Transmission

The transmission objective is formulated through ergodic rates and weighted sum-rate maximization, then transformed into a WMMSE problem with closed-form updates. These updates expose a compact recovery structure for high-dimensional receive and precoding vectors.

  • Optimization formulation: The multi-satellite objective is weighted sum-rate maximization under satellite transmit-power budgets.User rate weights and satellite power budgets enter the optimization formulation.
  • Optimization formulation: The weighted sum-rate problem is optimized through its weighted minimum mean square error equivalent.The WMMSE formulation introduces error weights and mean-square errors for each satellite-user link.
  • Closed-form updates: The closed-form updates use block-coordinate optimization and Lagrange multipliers associated with satellite transmit-power constraints.The derivation follows WMMSE equivalence and KKT-based updates.
  • Neural recovery structure: The WMMSE-induced structure recovers high-dimensional beamforming and precoding vectors from compact variables, weights, multipliers, and sCSI.This structure supports a neural recovery framework that predicts compact variables rather than directly learning full precoders.

III. MULTI-SATELLITE COOPERATION: FROM TEACHER NN TO STUDENT NNS

The framework uses centralized offline training to transfer cooperative precoding knowledge to lightweight decentralized students. A knowledge-guided compact-variable target further reduces the representation burden of precoder learning.

  • Centralized training and decentralized execution: During offline ground-based training, a centralized teacher learns from global-clean sCSI and transfers cooperative precoding knowledge to students trained with partial-noisy inputs.The trained students are uploaded for independent onboard inference using local sCSI and limited ISL information.
  • Knowledge-guided precoder recovery: Instead of directly learning high-dimensional precoders, the network predicts compact variables and reconstructs precoders using the closed-form solution.The compact variables include {w_s,k, u_s,k}, {λ_s}, {ϱ_s,k′}, and {b_s,k}.
  • Knowledge-guided precoder recovery: The compact-variable reformulation reduces representation complexity and provides a more structured, interpretable learning target.This target supports transmission-knowledge transfer between centralized discovery and decentralized precoder recovery.

B. Large Centralized Teacher NN for MSMS

The centralized teacher constructs global satellite-user features, tokenizes each feature vector into patches, and processes them with a patch-wise Transformer encoder before subsequent modules.

  • Teacher input construction: The centralized teacher input combines local features with transmitted features from all cooperative satellites into Htea.The feature set includes satellite and user angles, average channel power, Rician factor, spatial covariance, and receiver noise power.
  • Patch tokenization: Each satellite-user feature vector is partitioned into length-E patch tokens, with zero-padding applied to the final patch when necessary.The resulting raw patch tensor is Ftea ∈ R^S×K×T×E.
  • Patch embedding: Patch tokens are linearly projected to hidden dimension D_h, combined with patch-index positional embeddings, and normalized.The embeddings are shared across satellite-user pairs.
  • Large Transformer Encoder: A Transformer encoder applies self-attention only across the patch dimension for each satellite-user pair, followed by feed-forward processing.Each layer contains multi-head attention and a feed-forward network.
  • Teacher training and inference: Algorithm 1 trains the centralized teacher with WSR loss and stores its learned variables for inference.During inference, the teacher produces the teacher precoding variables.
  • Large Transformer Encoder: Patch-wise mean pooling produces one vector per satellite-user pair for subsequent teacher-network modules.The pooled representation is Xtea ∈ R^S×K×D_h.

2) Dual-Axis Attention (DAA):

Dual-axis attention captures both intra-satellite user coupling and inter-satellite information for the same user through separable attention sweeps.

  • User-axis attention: Each DAA stage first applies user-axis attention across users within each satellite.This encodes intra-satellite multi-user coupling.
  • Satellite-axis attention: The satellite-axis sweep then aggregates embeddings of the same user across satellites.It fuses inter-satellite information while holding the user index fixed.
  • Stacked DAA stages: Stacking L_axis two-axis stages yields Htea for subsequent teacher-network modules.DAA denotes compositions of the user-axis and satellite-axis updates.

3) Tensor-Equivariant Neural (TEN) Network:

The TEN network transforms inputs into latent representations while preserving permutation equivariance across satellite and user dimensions.

  • Tensor-equivariant representation: TEN aggregates mean features across subsets of equivariant dimensions and combines them with original features through learnable linear transformations.This captures original and multidimensional global features.
  • TEN architecture: TEN uses an input linear layer followed by stacked blocks containing an MDE module, ReLU, layer normalization, and output linear layer.The architecture is parameterized by Lten stacked blocks.
  • Multidimensional equivariance: Each MDE layer is a linear combination of mean-and-repeat patterns defined on subsets of the satellite and user dimensions.The two equivariant dimensions are {S, K}.

4) Parameter Decoder (PD) and Precoder Recovery (PR):

The student architecture uses a lightweight, weight-shared per-satellite network, while parameter decoding and fixed precoder recovery convert its outputs into local closed-form precoders under partial and noisy observations.

  • Parameter Decoder and Precoder Recovery: The parameter-decoding module recovers precoder-related parameters from the teacher-side representation before the fixed precoder-recovery module reconstructs receive and precoding vectors.The decoder uses a lightweight MLP with layer normalization, two linear transformations, and GELU activation.
  • Lightweight decentralized student NNs: The decentralized student is instantiated with shared weights across satellites and maps local features plus limited peer information to each satellite’s precoder.Only satellite state information and user-terminal positions are exchanged among peer satellites.
  • Lightweight decentralized student NNs: Student inputs omit other satellites’ average channel power and Rician factors, while all angle information is corrupted by practical estimation and hardware-related errors.The design substitutes outer products of line-of-sight steering vectors for missing correlation matrices.
  • Lightweight decentralized student NNs: The student reduces the teacher’s deep Transformer backbone to a single-layer Transformer encoder for patch tokens to satisfy onboard latency and signaling constraints.This reduction accompanies per-satellite inference rather than centralized aggregation.

IV. KNOWLEDGE DISTILLATION FRAMEWORK DESIGN

The framework distills a centralized teacher’s global-clean, system-level precoding policy into decentralized students that operate on partial, angle-noisy observations and produce local precoders.

  • Teacher–student asymmetry: The teacher receives global features across satellites and user terminals with accurate angles, whereas each student receives only its own partial and angle-noisy observation.This creates an explicit global-clean teacher-view domain and partial-noisy student-view domain.
  • Decentralized student operation: The student training and inference procedure constructs local and other-satellite feature tensors, stacks them into student inputs, and executes the shared network independently for each satellite.The resulting inference uses partial noisy inputs to obtain local precoders.
  • Teacher–student asymmetry: The teacher learns a system-level mapping to jointly coordinated multi-satellite precoders, while each student learns a node-level mapping to its corresponding satellite’s local precoder.The two mappings differ in both observation scope and precoding decision scope.
  • Knowledge transfer: Knowledge transfer spans both the global-clean-to-partial-noisy observation domains and the centralized system-level-to-decentralized satellite-level precoding policies.The framework is designed to preserve global interference-management and multi-satellite coordination knowledge under decentralized inference.

B. Angle and Phase Correction

To reduce degradation from angle noise, the hybrid KD design predicts bounded angle and phase residuals, reconstructs corrected steering vectors, and uses them during student distillation.

  • Correction modules: The angle and phase residual prediction and Kronecker steering-vector correction modules are placed after the student’s TEN network.These modules explicitly provide angle- and phase-correction capabilities within the decentralized student.
  • Angle correction: An angle-error prediction head produces a four-dimensional residual vector, which is converted into bounded elevation and azimuth corrections.The maximum correction magnitudes are controlled separately for the elevation and azimuth angles.
  • Phase correction: Per-antenna phase correction with normalization is then applied to the reconstructed steering vectors.The phase residuals include separate satellite-side and user-terminal-side components with bounded correction magnitudes.
  • Steering-vector reconstruction: The corrected angles are used to reconstruct satellite-side and user-terminal-side steering vectors through a Kronecker product.The corrected angle vector is formed by subtracting the predicted residual from the noisy angle observation.
  • KD training: During KD training, the teacher remains frozen while the student is updated using the overall hybrid loss and later performs inference from partial noisy inputs.The training loop feeds the global teacher input to the teacher and each student input to its corresponding student network.

C. Training Objective and Distillation Loss Design

Student training combines task optimization with precoder- and steering-vector-level distillation, using instantaneous-CSI WSR evaluation and samples spanning multiple transmit power levels.

  • Hybrid training objective: The hybrid objective combines a task-driven WSR loss with KD losses that transfer structured knowledge at the precoder and steering-vector levels.The teacher remains fixed and supplies supervision targets during student training.
  • WSR optimization: The WSR loss is the negative sample-averaged weighted sum rate, so minimizing it is equivalent to maximizing system WSR.Achievable rates are evaluated using instantaneous CSI from the student-generated precoders.
  • Precoder distillation: The precoder-level KD loss compares teacher and student precoder matrices using cosine similarity for complex-valued vectors or matrices.The cosine formulation includes a positive numerical-stability constant.
  • Steering-vector distillation: Steering-vector KD aligns student steering vectors reconstructed from corrected angles and phase with teacher vectors generated from clean angle information.The losses target geometry-consistent satellite-side and user-terminal-side steering vectors.
  • Training data: The dataset contains 10,000 samples split into 7,000 training, 2,000 validation, and 1,000 testing samples across five randomly selected transmit power levels.The power levels range from −10 to 10 dBW in 5 dB increments.

V. NUMERICAL RESULTS

The numerical results compare optimization-based and learning-based schemes across performance, robustness, scalability, and inference complexity. The proposed Dec-Student-HKD improves decentralized transmission under noisy partial sCSI while maintaining low inference cost.

  • Performance Comparison: Dec-Student-HKD consistently outperforms Dec-Student and Dec-Student-PKD in multi-satellite sum rate, with gains becoming more noticeable at larger P sat.The comparison uses S = 3 and K = 24.
  • Performance Comparison: About 7% higher precoder cosine similarity is achieved by the proposed distillation scheme across transmit-power levels.The gain indicates better directional alignment between recovered and teacher precoders.
  • Computational and Inference Complexity: Learning-based schemes avoid iterative optimization, while decentralized students use a single-layer structure and achieve nearly identical inference complexity across student variants.The dominant student complexity remains approximately KM^3 despite Dec-Student-HKD's additional low-order K(M+N) term.
  • Computational and Inference Complexity: Cen-Teacher achieves a running time of 0.044 s, whereas lightweight decentralized students reduce running time to below 0.01 s.The measurements were performed on an Intel i9 185H CPU @ 2.30 GHz.
  • Robustness to Angle Error Distributions: Dec-Student-HKD remains robust across uniform and Gaussian angle-error distributions and substantially outperforms Cen-Teacher (noisy) under large angle perturbations.With decreasing angle error, its performance becomes comparable to, and occasionally slightly better than, Cen-Teacher (noisy).
  • Scalability: Without retraining, Dec-Student-HKD remains closer to Cen-Teacher and Cen-WMMSE than Dec-Student as the numbers of satellites and UTs vary.The learning-based schemes are trained with S = 3 and K = 24 before testing under different network sizes.
Loading 2608.29620v1…