Source-linked AI summary
Optimal Power Allocation and AI Receiver Design for Superimposed DMRS and Data Transmission
Sha Hu, Zhongwang Fu
TL;DR
Superimposed DMRS and data can increase spectral efficiency, but their entangled channel-estimation and detection errors make receiver design challenging. The paper develops analytical power and pilot-allocation methods and a Transformer-based iterative receiver, which achieve higher throughput and spectral efficiency than conventional non-overlapped DMRS baselines.
Problem
Entangled channel-estimation and MIMO-detection errors on shared resources make effective SI-DMRS design challenging.
Method
The paper derives ICED MSE analyses for power and DMRS allocation and designs a Transformer-encoder AI receiver with iterative CE and detection.
Results
The AI-ICED receiver with SI-DMRS delivers higher throughput and spectral efficiency than conventional 5G-NR baselines, with rapid CE-MSE and BLER convergence.
Takeaways & Limitations
An allocation of around 0.8 to DMRS and 0.2 to data provides an effective SI-DMRS design with zero DMRS overhead.
Abstract
from arXiv · showhide
In this paper, we consider transmissions with superimposed (SI) demodulation-reference-symbol (DMRS) and data in orthogonal frequency-division multiplexing (OFDM) based multiple-input multiple-output (MIMO) systems. First, we derive an analytical framework to characterize the iterative behavior between the mean-square errors (MSEs) of channel estimation (CE) and MIMO detection (MD) within an iterative CE and detection (ICED) process. This framework is subsequently utilized to optimize power allocation and pilot patterns between the DMRS and data symbols for SI-DMRS transmission. Second, we design an artificial intelligence (AI) based receiver built upon Transformer encoders for SI-DMRS transmissions, which incorporates an iterative CE and detection (ICED) structure. Simulation results demonstrate that the proposed AI-ICED receiver, combined with SI-DMRS, effectively increases spectral efficiency (SE) compared to conventional systems using non-overlapped DMRS and data symbols.
I. INTRODUCTION
The paper addresses the spectral-efficiency and channel-estimation challenges of superimposed DMRS in MIMO-OFDM by jointly optimizing allocations and designing a Transformer-based AI-ICED receiver. It derives iterative CE/MD MSE analyses, integrates physics-guided refinement, and reports rapid convergence, near-optimal detection, and SE gains over standard 5G-NR.
- Motivation: AI-based receivers unify channel estimation and MIMO detection to harvest joint processing gains beyond independently optimized receiver modules.The introduction motivates AI-native transceiver designs for MIMO-OFDM and highlights unified processing as a path to joint gains.
- Problem formulation: Superimposed DMRS overlaps pilots with data to eliminate pilot overhead, but entangles channel estimation and detection and can compromise CSI accuracy and symbol detection.The paper frames the central trade-off: more pilots improve CE but reduce SE, whereas fewer pilots or SI-DMRS increase SE while risking unreliable detection.
- Proposed receiver: The proposed AI-ICED receiver uses Transformer encoders in an unfolded iterative cascade to jointly refine CE and MD for SI-DMRS MIMO-OFDM transmissions.Its backbone incorporates VWN, TPA, and MoE layers, while a differentiable PGFC module forwards explicit physical features between iteration stages.
- Optimization framework: Theoretical analysis derives CE and MD MSEs within ICED, proves an across-iteration equilibrium, and enables optimal power and DMRS resource allocation for general settings.MIESM combines the two RE subsets for data detection with and without SI-DMRS to assess final decoding performance.
- Results: 2 to 3 iterations yield rapid convergence, while simulations show near-optimal detection with genie-CSI and substantial SE gains over standard 5G-NR.The receiver is benchmarked against optimal MLD and genie-CSI-aided systems under SI-DMRS schemes.
Notations:: · II. SYSTEM MODEL AND ICED PROCESS · A. Received Signal Model
The paper defines notation and an SI-DMRS MIMO-OFDM received-signal model, then parameterizes ICED analysis through resource allocation and power factors. SI-DMRS avoids pilot overhead but makes channel estimation and detection more challenging, motivating power redistribution toward SI-DMRS-bearing resource elements.
- Notations::: SI-DMRS superimposes additional DMRS symbols onto data without pilot overhead, while splitting total transmit power between DMRS and data complicates CE and MD.The power-allocation factor between SI-DMRS and data is b, while data-only power uses a, with 0 < a, b ≤1.
- II. SYSTEM MODEL AND ICED PROCESS: The system is a MIMO-OFDM link with Nt transmit and Nr receive antennas using a 2D time-frequency resource grid of Nsym OFDM symbols and Nsc subcarriers.Transmit symbols use an M-QAM constellation and are indexed by OFDM symbol and subcarrier.
- A. Received Signal Model: Without DMRS, each received vector follows the frequency-domain MIMO channel plus AWGN, with noise distributed as CN (0, σ2I).The channel and noise entries are modeled as i.i.d. complex Gaussian variables, with no spatial correlation assumed.
- A. Received Signal Model: Unit-power data vectors yield SNR = 1/σ2, and stacking all resource elements produces received tensor Y and channel tensor H for subsequent processing.The observation tensors span receive antennas, transmit antennas, OFDM symbols, and subcarriers as specified by their dimensions.
- II. SYSTEM MODEL AND ICED PROCESS: Within a subframe, N resource elements carry data without DMRS and M carry SI-DMRS plus data, while total transmit power is normalized to N+M.These parameters establish the framework for analyzing iterative CE and MD behavior in ICED.
- A. Received Signal Model: Data-only resource elements use power factor a, SI-DMRS symbols use bk, and data superimposed with SI-DMRS uses (1−b)k.The allocation preserves total power through aN +kM = N +M.
- A. Received Signal Model: The parameter a enables borrowing power from data-only symbols to increase SI-DMRS-bearing resource-element SNR and potentially improve overall detection performance.This redistribution is introduced to analyze whether stronger SI-DMRS observations benefit detection.
B. CE with SI-DMRS · C. CE De-noising Among Pilots · D. MD with SI-DMRS
The paper models iterative channel estimation and MIMO detection for SI-DMRS, including pilot-domain de-noising and residual interference. It establishes a stable ICED equilibrium whose minimum CE-MSE is obtained from a third-order polynomial.
- B. CE with SI-DMRS: SI-DMRS channel estimation removes the estimated data component and tracks CE error and detection error across ICED iterations.The framework defines CE-MSE and MD-MSE from the channel and detection error quantities.
- B. CE with SI-DMRS: The next CE-MSE is computed from the SI-DMRS received-signal model using the previous iteration’s estimation and detection performance.This update forms the CE side of the iterative CE–detection analysis.
- C. CE De-noising Among Pilots: An LMMSE filter de-noises CE across multiple REs carrying DMRS, using time-frequency channel correlation for each Tx–Rx pair.The refined estimate uses CE over M/Nt REs for each link.
- C. CE De-noising Among Pilots: The effective noise power includes both thermal noise and residual data interference, enabling maximum coherent processing subject to correlation-dependent loss.The loss is represented by γ ≥1 relative to the maximum coherent gain.
- C. CE De-noising Among Pilots: Direct DMRS superimposition across transmitters produced no noticeable performance gains because it introduced spatial interference and reduced effective DMRS power per transmitter.The evaluated configuration used non-overlapped DMRS for different transmitters on separate REs.
- D. MD with SI-DMRS: After de-noising CE and removing the DMRS component, the receiver forms the SI-DMRS data-detection model and approximates MD-MSE analytically.The resulting CE-MSE and MD-MSE equations define the ICED equilibrium state.
- D. MD with SI-DMRS: A stable equilibrium exists for CE-MSE and MD-MSE, and the minimum CE-MSE solution p, with 0≤p≤1, is obtained as a root of a third-order polynomial.Property 1 formalizes the stability and polynomial-root characterization.
- D. MD with SI-DMRS: Property 1 determines converged CE-MSE and MD-MSE directly from equilibrium equations for general configurations of (N, M, a, b, Nt, γ, σ2).Overall subframe decoding analysis additionally requires MD-MSE on REs beyond those captured by the SI-DMRS equilibrium.
E. MD on REs without SI-DMRS … B. Optimal Power Allocations
The paper derives MD for REs without SI-DMRS after CE and MD convergence, then uses analytical MIESM-based evaluation to optimize SI-DMRS power allocation. Jointly optimizing the RE power factors improves effective MD-MSE, while the optimal DMRS allocation rises with SNR because CE becomes the bottleneck.
- E. MD on REs without SI-DMRS: After CE and MD iterations converge on REs with SI-DMRS, the accurate channel estimate is used for data detection on remaining REs without SI-DMRS.The paper separately models detection on REs without SI-DMRS using the resulting channel estimate.
- III. OPTIMAL POWER ALLOCATIONS FOR SI-DMRS TRANSMISSIONS: The optimization seeks power factors (a, b) that maximize effective SNR θeff after computing the equilibrium CE-MSE and MD-MSE.Equilibrium (p, q) and t values are obtained for different allocation strategies, followed by θeff computation from (q, t).
- B. Optimal Power Allocations: For a = 1, an optimal DMRS allocation factor b minimizes q, while CE-MSE decreases continuously as b increases.The examples use no power boost on REs carrying exclusively data symbols.
- B. Optimal Power Allocations: As SNR increases, the optimal b increases because CE-MSE becomes the bottleneck affecting MD-MSE.This identifies the reason for allocating more power to DMRS at higher SNR.
- B. Optimal Power Allocations: MIESM-derived effective MD-MSE is minimized at a higher b than when q is considered alone.The effective metric better reflects overall decoding performance, using β =1 and σ2=−15dB in the cited example.
- B. Optimal Power Allocations: Data symbols on REs without SI-DMRS require higher CE accuracy to achieve improved error performance.This conclusion follows from examining the MD-MSE curve for t.
- B. Optimal Power Allocations: −10.42dB to −12.2dB: effective MD-MSE decreases when a is reduced from 1 to 0.67 through joint optimization of a and b.The result demonstrates performance benefits from optimizing power allocation across REs.
C. Is Superimposing More DMRS with Data Beneficial? · IV. THE PROPOSED TRANSFORMER BASED AI-ICED RECEIVER DESIGN · A. Tokenization
The analysis finds that adding more superimposed DMRS is not universally beneficial, because performance depends on de-noising gains and system parameters. The proposed Transformer-based AI-ICED receiver exploits SI-DMRS through contextual DMRS inputs, iterative processing, and physics-aligned tokenization.
- C. Is Superimposing More DMRS with Data Beneficial?: Increasing M from 24 to 96 at constant γ = 2 yields gains, but the improvement is limited even when de-noising gain increases linearly with M.The evaluated configurations use M = 24, 48, and 96 and γ = 2, 4, and 6.
- C. Is Superimposing More DMRS with Data Beneficial?: 1.55dB separates M = 24, γ = 2 from M = 48, γ = 4 when γ/M is constant, showing that more superimposed DMRS is not always beneficial.The optimal choice depends heavily on de-noising gains and operational system parameters.
- C. Is Superimposing More DMRS with Data Beneficial?: The analytical framework provides a perspective for analyzing general cases and optimizing SI-DMRS configurations.Increasing the number of superimposed DMRS can also increase PAPR in MIMO-OFDM systems.
- IV. THE PROPOSED TRANSFORMER BASED AI-ICED RECEIVER DESIGN: The proposed Transformer-based AI-ICED receiver harnesses SE gains from SI-DMRS and processes received signals, DMRS symbols, and DMRS coordinate masks as tensor inputs.DMRS information acts as contextual prompts that map observations directly to symbols.
- IV. THE PROPOSED TRANSFORMER BASED AI-ICED RECEIVER DESIGN: The receiver uses a cascaded ICED loop that updates channel estimation and symbol probabilities, converts logits into bit LLRs, and feeds them to the channel decoder.Each stage includes tokenization, an ICED module, and a physics-guided feature construction module between iterations.
- A. Tokenization: Tokenization spatially flattens and concatenates physical tensors, separates real and imaginary components, and linearly projects them into a unified dvirtual-dimensional token sequence.The inputs are the observation tensor Y R, symbol tensor sR, and binary SI-DMRS mask Ωp.
- A. Tokenization: At i = 0, tokenization initializes from Y R with zero soft symbols, while later PGFC outputs expand the feature dimension before projection across L REs.The tokenizer outputs X(i)token ∈ R^B×L×dvirtual for the next ICED module.
- A. Tokenization: Flattening 2D time-frequency grids into length-L sequences enables RoPE-based attention to capture local channel coherence, while latent features let attention perform MD internally.Spatial antennas are collapsed into one dvirtual-dimensional feature vector per token, leaving spatial-correlation resolution to the network.
B. ICED Backbone
The AI-ICED receiver uses an unfolded cascaded architecture that fuses channel estimation and MIMO detection into one iterative Transformer-based pipeline. Its encoder blocks process channel features and symbol detection using VWN, TPA, and MoE designs.
- Architecture: AI-ICED unifies channel estimation and MIMO detection in a fused, unfolded cascaded pipeline rather than treating them as disjoint modules.The architecture operates on a tokenized sequence and is illustrated as an integrated CE–MD process.
- Channel estimation: Customized Transformer encoder blocks iteratively resolve time-frequency correlations and spatial interference to produce a channel latent representation.The channel-estimation branch uses NCE blocks incorporating VWN, TPA, and MoE techniques.
- Channel estimation: A projection head reconstructs the MIMO channel by mapping latent features back to physical antenna dimensions and reshaping them into a real-valued CE tensor.The reconstructed channel estimate is subsequently connected to the MIMO-detection input through an additive residual bridge.
- MIMO detection: NMD Transformer encoder blocks apply the same VWN-TPA-MOE design to detect symbols, after which a classification head produces real and imaginary PAM logits.Standard softmax normalization maps the logits to a probability tensor for loss computation.
C. Physics Guided Feature Construction (PGFC) · D. Training Loss Design · E. Inference Complexity
The receiver uses physics-guided features to connect iterative channel estimation and detection, jointly trains both objectives with deep supervision, and controls inference cost through stage-wise complexity reduction strategies.
- C. Physics Guided Feature Construction (PGFC): PGFC forms a differentiable bridge that converts current channel and symbol estimates into structured communication-theoretic features for subsequent VWN-TPA encoding.This reduces reliance on learning fundamental spatial projections solely from data and supports convergence and generalization across channel conditions.
- C. Physics Guided Feature Construction (PGFC): The augmented tokenizer input combines updated soft states and pilot-mask information with assembled PGFC tensors for the next ICED stage.The feature construction includes probability-based symbol estimates and matched-filter outputs after conversion to real-valued representations.
- C. Physics Guided Feature Construction (PGFC): During training, the channel projection head and re-embedding estimate H and measure CE-MSE, while the two modules can be combined during inference.A stop-gradient operation on the updated soft state isolates backward propagation within each iterative stage and stabilizes deep-supervision dynamics.
- D. Training Loss Design: The stage-wise training loss balances channel estimation and MIMO detection using hyper-parameter β1 ∈[0, 1].Detection uses cross-entropy, whereas channel-estimation loss uses MSE.
- D. Training Loss Design: Deep supervision injects gradients into all intermediate stages, with total-loss weighting controlled by β2 ∈[0, 1] across Nit+1 stages.This design targets vanishing gradients in the unfolded cascade.
- E. Inference Complexity: Inference complexity is dominated by attention-score computation, low-rank TPA projections, and MoE operations, while total cost scales linearly with the applied Nit +1 stages.The architecture addresses this challenge through VWN, TPA, and MoE design strategies.
- E. Inference Complexity: VWN separates representational capacity from active width, TPA reduces dense projection complexity with low-rank factorization, and MoE preserves dense-FFN FLOP equivalence through sparse expert activation.TPA retains the O(L2dmodel) attention-score term for global time-frequency dependencies.
- E. Inference Complexity: MoE activates exactly Ns shared experts and Kr routed experts per token, maintaining computational efficiency while expanding parameter capacity through expert specialization.The resulting complexity expressions are summarized in Table I in Appendix I.
V. NUMERICAL RESULTS · A. CE-MSE and BLER Convergences of the Proposed AI-ICED Receiver · B. Throughput Increments
Under SI-DMRS, the AI-ICED receiver converges rapidly and approaches genie-aided MLD performance, while incurring a 2–4 dB CE-MSE degradation versus 5G-NR. Its zero-DMRS-overhead advantage nevertheless yields consistently higher throughput than the 5G-NR baseline across tested configurations.
- V. NUMERICAL RESULTS: SI-DMRS experiments use two PRBs and two OFDM symbols, giving M =48, N =240, and a maximal SE gain of 20%.The power allocation factor is b=0.8, with no power boosting for comparison against 5G-NR.
- A. CE-MSE and BLER Convergences of the Proposed AI-ICED Receiver: Both CE-MSE and BLER converge within a single iteration for 4×4 MIMO, 16-QAM, and LDPC code-rate 2/3.The receiver is trained at 16dB SNR and evaluated across a wide SNR range.
- A. CE-MSE and BLER Convergences of the Proposed AI-ICED Receiver: Approximately 0.5dB in SNR is gained for both CE-MSE and BLER after two additional iterations, and performance approaches optimal MLD with genie-aided CSI.This result demonstrates the effectiveness of the iterative AI-ICED design.
- A. CE-MSE and BLER Convergences of the Proposed AI-ICED Receiver: 2–4 dB of CE-MSE degradation occurs under SI-DMRS relative to the 5G-NR baseline because DMRS and data interfere.The error remains insignificant compared with noise power, and dedicated models at regular SNR intervals address train-test mismatch.
- B. Throughput Increments: SI-DMRS with AI-ICED consistently outperforms the 5G-NR baseline, whose throughput saturates at high SNR because of DMRS overhead.With genie CSI, AI-ICED nearly matches genie-CSI-aided MLD and outperforms genie-CSI-aided LMMSE at middle-to-high SNR.
VI. SUMMARY … B. Derivations of MD-MSE
The paper summarizes an AI-ICED receiver for SI-DMRS MIMO-OFDM transmission and derives analytical CE-MSE and MD-MSE frameworks for iterative design. The analysis supports power allocation and performance evaluation against conventional baselines.
- VI. SUMMARY: The Transformer encoder based AI-ICED receiver integrates VWN, TPA, and MoE architectures to address DMRS overheads in MIMO-OFDM systems.It reformulates channel estimation and MIMO detection as joint contextual learning using a PGFC bridge and iterative unfolded cascade.
- VI. SUMMARY: CE-MSE and BLER converge rapidly within a few iterations, while detection performance approaches optimal MLD with genie CSI input.The receiver is described as resilient to pilot sparsity and capable of significant performance gains.
- VI. SUMMARY: SI-DMRS with the AI-ICED receiver delivers higher throughput and SE than conventional 5G-NR baselines using non-superimposed DMRS and data.These gains are validated through throughput envelopes.
- VI. SUMMARY: The analytical framework derives approximated closed forms for CE-MSE and MD-MSE and evaluates final ICED equilibrium states without numerical simulations.The framework addresses the entanglement between channel estimation and MIMO detection in SI-DMRS power allocation.
- VI. SUMMARY: Allocating a power factor of around 0.8 to DMRS and 0.2 to data provides an effective SI-DMRS design.The analysis also considers boosting SI-DMRS resource elements by borrowing power from remaining resource elements.
- VI. SUMMARY: MIESM combines MD-MSEs from resource elements containing SI-DMRS and those without it to evaluate final decoding performance.This provides a basis for combining the two resource-element groups in performance analysis.
- B. Derivations of MD-MSE: With an LMMSE estimator, the derivation gives the MD-MSE and distinguishes instantaneous channel-dependent MD-MSE from its ergodic counterpart.The ergodic MD-MSE averages over the channel realization using Jensen’s inequality and is bounded for iterative analysis.
- B. Derivations of MD-MSE: The general iterative analysis approximates E{∆x∆x†} by its bound.This approximation is used to analyze iterative behavior under the derived MD-MSE bound.
C. Proof of Property 1 … G. TPA
The paper derives the CE-MSE/MD-MSE equilibrium and MD-MSE approximation, then specifies MIESM mapping and introduces VWN with TPA to efficiently model complex MIMO-OFDM channels. VWN separates memory capacity from computational width, while TPA factorizes attention tensors and applies physical 2D positional adaptation.
- C. Proof of Property 1: The CE-MSE and MD-MSE equilibrium is characterized by equations that, after substitutions, produce a cubic polynomial with coefficients defined in Property 1.The derivation substitutes u = 1 − b, v = σ²_k, and ρ = γN_t before collecting powers of p.
- D. Derivations of MD-MSE on REs without DMRS: For REs without DMRS, the MD-MSE is derived using an LMMSE estimator and approximated following the argumentation in Appendix B.The approximation is explicitly based on the same reasoning used in Appendix B.
- E. MIESM: MIESM approximates the constrained mutual-information mapping for modulation order M using a widely used exponential fit with K = 3.The parameters for this exponential fit are listed in Table III.
- F. VWN: VWN separates memory capacity from computational width by expanding each input token to dvirtual = (n/m) · dmodel and reshaping it into n dimensions.The expansion uses n > m, while heavy nonlinear operations are restricted to an m-slot computational subspace.
- F. VWN: Within VWN, complex multi-path residuals traverse an n-dimensional highway while expensive dense computations remain isolated in the efficient m-dimensional subspace.Dynamic generalized hyper-connections manage the interaction between these pathways.
- G. TPA: TPA improves parameter efficiency and injects physical priors by factorizing queries, keys, and values into sums of contextual tensor products.This factorization replaces the monolithic parameterization associated with standard multi-head attention.
- G. TPA: TPA adapts attention to MIMO-OFDM by treating the sequence as a 2D physical grid, applying RoPE2D to feature bases, and reconstructing Q, K, and V through tensor contraction.The factorized tensors are then used for scaled dot-product attention before routing to the subsequent MoE block.
H. MoE … HYPERPARAMETER CONFIGURATION OF THE PROPOSED AI-ICED RECEIVER
The proposed AI-ICED receiver uses a DeepSeek-adapted MoE block with shared and selectively routed SwiGLU experts. Its configuration specifies VWN, TPA, MoE, optimization, and learning-rate hyperparameters, with Table IV providing the modulation configuration.
- H. MoE: The MoE block combines Ks always-active shared experts with Kr selectively routed experts, using SwiGLU-based feed-forward structures.This block follows the DeepSeek-MoE framework.
- H. MoE: A lightweight gating network selects the top-Kr routed experts using Sigmoid scores from a linear projection.The selected entries are normalized into final routing weights.
- I. NN Configuration of the AI-ICED Receiver: The VWN chassis uses m=2 memory slots and n=3 active computational slots per layer, yielding dvirtual = 768.The expanded virtual dimension is computed as (n/m) · dmodel = 768.
- I. NN Configuration of the AI-ICED Receiver: The TPA sub-layer uses query and key/value ranks rq = 16 and rk = 16.Both attention parameter ranks are set to 16.
- I. NN Configuration of the AI-ICED Receiver: The MoE sub-layer uses Ns = 1 shared expert, Nr = 4 routed experts, and Kr = 2 active experts per token.The per-expert hidden dimension is scaled by (8/3)/(Kr +Ns) to maintain equivalence with a dense SwiGLU FFN.
- I. NN Configuration of the AI-ICED Receiver: Training uses Muon for multi-dimensional weight matrices and AdamW for one-dimensional parameters such as layer normalizations and biases.This hybrid strategy is intended to optimize the deep cascaded Transformer backbone.
- Modulation: The learning rate warms up linearly for the first 10% of steps, then follows cosine annealing to a minimum ratio of 0.01 of the peak value.The modulation configuration is presented in Table IV.