Paper deep dive
Optimal Power Allocation and AI Receiver Design for Superimposed DMRS and Data Transmission
Sha Hu, Zhongwang Fu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/17/2026, 4:24:32 AM
Summary
This paper proposes an optimal power allocation framework and an AI-based receiver design for superimposed demodulation-reference-symbol (SI-DMRS) and data transmission in MIMO-OFDM systems. The authors derive an analytical framework to characterize the iterative behavior between channel estimation (CE) and MIMO detection (MD) mean-square errors (MSEs) within an iterative CE and detection (ICED) process. This framework is used to optimize power allocation and pilot patterns. Additionally, an AI receiver based on Transformer encoders is designed to perform joint CE and MD, demonstrating improved spectral efficiency compared to conventional non-overlapped DMRS systems.
Entities (8)
Relation Signals (5)
AI-ICED Receiver → uses → Transformer Encoder
confidence 97% · design an artificial intelligence (AI) based receiver built upon Transformer encoders for SI-DMRS transmissions
AI-ICED Receiver → performs → ICED
confidence 96% · AI-ICED receiver... incorporates an iterative CE and detection (ICED) structure
Analytical Framework → optimizes → Power Allocation
confidence 94% · This framework is subsequently utilized to optimize power allocation and pilot patterns between the DMRS and data symbols
SI-DMRS → increases → Spectral Efficiency
confidence 93% · proposed AI-ICED receiver, combined with SI-DMRS, effectively increases spectral efficiency (SE)
MIESM → combines → MD-MSE
confidence 92% · incorporated mutual information effective signal-to-noise ratio mapping (MIESM) to effectively combine the two RE subsets for data detections
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In this paper, we consider transmissions with superimposed (SI) demodulation-reference-symbol (DMRS) and data in orthogonal frequency-division multiplexing (OFDM) based multiple-input multiple-output (MIMO) systems. First, we derive an analytical framework to characterize the iterative behavior between the mean-square errors (MSEs) of channel estimation (CE) and MIMO detection (MD) within an iterative CE and detection (ICED) process. This framework is subsequently utilized to optimize power allocation and pilot patterns between the DMRS and data symbols for SI-DMRS transmission. Second, we design an artificial intelligence (AI) based receiver built upon Transformer encoders for SI-DMRS transmissions, which incorporates an iterative CE and detection (ICED) structure. Simulation results demonstrate that the proposed AI-ICED receiver, combined with SI-DMRS, effectively increases spectral efficiency (SE) compared to conventional systems using non-overlapped DMRS and data symbols.
Tags
Links
- Source: https://arxiv.org/abs/2608.13809v1
- Canonical: https://arxiv.org/abs/2608.13809v1
Trouble viewing inline? Open PDF directly →
Full Text
75,456 characters extracted from source content.
Expand or collapse full text
Optimal Power Allocation and AI Receiver Design for Superimposed DMRS and Data TransmissionThanks: The authors work with Nyquist Research Center, Huawei Technologies Sweden AB, Sweden, and Zhongwang Fu was an intern during the time of this work [1]. Email: hu.sha, zhongwang.fu@huawei.com. Sha Hu Zhongwang Fu Affiliation: Abstract In this paper, we consider transmissions with superimposed (SI) demodulation-reference-symbol (DMRS) and data in orthogonal frequency-division multiplexing (OFDM) based multiple-input multiple-output (MIMO) systems. First, we derive an analytical framework to characterize the iterative behaviour between the mean-square errors (MSEs) of channel estimation (CE) and MIMO detection (MD) within an iterative CE and detection (ICED) process. This framework is subsequently utilized to optimize power allocation and pilot patterns between the DMRS and data symbols for SI-DMRS transmission. Second, we design an artificial intelligence (AI) based receiver built upon Transformer encoders for SI-DMRS transmissions, which incorporates an iterative CE and detection (ICED) structure. Simulation results demonstrate that the proposed AI-ICED receiver, combined with SI-DMRS, effectively increases spectral efficiency (SE) compared to conventional systems using non-overlapped DMRS and data symbols. I Introduction Driven by evolutions towards the sixth-generation radio (6GR) networks, artificial intelligence (AI) and deep neural-network (N) driven paradigm shifts have became a broad consensus [2, 3, 5, 7, 6, 4]. Future wireless networks are steadily transitioning towards AI-native designs, including AI transceiver designs for multiple-input multiple-output (MIMO) and orthogonal frequency-division multiplexing (OFDM) systems. Unlike a conventional receiver design that relies on a cascade of independently optimized modules, AI-based receiver can unify multiple core functionalities such as channel estimation (CE) and MIMO detection (MD) [9, 8, 10, 12, 13, 16, 11, 14, 15] to harvest joint processing gains. Although classical algorithms such as linear minimum mean square error (LMMSE) estimator and quasi-maximum likelihood detector (MLD) [17, 18] demonstrate good performance, exploring AI design is still of interest. Firstly, the inference process naturally aligns with contemporary parallel computing hardware, offering a path to reduce computational latency. Secondly, a unified N architecture can mitigate information barriers among different modules, achieving joint processing gains through latent feature exchanges, thereby demonstrating a potential to surpass conventional baselines in challenge scenarios. A direction that AI could be beneficial in 6GR system is to optimize the overheads of pilots to increase spectral efficiency (SE). A classical dilemma is that to attain a good CE accuracy, more pilots should be transmitted but then the SE is decreased, and vice-visa, with fewer pilots the CE is not good enough and may degrade the data detection performance. In current fifth-generation new-radio (5G-NR) system, demodulation reference signal (DMRS) as pilots are transmitted on stand-alone resource elements (RE) to avoid interfering data transmissions. To increase SE, superimposed DMRS (SI-DMRS) has been proposed [20, 19, 21] which directly transmits DMRS overlapped with data symbols. This yields zero overheads from pilots, but the challenge is to attain an accurate CE from SI-DMRS and data11 1 Note that in our framework introduced later, the current DMRS patterns in 5G-NR can be seen as special cases of SI-DMRS by setting the power allocation of data symbols to zero on REs with SI-DMRS.. While reducing pilot overheads or adopting superimposed DMRS (SI-DMRS) directly increases SE, it degrades channel state information (CSI) accuracy, which compromises symbol detection. Although attention-based N architectures have been proposed to enhance CE under limited pilots [10], questions of how to optimize the balance between DMRS overheads and data detection, and design an AI receiver that is capable of reliable data recovery with a minimal pilot overheads, remain inadequately addressed. Although conventional iterative CE and detection (ICED) [22, 23] can be applied, the two tasks of CE and MD on REs with SI-DMRS are entangled with each other, and it is challenging to achieve satisfying performances for both. However, this type of problem is perfectly for AI to attack, which learns from training dataset to approach the performance bound of joint CE and MD that is infeasible with conventional methods. On the other hand, with the rapid iteration of large foundation models centred on Transformers (e.g., Qwen, DeepSeek, Kimi [25, 24, 26]), in-context learning (ICL) has attracted widespread attention [27, 29, 28]. ICL essentially provides an unprecedented few-shot adaptation paradigm. Mapping this to communications systems, a few pilot signals and their RE-locations can serve as contextual information for interpreting the channel state information (CSI) and data-detecting of the received MIMO-OFDM grids. Motivated by challenges with SI-DMRS and inspirations from ICL advances, we propose an AI-ICED receiver architecture for resolving CE and MD with SI-DMRS transmissions in MIMO-OFDM systems, which leverages advanced Transformer encoders to provide a unified design for diverse transmission schemes that achieves near-optimal performance. The architecture adopts an unfolded iterative cascade with a Transformer-encoder based backbone that incorporates Virtual Width Networks (VWN) [30], Tensor Product Attention (TPA) [31], and Mixture-of-Experts (MoE) [32] layers. Between iteration stages, a differentiable physics-guided feature construction (PGFC) module converts intermediate predictions into explicit physical features and are forwarded to subsequent stages. The main contributions of this paper are summarized as follows: • Theoretical analysis of power and DMRS allocations: We have derived the MSEs for CE and MD within the ICED structure, proving the existence of an equilibrium state across iterations. Consequently, determining optimal power and DMRS resource allocations for general settings becomes straightforward. Furthermore, we incorporated mutual information effective signal-to-noise ratio mapping (MIESM) [42] to effectively combine the two RE subsets for data detections, with and without superimposed DMRS, to evaluate final decoding performance, thereby enhancing the practical applicability of the proposed optimizations. • Transformer unified AI-ICED receiver design: We have designed an AI-ICED receiver using Transformer encoder blocks as its core backbone. The AI receiver processes the received signal, DMRS symbols, and a pilot mask to output symbol likelihoods at each stage. In the final stage, bit log-likelihood ratios (LLRs) are computed and passed to the decoder to evaluate block error rate (BLER) performance. The key features of the proposed AI-ICED receiver are twofold: joint CE and MD at each iteration stage, and a physics-guided feature construction (PGFC) mechanism between iterations that enables progressive and joint refinements. • Thorough simulation results and benchmarks: We have evaluated the proposed AI-ICED receiver against optimal MLD and genie-CSI-aided benchmarks. The results quantify the effectiveness of AI receiver under superimposed DMRS (SI-DMRS) schemes, demonstrating near-optimal MD performance with genie-CSI. Furthermore, the proposed receiver converges rapidly within 2 to 3 iterations and achieves substantial SE gains by mitigating DMRS overhead compared to a standard 5G-NR baseline. The remainder of this paper is organized as follows. Sec. I introduces the MIMO-OFDM system model with SI-DMRS, and a theoretical framework for analysing the MSEs of CE and MD in conventional ICED process is developed. Sec. I investigated the optimal power allocation of SI-DMRS based on the analytical framework established in Sec. I. Sec. IV presents the proposed AI-ICED receiver design in detail. Sec. V provides the simulation results, and Sec VI concludes the paper. Notations: The Hermitian of a matrix A is denoted as † A , and the identity matrix is denoted as I. The expectation operator is ⋅E\·\. The variables B, NrN_r, NtN_t, NsymN_sym, and NscN_sc denote the batch size, number of receive antennas, number of transmit antennas, number of OFDM symbols within one subframe that carries data (excludes OFDM symbols carrying control channel), and number of subcarriers for the considered bandwidth, respectively. The variables N and M denotes the number of REs that only carries data symbols, and the number of REs carrying both SI-DMRS and data, respectively. Furthermore, L=NsymNsc=N+ML\!=\!N_symN_sc\!=\!N\!+\!M denote the input sequence length for AI receiver for a CE and MD occurrence. The power allcoation factor between SI-DMRS and data is denoted as b, and the power factor for data without SI-DMRS is set to a, both <a,b≤10\!<\!a,b\!≤\!1. Furthermore, dmodeld_model is model-size in Transformer encoder blocks, and dffd_f is (feed-forward network) FFN dimension. Fig. 1: Example patterns of DMRS and data transmission in standard 5G-NR (top) versus with SI-DMRS (bottom). The bandwidth is 2 physical resource block (2PRB) with Nsc=24N_sc\!=\!24, and two OFDM symbols are used for control channel such that Nsym=12N_sym\!=\!12 for the SI-DMRS case. In both cases, the transmit power per RE is normalized to unity. Under the SI-DMRS scheme, superimposing additional DMRS symbols onto data introduces no pilot overhead, however, because total transmit power is split between the DMRS and data components, both CE and MD become more challenging I System Model and ICED Process I-A Received Signal Model We consider a MIMO-OFDM system consisting of NtN_t transmit-antennas (Tx) and NrN_r receiving-antennas (Rx). The transmitted signals are structured into a two-dimensional (2D) time-frequency resource grid, where transmission time interval (TTI) comprises NsymN_sym OFDM symbols in the time domain and NscN_sc subcarriers in the frequency domain [33]. Let n,k x_n,k denote the transmit symbol vector mapped to the nnth OFDM symbol and the kkth subcarrier, with its entries drawn from an M-ary quadrature amplitude modulation (M-QAM) constellation. The corresponding received signal vector n,k∈ℂNr×1 y_n,k\!∈\!C^N_r× 1 without DMRS is n,k=1Ntn,kn,k+n,k, y_n,k= 1N_t H_n,k x_n,k+ w_n,k, (1) where n,k∈ℂNr×Nt H_n,k\!∈\!C^N_r\!×\!N_t is the complex-valued frequency-domain MIMO channel matrix, and n,k∈ℂNr×1 w_n,k\!∈\!C^N_r\!×\!1 represents the additive white Gaussian noise (AWGN) at the receiver. The entries of n,k w_n,k are independent and identically distributed (i.i.d.) complex Gaussian random variables distributed as n,k∼(,σ2) w_n,k\! \!CN( 0,σ^2 I), where σ2σ^2 denotes noise denotes the noise variance per receive antenna. For simplicity, we assume no spatial correlation, meaning the entries of n,k H_n,k are also i.i.d. complex Gaussian random variables satisfying n,kn,k†=NtE\ H_n,k H_n,k \\!=\!N_t I. Furthermore, the data vectors are transmitted with unit-power such that n,kn,k†=E\ x_n,k x_n,k \\!=\! I, and the received signal-to-noise ratio (SNR) is defined as SNR=1/σ2SNR\!=\!1/σ^2. By stacking all REs within one subframe, the received signal tensor ∈ℂNr×Nsym×Nsc Y\!∈\!C^N_r\!×\!N_sym\!×\!N_sc and the corresponding channel tensor ∈ℂNr×Nt×Nsym×Nsc H\!∈\!C^N_r\!×\!N_t\!×\!N_sym\!×\!N_sc constitute the complete observation space for the subsequent processing tasks. To establish a framework for the iterative behaviours in ICED, we define some parameters as follows. Within one subframe, N REs are allocated for data transmission without DMRS, and M REs are allocated for SI-DMRS and data transmission. The total transmit power within one subframe is normalized to N+MN\!+\!M. Further, denote μ=NMμ\!=\! NM and k=+μ(1−a)k\!=\!1\!+\!μ(1-a). The power allocation is structured as follows: • Data symbol vectors transmitted on REs without DMRS are with a power factor a. • SI-DMRS symbols are transmitted with power bkbk. • Data symbol vectors with SI-DMRS are transmitted with power (−b)k(1\!-\!b)k. Accordingly, the total power is fixed since aN+kM=N+MaN\!+\!kM\!=N\!+\!M. The purpose of introduce the parameter a is to analyse the potential benefits of borrowing powers from data symbols to boost the SNRSNR on REs with SI-DMRS, thereby the overall detection performance can be improved. I-B CE with SI-DMRS On REs containing SI-DMRS, after removing the estimated data component, the received signal model for estimating the channel vector h for each Tx reads 1=bksDMRS+(1−b)kNt(−~~)+, y_1= bk hs_DMRS+ (1-b)kN_t( H x- H x)+ w, (2) where ~ H and ~ x are the estimated MIMO channel and data symbols, respectively, and s is the transmitted DMRS symbol for a given Tx22 2 We consider the case where DMRS for different Tx are non-overlapped and transmitted on separate REs. We also evaluated direct DMRS superimposition across Tx on the same RE (without orthogonal cover codes (OCC) as in 5G-NR), but found no noticeable performance gains, as it introduces spatial interference and reduces the effective DMRS power per Tx.. Denote the CE error matrix as =−~ \!H\!=\! H\!-\! H, with h denoting the column corresponding to a single transmit antenna. Likewise, define the detection error vector as =−~ \!x\!=\! x\!-\! x. Assuming the CE-MSE and MD-MSE from the previous iteration are respectively defined as ()† \ \!h( \!h) \\! = \!\!=\!\! p, 0≥p≤1, \!p I,\;\;0≥ p≤ 1, (3) ()† \ \!x( \!x) \\! = \!\!=\!\! q, 0≥d<≤1. \!q I,\;\;0≥ d<≤ 1. (4) In the next CE step in an ICED process, the update CE-MSE based on (2) can be computed as (see Appendix-A) ()†=k(1−b)(p+q−pq)+σ2kb+k(1−b)(p+q−pq)+σ2. \ \!h( \!h) \= k(1-b)(p+q-pq)+σ^2kb+k(1-b)(p+q-pq)+σ^2 I. (5) I-C CE De-noising Among Pilots Note that the CE-MSE in (5) is evaluated for a single RE, and typically an LMMSE filter is subsequently applied to de-noise the to the CE over multiple REs across multiple REs carrying DMRS. In this case, the refined channel estimate is given by LMMSE=h(h+σ~2)−1¯, h_LMMSE= R_h( R_h+ σ^2 I)^-1 h, (6) where h R_his the time-frequency channel correlation matrix for each Tx–Rx pair, the vector ¯ h comprises the CE on M/NtM/N_t REs for each link, and σ~2 σ^2 is the effective noise power that incorporates both the noise and the residual data interference components. This formulation enables a maximum coherent processing gain of M/NtM/N_t. Although the practical gain can be analysed depends on h R_h, for simplicity we assume that ()†E\ \!h( \!h) \ is decreased by a factor γNtM γ N_tM, where γ≥1γ\!≥\!1 accounts for any loss relative to the maximum coherent gain. Consequently, after de-noising, the CE-MSE is modelled as pp I, where p=(γNtM)k(1−b)(p+q−pq)+σ2kb+k(1−b)(p+q−pq)+σ2. p= ( γ N_tM )\! k(1-b)(p+q-pq)+σ^2kb+k(1-b)(p+q-pq)+σ^2. (7) I-D MD with SI-DMRS After CE is obtained and de-noised, after removing the DMRS component, the received signal model for detecting x is 2 y_2\! = \!\!\!\!=\!\!\!\! bksDMRS+(−b)kNt(~+)+ \!\! bk \!hs_DMRS\!+\! (1\!-\!b)kN_t( H\!+\! \!H) x\!+\! w = \!\!\!\!=\!\!\!\! (−b)kNt~+(−b)kNt+bksDMRS+. \!\! (1\!-\!b)kN_t H x\!+\! (1\!-\!b)kN_t \!H x\!+\! bk \!hs_DMRS\!+\! w\!. For analytical tractability, the MD-MSE ()†E\ x( x) \ based on (I-D) can be approximated as qq I (see Appendix-B), where q=kp+σ2k(1−b+pb)+σ2. q= kp+σ^2k(1-b+pb)+σ^2. (9) As seen, equations (7) and (9) define the equilibrium state of CE-MSE and MD-MSE. This is formalized in Property 1, with its proof provided in Appendix C. Property 1 There exists a stable equilibrium state for CE-MSE and MD-MSE within the ICED framework. Furthermore, the minimum CE-MSE solution p (≤p≤10\!≤\!p\!≤\!1) can be obtained as the root of a third-order polynomial equation: Ap3+Bp2+Cp+D=0, Ap^3+Bp^2+Cp+D=0, (10) where u=−bu\!=\!1\!-\!b, v=σ2kv\!=\! σ^2k, and ρ=γNtMρ= γ N_tM, and A A = \!\!\!\!=\!\!\!\! u2, u^2, (11) B B = \!\!\!\!=\!\!\!\! −((ρ+2)u2+(1−u)(1+v)), -((ρ\!+\!2)u^2\!+\!(1-u)(1+v)), (12) C C = \!\!\!\!=\!\!\!\! ρ(u(+u)+(1−u)v)−(u(1−u)+(1+u)v+v2), ρ(u(1\!+\!u)\!+\!(1-u)v)\!-\!(u(1-u)\!+\!(1+u)v\!+\!v^2), (13) D D = \!\!\!\!=\!\!\!\! ρv(2u+v). ρ v(2u\!+\!v). (14) Property 1 provides a theoretical framework to analyse the iterative behaviour of CE and MD on REs with SI-DMRS for general configurations of (N,M,a,b,NtN,M,a,b,N_t, γ, σ2σ^2), enabling the converged CE-MSE and MD-MSE to be determined directly by solving the equilibrium equations. However, to analyse the overall decoding performance across the entire subframe, we must also account for the MD-MSE on REs without DMRS. I-E MD on REs without SI-DMRS Once the iterative steps between CE and MD on REs with SI-DMRS have converged, the CE ~ H achieves high accuracy and can be utilized for data-detection on the remaining REs33 3 In AI-based receivers, it is feasible and beneficial to detect data symbols across all REs and leverage all data estimates to further enhance CE by fully exploiting the correlations between the data and channel. This methodology is adopted in our AI receiver design, which is elaborated later.. On REs without SI-DMRS, the received signal model for detecting x reads 3 y_3\! = \!\!\!\!=\!\!\!\! aNt(~+)+ \! aN_t( H+ \!H) x\!+\! w (15) = \!\!\!\!=\!\!\!\! aNt~+aNt+. aN_t H x\!+\! aN_t \!H x\!+\! w. In this case, the MD-MSE as approximated as tt I (see Appendix-D), where t=ap+σ2a(1−b)+σ2. t= ap+σ^2a(1-b)+σ^2. (16) I Optimal Power Allocations for SI-DMRS Transmissions With the CE-MSE and MD-MSE for the ICED process derived above, we now seek to address the following fundamental question: How can we optimize the power allocation factors (a,b)(a,b) to maximize the final performance? I-A MIESM Before addressing this question, we first introduce the mutual information effective SNR mapping (MIESM) to quantify the combined MD-MSE. The basic principles are as follows: • Convert the MD-MSE of the two subsets into effective SNR [PK08] values: θ1=1q−1,θ2=1t−1. _1= 1q-1, _2= 1t-1. (17) • Map SNR to MI using a constellation-constrained MI function: Ii=fMI(θiβ), I_i=f_MI\! ( _iβ ), (18) where fMI(⋅)f_MI(·) is the MI mapping for the chosen modulation and the parameter β is a calibration factor. Typical functions of fMI(⋅)f_MI(·) are shown in Appendix-E. • Compute the weighted average MI: Iavg=MI1+NI2M+N. I_avg= MI_1+NI_2M+N. (19) • Map the average MI back to an equivalent SNR: θeff=βfMI−1(Iavg), _eff=β\,f_MI^-1\! (I_avg ), (20) At last, the combined MD-MSE is calculated as 1/(+θeff)1/(1\!+ _eff). I-B Optimal Power Allocations Fig. 2: The equilibrium state of CE-MSE and MD-MSE based on Property 1 with (a=1a\!=\!1, γ=2γ\!=\!2, Nt=4N_t\!=\!4, N=264N\!=\!264, M=24M\!=\!24). Fig. 3: The effective MD-MSE based on MIESM mapping of p and t with (a=1a\!=\!1, γ=2γ\!=\!2, Nt=4N_t\!=\!4, N=264N\!=\!264, M=24M\!=\!24) and σ2=−15σ^2\!=\!-15dB. From Property 1, the equilibrium state of (p,q)(p,q) and t can be directly computed for different power allocation strategies. Subsequently, the effective SNR θeff _eff is obtained from (q,t)(q,t). Thus, the goal of optimal power allocation is to maximize θeff _eff. This provides a clean analytical approach to assess the performance of SI-DMRS, with several illustrative examples shown below. In Fig. 2, the values of p and q are illustrated for different noise powers (σ2=−10σ^2\!=\!-10dB, −20-20dB, and −30-30dB) with a=1a\!=\!1 (i.e., without power-boosting from REs carrying exclusively data symbols). As observed, there exists an optimal value of b-the power allocation factor for DMRS-that minimizes q (the MD-MSE), whereas the CE-MSE continues to decrease as b increases. Furthermore, as the SNR increases, the optimal b also increases. This behaviour occurs because, at higher SNRs, the CE-MSE becomes the bottleneck affecting the MD-MSE. In Fig. 3, the effective MD-MSE is presented using β=1β\!=\!1 in (18) and σ2=−15σ^2\!=\!-15dB, which aligns more closely with overall decoding performance. As seen, the effective MD-MSE-derived via MIESM from (q, t)-is minimized at a higher value of b compared to the case that considers q in isolation. Furthermore, examining the curve for t indicates that the MD-MSE of data symbols on REs without SI-DMRS requires higher CE accuracy to achieve improved error performance. Note that in both Fig. 2 and Fig. 3, the data power allocation factor is fixed to a=1a\!=\!1. Fig. 4 illustrates the joint optimization over both a and b. As demonstrated, the effective MD-MSE is reduced from −10.42-10.42dB to −12.2-12.2dB when a is decreased from 11 to 0.670.67. This highlights the significant performance benefits of jointly optimizing power allocation across REs44 4 A potential drawback of this approach, however, is an increase in the peak-to-average-power ratio (PAPR) for the MIMO-OFDM system.. I-C Is Superimposing More DMRS with Data Beneficial? Fig. 4: 2D contour of effective MD-MSE for joint power allocations over (a,b)(a,b) with (γ=2γ\!=\!2, Nt=4N_t\!=\!4, N=264N\!=\!264, M=24M\!=\!24) and σ2=−15σ^2\!=\!-15dB. Fig. 5: The minimal effective MIESM with different configurations of M, i.e., for superimposing more DMRS with data. The gains are limited even under the idea assumption that the de-nosing gain of CE linearly increases in the number of DMRS. Another interesting question is whether we should superimpose more DMRS with data. In the preceding figures, the number of DMRS is set to M=24M\!=\!24, with each Tx occupying 6 REs out of a total grid of 240 REs.In Fig. 5, we evaluated different configurations of M (2424, 4848, and 9696) and plotted the minimum effective MD-MSE. For each configuration, we also tested different factors (γ=2γ\!=\!2, 44, and 66) representing varying de-noising gains. As shown, for a constant γ=2γ\!=\!2, increasing M from 2424 to 9696 yields gains. However, keeping γ constant implies that the de-noising gain increases linearly with M. When compared to the extreme case where the de-noising gain remains unchanged (i.e., when γ/Mγ/M is constant), the configuration with M=24M\!=\!24 and γ=2γ\!=\!2 outperforms the M=48M\!=\!48 and γ=4γ\!=\!4 configuration by 1.551.55dB. Therefore, it is not always beneficial to superimpose more DMRS with data, and the optimal choice depends heavily on the de-noising gains and operational system parameters. Nevertheless, this analytical framework provides a valuable perspective for analysing general cases to optimize SI-DMRS. IV The Proposed Transformer based AI-ICED Receiver Design With the theoretical framework established, we next present the AI-ICED receiver design, which harnesses the spectral efficiency (SE) gains enabled by SI-DMRS transmission. The overall architecture is illustrated in Fig. 6. The AI receiver processes three tensor inputs: the received signal tensor Y, the DMRS symbols ps_p, and the DMRS coordinate mask Ωp _p. By treating the DMRS information as contextual prompts, the AI architecture establishes a direct mapping from the observation space to the symbol space. Meanwhile, the AI receiver applies an ICED design, and at iteration stage i, it updates the CE and the a posteriori probabilities of symbol vectors, pi(∣,~(i−1),~(i−1))p^i( x\! \! Y, H^(i-1), x^(i-1)), by capturing the dependencies between the intermediate CE and soft symbol estimates from the previous iteration. Ultimately, the logits generated by the AI-ICED receiver are used to calculate bit log-likelihood ratios (LLRs), which are then fed into the channel decoder. As depicted in Fig. 6, the overall design is structured as a cascaded loop where the inputs at each stage are processed through three primary modules: the tokenizer, the ICED module, and the physics-guided feature construction (PGFC) module between iteration stages. Fig. 6: The structure of the MIMO-OFDM system and the proposed AI-ICED receiver, along with details regarding the N construction and configurations provided in Appendix F through I and Table I. IV-A Tokenization This module serves as a tensor pre-processor that performs a spatial flattening operation, concatenates the tensors along their feature axes, and applies a linear projection. This step maps the physical tensors into a unified dvirtuald_virtual-dimensional token sequence. Note that all defined complex-valued tensors include a batch dimension B, with the real and imaginary components decoupled and stacked along a dedicated feature axis of size two. The module receives the real-valued observation tensor ℝ∈ℝB×Nr×Nsym×Nsc×2 Y_R\!∈\!R^B× N_r× N_sym× N_sc× 2, the symbol tensor ℝ∈ℝB×Nt×Nsym×Nsc×2 s_R\!∈\!R^B× N_t× N_sym× N_sc× 2, and the binary mask p∈0,1B×Nt×Nsym×Nsc _p\!∈\!\0,1\^B× N_t× N_sym× N_sc indicating where an SI-DMRS is transmitted. The raw feature vector at a specific RE coordinate (n,k)(n,k) is constructed by concatenating the flattened antennas and the real-imaginary components, and the mapping is formulated as raw(i)[n,k]=Concat((i−1)[n,k],~(i−1)[n,k],p[n,k]), x_raw^(i)[n,k]=Concat ( ^(i-1)[n,k], s^(i-1)[n,k], _p[n,k] ), (21) where the symbol vector is reconstructed according to the SI-DMRS scheme and power allocations as ~n,k(i)=ρsDMRS+1−b~n,k(i−1),if p(n,k)=1a~n,k(i−1),Otherwise. s^(i)_n,k= cases ρ\,s_DMRS\!+\! 1-b\, x^(i-1)_n,k,&if _p(n,k)=1\\ a x^(i-1)_n,k,&Otherwise cases\!. (22) At initial stage (i=0i\!=\!0), the tensor (−1)[n,k] ^(-1)[n,k] is initialized by ℝ[n,k] Y_R[n,k], and ~(−1)[n,k]= x^(-1)[n,k]\!=\! 0 such that ~n,k(0) s^(0)_n,k only contains known DMRS. The dimension of raw(0)[n,k] x_raw^(0)[n,k] is Din(0)=2Nr+3NtD_in^(0)\!=\!2N_r\!+\!3N_t. In subsequent stages, the dimension Din(i)D_in^(i) expands to accommodate more features generated by the PGFC module. A projection matrix (i)∈ℝDin(i)×dvirtualW^(i)\!∈\!R^D_in^(i)× d_virtual is applied to map the raw feature vector into a dvirtuald_virtual-dimensional latent space as raw(i)[n,k](i) x_raw^(i)[n,k]W^(i). By aggregating all projected vectors across L REs, the tokenizer outputs a sequence token(i)∈ℝB×L×dvirtual X_token^(i)\!∈\!R^B× L× d_virtual as input to the next ICED module. This tokenization strategy closely aligns with the physical characteristics of MIMO-OFDM systems. By flattening the 2D time-frequency grids into a sequence of length L, the self-attention mechanism is augmented with Rotary Position Embedding (RoPE) [36] and effectively captures the local coherence of MIMO channels. Furthermore, by collapsing spatial antennas into a single dvirtuald_virtual-dimensional feature vector per token, the architecture delegates the resolution of spatial correlations to the network. This structural design enables the attention layers to perform MD internally within the high-dimensional latent representation. Fig. 7: The VWN-TPA-MoE architecture enhanced Transformer encoder design. IV-B ICED Backbone The ICED operating on the tokenized sequence is designed as an unfolded cascaded architecture that performs CE and MD. Rather than treating them as disjoint modules, the AI-ICED design unifies them into a fused pipeline, as illustrated in Fig. 6. The Transformer encoder architecture incorporates with the latest VWN, TPA and MoE techniques to optimize feature representation. The prompt token(i)X_token^(i) received from the tokenizer is firstly processed by NCEN_CE stacked customized Transformer encoder blocks that incorporates VWN, TPA, and MOE. These specialized encoder blocks iteratively resolve the time-frequency correlations and spatial interference, yielding a channel latent representation feature(i)∈ℝB×L×dvirtual H_feature^(i)\!∈\!R^B× L× d_virtual. To reconstruct the MIMO channel, a projection head (a linear fully-connected layer) CE(i)W_CE^(i) maps feature(i)H_feature^(i) back to physical antenna dimensions ^flat(i)=feature(i)CE(i)∈ℝB×L×(2NrNt). H_flat^(i)= H_feature^(i) W_CE^(i) ^B× L×(2N_rN_t). (23) It is subsequently reshaped to the real-valued CE tensor ^(i)∈ℝB×Nr×Nt×Nsym×Nsc×2 H^(i)\!∈\!R^B× N_r× N_t× N_sym× N_sc× 2. Instead of solely using ^flat H_flat, the input to MD module is formulated via an additive residual bridge d(i)=token(i)+feature(i)+^flat(i)re−embed(i), Z_d^(i)= X_token^(i)+ H_feature^(i)+ H_flat^(i) W_re-embed^(i), (24) where re−embed(i) W_re-embed^(i) acts as a re-embedding layer that projects the CE back into the high-dimensional latent space. This design linearly aggregates the original input prompt token(i)X_token^(i), the uncompressed high-dimensional channel memory feature(i)H_feature^(i), ensuring maximum information retention55 5 However, the purpose of channel projection head and re-embedding is just to obtain the estimate of H and measure the CE-MSE for training loss optimization during training, once the N is trained, these two modules can be combined together as one in inference stage.. The fused sequence d(i)Z_d^(i) is subsequently processed by NMDN_MD transformer encoder blocks, which apply the same VWN-TPA-MOE design to detect the symbols. Finally, a classification head maps the output features into logits for the real and imaginary pulse amplitude modulation (PAM) symbols: PAM(i)=EncoderMD(d(i))MD(i), Z_PAM^(i)=Encoder_MD ( Z_d^(i) ) W_MD^(i), (25) where MD(i) W_MD^(i) is the projection weight, and M M represents the PAM alphabet size converted from the M-QAM modulation. The output tensor PAM(i)∈ℝB×L×(2NtM) Z_PAM^(i)\!∈\!R^B× L×(2N_t M) encompasses the real and imaginary logits, Re(i) Z_ Re^(i) and Im(i)Z_ Im^(i), which are utilized for loss computation. By applying standard softmax normalization, these logits are mapped to the probability tensor ^(i) P^(i). IV-C Physics Guided Feature Construction (PGFC) While the VWN-TPA-MoE encoders provide a powerful computational engine, a purely data-driven approach faces limitations in modeling the high-dimensional state space of MIMO-OFDM detection. Rather than propagating raw outputs across refinement stages, PGFC functions as a differentiable bridge. It constructs communication-theoretic tensors derived from current channel and symbol estimates and supplies the subsequent VWN-TPA encoder with structured features. This relieves the network from learning fundamental spatial projections strictly from data, thereby accelerating convergence and improving generalization across varying channel conditions. Using the probability tensor ^(i) P^(i) output from the ICED module, the soft estimates ~(i) x^(i) of the transmitted data can be computed. Based on these estimates, several features are assembled for the next stage: an interference cancellation (IC) result, (i)=−^(i)~(i) R^(i)= Y- H^(i) S^(i); matched-filter (MF) outputs, MF(i)=(^(i)) Z_MF^(i)=( H^(i)) H Y and MF(i)=(^(i))(i) R_MF^(i)=( H^(i)) H R^(i); and the Gram matrix (i)=(^(i))^(i) G^(i)=( H^(i)) H H^(i). After converting them into their real-valued versions, the tensor (i) ^(i) in (21) sent to the tokenizer is augmented as follows: (i)=Concat(ℝ,ℝ(i),MF,ℝ(i),ℝ(i),MF,ℝ(i)). ^(i)=Concat ( Y_R,\; R_R^(i),\; Z_MF,R^(i),\; G_R^(i),\; R_MF,R^(i) ). (26) Combined with the updated soft state ~(i) s^(i) and the pilot mask p _p are forwarded to the tokenizer for the next ICED stage66 6 A stop-gradient operation is applied to ~(i) s^(i) to isolate the backward pass within each iterative stage, stabilizing the deep supervision dynamics across the unfolded cascade. IV-D Training Loss Design To simultaneously optimize CE and MD, the training loss at iteration stage i balances these objectives using a hyper-parameter β1∈[0,1] _1∈[0,1]: ℒstage(i)=β1ℒ(i)+(1−β1)ℒ(i), _stage^(i)= _1L_ H^(i)+(1- _1)L_ S^(i), (27) where the detection loss is measured via cross-entropy as ℒ(i)=CrossEntropy(PAM(i),PAM,Label), _ S^(i)=CrossEntropy ( Z_PAM^(i), Z_PAM,Label ), (28) and the CE loss is computed using the MSE: ℒ(i)=MSE(^(i),Label). _ H^(i)=MSE ( H^(i),H_Label ). (29) To further mitigate vanishing gradients across the unfolded cascade, deep supervision injects gradients into all intermediate stages. Controlled by a weighting factor β2∈[0,1] _2∈[0,1], the total loss over Nit+1N_it\!+\!1 stages is defined as ℒtotal=(1−β2)ℒstage(Nit)+β2Nit∑i=0Nit−1ℒstage(i). _total=(1- _2)L_stage^(N_it)+ _2N_it _i=0^N_it-1L_stage^(i). (30) IV-E Inference Complexity A critical challenge for AI-based receivers is inference complexity. The per-stage computational complexity is primarily bounded by attention-score computation, low-rank TPA projections, and MoE operations. As the unfolded receiver iterates over Nit+1N_it\!+\!1 stages, the total inference cost scales linearly with the number of applied stages .To manage this complexity, the architecture incorporates three key design strategies: VWN expands the latent memory capacity to dvirtual=(n/m)dmodeld_virtual\!=\!(n/m)d_model, but restricts dense nonlinear operations to m out of n virtual slots. This effectively decouples representational capacity from active computational width. TPA factorizes the Query, Key, and Value projections into low-rank tensor products. Given a query rank rqr_q and a shared key/value rank r, the dense projection complexity drops from (Ldmodel2)O(Ld_model^2) to (Ldmodel(rq+2r))O(Ld_model(r_q\!+\!2r)). Meanwhile, the (L2dmodel)O(L^2d_model) attention-score term is retained to capture global time-frequency dependencies. The MoE feed-forward block expands parameter capacity via expert specialization. By activating exactly NsN_s shared experts and KrK_r routed experts per token, the active computation maintains FLOP equivalence with a standard dense Transformer FFN. A list of the complexity is summarized in Table I in Appendix I. V Numerical Results Next we analyse the performance of the AI-ICED receiver under SI-DMRS and 5G-NR settings. The configurations and channel models are aligned with 3GPP specifications and are summarized in Table I, and the hyper-parameter configurations for the AI-ICED receiver construction are summarized in Table 3, respectively. The basic configurations are the same as depicted in Fig. 1, where we assume a bandwidth of two PRBs, and two OFDM symbols are used for SI-DMRS transmissions (so M=48M\!=\!48), regardless of NtN_t, and N=240N\!=\!240. So the maximal SE gain can be obtained is 20%. Performance is characterized in terms of CE-MSE, BLER, and throughput for different SNRs and code-rates. For all SI-DMRS configurations, the DMRS pattern adopts the layout shown in Fig. 1. Based on the theoretical results in Sec. I, which are also validated by our tests, the power allocation factor is set to b=0.8b\!=\!0.8, assigning a larger proportion of the transmit power to the pilot symbols. For a fair comparison with the 5G-NR baseline, no power boosting is applied (i.e., a=1a\!=\!1). The multipath fading channels are modelled as Extended Typical Urban (ETU70) [34]. For evaluation, the generated datasets comprise 800,000 frames for the ×22\!×\!2 MIMO setup, alongside 200,000 subframes for the ×44\!×\!4 MIMO configuration. All datasets are partitioned into 81% for training, 9% for validation, and 10% for testing. Finally, the block error rate (BLER) and throughput of the proposed architecture are benchmarked against classical baselines including conventional LMMSE and the optimal log-map based maximum likelihood detector (MLD). Furthermore, to decouple the errors induced by CE from detection, a genie-aided variant of the proposed architecture is tested. It bypasses the estimated channel tensor ^flat H_flat and directly injects the ground-truth H into the latent residual fusion bridge, which reflects the MIMO detection performance under genie-aided channel state information (CSI). Note that the same AI-ICED design is also applied to the 5G-NR baseline simulations using separate training sessions, which corresponds to a special case of SI-DMRS where b=1b\!=\!1. Consequently, the overarching design principles and AI receiver architecture apply equally to standard 5G-NR configurations. V-A CE-MSE and BLER Convergences of the Proposed AI-ICED Receiver Fig. 8: CE-MSE over iteration stages for ×44\!×\!4 MIMO and 16QAM modulation. Fig. 9: BLER performance with the same configurations in Fig. 8. Fig. 10: CE-MSE under ×22\!×\!2 MIMO and 16-QAM across a wide SNR range. The CE-MSE and BLER performance of the proposed AI-ICED receiver under a ×44\!×\!4 MIMO and 16-QAM modulation (4-PAM) are presented in Fig. 8 and Fig. 9, respectively, using an LDPC code-rate of 2/32/3. The AI-ICED receiver is trained at an SNR of 1616dB, but evaluated across a wide SNR range. As shown, both the CE-MSE and BLER converge rapidly within a single iteration. With two additional iterations, the performance gain is approximately 0.50.5dB in SNR for both metrics, highlighting the effectiveness of the iterative AI-ICED design. Furthermore, as illustrated in Fig. 9, the AI-ICED receiver approaches the performance of the optimal MLD provided with genie-aided CSI. In Fig. 10, the CE-MSE across a wide SNR range is presented, along with the corresponding required operational SNR for different code-rates. Although the AI-ICED receiver trained at a fixed SNR can generalize to adjacent SNRs, train-test SNR mismatch prevents optimal detection over the entire operating range. To handle this mismatch, dedicated models are trained at regular SNR intervals to capture variations and maintain optimality. As shown, due to the interference introduced between DMRS and data symbols under SI-DMRS, there is a CE-MSE degradation of approximately 2–4 dB2--4 dB compared to the 5G-NR baseline. However, the CE-MSE error remains insignificant compared to the noise power. Furthermore, this performance loss is well justified by the increased SE and higher throughput achieved by the AI-ICED receiver as shown later. V-B Throughput Increments Fig. 11: Throughput envelopes under the ×22\!×\!2 and 16-QAM modulation. Fig. 12: Throughput envelopes under the ×44\!×\!4 and 16-QAM modulation. Although both the CE-MSE and BLER can be worse than 5G-NR baselines without superimposed DMRS, the primary advantage of SI-DMRS is the increased SE resulting from its zero DMRS overhead. The throughput envelopes are illustrated in Fig. 11 and Fig. 12 for the ×22\!×\!2 and ×44\!×\!4 cases for LDPC code-rates (1/3, 1/2, 2/3) corresponding to low, medium, and high code-rates, respectively, The proposed AI-ICED receiver is trained at SNR points (10dB, 14dB, 16dB) to cover different SNR range and code-rates. As shown, the throughputs achieved using SI-DMRS (represented by the yellow curves) combined with the AI-ICED receiver consistently outperform the 5G-NR baseline (the green curves), even though both utilize the proposed AI-ICED receivers. Notably, as the SNR increases, the throughput of the 5G-NR approach saturates due to DMRS overheads, whereas transmissions utilizing SI-DMRS achieve full throughput. Conversely, when using genie CSI, the AI-ICED receiver simplifies to an AI-based MIMO detector, performing (illustrated by the blue curves) nearly as well as the performance bound, namely the genie-CSI-aided MLD (the purple curves). This highlights the effectiveness of AI-ICED in approaching optimal detector performance with genie CSI. Additionally, compared against the baseline genie-CSI-aided LMMSE detector (the red curves), the proposed AI-ICED receiver demonstrates superior performance in middle-to-high SNR regions where detector capability is critical. VI Summary We have introduced a Transformer encoder based AI-ICED receiver that overcomes DMRS overheads in MIMO-OFDM systems. By integrating the latest VWN, TPA, and MoE architectures, the proposed AI-ICED receiver design reformulates CE and MD as a joint contextual learning process, utilizing a PGFC bridge and an iterative unfolded cascade that is highly resilient to pilot sparsity. With the AI-ICED receiver, both the CE-MSE and BLER converge rapidly within a few iterations, yielding significant performance gains. Furthermore, it achieves detection performance close to that of the optimal MLD when given genie CSI input. Validated by throughput envelopes, the SI-DMRS scheme, combined with the AI-ICED receiver, delivers higher throughput and SE compared to conventional 5G-NR baselines with non-superimposed DMRS and data transmission. A key question regarding the SI-DMRS scheme involves the power allocation between DMRS and data symbols, as CE and MD are entangled. We have derived the CE-MSE and MD-MSE in approximated closed forms and developed an analytical framework to directly evaluate the final equilibrium states under an ICED process, eliminating the need for numerical simulations. Studies indicate that allocating a power factor of around 0.80.8 to DMRS and 0.20.2 to data provides an effective design. Furthermore, we have extended the analysis to cases where the power on REs with SI-DMRS can be boosted by borrowing power from remaining REs, yielding substantial additional gains. To evaluate the final decoding performance, we have adopted MIESM to combine MD-MSEs from the two sets of REs: those containing SI-DMRS and those without. Appendices VI-A Derivations of CE-MSE Letting 1=(1−b)kNt(−~~)+, z_1\!=\! (1-b)kN_t( H x- H x)+ w, (31) and since the estimation error \!H is orthogonal to ~ H from the orthogonality principle, it holds that ~†=E\ \!H H \\!=\! 0. Similarily, ~†=E\ \!x x \\!=\! 0. Hence, it holds that 11† \ z_1 z_1 \ = \!\!\!\!=\!\!\!\! (1−b)kNt(†ΔΔ†CLOSE (1-b)kN_t (E\ H H \E\ \! x \! x \ (32) +ΔΔ†† \!\!\!\!+E\ \! H \! H \E\ x x \ OPEN−ΔΔ†ΔΔ†)+σ2 \!\!\!\!-E\ \! H \! H \E\ \! x \! x \ )\!+\!σ^2 I = \!\!\!\!=\!\!\!\! (k(1−b)(p+q−pq)+σ2), (k(1-b)(p+q-pq)+σ^2 ) I, where Δ† \ \! H \! H \\ = \!\!\!\!=\!\!\!\! NtΔ†=Ntp N_tE\ \! h \! h \\=N_tp I (33) Δ† \ \! x \! x \\ = \!\!\!\!=\!\!\!\! q. q I. (34) With the LMMSE estimator, the CE-MSE equals † \ \!h \!h \ = \!\!\!\!=\!\!\!\! (bk†11†−1+)−1 \! (\!bk E\ z_1 z_1 \^-1\!+\! I\! )^-1 (35) = \!\!\!\!=\!\!\!\! (bk(1−b)(p+q−pq)+σ2+)−1 (\! bkk(1-b)(p+q-pq)+σ^2\!+\!1\! )^-1\! I = \!\!\!\!=\!\!\!\! k(1−b)(p+q−pq)+σ2bk+k(1−b)(p+q−pq)+σ2. k(1-b)(p+q-pq)+σ^2bk+k(1-b)(p+q-pq)+σ^2 I. VI-B Derivations of MD-MSE Similar, letting 2=(1−b)kNt+bks+, z_2\!=\! (1-b)kN_t \!H x\!+\! bk \!hs\!+\! w, (36) it holds that 22† \ z_2 z_2 \\!\! = \!\!\!\!=\!\!\!\! (−b)kNtΔΔ†† \!\! (1\!-\!b)kN_tE\ \! H \! H \E\ x x \ (37) +bkΔΔ†+σ2 \!\!+bkE\ \! h \! h \\!+\!σ^2 I = \!\!\!\!=\!\!\!\! (−b)kΔΔ†+bkΔΔ†+σ2 \!(1\!-\!b)kE\ \! h \! h \+bkE\ \! h \! h \\!+\!σ^2 I = \!\!\!\!=\!\!\!\! (kp+σ2). (kp+σ^2) I. With an LMMSE estimator, the MD-MSE equals † \ \!x \!x \ = \!\!\!\!=\!\!\!\! ((1−b)kNt~†22†−1~+)−1 \! (\! (1-b)kN_t H E\ z_2 z_2 \^-1 H\!+\! I\! )^-1 (38) = \!\!\!\!=\!\!\!\! ((1−b)kNt(kp+σ2)~†~+)−1. (\! (1-b)kN_t(kp+σ^2) H H\!+\! I\! )^-1. Note that an instantaneous MD-MSE depends on the MIMO channel realization H, although the expectation has been taken over x, w and h. By Jensen’s inequality, the ergodic MD-MSE with taking expectation over H holds that †≥((1−b)kNt(kp+σ2)~†~+)−1. _ H\E\ \!x \!x \\\!≥\! (\! (1-b)kN_t(kp+σ^2)E\ H H\\!+\! I\! )^-1. (39) Noting that †=†−†=Nt(1−q), \ H H\\!=\!E\ H H\\!-\!E\ \!H \!H\=N_t(1-q) I, (40) the ergodic MD-MSE is bounded as †≥kp+σ2k(1−b+bp)+σ2. _ H\E\ \!x \!x \\\!≥\! kp+σ^2k(1-b+bp)+σ^2 I. (41) In order to analyse the general iterative behaviour, we approximate †E\ x x \ by its bound, and assuming †≈kp+σ2k(1−b+bp)+σ2. \ \!x \!x \≈ kp+σ^2k(1-b+bp)+σ^2 I. (42) VI-C Proof of Property 1 Note that the equilibrium state of CE-MSE and MD-MSE is defined by the equations p p = \!\!\!\!=\!\!\!\! γNtMk(1−b)(p+q−pq)+σ2kb+k(1−b)(p+q−pq)+σ2, γ N_tM k(1-b)(p+q-pq)+σ^2kb+k(1-b)(p+q-pq)+σ^2, (43) q q = \!\!\!\!=\!\!\!\! kp+σ2k(1−b+pb)+σ2. kp+σ^2k(1-b+pb)+σ^2. (44) With substitutions u=1−bu=1-b, v=σ2kv= σ^2k, and ρ=γNtMρ= γ N_tM, it holds that p+q−pq=p(u+1)−up2+vu+pb+v. p+q-pq= p(u+1)-up^2+vu+pb+v. (45) Inserting it into the p-equation and collecting powers of p yields the cubic polynomial Ap3+Bp2+Cp+D=0, Ap^3+Bp^2+Cp+D=0, (46) with coefficients A,B,C,DA,B,C,D defined in Property 1. VI-D Derivations of MD-MSE on REs without DMRS Letting 3=aNt+, z_3\!=\! aN_t \!H x\!+\! w, (47) it holds that 33† \ z_3 z_3 \\!\! = \!\!\!\!=\!\!\!\! aNtΔΔ††+σ2 \!\! aN_tE\ \! H \! H \E\ x x \\!+\!σ^2 I (48) = \!\!\!\!=\!\!\!\! (ap+σ2). (ap+σ^2) I. With an LMMSE estimator, the MD-MSE equals † \ \!x \!x \ = \!\!\!\!=\!\!\!\! (aNt~†33†−1~+)−1 \! (\! aN_t H E\ z_3 z_3 \^-1 H\!+\! I\! )^-1 (49) = \!\!\!\!=\!\!\!\! (aNt(ap+σ2)~†~+)−1. (\! aN_t(ap+σ^2) H H\!+\! I\! )^-1. Following the same argumentations in Appendix-B, we approximate it as †≈ap+σ2a(1+p)+σ2. \ \!x \!x \\!≈\! ap+σ^2a(1+p)+σ^2. (50) VI-E MIESM For a modulation order M (number of bits carried by each constellation symbol), the constellation constrained MI mapping function f(θ)f(θ) is commonly approximated as f(γ)≈log2(M)(1−∑k=1Kake−bkθ), f(γ)≈ _2(M) (1- _k=1^Ka_ke^-b_kθ ), (51) where ∑k=1Kak=1. _k=1^Ka_k=1. (52) For the widely used exponential fit with K=3K\!=\!3, the paramters are listed in Table I. VI-F VWN Standard Transformer encoders use a fixed hidden dimension (dmodeld_model) across their layers. Increasing this dimension to handle complex MIMO-OFDM wireless channels causes a quadratic rise in computing cost. As shown in Fig. 7, VWN separate memory capacity from computational width. The input token token(i) X_token^(i) is expanded to a larger dimension dvirtual=(n/m)⋅dmodeld_virtual=(n/m)· d_model (where n>mn>m) and reshaped into n blocks: block(i)∈ℝB×L×n×db X_block^(i) ^B× L× n× d_b Here, the block dimension is db=dmodel/md_b=d_model/m. Inside each VWN layer, heavy non-linear operations ℱ(⋅)F(·) (such as TPA or MoE) are restricted to a narrow m-slot computational subspace. Dynamic generalized hyper-connections (DGHC) [30] manage this interaction. A lightweight linear projection and Tanh-activation on the normalized input dynamically generate a write matrix ℬ∈ℝm×nB\!∈\!R^m× n and a joint transformation matrix ∈ℝ(m+n)×nA\!∈\!R^(m+n)× n. Horizontally partitioning A produces the read component ̊∈ℝm×n A\!∈\!R^m× n and the decay residual ^∈ℝn×n A\!∈\!R^n× n. The core read-compute-write cycle is executed as out(i)=ℬ⊤ℱ(̊block(i))+^block(i). _out^(i)=B F ( AX_block^(i) )+ AX_block^(i). (53) It outlines the data flow: the network tracks complex multi-path residuals across the n-dimensional highway, while isolating expensive dense computations to the efficient m-dimensional subspace through dynamic gating. VI-G TPA While the VWN architecture manages macroscopic memory capacity, each layer’s core process relies on TPA. Standard Multi-Head Attention (MHA) [35] uses monolithic parameter matrices, limiting its inductive bias. TPA addresses this by factorizing queries, keys, and values into sums of contextual tensor products [31], improving parameter efficiency and injecting physical priors. Contextual Factorization: Given the intermediate state ∈ℝB×L×dmodelX\!∈\!R^B× L× d_model inside a VWN computational slot, TPA decomposes query, key, and value generation into independent latent factor maps. Let h denote the number of attention heads, dh=dmodel/hd_h\!=\!d_model/h the per-head dimension, and rq,rkr_q,r_k the rank constraints for queries and keys/values, respectively. For queries, two linear projections produce head-specific mixing weights Q∈ℝB×L×h×rqA_Q\!∈\!R^B× L× h× r_q and a shared feature basis Q∈ℝB×L×rq×dhB_Q\!∈\!R^B× L× r_q× d_h such that Q() _Q(X)\!\! = = AQ, \!\!XW_A^Q, (54) Q() _Q(X)\!\! = = BQ, \!\!XW_B^Q, (55) where AQ∈ℝdmodel×(h⋅rq)W_A^Q\!∈\!R^d_model×(h· r_q) and BQ∈ℝdmodel×(rq⋅dh)W_B^Q\!∈\!R^d_model×(r_q· d_h) are factor weight matrices. Analogous projections with rank rkr_k yield (K,KA_K,B_K) and (V,VA_V,B_V). Physical Adaptation with 2D-RoPE and Rank Scaling: Standard 1D rotary position embeddings [36] fail to preserve the 2D coherence patterns of MIMO-OFDM channels. To fix this, we treat the sequence axis as a 2D physical grid and apply RoPE2DRoPE_2D to the feature bases before tensor contraction: ~Q B_Q\!\! = = RoPE2D(Q), \!\!RoPE_2D (B_Q ), (56) ~K B_K\!\! = = RoPE2D(K). \!\!RoPE_2D (B_K ). (57) The full query tensor ∈ℝB×L×h×dhQ\!∈\!R^B× L× h× d_h is reconstituted via tensor contraction =Q()~Q.Q=A_Q(X) B_Q. (58) For the i-th head at the t-th token, this represents a sum of outer products t(i)=∑j=1rqt,j(i)⊗t,jQ_t^(i)\!=\! _j=1^r_qa_t,j^(i) _t,j. Scaled Dot-Product Attention: Using the factorized tensors ,,∈ℝB×L×h×dhQ,K,V\!∈\!R^B× L× h× d_h, attention is computed per head as headi=Softmax(ii⊤dh)i, _i=Softmax\! ( Q_iK_i d_h )\!V_i, (59) where i,i,i∈ℝB×L×dhQ_i,K_i,V_i ^B× L× d_h denote the slices along the head dimension. The outputs are concatenated and projected via O∈ℝ(h⋅dh)×dmodelW_O\!∈\!R^(h· d_h)× d_model as TPA(,,)=Concat(head1,…,headh)O. (Q,K,V)=Concat (head_1,…,head_h )\,W_O. (60) This output completes the TPA layer before routing to the subsequent MoE block. VI-H MoE After the TPA process, the outputs pass through a MoE feed-forward block adapted from the DeepSeekMoE framework [24]. The expert pool is divided into KsK_s always-active shared experts and KrK_r selectively routed experts, both utilizing a SwiGLU-based feed-forward structure [37]. For an input token ∈ℝdmodel x\!∈\!R^d_model, the MoE forward pass is given by =∑k=1KsMLPs(k)()+∑k∈gk()MLPr(k)(), = _k=1^K_sMLP_s^(k)(x)+ _k g_k(x)\,MLP_r^(k)(x), (61) where K represents the set of top-KrK_r experts selected by a lightweight gating network, with gating scores computed via a linear projection followed by a Sigmoid activation =Sigmoid(route)s\!=\!Sigmoid(xW_route), and the top-KrK_r entries are normalized to yield the final routing weights: gk()=sk∑k∈sk+ϵ,k∈, g_k(x)= s_k _k s_k+ε, k , (62) where ϵε is a small constant ensuring numerical stability. VI-I N Configuration of the AI-ICED Receiver The proposed AI-ICED receiver is configured using the hyperparameters listed in Table IV. The VWN chassis is configured with m=2m\!=\!2 memory slots and n=3n\!=\!3 active computational slots per layer, producing an expanded virtual dimension dvirtual=(n/m)⋅dmodel=768d_virtual\!=\!(n/m)· d_model\!=\!768. The TPA sub-layer operates with a query rank rq=16r_q\!=\!16 and a key/value rank rk=16r_k\!=\!16. The MoE sub-layer employs Ns=1N_s\!=\!1 shared expert, Nr=4N_r\!=\!4 routed experts, and Kr=2K_r\!=\!2 active experts per token. The per-expert hidden dimension scaled by a factor of (8/3)/(Kr+Ns)(8/3)/(K_r\!+\!N_s) to maintain equivalence with a dense SwiGLU FFN [37]. To effectively optimize the deep cascaded Transformer backbone, a hybrid optimization strategy is employed: Multi-dimensional weight matrices are updated using the momentum-based Muon optimizer [38]; while one-dimensional parameters, such as layer normalizations and biases, are routed to AdamW [39]. The learning rate follows a linear warmup at first 10%10\% steps, followed by cosine annealing down to a minimum ratio of 0.010.01 of the peak value [40, 41]. References [1] Z. Fu, “Deep in-context learning (ICL) for wireless communications,” Master’s thesis, Dept. of electrical and information technology, Lund University, Lund, Sweden, Jun. 2026. [2] X. Wang, L. Lu, Q. Li, Q. Sun, N. Shi, Z. Chen, and T. Sun, “A task-driven design approach for 6G AI-native architecture,” Engineering, vol. 56, no. 1, p. 87–103, 2026. [3] Y. Wu et al., “A comprehensive review of AI-native 6G: Integrating semantic communications, reconfigurable intelligent surfaces, and edge intelligence for next-generation connectivity,” Front. Commun. Netw., vol. 5, p. 1655410, 2024. [4] Z. Qin, H. Ye, G. Y. Li, and B.-H. F. Juang, “Deep learning in physical layer communications,” IEEE Wireless Commun., vol. 26, no. 2, p. 93–99, Apr. 2019. [5] J. Sajid et al., “Empowering embodied AI in 6G networks: Architecture, enablers, and open challenges,” arXiv preprint arXiv:2606.20592, 2026. [6] T. O’Shea and J. Hoydis, “An introduction to deep learning for the physical layer,” IEEE Trans. Cogn. Commun. Netw., vol. 3, no. 4, p. 563–575, Dec. 2017. [7] X. Lin, “An overview of the 3GPP study on artificial intelligence for 5G new radio,” arXiv preprint arXiv:2308.05315, 2023. [8] X. Qin, S. Hu, J. Zhang, J. Qian, and H. Wang, “AI receiver design with deep learning based channel estimation and MIMO detection,” in Proc. IEEE PIMRC, 2024, p. 1–7. [9] X. Li, X. Zhou, Y. Cao, J. Zhang, C.-K. Wen, X. Li, and S. Jin, “Learning-aided iterative receiver for superimposed pilots: Design and experimental evaluation,” arXiv preprint arXiv:2507.10074, 2025. [10] X. Qin and S. Hu, “Dual-attention based 3D channel estimation,” arXiv preprint arXiv:2604.01769, 2026. [11] S. Cammerer et al., “A neural receiver for 5G NR multi-user MIMO,” in Proc. IEEE Globecom Workshops (GC Wkshps), Kuala Lumpur, Malaysia, 2023, p. 329–334. [12] S. Hu, “Invariant transformation and resampling based epistemic-uncertainty reduction,” arXiv preprint arXiv:2602.23315, 2026. [13] F. B. Saghezchi, M. Pourghasemian, B. Ding, A. Abdi, B. Lee, and A. Baron, “AI-native radio transceiver signal processing for next-generation mobile communication systems,” IEICE Trans. Commun., vol. E109-B, no. 4, p. 555–572, Apr. 2026. [14] H. He, C.-K. Wen, S. Jin, and G. Y. Li, “A model-driven deep learning network for MIMO detection,” in Proc. IEEE Glob. Conf. Signal Inf. Process. (GlobalSIP), Anaheim, CA, USA, 2018, p. 584–588. [15] A. Mazumdar, C. N. Manchon, O. E. Barbu, and R. O. Adeogun, “Comparative evaluation of model based deep learning receivers in coded MIMO systems,” in Proc. IEEE Veh. Technol. Conf. (VTC-Fall), 2024, p. 1–7. [16] M. -H. Hsieh and C. -H. Wei, “Channel estimation for OFDM systems based on comb-type pilot arrangement in frequency selective fading channels,” IEEE Trans. Consum. Electron., vol. 44, no. 1, p. 217–225, Feb. 1998. [17] D. Wubben, R. Bohnke, V. Kuhn, and K.-D. Kammeyer, “MMSE-based lattice-reduction for near-ML detection of MIMO systems,” in Proc. ITG Workshop Smart Antennas, 2004, p. 106–113. [18] S. Hu and F. Rusek, “A soft-output MIMO detector with achievable information rate based partial marginalization,” IEEE Trans. Signal Process., vol. 65, no. 6, p. 1622–1637, Mar. 2017. [19] K. Upadhya, S. A. Vorobyov, and M. Vehkapëä, “Superimposed pilots are superior for mitigating pilot contamination in massive MIMO,” IEEE Trans. Signal Process., vol. 65, no. 11, p. 2917–2932, Jun. 2017. [20] S. Rezaie, M. Honkala, D. Korpi, D. C. Melgarejo, T. Izydorczyk, D. Gold, and O.-E. Barbu, “Superimposed DMRS for spectrally efficient 6G uplink multi-user OFDM: Classical vs AI/ML receivers,” arXiv preprint arXiv:2506.20248, 2025. [21] X. Li, X. Zhou, J. Zhang, C.-K. Wen, and S. Jin, “AI-driven iterative receiver for superimposed pilot Schemes in MIMO-OFDM systems,” in Proc. IEEE WCNC, 2025, p. 1–6. [22] K. Ying et al., “Conditional diffusion model-driven sassive MIMO iterative detection,” in Proc. IEEE Int. Conf. Commun. (ICC), Glasgow, United Kingdom, 2026, p. 1–6. [23] M. Abuthinien, S. Chen, and L. Hanzo, “Semi-blind joint maximum likelihood channel estimation and data detection for MIMO systems,” IEEE Signal Process. Lett., vol. 15, p. 202–205, 2008. [24] DeepSeek-AI et al., “DeepSeek-V3.2: pushing the frontier of open large language models,” arXiv preprint arXiv:2512.02556, 2025. [25] A. Yang et al., “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025. [26] Kimi Team et al., “Kimi K2.5: Visual Agentic Intelligence,” arXiv preprint arXiv:2602.02276, 2026. [27] Z. Song, M. Zecchin, B. Rajendran, and O. Simeone, “Turbo-ICL: In-context learning-based turbo equalization,” arXiv preprint arXiv:2505.06175, 2025. [28] M. Shanmugam, S. Balaraman, and K. Ranganathan, “Improving channel equalization in cell-free MIMO networks using reconfigurable intelligent surfaces and in-context learning,” Ann. Telecommun., vol. 81, p. 259–273, 2026. [29] M. Zecchin, K. Yu, and O. Simeone, “In-context learning for MIMO equalization using Transformer-based sequence models,” in Proc. IEEE Int. Conf. Commun. Workshops (ICC Workshops), 2024, p. 1573–1578, [30] Seed et al., “Virtual Width Networks,” arXiv preprint arXiv:2511.11238, 2025. [31] Y. Zhang et al., “Tensor product attention is all you need,” arXiv preprint arXiv:2501.06425, 2026. [32] D. Dai et al., “DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models,” arXiv preprint arXiv:2401.06066, 2024. [33] 3GPP, “5G; NR; Physical channels and modulation,” 3GPP Technical Specification TS 38.211, V18.6.0, Apr. 2025. [34] 3GPP, ”Evolved Universal Terrestrial Radio Access (E-UTRA); User Equipment (UE) radio transmission and reception,” 3GPP Technical Specification TS 36.101, V18.9.0, Apr. 2025. [35] A. Vaswani et al., “Attention is all you need,” in Advances Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, p. 5998–6008. [36] J. Su et al., “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing, vol. 568, p. 127063, 2024. [37] N. Shazeer, “GLU variants improve Transformer,” arXiv preprint arXiv:2002.05202, 2020. [38] K. Jordan et al., “Muon: An optimizer for hidden layers in neural networks,” 2024. [Online] Available: https://kellerjordan.github.io/posts/muon/. [39] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2019. [40] P. Goyal et al., “Accurate, large Minibatch SGD: Training ImageNet in 1 Hour,” arXiv preprint arXiv:1706.02677, 2017. [41] I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” arXiv preprint arXiv:1608.03983, 2017. [42] J. Yang, H. Zhao, W. Wang, and C. Zhang, “An effective SINR mapping models for 256QAM in LTE-Advanced system,” in IEEE Annu. Int. Symp. Pers., Indoor, Mobile Radio Commun. (PIMRC), Washington, DC, USA, 2014, p. 343–347. [43] N. Kim, Y. Lee, and H. Park, “Performance analysis of MIMO system with linear MMSE receiver,” IEEE Trans. Wireless Commun., vol. 7, no. 11, p. 4474–4478, Nov. 2008. TABLE I: Dominant inference complexity per stage. Operation Dense Transformer block Proposed backbone Latent memory width dmodeld_model dvirtual=(n/m)dmodeld_virtual=(n/m)d_model Q/K/V projections (Ldmodel2)O(Ld_model^2) (Ldmodel(rq+2r))O\! (Ld_model(r_q+2r) ) FFN computation (Ldmodeldff)O(Ld_modeld_f) (Ldmodeldff)O(Ld_modeld_f) Attention score (L2dmodel)O(L^2d_model) (L2dmodel)O(L^2d_model) TABLE I: System parameters. Symbol Description Value Nr×NtN_r× N_t MIMO size 2×22× 2, 4×44× 4 A Modulation scheme 16QAM, 64QAM NsymN_sym Number of OFDM symbols carrying data 1212 NscN_sc Active subcarriers 2424 N Number of REs carrying only data 240240 M Number of REs carrying both SI-DMRS and data 4848 L Token lenght for the AI-ICED receiver 288288 Codes LDPC Code-rates 1/31/3, 1/21/2, 2/32/3 (a,b)(a,b) Power allocation factors (a=1,b=0.8)(a=1,b=0.8) Channels 3GPP channel models ETU-70Hz, EPA-5Hz TABLE I: Three-term exponential MI mapping parameters. Modulation a1a_1 b1b_1 a2a_2 b2b_2 a3a_3 b3b_3 QPSK 0.9810 1.0420 0.0190 4.2500 0.0000 0.0000 16-QAM 0.6053 0.2845 0.3541 1.2562 0.0406 5.3784 64-QAM 0.4076 0.0717 0.3951 0.3556 0.1973 1.8315 TABLE IV: Hyperparameter Configuration of the Proposed AI-ICED Receiver Symbol Description Value Transformer Backbone dmodeld_model Latent dimension 512512 NheadN_head Number of attention heads 88 m/nm/n VWN memory / compute slots 2/32/3 dvirtuald_virtual Virtual memory width (n/m)⋅dmodel(n/m)· d_model 768768 rq/r_q/r TPA query rank / key-value rank 16/1616/16 Ns/NrN_s/N_r Shared / routed experts (MoE) 1/41/4 KrK_r Active routed experts per token 22 Expert factor Per-expert hidden factor (8/3)/(Kr+Ns)(8/3)/(K_r+N_s) 8/98/9 Dropout Dropout probability 0.10.1 Cascade Architecture NCE(0)/NMD(0)N_CE^(0)/N_MD^(0) Bootstrap CE / MD encoder layers 3/33/3 NCE(i)/NMD(i)N_CE^(i)/N_MD^(i) Refinement CE / MD encoder layers 6/66/6 NitN_it Number of ICED iteration 11 Training Protocol Epochs Total training epochs 100100 Batch size Samples per batch 1616 Optimizer Hybrid Muon + AdamW LR (Muon) Peak learning-rate for multi-dimensional weights ×10−32\!×\!10^-3 LR (AdamW) Peak learning-rate for embeddings, norms, biases ×10−43\!×\!10^-4 Weight decay ℓ2 _2 regularization penalty 0.010.01 β1 _1 Loss balance between bit cross-entropy and CE-MSE 0.10.1 β2 _2 Auxiliary stage weight 0.30.3