Paper deep dive
Multi-layer MIMO Relay as Deep Physical Neural Networks: Power Amplifiers as Activation Functions
Meng Hua, Itsik Bergel, Deniz Gündüz
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/22/2026, 2:28:14 AM
Summary
This paper proposes a deep Wireless Physical Neural Network (WPNN) architecture using a multi-hop Multiple-Input Multiple-Output (MIMO) relay system. It leverages the intrinsic nonlinearity of Power Amplifiers (PAs) as activation functions for neural layers. Two transceiver designs are presented: a Least Squares (LS)-based scheme requiring only receiver-side Channel State Information (CSI), and a Singular-Value-Decomposition (SVD)-based scheme requiring both transmitter and receiver CSI. Simulations on Fashion-MNIST demonstrate that the LS-based scheme effectively exploits PA nonlinearity for accurate image classification, outperforming the SVD-based approach in deep cascades.
Entities (8)
Relation Signals (6)
Multi-hop MIMO Relay → implements → Wireless Physical Neural Network
confidence 96% · propose a deep WPNN in which nonlinear activations are realized by a multi-hop multiple-input multiple-output (MIMO) relay network
Power Amplifier → actsas → Activation Function
confidence 95% · power amplifier's intrinsic nonlinearity acting as an activation function
Multi-hop MIMO Relay → evaluatedon → Fashion-MNIST
confidence 94% · through simulations on the Fashion-MNIST dataset
Least Squares → requires → Receiver-side CSI
confidence 93% · a least squares (LS)-based scheme requiring only receiver-side CSI
Singular Value Decomposition → requires → Transmitter-side and Receiver-side CSI
confidence 93% · a singular-value-decomposition (SVD)-based scheme requiring both transmitter-side and receiver-side CSI
Least Squares → outperforms → Singular Value Decomposition
confidence 88% · the SVD-based scheme that decouples eigenmodes does not benefit from PA nonlinearity in deep cascades, whereas the LS-based scheme exploits nonlinearity and scales gracefully with depth
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Wireless physical neural networks (WPNNs) embed neural computation directly into analog hardware, offering lower energy consumption and latency than conventional digital implementations. In this paper, we propose a deep WPNN in which nonlinear activations are realized by a multi-hop multiple-input multiple-output (MIMO) relay network, in which each relay implements a trainable complex linear gain and bias, followed by the power amplifier's intrinsic nonlinearity acting as an activation function. The cascade of multiple relays therefore realizes an over-the-air fully connected network whose parameters can be trained end-to-end. We develop two transceiver designs for different channel state information (CSI) availability scenarios: a least squares (LS)-based scheme requiring only receiver-side CSI, and a singular-value-decomposition (SVD)-based scheme requiring both transmitter-side and receiver-side CSI. Simulation results show that the proposed architecture enables accurate over-the-air inference for image classification. In particular, the results highlight the advantage of exploiting hardware nonlinearity for enhanced inference capability.
Tags
Links
- Source: https://arxiv.org/abs/2607.18354v1
- Canonical: https://arxiv.org/abs/2607.18354v1
Trouble viewing inline? Open PDF directly →
Full Text
32,985 characters extracted from source content.
Expand or collapse full text
Multi-layer MIMO Relay as Deep Physical Neural Networks: Power Amplifiers as Activation Functions Meng Hua, Itsik Bergel, , and Deniz Gündüz This work was supported by UKRI under the projects AI-R (EP/X030806/1) and INFORMED-AI (EP/Y028732/1), and by the SNS JU project 6G-GOALS under the EU Horizon program (Grant Agreement No. 101139232). M. Hua and D. Gündüz are with the Department of Electrical and Electronic Engineering, Imperial College London, London SW7 2AZ, U.K. (e-mail: m.hua,d.gunduz@imperial.ac.uk). I. Bergel is with the Faculty of Engineering, Bar-Ilan University, Ramat Gan 5290002, Israel (e-mail: itsik.bergel@biu.ac.il). Abstract Wireless physical neural networks (WPNNs) embed neural computation directly into analog hardware, offering lower energy consumption and latency than conventional digital implementations. In this paper, we propose a deep WPNN in which nonlinear activations are realized by a multi-hop multiple-input multiple-output (MIMO) relay network, in which each relay implements a trainable complex linear gain and bias, followed by the power amplifier’s intrinsic nonlinearity acting as an activation function. The cascade of multiple relays therefore realizes an over-the-air fully connected network whose parameters can be trained end-to-end. We develop two transceiver designs for different channel state information (CSI) availability scenarios: a least squares (LS)-based scheme requiring only receiver-side CSI, and a singular-value-decomposition (SVD)-based scheme requiring both transmitter-side and receiver-side CSI. Simulation results show that the proposed architecture enables accurate over-the-air inference for image classification. In particular, the results highlight the advantage of exploiting hardware nonlinearity for enhanced inference capability. I Introduction The remarkable success of artificial intelligence (AI) has been largely driven by deep learning models executed on graphics processing units (GPUs), whose massive parallelism and numerical precision enable efficient training and inference. However, GPUs inherit the von Neumann architecture, in which the physical separation between memory and computation incurs substantial energy and latency costs, particularly for large-scale or real-time deployments. The emerging concept of the physical neural network (PNN) offers a promising alternative [14, 10, 7]: linear mappings and nonlinear activations are realized through controllable analog mechanisms, including electronic, optical, or wireless, whose intrinsic dynamics and nonlinearity enable in-situ inference with significantly lower energy and delay. Recently, research on implementing wireless PNNs (WPNNs) over the air has gained increasing attention [4]. First, physical substrates such as the phase shifts of reconfigurable intelligent surfaces (RISs) and the amplification coefficients of relays can be regarded as trainable neurons within the wireless medium. Second, fundamental algebraic operations, such as matrix multiplication and summation, can be inherently realized over the air by exploiting the superposition property of the wireless multiple-access channel [1], thereby significantly reducing computation energy consumption and processing latency. Nevertheless, research on over-the-air WPNNs remains in its early stage, with only a limited number of studies reported to date [12, 5, 16, 8, 9, 15, 6, 11, 2, 13, 3]. In particular, [12, 5, 16, 8, 9, 15, 6, 11] investigated the use of RISs as neural network neurons, where the adjustable phase shifts of RIS elements are treated as trainable parameters of the network. For instance, in [11], the authors proposed leveraging RISs to realize one-dimensional convolution operations by exploiting multipath propagation delays, wherein each channel impulse response acts as an individual finite impulse response filter that convolves with the transmitted signal to emulate a digital convolutional neural network. Relays have also been explored as WPNN hardware in [2, 13, 3], where the relay amplification coefficients are interpreted as analog neuron weights. When multiple relays are employed, the combined effects of wireless channels and amplification gains form an equivalent virtual multiple-input multiple-output (MIMO) system. For example, [2] demonstrated that by exploiting both the amplification gain and the inherent nonlinearity of power amplifiers (PAs), significant communication performance gains can be achieved. However, the aforementioned studies have not fully exploited the spatial degrees of freedom or the intrinsic nonlinearity of physical devices, which fundamentally limit their expressive capacity. In this paper, we study deep physical neural networks via a multi-hop MIMO relay architecture, as illustrated in Fig. 1. Our main contributions are as follows. First, we propose a multi-hop MIMO relay architecture for realizing a deep WPNN, formally associating the trainable amplification, bias, and PA nonlinearity at each relay with the weight, bias, and activation function of one FC layer. Second, we propose two transceiver schemes for different channel state information (CSI) availability scenarios: a least squares (LS)-based scheme requiring only receiver-side CSI (CSIR), and a singular-value-decomposition (SVD)-based scheme requiring both transmitter-side and receiver-side CSI (CSIR/T). Both schemes support end-to-end training with standard backpropagation through the PA model. Third, through simulations on the Fashion-MNIST dataset, we show that, contrary to expectation, the SVD-based scheme that decouples eigenmodes does not benefit from PA nonlinearity in deep cascades, whereas the LS-based scheme exploits nonlinearity and scales gracefully with depth. Figure 1: The architecture of multi-hop MIMO relay systems. I System model As illustrated in Fig. 1, we consider a multi-hop MIMO relay system, where information is transmitted from a source to a user through M MIMO relays R1,…,RMR_1,…,R_M. For notational convenience, we denote the source node and the user as R0R_0 and RM+1R_M+1, respectively. All nodes in the network are equipped with equal number of transmit and receive antennas. Specifically, let NmN_m denote the number of transmit or receive antennas at node RmR_m, for m∈0,…,M+1m∈ \0,…,M+1 \. The objective of this relay network is to perform a generic inference task on the signal ∈ S∈ S available at R0R_0, and conveyed to RM+1R_M+1 through M MIMO relays. The task is represented by :→T:S , where Y denotes the task-specific output space. A concrete example of an image classification task will be introduced in Section IV. Let m∈ℂNm×Nm−1H_m∈C^N_m×N_m-1 denote the complex baseband equivalent MIMO channel from node Rm−1R_m-1 to node RmR_m for m∈1,…,M+1m∈ \1,…,M+1 \. We adopt a Rician fading model for all wireless links. Specifically, the channel matrix of the m-th hop, m H_m, is expressed as m = αm(K+1m,LoS+1K+1m,NLoS), H_m = _m ( KK+1H_m,LoS+ 1K+1H_m,NLoS ), (1) where αm _m captures the large-scale fading of the m-th hop, and K represents the Rician factor. The line-of-sight (LoS) component m,LoSH_m,LoS is modeled as a deterministic rank-one matrix and is given by m,LoS=r(θmr;Nm)tH(θm−1t;Nm−1), H_m,LoS=a_r ( _m^r;N_m )a_t^H ( _m-1^t;N_m-1 ), (2) where r(θmr;Nm)a_r ( _m^r;N_m ) and t(θm−1t;Nm−1)a_t ( _m-1^t;N_m-1 ) denote the receive and transmit array response vectors at nodes RmR_m and Rm−1R_m-1, respectively. The angles of arrival (AoA) and departure (AoD) associated with the LoS path are denoted by θmr _m^r and θm−1t _m-1^t, respectively, which are assumed to be independently and uniformly distributed over [0,π] [0,π ]. For a uniform linear array with N antennas, its array response is given by c(θ;N)=[1,ejπsinθ,…,ejπsinθ(N−1)]T,c∈r,t. a_c (θ;N )= [1,e^jπ θ,…,e^jπ θ (N-1 ) ]^T,c∈ \r,t \. (3) The non-LoS (NLoS) component m,NLoSH_m,NLoS models the small-scale scattering, where each entry is independently distributed as [m,NLoS]i,j∼(0,1) [H_m,NLoS ]_i,j CN (0,1 ). Let m∈ℂNm×LX_m∈C^N_m× L and m∈ℂNm×LY_m∈C^N_m× L denote the transmitted and received symbols at node RmR_m for m∈0,1,…,M+1m∈ \0,1,…,M+1 \, respectively, where L denotes the number of transmitted symbols, which will be specified later. We have m=mm−1+m,m∈1,…,M+1, Y_m= H_m X_m-1+ N_m,~m∈ \1,…,M+1 \, (4) where m∈ℂNm×LN_m∈C^N_m× L denotes the additive white Gaussian noise, with each entry in m N_m independently distributed as [m]i,j∼(0,σ2) [ N_m ]_i,j CN (0,σ^2 ). We consider two cases, namely CSIR MIMO and CSIR/T MIMO, depending on the availability of CSI at each node. I-A CSIR MIMO System In a CSIR MIMO system, the CSI is available at the receiver for each hop, which allows it to equalize the transmitted signals. According to the input–output relationship in (4), with the knowledge of mH_m, RmR_m can perform receiver-side signal processing to estimate the transmitted symbols. We adopt the LS estimator to recover an estimate ^m−1 X_m-1 of m−1 X_m-1. By applying a node-specific processing function parameterized by m θ_m, the transmitted signal at RmR_m can be expressed as m=gm(m,m,σ2),m∈1,…,M+1, X_m=g_ θ_m ( Y_m, H_m,σ^2 ),~m∈ \1,…,M+1 \, (5) where gm(⋅)g_ θ_m (· ) denotes the signal processing function implemented at RmR_m. I-B CSIR/T MIMO System In a CSIR/T MIMO system, the CSI is available at both the transmitter and receiver for each hop. Therefore, the encoding function parameterized by WPNN parameters ϕm φ_m at RmR_m can be expressed as m=fϕm(m,m,m+1,σ2),m∈1,…,M, \!\!\! X_m=f_ φ_m ( Y_m, H_m, H_m+1,σ^2 ),~m∈ \1,…,M \, (6) where fϕm(⋅)f_ φ_m (· ) denotes the signal processing function implemented at RmR_m. The WPNN realized by the multi-hop MIMO relay network is trained end-to-end to approximate the task mapping T. Denoting by ^() T_ ( S) the mapping from the input sample S to the user’s decision, parameterized by all trainable physical-layer parameters (amplification matrices, biases, transmit/receive processing, and final read-out layer), the objective is to minimize a task-specific loss: min(,y)∼P,y[ℒ(^(S),y)] _ E_( S,y) P_ S,y[ L( T_ (S),y)], where y∈y is the ground-truth label. I Role of PA Nonlinearity in Multi-Hop MIMO Relay Systems To motivate the role of PA nonlinearity, we contrast linear and nonlinear PA regimes. Linear PA: Assume that the PA at each relay operates strictly in its linear region so that m=mmX_m=W_mY_m, where m∈ℂNm×NmW_m∈C^N_m× N_m is the linear amplification matrix. Substituting it into (4) recursively yields M+1=M+1MMM−1⋯10+eff, Y_M+1=H_M+1W_MH_MW_M-1·sH_1X_0+N_eff, (7) where effN_eff aggregates the per-hop noise and is independent of 0 X_0. The end-to-end map 0→M+1X_0→Y_M+1 is therefore linear and collapses to a single composite FC layer for any M, fundamentally limiting the expressive capability of the WPNN. Basically, the achievable function class is the set of complex linear maps of limited rank minmrank(mm) _mrank (H_mW_m ). Nonlinear PA: When the PA exhibits a nonlinear transfer characteristic, the relay output becomes m=PA(mm), X_m=PA (W_mY_m ), (8) where PA(⋅) PA (· ) denotes the element-wise nonlinear input–output map, whose specific form will be detailed in Section IV. Each relay hop realizes linear mixing followed by a pointwise nonlinear feature transformation, so the M-hop relay chain is equivalent to an M-layer FC network with implicit activations. This breaks the limited rank bound of linear PA, making the achievable function class grow with M. This is the structural basis for interpreting the multi-hop MIMO relay system as a deep WPNN. IV Deep Learning Optimization with Nonlinear PA In this section, we instantiate the general framework developed in Section I for an image classification task, while noting that the proposed design can be readily extended to other inference tasks. We then present the training strategy and the corresponding loss function used to train the WPNN for this task. IV-A Transmission Architecture Design IV-A1 CSIR Design Let the input sample be an image ∈ℝC×H×W S∈R^C× H× W, where C, H, and W represent the number of color channels, height, and width, respectively, and the output space Y is the discrete set of image classes. The source node R0R_0 first normalizes S so that its pixel values lie in [0,1] [0,1 ]. The normalized image is then vectorized and reshaped into a complex-valued matrix c∈ℂN0×L S_c∈C^N_0× L with L=C×H×W2N0L= C× H× W2N_0. Then, the signal transmitted by R0R_0 is given by 0=PA(0c+0LT), X_0=PA (F_0S_c+b_01_L^T ), (9) where 0∈ℂN0×N0 F_0∈C^N_0×N_0 and 0∈ℂN0×1 b_0∈C^N_0× 1 denote the precoder matrix and the bias vector, respectively, and L1_L is the L-length vector of all ones. The bias vector 0 b_0 plays a role analogous to that in digital neural networks and can be practically realized by injecting a direct current offset, which is trainable. The Rapp PA model, denoted by PA(⋅) PA (· ), can be modeled as [13] PA(x)=x(1+(|x|/xsat)2p)1/(2p), PA (x )= x (1+ ( |x | / |x |x_sat . -1.2ptx_sat )^2p )^1 / 1 (2p ) . -1.2pt (2p ), (10) where p and xsatx_ sat represent the PA parameters. Following [2], we set p=2p=2 and xsat=1x_ sat=1, under which the amplitude response of the nonlinear power amplifier exhibits a smooth saturation behavior similar to that of a tanhtanh function. Accordingly, the signal received at relay R1R_1 can be represented as 1=10+1. Y_1= H_1 X_0+ N_1. (11) Then, a LS MIMO estimator is employed to exploit the CSI to decouple the entangled signal 1 Y_1 as ^0 X_0: ^0=1+1=1+(10+1), X_0= H_1^+ Y_1= H_1^+ ( H_1 X_0+ N_1 ), (12) where (⋅)+ (· )^+ denotes the the Moore–Penrose pseudo-inverse. Next, a trainable relay amplification matrix, denoted by 1∈ℂN1×N0 F_1∈C^N_1× N_0, is applied to scale ^0 X_0, and a trainable bias vector 1∈ℂN1×1 b_1∈C^N_1× 1 is added, and the result passes through the nonlinear PA. The output signal at relay R1R_1 can thus be expressed as 1=PA(1^0+1LT). X_1=PA (F_1 X_0+b_11_L^T ). (13) It can be seen that this nonlinear transformation plays the role of an activation function, enabling the relay to realize both amplification and nonlinear feature mapping over the air. Therefore, the relay effectively performs one FC layer, where the amplification matrix 1 F_1 and the bias vector 1 b_1 correspond to the trainable weights and bias of a conventional digital neural network, respectively. At the subsequent relay nodes R2,…,RMR_2,…,R_M, similar signal processing operations are performed, where each relay employs its own trainable amplification matrix and bias vector, followed by the inherent nonlinear amplifier characteristic. Consequently, the entire multi-hop relay chain can be viewed as a cascade of over-the-air FC layers, thereby forming a multi-layer WPNN. In this sense, the CSIR MIMO relay network inherently realizes a deep neural architecture in the analog domain, with each relay acting as one neural layer that jointly contributes to the end-to-end inference process. IV-A2 CSIR/T Design Given m∈ℂNm×Nm−1H_m∈C^N_m×N_m-1, we first decompose the channel matrix m H_m by SVD, yielding m=mmmH H_m= U_m _m V_m^H, where m∈ℂNm×NmU_m∈C^N_m×N_m and m∈ℂNm−1×Nm−1V_m∈C^N_m-1×N_m-1 are unitary matrices, and m∈ℂNm×Nm−1 _m∈C^N_m×N_m-1 is a diagonal matrix whose singular values are real and sorted in a descending order. For a CSIR/T MIMO system, the CSI can be leveraged at the transmitter side, and the output signal at source R0R_0 is given by 0=PA(10c+0LT). X_0=PA (V_1F_0S_c+b_01_L^T ). (14) At relay R1R_1, the received signal is first processed by the combiner 1H U_1^H to decouple the spatial streams according to the SVD structure of 1 H_1. The resulting signal is then scaled by the pseudo-inverse of the singular-value matrix 1+ _1^+ to normalize the power across the eigenmodes. Subsequently, an amplification matrix 1 F_1 is applied to control the relay gain and adapt the transmitted power level. Finally, the processed signal is multiplied by the right singular matrix 2 V_2, which serves as a pre-processing operation aligned with the channel 2 H_2 toward the next hop. This process at R1R_1 can be written as 1=PA(211+1H1+1LT). X_1=PA (V_2F_1 _1^+U_1^HY_1+b_11_L^T ). (15) This sequential combination of receive combining, singular-value equalization, amplification, and transmit pre-processing can be extended to subsequent relays, yielding an end-to-end mapping whose parameters m,mm=1M \F_m,b_m \_m=1^M are jointly trained. IV-B Communication Design Based on Subsection IV-A, the signal received at the user after signal processing is given by M+1=M+1^M+M+1LT,CSIR,M+1M+1+M+1HM+1+M+1LT,CSIR/T, X_M+1= \ array[]*20lF_M+1 X_M+b_M+11_L^T, 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt 1.0pt CSIR,\\ F_M+1 _M+1^+U_M+1^HY_M+1+b_M+11_L^T, 1.0pt 1.0pt 1.0ptCSIR/T, array . (18) where M+1∈ℂNM+1×NMF_M+1∈C^N_M+1×N_M and M+1∈ℂNM+1×1 b_M+1∈C^N_M+1× 1 denote the combiner and bias at user, respectively, and ^M X_M denotes the estimated signals transmitted from RMR_M, which is similarly defined in (12). After obtaining the processed signal M+1 X_M+1, it is first converted into its real-valued representation by concatenating the real and imaginary parts. The resulting real-valued tensor is then flattened into a one-dimensional feature vector, which is fed into a task-specific read-out layer to produce the final inference output. For the image classification example, the read-out layer consists of an FC layer followed by a softmax activation producing the class probability vector p over the C classes. IV-C Training Loss For the image classification task, a standard cross-entropy loss is adopted to train the WPNN, expressed as ℒloss=−∑i=1Cpilog(p^i), L_ loss=-Σ _i=1^Cp_i ( p_i ), (19) where C denotes the number of classes, pip_i is the one-hot true label, and p^i p_i is the predicted probability of the iith class. For other tasks, ℒloss L_ loss would be replaced by the corresponding task-specific loss function without changing the underlying architecture or the training procedure. V Numerical results We evaluate the proposed multi-hop MIMO relay-based WPNN in terms of classification accuracy on the Fashion-MNIST dataset, which contains 60,000 training examples and 10,000 test examples across 10 categories, where each example is a 28×2828× 28 grayscale image. The path loss between the source node and the user is normalized to one. The relay nodes are uniformly placed along the line connecting the source node and the user, yielding αm=(M+1)2 _m= (M+1 )^2. The signal-to-noise ratio (SNR) is defined as SNR=10log101σ2 SNR=10lo g_10 1σ^2 in dB. Moreover, the average transmit power at the linear PA is set to 11 W. Unless specified otherwise, we set N0=28N_0=28, L=14L=14, and Nm=32N_m=32, m∈1,…,M+1m∈ \1,…,M+1 \. The proposed WPNN is trained in an end-to-end manner using the Adam optimizer with a learning rate of 10−410^-4 and a batch size of 64. During training, one independent channel realization is generated for each image transmission. The following schemes are considered: • Upper bound: This scheme employs an ideal digital neural network with the same layer dimensions and depth as the proposed WPNN. Each relay-associated physical layer is replaced by a perfect digital FC layer, without wireless-channel distortion. The nonlinear activation is implemented by the standard tanh(⋅) (·) function. • LS, LPA: This scheme corresponds to the CSIR case, requiring only receiver CSI for each hop. All PAs are modeled as ideal linear amplifiers. • LS, NPA: This scheme adopts the same LS-based transceiver architecture as “LS, LPA”, except that the nonlinear Rapp PA model is employed. • SVD, LPA: This scheme corresponds to the CSIR/T case, exploiting both transmitter- and receiver-side CSI for each hop. All PAs are modeled as ideal linear amplifiers. • SVD, NPA: This scheme adopts the same SVD-based transceiver architecture as “SVD, LPA”, except that the nonlinear Rapp PA model is employed. • Training-testing PA mismatch (TPM): The WPNN is trained assuming ideal linear PAs but is evaluated using nonlinear PAs, without retraining or fine-tuning. This mismatch setting is considered for both the LS- and SVD-based transceiver architectures. Figure 2: Classification accuracy versus SNR. Fig. 2 compares the classification accuracy of different transceiver schemes under varying SNRs for M=1M=1 relay and K=−∞K=-∞ (dB). We can observe that for SNRs above 0 dB, the LS scheme with nonlinear PA outperforms linear PA due to its superior expressiveness. In particular, at an SNR of 2020 dB, the performance of “LS, NPA” scheme is already close to the upper bound. For the SVD-based scheme, the linear PA consistently performs better since the nonlinearity prevents full eigenmode decoupling, leaving residual inter-stream interference. Moreover, the TPM schemes under both the LS- and SVD-based architectures exhibit clear performance losses relative to their nonlinear-PA counterparts. These losses become more pronounced as the SNR increases, since the inconsistency between the assumed and actual PA models becomes the dominant performance-limiting factor when noise is weak. These results illustrate that by appropriately tuning the hardware parameters at each node, the proposed WPNN can closely approximate the performance of a digital neural network. Figure 3: Classification accuracy versus number of relays M. Fig. 3 illustrates the impact of the number of relay nodes M on the classification accuracy for different transceiver schemes under SNR=20 SNR=20 dB and K=−∞K=-∞ (dB). For the “LS, NPA” scheme, the classification accuracy increases with M, from 0.8988 at M=1M=1 to 0.9258 at M=5M=5, and consistently outperforms its linear counterpart across all M. In contrast, the “SVD, NPA” scheme exhibits noticeable degradation as M increases, since the nonlinear PA distorts the SVD-based beamforming structure and such distortion accumulates across multiple hops. Also, for both linear PA schemes, the accuracy remains the same with increasing M due to the absence of nonlinear activation, which limits the expressive capability of the cascaded architecture. It is further observed that the accuracy of TPM schemes under both the LS- and SVD-based architectures diminishes with M. This is due to the accumulation of errors caused by PA-model mismatch over multiple hops. We further consider two transmit-power constraints for the TPM scheme. Specifically, the “LS, TPM (0.7 W)” scheme corresponds to the case in which the relays are trained with an average transmit power of 0.7 W. Interestingly, “LS, TPM (0.7 W)” outperforms “LS, TPM” because the lower training power keeps the PAs closer to their linear operating region, thereby reducing the mismatch between the linear PA model assumed during training and the nonlinear PA behavior encountered during testing. Figure 4: Classification accuracy versus Rician factor K. Fig. 4 illustrates the impact of the Rician factor K on the classification accuracy for M=1M=1 under SNR=−20dBSNR=-20~dB and SNR=10dBSNR=10~dB. It can be observed that the classification accuracy generally degrades as K increases for both LS- and SVD-based schemes. This is because lower-rank channels destroy the spatial degrees-of-freedom on which the WPNN relies. At SNR=10dBSNR=10~dB, the accuracy still decreases with K, but the drop is much less drastic since the noise is no longer the dominating impairment during per-hop equalization. For instance, even at K=20dBK=20~dB, the LS- and SVD-based schemes achieve accuracies of 0.69570.6957 and 0.77280.7728, respectively. VI Conclusion This paper has proposed a deep WPNN realized through a multi-hop MIMO relay network, in which each relay implements a trainable linear precoding stage followed by the intrinsic nonlinear activation of its PA. Cascading M such relays yields an over-the-air multi-layer FC network that unifies communication and computation over the same wireless infrastructure. Two transceiver designs were developed: an LS-based scheme requiring only receiver CSI, and an SVD-based scheme exploiting joint transmitter–receiver CSI. Three main findings emerged from our study: First, PA nonlinearity is a resource, rather than merely an impairment, for over-the-air computing. The LS scheme with a nonlinear PA monotonically improves with M. Second, this improvement is architecture-dependent: nonlinearity disrupts SVD eigenmode decoupling, so a linear-PA design is preferable for CSIR/T systems. Third, hardware-model mismatch and increasing channel rank deficiency (large Rician K) cause errors that compound over hops, motivating mismatch-aware training. Two natural extensions are: (i) more complex over-the-air architectures beyond FC layers, e.g., convolutional or attention-based WPNNs; and (i) robust, mismatch- and CSI-uncertainty-aware end-to-end training for deployment under realistic hardware imperfections. References [1] M. M. Amiri and D. Gündüz (2020-05) Federated learning over wireless fading channels. IEEE Trans. Wireless Commun. 19 (5), p. 3546–3557. Cited by: §I. [2] I. Bergel (2024-Dec.) Non-linear relay optimization using deep-learning tools. IEEE Trans. Wireless Commun. 23 (12), p. 19289–19301. Cited by: §I, §IV-A1. [3] C. Bian, M. Hua, and D. Gündüz (2025-Nov.) Over-the-air inference through analog computation over multi-hop MIMO networks. IEEE Wireless Commun. Lett. 14 (11), p. 3739–3743. Cited by: §I. [4] M. Hua, I. Bergel, T. Girici, M. Di Renzo, and D. Gunduz (2026) Wireless physical neural networks (WPNNs): opportunities and challenges. External Links: Link Cited by: §I. [5] M. Hua, C. Bian, H. Wu, and D. Gunduz (2025) Implementing neural networks over-the-air via reconfigurable intelligent surfaces. IEEE Trans. Wireless Commun. 25, p. 11562–11576. Cited by: §I. [6] M. Hua, H. Wu, and D. Gündüz (2026-Mar.) CNNs in the air via reconfigurable intelligent surfaces. IEEE Wireless Commun. Lett. 15, p. 2124–2128. Cited by: §I. [7] R. Iten, T. Metger, H. Wilming, L. del Rio, and R. Renner (2020-01) Discovering physical concepts with neural networks. Phys. Rev. Lett. 124, p. 010508. Cited by: §I. [8] C. Liu, Q. Ma, Z. J. Luo, Q. R. Hong, Q. Xiao, H. C. Zhang, L. Miao, W. M. Yu, Q. Cheng, L. Li, et al. (2022-Feb.) A programmable diffractive deep neural network based on a digital-coding metasurface array. Nat. Electron. 5 (2), p. 113–122. Cited by: §I. [9] M. Liu, J. An, C. Huang, and C. Yuen (2026) Over-the-air ODE-inspired neural network for dual task-oriented semantic communications. IEEE Trans. Cogn. Commun. Netw. 12, p. 805–819. Cited by: §I. [10] A. Momeni et al. (2025) Training of physical neural networks. Nature 645 (8079), p. 53–61. Cited by: §I. [11] G. Sanchez et al. (2023-Dec.) AirNN: over-the-air computation for neural networks via reconfigurable intelligent surfaces. IEEE/ACM Tran. Netw. 31 (6), p. 2470–2482. Cited by: §I. [12] K. Stylianopoulos, P. Di Lorenzo, and G. C. Alexandropoulos (2026-Mar.) Over-the-air edge inference via end-to-end metasurfaces-integrated artificial neural networks. IEEE Trans. Wireless Commun. 25, p. 13818–13834. Cited by: §I. [13] R. Wang, Y. Jiang, and W. Zhang (2022-Apr.) Distributed learning for MIMO relay networks. IEEE J. Sel. Topics Signal Process. 16 (3), p. 343–357. Cited by: §I, §IV-A1. [14] L. G. Wright, T. Onodera, M. M. Stein, T. Wang, D. T. Schachter, Z. Hu, and P. L. McMahon (2022) Deep physical neural networks trained with backpropagation. Nature 601 (7894), p. 549–555. Cited by: §I. [15] Y. Yang, Z. Zhang, Y. Tian, Z. Yang, R. Jin, L. Liu, and C. Huang (2024) Realizing over-the-air neural networks in RIS-assisted MIMO communication systems. In IEEE WCNC, Dubai, UAE, p. 1–5. Cited by: §I. [16] J. Zhang, H. Chen, and D. M. Blough (2024) A radio-frequency-based 2-D convolutional layer using transmissive intelligent surfaces. In IEEE VTC, Washington, DC, USA, p. 1–7. Cited by: §I.