Paper deep dive
A Coherence Law for Trainability in Noisy Equivariant Quantum Neural Networks
Hassan Ugail, Newton Howard
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 7/5/2026, 2:06:10 AM
Summary
The paper introduces a coherence law for the trainability of U(1)-equivariant quantum neural networks under noise. It identifies that while symmetry protects against barren plateaus, the survival of gradients depends on 'readout-visible aligned sector coherence'. The authors prove a light-cone reduction theorem showing that active gradients are confined to the backward light cone of the readout and possess a lower bound independent of system size. They define a training law where gradient degradation is governed by the Rayleigh quotient of the noise generator along the gradient-carrying mode. Density-matrix simulations confirm that this aligned coherence rate is a superior diagnostic for noisy trainability compared to standard channel diagnostics, particularly for correlated-dephasing channels.
Entities (6)
Relation Signals (3)
Backward light cone → confines → Active Gradients
confidence 100% · Causality fixes where the gradient can live, confining it to the backward light cone of the readout inside the active charge sector.
U(1)-equivariant brickwork circuits → isgovernedby → Readout-visible aligned coherence rate
confidence 95% · The law predicts no gradient loss for this channel, and none is seen. Sector coherence outperforms every standard channel diagnostic...
Correlated-dephasing channel → demonstrates → Readout-visible aligned coherence rate
confidence 90% · The sharpest test comes from a correlated-dephasing channel that has a large worst-case rate but a near-zero aligned rate.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Symmetry provides a quantum neural network structure, but on its own it does not keep the network trainable once noise is present. We ask which physical quantity decides whether the gradients of an equivariant circuit survive decoherence, and we answer with a compact training law. Working with U(1)-equivariant brickwork circuits that conserve a charge, we find that two distinct effects govern a trainable gradient. Causality fixes where the gradient can live, confining it to the backward light cone of the readout inside the active charge sector. Coherence then determines how fast it decays through the contraction of the off-diagonal sector modes that the projected readout can actually observe. We prove a light-cone reduction that pins the noiseless gradient to the sector-restricted cone with a lower bound independent of the total qubit number, and we define a readout-visible aligned coherence rate as a Rayleigh quotient of the noise generator along the gradient-carrying mode. A perturbative open-system analysis turns this rate into a leading-order training law. Density-matrix simulations then confirm that the finite-noise degradation follows a single accumulated variable built from noise depth and coherence contraction, with a coefficient of determination of 0.979. The sharpest test comes from a correlated-dephasing channel that has a large worst-case rate but a near-zero aligned rate. The law predicts no gradient loss for this channel, and none is seen. Sector coherence outperforms every standard channel diagnostic we compare it against, and the analysis identifies readout-visible sector coherence as the quantity that links equivariant architecture, open-system dynamics and noisy trainability.
Tags
Links
- Source: https://arxiv.org/abs/2606.30688v1
- Canonical: https://arxiv.org/abs/2606.30688v1
Trouble viewing inline? Open PDF directly →
Full Text
60,833 characters extracted from source content.
Expand or collapse full text
A Coherence Law for Trainability in Noisy Equivariant Quantum Neural Networks Hassan Ugail Centre for Visual Computing and Intelligent Systems University of Bradford United Kingdom &Newton Howard School of Individualized Study Rochester Institute of Technology United States Abstract Symmetry provides a quantum neural network structure, but on its own it does not keep the network trainable once noise is present. We ask which physical quantity decides whether the gradients of an equivariant circuit survive decoherence, and we answer with a compact training law. Working with U(1)U(1)-equivariant brickwork circuits that conserve a charge, we find that two distinct effects govern a trainable gradient. Causality fixes where the gradient can live, confining it to the backward light cone of the readout inside the active charge sector. Coherence then determines how fast it decays through the contraction of the off-diagonal sector modes that the projected readout can actually observe. We prove a light-cone reduction that pins the noiseless gradient to the sector-restricted cone with a lower bound independent of the total qubit number, and we define a readout-visible aligned coherence rate as a Rayleigh quotient of the noise generator along the gradient-carrying mode. A perturbative open-system analysis turns this rate into a leading-order training law. Density-matrix simulations then confirm that the finite-noise degradation follows a single accumulated variable built from noise depth and coherence contraction, with a coefficient of determination of 0.9790.979. The sharpest test comes from a correlated-dephasing channel that has a large worst-case rate but a near-zero aligned rate. The law predicts no gradient loss for this channel, and none is seen. Sector coherence outperforms every standard channel diagnostic we compare it against, and the analysis identifies readout-visible sector coherence as the quantity that links equivariant architecture, open-system dynamics and noisy trainability. Keywords quantum neural networks ⋅· equivariant quantum machine learning ⋅· trainability ⋅· quantum coherence ⋅· noisy variational circuits ⋅· quantum information 1 Introduction Variational quantum learning rests on a simple operational requirement. A parametrised circuit can be trained only if the gradients of its output stay measurable at a realistic sampling cost [1, 2, 3, 4]. Quantum neural networks and the wider family of variational quantum algorithms encode information into a quantum state, process it through trainable gates, and read out an expectation value whose parameter derivatives drive optimisation [5, 6, 7]. When those derivatives fall below the sampling floor, the model cannot be optimised at all, and so the survival of the gradient signal is the question that governs near-term quantum learning. Two distinct mechanisms can destroy this signal. One is the barren-plateau phenomenon, in which gradient second moments concentrate exponentially around zero as circuits grow wider, deeper, or overly expressive [8, 9, 10, 11]. The other is noise. Markovian decoherence acting throughout the circuit suppresses the gradient, drives noise-induced plateaus of its own, and caps the performance reachable by a noisy optimisation loop [12, 13, 14]. The two are physically separate. One concerns the geometry of the noiseless landscape; the other, the open-system dynamics layered on top of it; and a model can be safe from the first while remaining acutely vulnerable to the second. What survives the noise is therefore the operational trainability question, and the natural way in is to ask which physical quantity governs that survival. Symmetry has become one of the most productive ways to structure quantum neural networks against the first mechanism. Following the lead of classical geometric deep learning [15, 16], equivariant quantum architectures restrict the trainable gates to operations that commute with a symmetry group of the task [17, 18, 19, 20, 21]. Under a U(1)U(1) symmetry generated by the total charge, the Hilbert space splits into charge sectors, and an equivariant ansatz keeps a sector-supported input inside its sector throughout the noiseless evolution. Such circuits have reduced dynamical Lie algebras and can provably escape the exponential gradient concentration that afflicts generic deep circuits [11, 19, 22], while sharpening generalisation from few samples [23, 24]. A broader principle is at work here. Preserving the structure a model is built on is often what protects its useful signal, a pattern that recurs across machine learning. It appears in deep face recognition under degraded or imperfect data [25, 26], in transfer-learned visual attribution and interpretable prediction [27, 28], and in the dynamical modelling of organised complexity in physical systems [29]. For the present circuits, the charge sectors supply both the physical structure and the natural arena for analysing trainability. Symmetry preservation, however, is not noise protection. A channel can commute with the U(1)U(1) action and still erase the quantum information that carries the gradient. The reason is that the gradient of a sector-projected readout lives not in sector populations but in coherences, the off-diagonal matrix elements within the active charge sector. Single-qubit dephasing makes the point concrete. It is phase-covariant; it moves no population between sectors, and a benchmark reporting only covariance defects and sector leakage would deem it harmless. Yet it contracts the intra-sector off-diagonal elements at a rate set by its strength, eroding the very matrix elements from which the gradient is built [30, 31, 32]. Three questions about a noisy equivariant circuit therefore need to be kept apart. One asks whether the channel respects the symmetry, a question of covariance. Another asks whether the population stays in the active sector, a question of sector retention. The last asks whether the off-diagonal signal that supports the gradient survives, a question of sector coherence. For common hardware channels the three answers diverge [33], and it is the third that decides trainability. The central claim of the paper is as follows. Gradients of a variational circuit are local response functions, and what governs them divides cleanly in two. Their support is fixed by causality, because in a local brickwork circuit a parameter can reach the readout only from inside the backward light cone of the measured observable, restricted to the relevant charge sector. Their amplitude is carried by intra-sector off-diagonal coherence, because the commutator structure behind a parameter derivative lives on the off-diagonal block of the sector. The trainability resource for a noisy equivariant quantum neural network is therefore neither symmetry covariance nor sector population. It is readout-visible sector coherence, the coherence on the off-diagonal modes of the active sector that the projected readout can actually see, within the light cone the readout can actually reach. The training law is built around three entities. The first is a light-cone reduction for active gradients in U(1)U(1)-equivariant brickwork circuits. It shows that the active-gradient second moment is supported only inside the readout backward light cone, and that, under a non-degeneracy condition and a light-cone sector-weight condition on the fixed input, it carries a positive lower bound independent of the total qubit number at fixed depth and sector. The second is the readout-visible aligned coherence rate, a Rayleigh quotient of the noise generator taken along the gradient-carrying mode. We prove that this rate is bounded above by the worst-case scalar rate, with equality for restricted-isotropic noise, and from it we derive a perturbative coherence-loss law for phase-covariant Markovian noise. The third is empirical. Density-matrix simulations confirm that the finite-noise degradation obeys M∝(γL)1.002λcoh0.964D_M (γ L)^1.002 _coh^0.964 with R2≃0.979R^2 0.979, that sector coherence outperforms standard channel diagnostics, and that a correlated-dephasing control with a large worst-case rate but a near-zero aligned rate produces no measurable gradient loss. That last control is decisive evidence that the aligned rate, rather than the worst-case rate, is the operative quantity. It is worth fixing the scope before going further. The clean scalar law holds for the single-excitation sector under restricted-isotropic noise. Higher sectors and structured noise explicitly call for an aligned diagnostic, and coherent unitary miscalibration sits outside the dissipative family altogether. Within these limits, readout-visible aligned sector coherence is the quantity that bridges equivariant architecture, open-system dynamics and noisy trainability. 2 Methods 2.1 Model, readout and gradient observables We study a U(1)U(1)-equivariant quantum neural network on a cycle graph CnC_n with periodic boundary conditions. The conserved charge is N^=∑j(I−Zj)/2 N= _j(I-Z_j)/2, and the circuit acts on a fixed input that lies in a single charge sector r, with PrP_r the orthogonal projector onto that sector. The input is the localised computational-basis state ρ0=|ψ0⟩⟨ψ0| _0= _0 _0 with |ψ0⟩=|1r0n−r⟩ _0 = 1^r0^n-r . The variational circuit is a depth-L brickwork ansatz built from two gate families that commute with N N [18, 19, 17]. Each layer applies trainable single-qubit rotations Rz(θℓ,j)R_z( _ ,j) on every site, followed by nearest-neighbour XYXY hopping gates exp[−iβℓ,e(Xe1Xe2+Ye1Ye2)] [-i _ ,e(X_e_1X_e_2+Y_e_1Y_e_2)] on alternating edges of the cycle. The hopping parameters are fixed and treated as part of the architecture, so the only trainable degrees of freedom are the single-qubit rotation angles. This study examines how noise degrades the parameter gradients of a fixed-state equivariant response rather than learning from a distribution of encoded data. The state the circuit processes is held fixed, and the object of interest is the second moment of the readout gradient with respect to the trainable angles. This is the quantity whose survival determines whether the circuit can be optimised at all, and it is well defined independently of any data-encoding map. The intra-sector off-diagonal coherence that carries the gradient is generated dynamically by the XYXY hopping gates, which mix the in-sector basis states as the excitation propagates through the variational layers. The readout-visible coherence studied throughout the paper is therefore a property of the variational evolution acting on the fixed input, and the sector-supported state remains in the same charge sector throughout the noiseless circuit. The output is the sector-projected local readout, OB=PrZ0Pr,O_B=P_rZ_0P_r, (1) and the noisy response is, fθγ=Tr[OBΦθγ(ρ0)],f^γ_θ=Tr\! [O_B\, ^γ_θ( _0) ], (2) where Φθγ ^γ_θ denotes the layered circuit channel with single-qubit Markovian noise of rate γ applied once per layer [34, 35, 36]. Trainability is quantified by the prediction-gradient second moment, ℳ2(θi;γ)=θ[(∂θifθγ)2],M_2( _i;γ)=E_θ [ ( _ _if^γ_θ )^2 ], (3) with expectations taken over a small-box parameter prior of width δinit _init [37] and derivatives evaluated by the parameter-shift rule [4, 38, 39]. The central response variable is the relative degradation, M(γ)=log[ℳ2(θi;0)/ℳ2(θi;γ)],D_M(γ)= \! [M_2( _i;0)\,/\,M_2( _i;γ) ], (4) which measures on a logarithmic scale how much active gradient strength the noise has removed. 2.2 Active gradients and the backward light cone Before asking how noise damages a gradient, it helps to ask where a gradient can live at all. The parameter-shift identity expresses a derivative as a commutator response, ∂θifθ0=−i2Tr[OB(ℓi)[Zqi,ρℓiθ]], _ _if^0_θ=- i2\,Tr\! [O_B( _i)\,[Z_q_i,ρ^θ_ _i] ], (5) where OB(ℓi)O_B( _i) is the readout evolved backwards in the Heisenberg picture. Each brickwork layer propagates operator support by at most one bond, so OB(ℓi)O_B( _i) is supported on the backward light cone LC(OB)LC(O_B) of the readout qubit, and outside that cone the commutator vanishes by causal locality. Because every gate commutes with N N, the trace restricts further to the charge sectors within the cone. A theorem makes this precise, subject to two conditions on the fixed background. Assumption 1 (Light-cone algebra non-degeneracy). For the fixed hopping background β†β , the commutator between the light-cone-evolved readout operator and the active generator is not identically zero on the sector-restricted light-cone subspace. Assumption 2 (Light-cone sector weight). The fixed input places a non-vanishing charge weight inside the readout light cone, uniformly in system size. There is a constant w0>0w_0>0, independent of n at fixed L and r, such that the light-cone reduced state ρLC _LC satisfies Tr[Pq≥1LCρLC]≥w0Tr[P^LC_q≥ 1\, _LC]≥ w_0, where for the single-excitation sector Pq≥1LC=Pq=1LCP^LC_q≥ 1=P^LC_q=1 projects onto the one-excitation subspace of the cone. The condition excludes inputs in which the excitation delocalises so strongly that its weight inside a fixed-size cone vanishes as 1/n1/n. Theorem 1 (Light-cone reduction of active gradients). Let LC(OB)LC(O_B) be the backward light cone of the readout, let |LC||LC| be the number of qubits in it, and let dLC=∑q=0r(|LC|q)d_LC= _q=0^r |LC|q be the sector-restricted light-cone dimension. Active gradients are supported only inside the cone, since ∂θifθ0=0 _ _if^0_θ=0 whenever θi∉LC(OB) _i (O_B). Under Assumptions 1 and 2, for any active parameter θi∈LC(OB) _i (O_B) and the small-box prior of width δinit _init, ℳ2(θi;0)≥cLC(L,r,δinit,β†,OB,w0)>0,M_2( _i;0)≥ c_LC (L,r, _init,β ,O_B,w_0 )>0, (6) where cLCc_LC is independent of the total system size n at fixed L and r. The locality statement that gradients vanish outside the cone needs no assumption on the input. The n-independent lower bound is the stronger claim and rests on Assumption 2, which guarantees that the excitation keeps finite weight inside the fixed-size cone as n grows. The theorem frames everything that follows. Only the light-cone parameters carry signal, so the question becomes what noise does to the gradient amplitude inside the cone, and that amplitude is carried by intra-sector coherence. Proof of Theorem 1. We adopt the generator convention Rz(θ)=exp(−iθZ/2)R_z(θ)= (-iθ Z/2), so the half-angle generator is Z/2Z/2 and the parameter-shift rule uses shifts of π/2π/2. If parameter θi _i acts on qubit qiq_i in layer ℓi _i, the parameter-shift identity gives Equation (5). Parameters for which the commutator [Zqi,OB(ℓi)][Z_q_i,O_B( _i)] vanishes identically on the sector-restricted Hilbert space are classified as inactive and excluded from Theorem 1. The proof begins with the sector-restricted light-cone reduction. Each brickwork layer propagates operator support by at most one bond, so the Heisenberg-evolved readout OB(ℓi)O_B( _i) is supported on the backward light cone LC(OB)LC(O_B). Outside the cone the commutator vanishes by causal locality, and because all gates commute with N N the trace further restricts to the charge sectors q≤rq≤ r within the cone, whose dimension is dLC=∑q=0r(|LC|q)d_LC= _q=0^r |LC|q, replacing the full sector dimension (nr) nr. Next, the prior factorises over inactive coordinates. Writing θ=(θLC,θLC¯)θ=(θ^LC,θ LC) for parameters inside and outside the cone, the derivative for an active parameter depends only on θLCθ^LC. The small-box prior is a product measure, so the coordinates outside the cone integrate to unit mass and θ[|∂θifθ0|2]=∫[−δinit,δinit]pLC|∂θifθLC0|2dpLCθLC(2δinit)pLC,E_θ\! [| _ _if^0_θ|^2 ]= _[- _init, _init]^p_LC| _ _if^0_θ^LC|^2\, d^p_LCθ^LC(2 _init)^p_LC, (7) where pLCp_LC, the number of light-cone parameters, depends on the local circuit geometry but not on the total number of qubits at fixed L and r. By Assumption 1 the integrand is a continuous function of θLCθ^LC that is not identically zero, so there is an open set ULCU_LC on which |∂θifθLC0|2≥μ(β†)>0| _ _if^0_θ^LC|^2≥μ(β )>0. The in-cone signal μ(β†)μ(β ) is proportional to the active-sector light-cone weight Tr[Pq≥1LCρLC]Tr[P^LC_q≥ 1 _LC], so Assumption 2 keeps it from vanishing as 1/n1/n when the excitation delocalises. Combining the factorisation with the non-degeneracy and sector-weight floor gives ℳ2(θi;0)≥vol(ULC)(2δinit)pLCw0μ~(β†) . . =cLC(L,r,δinit,β†,OB,w0)>0,M_2( _i;0)≥ vol(U_LC)(2 _init)^p_LC\,w_0\, μ(β ) . .=c_LC(L,r, _init,β ,O_B,w_0)>0, (8) a constant independent of n at fixed L and r, where μ~ μ is the in-cone signal normalised by the sector weight. The bound is conditional on the fixed non-degenerate background β†β and on Assumption 2, and is not claimed to be uniform over degenerate backgrounds or maximally delocalised inputs. ∎ 2.3 The readout-visible aligned coherence rate To make the coherence dependence quantitative, we need an operator-level rate. Let Φγ=I+γℒ+O(γ2) ^γ=I+ +O(γ^2) be the small-noise expansion of the single-qubit channel and let Πoff _off project onto operators on the off-diagonal block of sector r in the eigenbasis of N N. The worst-case coherence-contraction rate is, λcoh . . =supA=ΠoffA‖A‖F=1−Re⟨A,Πoffℒ(A)⟩F, _coh . .= _ subarraycA= _offA\\ \|A\|_F=1 subarray-Re\, A, _offL(A) _F, (9) the largest decay rate of the restricted generator on the intra-sector off-diagonal subspace in the Frobenius inner product. This worst-case rate is convenient but not fundamental, because it maximises over all off-diagonal directions while the gradient occupies only one. The active gradient of parameter θi _i is generated by the commutator direction Gi(θ)=Πoff[Zqi,OB(ℓi;θ)]G_i(θ)= _off[Z_q_i,O_B( _i;θ)], the readout-visible mode through which the projected readout actually responds. The fundamental object is the readout-visible aligned rate, the ensemble-weighted decay rate of the restricted generator along this direction, λvis(θi) . . =θ[−Re⟨Gi(θ),Πoffℒ(Gi(θ))⟩F⟨Gi(θ),Gi(θ)⟩F]. _vis( _i) . .=E_θ\! [ -Re\, G_i(θ), _offL(G_i(θ)) _F G_i(θ),G_i(θ) _F ]. (10) By construction, the aligned rate is bounded above by the worst-case rate, since for every θ the supremum in Equation (9) runs over a set that contains the normalised gradient direction, so λvis(θi)≤λcoh _vis( _i)≤ _coh (11) for every active parameter. The two coincide when the restricted generator acts isotropically on the off-diagonal block, because every normalised direction then decays at the common rate. The distinction is invisible for an isotropic family but decisive for structured noise, where a channel can contract directions orthogonal to GiG_i and so carry a large λcoh _coh while leaving λvis _vis near zero. Applying the noise once per layer accumulates the aligned contraction into the layered coherence defect Δcohlay(γ)=γLλvis+O((γL)2) _coh^lay(γ)=γ L\, _vis+O((γ L)^2), which reduces to γLλcohγ L\, _coh in the restricted-isotropic case. For the four restricted-isotropic channels used in the main sweep, the worst-case and aligned rates agree to four significant figures, with λcoh=1.000 _coh=1.000 for amplitude damping, 3.9963.996 for dephasing, 6.6486.648 for depolarising and 7.9737.973 for the U(1)U(1)-breaking X-error control. Here, restricted-isotropic refers to the action on the intra-sector off-diagonal block rather than to ordinary channel isotropy. Theorem 2 (Perturbative coherence-loss law). For an active parameter θi _i and a phase-covariant Markovian noise channel of rate γ with Frobenius-dissipative restricted generator acting once per layer [40, 34, 35], and under the first-order alignment condition that the noisy derivative acquires no leading-order component along readout-visible directions orthogonal to GiG_i, the leading-order degradation is a sum of layer-resolved aligned contractions, ℳ2(θi;γ)=ℳ2(θi;0)[1−2γ∑ℓ=1Lλvis(ℓ)(θi)+O((γLλcoh)2)].M_2( _i;γ)=M_2( _i;0) [1-2\,γ _ =1^L _vis^( )( _i)+O ((γ L\, _coh)^2 ) ]. (12) In the restricted-isotropic family, where every off-diagonal direction decays at the common rate, the layer sum collapses to a scalar accumulation, ℳ2(θi;γ)=ℳ2(θi;0)[1−2γLλvis(θi)+O((γLλcoh)2)],M_2( _i;γ)=M_2( _i;0) [1-2\,γ L\, _vis( _i)+O ((γ L\, _coh)^2 ) ], (13) with the bounded alignment factor ωi=λvis(θi)/λcoh∈[0,1] _i= _vis( _i)/ _coh∈[0,1] equal to one for isotropic generators and strictly less than one when the gradient mode is misaligned with the fastest-decaying direction. Because three closely related alignment quantities appear in the paper, Table 1 collects their definitions. The theoretical ratio ωi=λvis/λcoh _i= _vis/ _coh is the object the theory identifies as fundamental, its perturbative empirical estimate ω ω is read off the small-noise experiments, and the finite-noise multiplier a a is an a posteriori diagnostic measured in the presence of finite noise. Table 1: The three alignment quantities used in the paper. The first is a theoretical ratio, the second its perturbative empirical estimate, and the third a finite-noise empirical multiplier. Symbol Definition Role ωi=λvis/λcoh _i= _vis/ _coh ratio of aligned to worst-case rate theoretical alignment factor, in [0,1][0,1] ω ω Δ~/(2γLλcoh) /(2γ L\, _coh) in the perturbative window empirical estimate of ωi _i at small noise a a finite-noise weight fitted to paired degradation a posteriori diagnostic of residual alignment Proof of Theorem 2. Let Φγ=I+γℒ+O(γ2) ^γ=I+ +O(γ^2) be the small-noise expansion of the layer channel and Πoff _off the projector onto the off-diagonal block of sector r. The worst-case rate is Equation (9) and the aligned rate is the Rayleigh quotient of Equation (10) with respect to the gradient mode GiG_i. Inserting the expansion into the commutator form of the derivative and projecting onto Gi(θ)=Πoff[Zqi,OB(ℓi;θ)]G_i(θ)= _off[Z_q_i,O_B( _i;θ)] gives, at each layer at which the noise acts, a first-order contraction of the gradient mode propagated to that layer at the pointwise rate λvis(ℓ)(θ) _vis^( )(θ). The single-layer contraction along the mode is the Rayleigh quotient by definition, and for a Frobenius-dissipative restricted generator, each such quotient is a non-negative spectral decay rate. Because the noise acts once per layer and the first-order terms add, the leading-order degradation is the layer sum 2γ∑ℓ=1Lλvis(ℓ)(θi)2γ _ =1^L _vis^( )( _i) of Equation (12) after the θ ensemble average. Writing the first-order noisy derivative as ∂θifθγ=∂θifθ0(1−γ∑ℓλvis(ℓ))+γδi _ _if^γ_θ= _ _if^0_θ(1-γ _ _vis^( ))+γ\, _i, where δi _i collects the first-order response along readout-visible directions orthogonal to GiG_i, the square produces a diagonal term and cross terms linear in δi _i. The first-order alignment condition is exactly the statement that the ensemble average of these cross terms vanishes at leading order, which holds identically in the restricted-isotropic family where δi=0 _i=0 because every off-diagonal direction decays at the common rate. The collapse to the scalar accumulation γLλvisγ L\, _vis of Equation (13) requires the aligned rate to be layer-independent, which holds in that family because the interleaved unitaries rotate the gradient mode without changing its decay rate. Outside it, the interleaved unitaries can carry the mode through directions of differing aligned rate, the layer-resolved rates no longer coincide, and the scalar λcoh _coh can no longer stand in for the layer-averaged aligned contraction. The alignment ratio ωi=λvis/λcoh _i= _vis/ _coh inherits the bounds [0,1][0,1] from Equation (11). ∎ 2.4 Noise channels and structural diagnostics The restricted-isotropic family comprises four single-qubit channels applied independently to every qubit once per layer, each parameterised so the per-layer rate γ enters the Kraus probabilities linearly at small noise. The term isotropic refers to the induced action on the intra-sector off-diagonal block in the r=1r=1 sector, on which the restricted generator contracts every normalised direction at the same rate so that λvis=λcoh _vis= _coh. Amplitude damping models energy relaxation, dephasing applies a Z flip and models transverse decoherence, depolarising replaces the qubit state by the maximally mixed state, and the X-error channel applies a bit flip and serves as the U(1)U(1)-breaking control. The structured family comprises inhomogeneous dephasing, site-dependent amplitude damping, biased Pauli noise, a coherent-dissipative mixed channel, and correlated two-site dephasing, the last built so its decaying coherence direction is orthogonal to the gradient mode, making it the zero-alignment control. The worst-case rates used in the regressions are λcoh=1.000 _coh=1.000 for amplitude damping, 3.9963.996 for dephasing, 6.6486.648 for depolarising, 7.9737.973 for the X-error control, 3.9963.996 for inhomogeneous dephasing, 1.0001.000 for site-dependent amplitude damping, 6.0006.000 for biased Pauli noise, 1.5001.500 for the coherent-dissipative mix, and 4.5004.500 for correlated dephasing. The aligned rate is evaluated directly from the generator only for the correlated-dephasing control, where the calculation confirms λvis≈0 _vis≈ 0 against λcoh=4.5 _coh=4.5. The comparison metrics, average gate infidelity 1−Favg1-F_avg, unitarity loss 1−u1-u, the diamond-distance upper bound, purity loss, and state off-diagonal loss, each enter the predictor comparison paired with log(γL) (γ L) exactly as the coherence rate does. 2.5 Numerical protocol and regression analysis All simulations evolve full density matrices, with charge-sector preservation verified to floating-point precision before noise is applied. Noise acts once per layer as a Markovian Kraus map [34, 35, 36, 41]. The small-box prior draws every trainable angle independently and uniformly from [−δinit,δinit][- _init, _init] with δinit=0.5 _init=0.5, while the hopping parameters remain fixed. Derivatives use the parameter-shift rule with shifts of π/2π/2. For the small-noise experiments the noiseless and noisy evaluations share common random numbers, which removes the dominant sampling variance from the ratio defining MD_M and is essential in the regime γLλcoh≪1γ L\, _coh 1. The main sweep uses the noise grid γL∈0.01,0.03,0.05,0.1,0.2,0.3γ L∈\0.01,0.03,0.05,0.1,0.2,0.3\ per channel-depth pair, the perturbative experiment uses γL∈0.001,0.003,0.005,0.01,0.02,0.03,0.05γ L∈\0.001,0.003,0.005,0.01,0.02,0.03,0.05\ with window cutoffs γLλcoh≤0.01,0.03,0.1,0.3,1.0γ L\, _coh≤\0.01,0.03,0.1,0.3,1.0\, and the size study spans n∈6,8,10n∈\6,8,10\ at L=3L=3, the largest size tractable for full-density-matrix evolution. Activity classification uses a noiseless preflight with a threshold of 10−1010^-10 on ℳ2M_2, far above the inactive floor near 10−3010^-30, which is the square of a double-precision gradient amplitude near 10−1510^-15. Bootstrap confidence intervals use 10001000 resamples. The statistical analysis fits ordinary least squares to logM _M, with the main model regressing on log(γL) (γ L) and logλcoh _coh and the reduced model on log(γL) (γ L) alone. Model comparison uses the Akaike information criterion, reported as differences relative to the best model. Robust uncertainty is assessed through HC0 standard errors, and a cluster bootstrap over channel-depth cells with 10001000 replicates, a mixed-effects model with random intercepts for channel and depth checks that the exponents are not group-structure artefacts, and a permutation test reassigns the channel rates to labels across 10001000 permutations. Cross-validation uses leave-one-channel-out, leave-one-depth-out and leave-one-noise-level-out folds. 3 Results 3.1 Active gradients localise to the light cone Figure 1 confirms the prediction numerically. The noiseless activity map log10ℳ2(θℓ,j;0) _10M_2( _ ,j;0) at n=8n=8 and r=1r=1 shows that active parameters sit inside the backward light cone of the readout, the active set grows with depth, and every parameter outside the cone rests at the numerical floor near 10−3010^-30. The red outline is the analytic light cone inferred from the brickwork causal structure, and it coincides cell for cell with the numerically active region. Across the maps the median active-to-inactive ratio of ℳ2M_2 exceeds 102710^27 at both depths, a structural separation rather than a faint signal lifted above a sampling floor. Layer ℓ=0 =0 stays inactive because the frozen initial state and the small-box prior make the first rotation layer act trivially on the gradient at leading order. Figure 1: Backward-light-cone localisation of active gradients. Numerically computed prediction-gradient second moment log10ℳ2(θℓ,j;0) _10M_2( _ ,j;0) for n=8n=8 qubits in charge sector r=1r=1 at depths L=3L=3 (a) and L=6L=6 (b). The red outline marks the analytic readout backward light cone inferred from the brickwork causal structure, a theoretical prediction rather than a contour fitted to the heatmap. Active parameters lie inside the cone while inactive parameters rest at the numerical floor near 10−3010^-30, giving a median active-to-inactive ratio above 102710^27. The L=6L=6 cone wraps around the cycle because of the periodic boundary. 3.2 The perturbative law holds at small noise Figure 2 validates the perturbative prediction using paired small-noise simulations at n=8n=8, L=4L=4 and r=1r=1, in which the noiseless and noisy second moments share common random numbers so that the ratio defining MD_M is far less sensitive to sampling-floor effects. The paired degradation Δ~=1−ℳ2(θi;γ)/ℳ2(θi;0) =1-M_2( _i;γ)/M_2( _i;0) is linear in γLγ L for all four channels, with per-channel R2R^2 above 0.9860.986. The inferred alignment estimate ω ω stays within [0,1][0,1] and approaches channel-specific limits, near one for amplitude damping and near one-half for dephasing. Pooled fits over increasingly tight perturbative windows recover exponents close to the target value of 1. Figure 2: Perturbative sector-coherence loss. Paired small-noise simulations at n=8n=8, L=4L=4, r=1r=1 for the four restricted-isotropic channels. (a) Per-channel linearity of Δ~=1−ℳ2(γ)/ℳ2(0) =1-M_2(γ)/M_2(0) in the noise-depth product γLγ L, with dashed small-noise tangents and per-channel R2R^2 above 0.9860.986. (b) The perturbative alignment estimate ω ω approaches channel-specific limits at 11 and 1/21/2 and remains in [0,1][0,1]. (c) Pooled exponents under tightening window cutoffs on γLλcohγ L\, _coh, with the exponent on log(γL) (γ L) contracting towards one and the coherence exponent between 0.890.89 and 0.950.95. 3.3 Finite-noise degradation follows accumulated coherence loss The perturbative law predicts the leading-order behaviour, but the operationally relevant regime extends to finite noise where higher-order terms are no longer negligible. The main sweep covers n=8n=8, r=1r=1, depths L∈3,4,5,6L∈\3,4,5,6\, the four restricted-isotropic channels and six noise levels per channel-depth pair, giving 9696 settings. Each setting is itself an ensemble estimate of MD_M with bootstrap confidence intervals, so the regression sample size of 9696 counts independent physical settings. The central regression is the log-linear model, logM=c+αlog(γL)+βlogλcoh+ϵ, _M=c+α (γ L)+β _coh+ε, (14) and on the pooled data it gives α=1.002α=1.002, β=0.964β=0.964 and R2=0.9789R^2=0.9789, so that, M∝(γL)1.002λcoh0.964.D_M (γ L)^1.002\, _coh^0.964. (15) Both exponents sit close to one, which is the accumulated form of the perturbative law specialised to the restricted-isotropic family, where λvis=λcoh _vis= _coh for every channel so that regressing on λcoh _coh is the same as regressing on the aligned rate. Because Theorem 2 is stated for phase-covariant noise, the primary theorem-supporting fit uses only the three symmetry-preserving channels, giving α=1.006α=1.006, β=0.889β=0.889 and R2=0.978R^2=0.978, statistically indistinguishable from the pooled result. When the symmetry-breaking X-error channel is held out and predicted from the three-channel fit, the held-out coefficient of determination is 0.9490.949, comparable to the within-family folds, which shows that it falls on the same scalar law without being part of the theorem-supporting regression. The finite-noise degradation is therefore organised, to good accuracy, by the single accumulated variable γLλcohγ L\, _coh, and Figure 3 displays the collapse. All four depths fall on a common line, so the relevant variable is the accumulated product γLγ L rather than depth itself. Adding an explicit logL L term changes nothing of substance, with a fitted coefficient of −0.031-0.031, a p-value of 0.710.71 and no improvement in R2R^2. Figure 3: Finite-noise accumulated coherence collapse. Gradient degradation MD_M against the coherence-weighted noise rate γLλcohγ L\, _coh for the main sweep at n=8n=8, r=1r=1, L∈3,4,5,6L∈\3,4,5,6\, four restricted-isotropic channels and six noise levels per channel-depth pair, giving n=96n=96 settings. (a) Points coloured by circuit depth. (b) The identical points coloured by the noise channel. The pooled fit gives M∝(γL)1.002λcoh0.964D_M (γ L)^1.002 _coh^0.964 with R2=0.979R^2=0.979. The dephasing offset band in (b) reflects residual channel structure absorbed by the alignment-weighted refinement. A fair reading of the headline fit asks how much of the explained variance is already captured by noise depth alone. The noise-depth-only regression on log(γL) (γ L) explains R2=0.670R^2=0.670. Adding the coherence rate raises R2R^2 to 0.9790.979 and reduces the root-mean-square error from 0.8100.810 to 0.2050.205. The unique contribution of the coherence rate is captured by the partial coefficient of determination, Rpartial2=Rfull2−RγL-only21−RγL-only2=0.936,R^2_partial= R^2_full-R^2_γ L-only1-R^2_γ L-only=0.936, (16) so after accounting for noise depth the sector-coherence rate explains roughly 94%94\% of the variance that noise depth alone leaves unresolved. The fit is stable under heteroscedasticity-robust standard errors, a cluster bootstrap over channel-depth cells, mixed-effects modelling with random intercepts for channel and depth, and leave-one-depth-out and leave-one-noise-level-out validation, with held-out R2R^2 between 0.930.93 and 0.990.99. 3.4 Sector coherence outperforms standard diagnostics The coherence rate would be of limited interest if any reasonable channel-strength metric organised the same data equally well. Figure 4 compares the main model against the standard alternatives, each pairing log(γL) (γ L) with one channel metric on the identical pooled response. Pairing γLγ L with λcoh _coh gives R2=0.979R^2=0.979, against 0.9610.961 for purity loss, 0.9490.949 for state off-diagonal loss, roughly 0.8080.808 for average infidelity, unitarity loss and the diamond-distance proxy, and 0.6700.670 for noise depth alone. The same ordering appears in the Akaike information criterion, where the coherence-rate model improves on the gate-level metrics by more than 210210 units. The advantage is not that the model carries an extra predictor. Average infidelity, unitarity, purity and the generic off-diagonal loss of the full state all average over degrees of freedom that play no part in the active-gradient response, whereas the coherence rate is the contraction rate on the intra-sector off-diagonal subspace through which the projected readout actually responds. A diagnostic matched to the response-carrying subspace outperforms diagnostics that dilute the same physics across the whole space, and the result survives a permutation null with one-sided p≃0.045p 0.045. Figure 4: Sector coherence outperforms standard noise diagnostics. Predictor comparison for the logM _M regression on the pooled dataset with n=96n=96 settings. (a) Coefficient of determination R2R^2 by predictor set, each alternative pairing log(γL) (γ L) with one channel metric, with the coherence-rate model highlighted. (b) The difference in the Akaike information criterion from the best model, with lower being better. The coherence-rate model improves on the gate-level metrics by more than 210210 AIC units and on the noise-depth baseline by more than 260260. 3.5 Structured noise reveals the aligned rate The restricted-isotropic family is the reference case, because its restricted generators contract every off-diagonal direction at the same rate. Structured noise breaks this degeneracy and tests whether the bare scalar λcoh _coh is still the right object. The structured study uses five anisotropic channels at n=8n=8, L=3L=3 and r=1r=1. These are inhomogeneous dephasing, site-dependent amplitude damping, biased Pauli noise, a coherent-dissipative mixed channel, and correlated two-site dephasing. Figure 5 presents the analysis. Four of the five structured channels follow the same power law in γLγ L as the restricted-isotropic family, each along its own coherence-weighted line. The fifth, correlated dephasing, is the decisive control. It is built so that its decaying coherence direction runs orthogonal to the gradient-carrying mode Gi(θ)G_i(θ) across the ensemble, which sends the aligned rate to λvis≈0 _vis≈ 0 even as the worst-case rate stays large at λcoh=4.5 _coh=4.5. The aligned theory predicts negligible degradation in advance, and the simulation bears this out, with the measured degradation pinned at the numerical floor, |M|<3×10−14|D_M|<3× 10^-14 at every noise level. This is a direct test of the inequality λvis≤λcoh _vis≤ _coh at its extreme, where a channel saturates a large worst-case rate while leaving the readout-visible rate near zero. As a zero-alignment control, it shows that the bare worst-case rate cannot, on its own, be the fundamental variable. Pooling the isotropic and anisotropic data with the control left out of the fit yields α=1.001α=1.001, β=1.040β=1.040 and R2=0.964R^2=0.964 across 126126 settings, with noticeably more residual scatter than the isotropic collapse. The drift of the fitted exponent from the isotropic 0.9640.964 towards 1.0401.040, together with the wider scatter, is the signature one expects when the regression is written in λcoh _coh but the response is governed by λvis _vis. The structured channels carry λvis<λcoh _vis< _coh by varying amounts that the scalar regressor cannot resolve. The finite-noise empirical multiplier a a, an a posteriori estimate of ωi=λvis/λcoh _i= _vis/ _coh measured directly from the paired degradation data and bounded by one across all settings, absorbs that gap. Weighting the rate by it raises the main pooled regression from R2=0.979R^2=0.979 to 0.9910.991 and cuts the root-mean-square error of logM _M from 0.2050.205 to 0.1310.131. Because a a is read off the response, it is a post hoc diagnostic rather than an independent predictor, so the decisive evidence for the aligned rate remains the zero-alignment control. In matched settings, the scalar coherence model gives weaker organisation outside the single-excitation sector, with R2R^2 of 0.710.71 at r=2r=2 and 0.890.89 at r=3r=3, so the scalar law is treated as a single-excitation, restricted-isotropic result, with direct evaluation of the aligned rate the appropriate route for higher sectors. Figure 5: Structured noise and readout-visible aligned coherence. Anisotropic-noise study at n=8n=8, L=3L=3, r=1r=1, pooled with the isotropic dataset. (a) Four structured channels follow the same power law in γLγ L, each along its own coherence-weighted line. (b) Correlated dephasing shown separately as a zero-alignment control. Despite a non-zero worst-case rate λcoh=4.5 _coh=4.5, its aligned rate satisfies λvis≈0 _vis≈ 0, so the measured degradation satisfies |M|<3×10−14|D_M|<3× 10^-14 at every noise level. (c) Pooled isotropic and anisotropic collapse with the control excluded, giving α=1.001α=1.001, β=1.040β=1.040 and R2=0.964R^2=0.964 over n=126n=126 settings. (d) Weighting the rate by the finite-noise multiplier a a improves the regression from R2=0.979R^2=0.979 to 0.9910.991 and reduces the root-mean-square error of logM _M from 0.2050.205 to 0.1310.131. 3.6 Size and topology checks Theorem 1 predicts that at fixed depth and charge sector, the degradation law is insensitive to the boundary conditions of the graph and to the total number of qubits. Both predictions hold. Replacing the cycle with an open chain at n=8n=8 leaves the per-channel degradation unchanged, and a topology-pooled regression across 4545 settings yields α=1.016α=1.016, β=0.943β=0.943 and R2=0.991R^2=0.991. A size check over n∈6,8,10n∈\6,8,10\ at L=3L=3 collapses onto the common line with α=1.007α=1.007, β=0.909β=0.909 and R2=0.983R^2=0.983, and an explicit logn n term yields a negligible size exponent η≃0.043η 0.043. The size sweep also bears on Assumption 2, since the noiseless active second moment shows no 1/n1/n suppression across the three sizes, remaining of order 10−510^-5 to 10−410^-4, consistent with the fixed input keeping finite light-cone sector weight as n grows. The degradation is exactly n-invariant for amplitude damping and dephasing, while the only size trend in depolarising noise is explained by its worst-case rate rising with n, evaluated per size, from λcoh=5.32 _coh=5.32 at n=6n=6 to 7.977.97 at n=10n=10. 4 Discussion Equivariance determines where gradients can live, and readout-visible sector coherence determines whether those gradients survive noise. The light-cone reduction pins the support of the active response to the sector-restricted backward light cone of the projected readout and keeps its noiseless second moment away from zero independently of system size. The perturbative law then names the off-diagonal mode GiG_i that carries the amplitude of that response, together with the aligned rate λvis _vis that controls its contraction. The finite-noise simulations show that this mechanism stays predictive well beyond the first-order window, with near-linear exponents in both accumulated noise depth and sector-coherence contraction. Nothing in the finite-noise regime was guaranteed by the first-order expansion, so the persistence of the near-linear law is a genuine empirical finding about this model family rather than a corollary of the theory. The same reasoning explains why standard channel diagnostics underperform. Average infidelity, unitarity loss, purity loss, and diamond-distance proxies each measure a legitimate notion of channel disturbance, but they average over degrees of freedom that need not participate in the active-gradient response. The sector-coherence rate is tied to the intra-sector off-diagonal block through which the projected readout responds, so its predictive advantage comes from a physical match between the diagnostic and the response subspace rather than from a statistical artefact of fitting. The point generalises beyond this model. A trainability diagnostic should be matched to the operator subspace of the response it is meant to predict. Structured noise shows where the scalar diagnostic must be refined. Correlated dephasing has a nonzero worst-case contraction yet produces no measurable gradient loss because its decaying coherence direction is orthogonal to the gradient mode, and its aligned rate is therefore near zero. This zero-alignment control is the decisive evidence that the aligned rate, not the bare worst-case rate, is the operator-level object that controls degradation, with the inequality λvis≤λcoh _vis≤ _coh becoming an equality only for restricted-isotropic channels. The definition of λvis _vis as a Rayleigh quotient supplies the route to an independent operator-level metric, and confirming that the directly computed λvis _vis reproduces the empirical alignment across the structured family is the clearest next step, since the naive generator-spectrum estimator misassigns alignment for correlated dephasing and a correct estimator needs the readout projection built in. Two practical consequences follow for noisy equivariant quantum models. One is a matter of benchmarking. Reporting symmetry covariance or sector leakage is not enough, since a channel can preserve both while erasing the coherence that carries the gradient, so a useful benchmark should also report how much sector coherence survives within the readout light cone. The other is a matter of design. Given a characterised noise model, hardware-aware ansatz design can weigh the readout light cone, the active sector, and the noise directions that contract the corresponding coherence modes, and so place the gradient-carrying coherence in the most slowly contracting directions on offer. A few limitations bound what has been shown. The study is scoped to the single-excitation U(1)U(1) sector and the restricted-isotropic family, which is the regime where the aligned rate reduces to the worst-case scalar. Higher sectors sit at the edge of that regime and show weaker single-scalar organisation, consistent with their richer off-diagonal structure needing the aligned diagnostic in place of the worst-case scalar. Full density-matrix simulation limits the accessible system sizes, with the size check reaching n=10n=10. Coherent unitary miscalibration, which rotates rather than contracts off-diagonal modes, calls for a different treatment. Experimental validation of the full coherence-rate law on hardware is left to future work, and no hardware results are reported here. 5 Conclusion We have set out a training law for noisy equivariant quantum neural networks that ties the survival of trainable gradients to two physical effects, causality and coherence. Causality confines the active gradient to the sector-restricted backward light cone of the projected readout. A reduction theorem makes this confinement quantitative, with a lower bound independent of the total qubit number, so that adding idle qubits outside the cone cannot dilute the signal. Coherence then sets the rate at which the gradient fades, through the readout-visible aligned contraction of the off-diagonal sector mode the readout can see. A perturbative open-system analysis turns this into a leading-order law, and density-matrix simulations confirm that it holds well into the finite-noise regime as M∝(γL)1.002λcoh0.964D_M (γ L)^1.002 _coh^0.964 with R2≃0.979R^2 0.979. The single most informative result is the negative control. A correlated-dephasing channel with a large worst-case contraction but a near-zero aligned rate produces no measurable gradient loss, exactly as the aligned theory predicts in advance, and against what any whole-channel strength measure would suggest. Together with the systematic advantage of sector coherence over every standard channel diagnostic, this control identifies readout-visible, aligned sector coherence as the operator-level quantity governing noisy trainability, in place of symmetry covariance or sector population. The wider lesson is that a trainability diagnostic should be matched to the operator subspace of the response it predicts. Within the single-excitation, restricted-isotropic regime that matching yields a clean and reproducible law, and the Rayleigh-quotient definition of the aligned rate marks the route towards higher sectors, structured noise, and eventually hardware, where the same light-cone and coherence structure should continue to organise how gradients survive under noise. Acknowledgments The authors acknowledge the computational resources provided by the Centre for Visual Computing and Intelligent Systems at the University of Bradford. Data availability The complete codebase and the raw and processed data supporting the findings of this study, including all CSV files, are available in the GitHub repository at https://github.com/ugail/Readout-Visible-Coherence-QNN. The archive contains the complete code along with the processed data required to reproduce all results reported here. Funding No funding was received for this work. Competing interests The authors declare no competing interests. References [1] J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, “Quantum machine learning,” Nature, vol. 549, p. 195–202, 2017. https://doi.org/10.1038/nature23474 [2] J. Preskill, “Quantum computing in the NISQ era and beyond,” Quantum, vol. 2, art. 79, 2018. https://doi.org/10.22331/q-2018-08-06-79 [3] M. Cerezo et al., “Variational quantum algorithms,” Nature Reviews Physics, vol. 3, p. 625–644, 2021. https://doi.org/10.1038/s42254-021-00348-9 [4] K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, “Quantum circuit learning,” Physical Review A, vol. 98, art. 032309, 2018. https://doi.org/10.1103/PhysRevA.98.032309 [5] V. Havlíček et al., “Supervised learning with quantum-enhanced feature spaces,” Nature, vol. 567, p. 209–212, 2019. https://doi.org/10.1038/s41586-019-0980-2 [6] M. Schuld, R. Sweke, and J. J. Meyer, “Effect of data encoding on the expressive power of variational quantum machine learning models,” Physical Review A, vol. 103, art. 032430, 2021. https://doi.org/10.1103/PhysRevA.103.032430 [7] A. Abbas et al., “The power of quantum neural networks,” Nature Computational Science, vol. 1, p. 403–409, 2021. https://doi.org/10.1038/s43588-021-00084-1 [8] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, “Barren plateaus in quantum neural network training landscapes,” Nature Communications, vol. 9, art. 4812, 2018. https://doi.org/10.1038/s41467-018-07090-4 [9] M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, “Cost function dependent barren plateaus in shallow parametrized quantum circuits,” Nature Communications, vol. 12, art. 1791, 2021. https://doi.org/10.1038/s41467-021-21728-w [10] Z. Holmes, K. Sharma, M. Cerezo, and P. J. Coles, “Connecting ansatz expressibility to gradient magnitudes and barren plateaus,” PRX Quantum, vol. 3, art. 010313, 2022. https://doi.org/10.1103/PRXQuantum.3.010313 [11] M. Ragone et al., “A Lie algebraic theory of barren plateaus for deep parameterized quantum circuits,” Nature Communications, vol. 15, art. 7172, 2024. https://doi.org/10.1038/s41467-024-49909-3 [12] S. Wang et al., “Noise-induced barren plateaus in variational quantum algorithms,” Nature Communications, vol. 12, art. 6961, 2021. https://doi.org/10.1038/s41467-021-27045-6 [13] D. Stilck França and R. García-Patrón, “Limitations of optimization algorithms on noisy quantum devices,” Nature Physics, vol. 17, p. 1221–1227, 2021. https://doi.org/10.1038/s41567-021-01356-3 [14] K. Sharma, S. Khatri, M. Cerezo, and P. J. Coles, “Noise resilience of variational quantum compiling,” New Journal of Physics, vol. 22, art. 043006, 2020. https://doi.org/10.1088/1367-2630/ab784c [15] T. Cohen and M. Welling, “Group equivariant convolutional networks,” Proceedings of the 33rd International Conference on Machine Learning, p. 2990–2999, 2016. https://proceedings.mlr.press/v48/cohenc16.html. arXiv DOI: https://doi.org/10.48550/arXiv.1602.07576 [16] M. M. Bronstein, J. Bruna, T. Cohen, and P. Veličkovi’c, “Geometric deep learning: grids, groups, graphs, geodesics, and gauges,” arXiv:2104.13478, 2021. https://doi.org/10.48550/arXiv.2104.13478 [17] M. Larocca et al., “Group-invariant quantum machine learning,” PRX Quantum, vol. 3, art. 030341, 2022. https://doi.org/10.1103/PRXQuantum.3.030341 [18] Q. T. Nguyen et al., “Theory for equivariant quantum neural networks,” PRX Quantum, vol. 5, art. 020328, 2024. https://doi.org/10.1103/PRXQuantum.5.020328 [19] L. Schätzki, M. Larocca, Q. T. Nguyen, F. Sauvage, and M. Cerezo, “Theoretical guarantees for permutation-equivariant quantum neural networks,” npj Quantum Information, vol. 10, art. 12, 2024. https://doi.org/10.1038/s41534-024-00804-1 [20] J. J. Meyer et al., “Exploiting symmetry in variational quantum machine learning,” PRX Quantum, vol. 4, art. 010328, 2023. https://doi.org/10.1103/PRXQuantum.4.010328 [21] H. Ugail and N. Howard, “Symmetry-organised complexity in quantum neural networks,” Symmetry, vol. 18, no. 6, art. 912, 2026. https://doi.org/10.3390/sym18060912 [22] A. Pesah et al., “Absence of barren plateaus in quantum convolutional neural networks,” Physical Review X, vol. 11, art. 041011, 2021. https://doi.org/10.1103/PhysRevX.11.041011 [23] M. C. Caro et al., “Generalization in quantum machine learning from few training data,” Nature Communications, vol. 13, art. 4919, 2022. https://doi.org/10.1038/s41467-022-32550-3 [24] S. Jerbi et al., “Quantum machine learning beyond kernel methods,” Nature Communications, vol. 14, art. 517, 2023. https://doi.org/10.1038/s41467-023-36159-y [25] H. Ugail, H. M. Alawar, A. A. Zehi, A. M. Alkendi, and I. L. Jaleel, “Evaluation of latent diffusion enhanced face recognition under forensic image degradations,” Discover Computing, vol. 29, art. 193, 2026. https://doi.org/10.1007/s10791-026-10082-4 [26] A. Elmahmudi and H. Ugail, “Deep face recognition using imperfect facial data,” Future Generation Computer Systems, vol. 99, p. 213–225, 2019. https://doi.org/10.1016/j.future.2019.04.025 [27] H. Ugail, D. G. Stork, H. G. M. Edwards, S. C. Seward, and C. Brooke, “Deep transfer learning for visual analysis and attribution of paintings by Raphael,” Heritage Science, vol. 11, article 268, 2023. https://doi.org/10.1186/s40494-023-01094-0 [28] A. A. Ibrahim, N. H. Ugail, and H. Ugail, “Is facial beauty in the eyes? A multi-method approach to interpreting facial beauty prediction in machine learning models,” Discover Artificial Intelligence, vol. 5, p. 16, 2025. https://doi.org/10.1007/s44163-025-00226-8 [29] H. Ugail and N. Howard, “Quantifying the dynamics of consciousness using hierarchical integration, organised complexity and metastability,” arXiv:2512.10972, Dec. 2025. https://doi.org/10.48550/arXiv.2512.10972 [30] I. Marvian and R. W. Spekkens, “Extending Noether’s theorem by quantifying the asymmetry of quantum states,” Nature Communications, vol. 5, art. 3821, 2014. https://doi.org/10.1038/ncomms4821 [31] M. Piani et al., “Robustness of asymmetry and coherence of quantum states,” Physical Review A, vol. 93, art. 042107, 2016. https://doi.org/10.1103/PhysRevA.93.042107 [32] T. Baumgratz, M. Cramer, and M. B. Plenio, “Quantifying coherence,” Physical Review Letters, vol. 113, art. 140401, 2014. https://doi.org/10.1103/PhysRevLett.113.140401 [33] H. Ugail and N. Howard, “A channel-level diagnostic for symmetry breaking in noisy equivariant quantum neural networks,” IEEE Access, vol. 14, 2026. https://doi.org/10.1109/ACCESS.2026.3706394 [34] G. Lindblad, “On the generators of quantum dynamical semigroups,” Communications in Mathematical Physics, vol. 48, p. 119–130, 1976. https://doi.org/10.1007/BF01608499 [35] V. Gorini, A. Kossakowski, and E. C. G. Sudarshan, “Completely positive dynamical semigroups of N-level systems,” Journal of Mathematical Physics, vol. 17, p. 821–825, 1976. https://doi.org/10.1063/1.522979 [36] H.-P. Breuer and F. Petruccione, The Theory of Open Quantum Systems. Oxford University Press, Oxford, 2002. https://doi.org/10.1093/acprof:oso/9780199213900.001.0001 [37] E. Grant, L. Wossnig, M. Ostaszewski, and M. Benedetti, “An initialization strategy for addressing barren plateaus in parametrized quantum circuits,” Quantum, vol. 3, art. 214, 2019. https://doi.org/10.22331/q-2019-12-09-214 [38] M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Killoran, “Evaluating analytic gradients on quantum hardware,” Physical Review A, vol. 99, art. 032331, 2019. https://doi.org/10.1103/PhysRevA.99.032331 [39] D. Wierichs, J. Izaac, C. Wang, and C. Y.-Y. Lin, “General parameter-shift rules for quantum gradients,” Quantum, vol. 6, art. 677, 2022. https://doi.org/10.22331/q-2022-03-30-677 [40] A. S. Holevo, “Covariant quantum Markovian evolutions,” Journal of Mathematical Physics, vol. 37, p. 1812–1832, 1996. https://doi.org/10.1063/1.531481 [41] M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information. Cambridge University Press, Cambridge, 2010. https://doi.org/10.1017/CBO9780511976667