Paper deep dive
Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support
Peiyong Wang, Udaya Parampalli, Casey R. Myers
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:A central design principle in modern machine learning and artificial intelligence is to align a model's inductive bias with the structure of its input data. For matrix-valued inputs, relevant matrix-level relationships can be characterised through spectral values and spectral subspaces; however, common coordinate-wise rotation-gate data-encoding unitaries used in most quantum machine learning models do not explicitly construct such a matrix-level representation. We introduce Quantum Spectral Models (QSMs), in which we construct the generator of the data-encoding unitary directly from each input matrix. We study three QSM variants based on symmetric, global block, and non-overlapping patch-local block Hamiltonians. Their outputs admit truncated Fourier representations in which input-dependent spectral gaps supply candidate phase carriers, while spectral subspaces help determine their coefficients. We evaluate the QSMs and comparison quantum models on two matrix representations of Pendigits and two controlled synthetic tasks defined by spectral statistics. At the largest evaluated circuit depth, QSM variants lead the tested quantum models in mean test accuracy across all four benchmarks. The patch-local QSM leads on Pendigits, whereas the global block-Hamiltonian QSM leads on the controlled spectral tasks. Ablations show a task-dependent reversal: subspace-preserving controls perform better on Pendigits, whereas spectral-value-only controls lead among the tested ablations on the synthetic tasks. Together, these results shed new light on quantum machine-learning model design by showing how input-conditioned spectral representations can provide an analysable inductive bias, while offering a broader perspective on structure-aware model design in machine learning and artificial intelligence.
Tags
Links
- Source: https://arxiv.org/abs/2607.22516v1
- Canonical: https://arxiv.org/abs/2607.22516v1
Trouble viewing inline? Open PDF directly →
Full Text
169,549 characters extracted from source content.
Expand or collapse full text
Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support Peiyong Wang CSIRO Technology Research Way, Clayton VIC 3168, Australia Peiyong.Wang at csiro.au Udaya Parampalli School of Computing and Information Systems The University of Melbourne Grattan Street, Parkville, VIC 3010, Australia udaya at unimelb.edu.au Casey R. Myers School of Physics, Mathematics and Computing The University of Western Australia 35 Stirling Hwy, Crawley WA, 6009, Australia casey.myers at uwa.edu.au Corresponding author. Also available at addwater0315 at gmail.comAlso with Pawsey Supercomputing Centre, 1 Bryce Avenue, Kensington WA, 6151, Australia Abstract A central design principle in modern machine learning and artificial intelligence is to align a model's inductive bias with the structure of its input data. For matrix-valued inputs, relevant matrix-level relationships can be characterised through spectral values and spectral subspaces; however, common coordinate-wise rotation-gate data-encoding unitaries used in most quantum machine learning models do not explicitly construct such a matrix-level representation. We introduce Quantum Spectral Models (QSMs), in which we construct the generator of the data-encoding unitary directly from each input matrix. We study three QSM variants based on symmetric, global block, and non-overlapping patch-local block Hamiltonians. Their outputs admit truncated Fourier representations in which input-dependent spectral gaps supply candidate phase carriers, while spectral subspaces help determine their coefficients. We evaluate the QSMs and comparison quantum models on two matrix representations of Pendigits and two controlled synthetic tasks defined by spectral statistics. At the largest evaluated circuit depth, QSM variants lead the tested quantum models in mean test accuracy across all four benchmarks. The patch-local QSM leads on Pendigits, whereas the global block-Hamiltonian QSM leads on the controlled spectral tasks. Ablations show a task-dependent reversal: subspace-preserving controls perform better on Pendigits, whereas spectral-value-only controls lead among the tested ablations on the synthetic tasks. Together, these results shed new light on quantum machine-learning model design by showing how input-conditioned spectral representations can provide an analysable inductive bias, while offering a broader perspective on structure-aware model design in machine learning and artificial intelligence. 1 Introduction Figure 1: Synthesiser view of the quantum spectral model. Top, conceptual signal flow in a music synthesiser; bottom, the corresponding flow in a quantum spectral model. Reading the panels from left to right, the control signal corresponds to the input matrix, the oscillator bank to the input-derived Hamiltonian, the available tones to its spectrum and spectral gaps, the mixer and filter to the trainable quantum block, and the output sound to quantum measurement and prediction. A matrix input M defines a Hamiltonian H(M)H(M) whose spectrum supplies sample-conditioned candidate phase carriers and whose spectral projectors help determine their coefficients through the trainable mixer and quantum measurement. Unlike a coordinate-wise rotation-gate data-encoding unitary with globally determined candidate frequency support, the quantum spectral model conditions both its candidate carriers and coefficient structure on each input matrix. The design of modern deep neural networks often reflects inductive biases about the structure of their input data. Convolutional neural networks encode spatial locality [27], while graph neural networks encode relational structure [3]. Decoder-only Transformers impose an autoregressive ordering through masked self-attention, preventing each token from accessing future tokens during language modelling [42]. Operator-learning architectures, meanwhile, learn maps between function spaces rather than between fixed-dimensional vectors [24]. The same representational decision arises in quantum machine learning before training begins: classical data must first be mapped to data-dependent quantum operations [46]. This encoding specifies the quantum representation on which the trainable circuit acts. Quantum models represent data through unitary evolution, whose dependence on a Hermitian generator can be understood through its spectral decomposition. Data reuploading circuits alternate data-encoding unitaries with trainable quantum layers [41]. A quantum Fourier model (QFM) is a data-reuploading model whose observable outputs can be analysed as truncated Fourier series: the spectra of the data-encoding generators determine the candidate frequencies, while the trainable circuit controls their coefficients and interference at measurement [46]. From this perspective, encoder design determines which frequencies, modes or basis functions are made available to the model. Common data-encoding unitaries in reuploading circuits inject individual scalar features through single-qubit rotation gates [46]. Trainable-frequency rotation-gate data-encoding unitaries can adjust their globally shared candidate frequency support during optimisation [21]. Once trained, that support remains shared across samples. These encoders process matrix entries through coordinate-wise rotations rather than explicitly constructing a generator from matrix-level spectral values, spectral subspaces, trajectory order or local organisation. In this work, we instead construct the generator of the data-encoding unitary directly from each input matrix. We call the resulting model class a Quantum Spectral Model (QSM). A QSM maps a matrix-structured input to a Hamiltonian and uses the corresponding time evolution as the data-encoding unitary. The resulting candidate frequency support is therefore conditioned on the input matrix rather than determined solely by the generator family or globally learned frequency parameters. In the synthesiser analogy in Figure 1, a conventional rotation-gate data-encoding unitary provides a fixed oscillator bank, whereas a QSM allows each input to determine the available components; the trainable mixer then combines these components before measurement. Hamiltonian embedding changes both the candidate support and the coefficients of the resulting truncated Fourier series. We refer to each application of a data-encoding unitary as an upload. The one-upload candidate support comprises eigengaps of the symmetrised input matrix for the symmetric-Hamiltonian QSM and signed singular-value gaps, including gaps involving zero modes, for the block-Hamiltonian QSM. In both cases, the candidate support varies across samples, while the Fourier coefficients depend on input-derived spectral projectors or subspaces together with the initial state, trainable mixer and quantum measurement observable (see Equations 19, 20, 24 and 23). A particular observable output may realise only a subset of the candidate support when coefficients vanish or cancel. A QSM therefore does more than pass eigenvalues or singular values to a classifier: its distinguishing feature is the coupling of sample-conditioned candidate phase carriers and sample-conditioned spectral subspaces through the trainable mixer and quantum measurement. This construction makes quantum data encoding a choice of inductive bias, because the input-derived operator determines which spectral information is made available to the trainable circuit. We therefore focus on the representational consequences of this operator choice within a common upload–mixer framework. Frequency-related learning phenomena have been studied in classical neural networks, where spectral bias interacts with data-manifold geometry [51, 43]. In quantum neural networks, trainable-frequency QFMs make the candidate frequency support itself optimisable [21]. Spectral representations also depend on the choice of transform, operator and basis [4]. Classical architectures obtain spectral structure from fixed Fourier transforms [28], spectral parameterisations of convolutional networks [44], graph-Laplacian bases [5] or Fourier-space operators learned across a training distribution [30, 17]. In a QSM, the current input matrix instead defines the Hamiltonian, whose spectral decomposition supplies sample-conditioned candidate phase carriers and coefficient-forming subspaces for unitary evolution. We study two global QSM variants. The symmetric-Hamiltonian QSM follows our earlier Hamiltonian embedding, Hsym(M)=(M+MT)/2H_sym(M)=(M+M^T)/2 [49]. We additionally study the block-Hamiltonian QSM, which uses Hblock(M)=(0MMT0)H_block(M)= pmatrix0&M\\ M^T&0 pmatrix (see Equations 14 and 15). We also evaluate fixed and trainable patch-SU(4)SU(4) data-encoding unitaries and a quantum spectral model with a non-overlapping patch-local block-Hamiltonian-based data-encoding unitary to test the effect of local structure; their definitions are given in Appendix Appendix E. We investigate how sample-conditioned candidate frequency support changes representation, optimisation and final-state geometry within a common upload–mixer modelling framework. We compare the three QSM variants with fixed rotation-gate, trainable-frequency rotation-gate, fixed patch-SU(4)SU(4) and trainable patch-SU(4)SU(4) data-encoding unitaries on two matrix representations of Pendigits [1] and two controlled spectral tasks. At depth 32, a QSM variant attains the largest mean among the tested quantum models in all four cases: the patch-local block-Hamiltonian QSM leads on both Pendigits representations, while the block-Hamiltonian QSM leads on both synthetic tasks (see Figures 3 and 1). The gradient diagnostics show that the coordinate-wise rotation-gate data-encoding unitaries have much more strongly suppressed mixer-gradient variance at high-depth initialisation, whereas several high-performing QSMs retain larger final gradient scales (see Figure 4). Their final-state fidelity kernels also show more label-aligned geometry in the fixed diagnostic batch, while the patch-SU(4)SU(4) controls demonstrate that gradient magnitude or fidelity-kernel spectral spread alone does not determine accuracy (see Figure 5). We further use ablations to distinguish the roles of spectral values and subspaces. On Pendigits, spectrum-only controls perform poorly, whereas controls that preserve the input-dependent spectral subspaces retain substantially higher accuracy. The ordering reverses on the synthetic tasks, whose labels are defined by eigenvalue or singular-value statistics: spectral-value-only controls are strong, while vector-only controls are weaker (Tables 2 and 7). This reversal establishes a task-dependent division of labour within the sample-conditioned representation: Pendigits relies more strongly on input-dependent subspace geometry and its interaction with the trainable mixer, whereas the controlled tasks reward the relevant spectral values directly. We make four contributions that connect the QSM construction, its Fourier analysis and the empirical comparisons. First, we formulate QSMs as matrix-conditioned models with Hamiltonian-based data-encoding unitaries and identify their representational distinction from encoders with globally shared frequency support. Second, we derive the eigengap and signed-singular-gap candidate supports of the symmetric- and block-Hamiltonian QSMs, together with the dependence of their coefficients on input-derived spectral projectors or subspaces. Third, controlled quantum-model comparisons across reuploading depths, complemented by gradient and fidelity diagnostics, characterise depth-dependent performance, optimisation behaviour and final-state geometry across four benchmark cases. Fourth, value-versus-subspace and structure-destroying ablations identify which component of the sample-conditioned QSM representation is informative for each task. 2 Related Work Variational quantum learning. Variational quantum circuits (VQCs), trained with classical optimisation algorithms, are widely studied as models for near-term quantum applications [8]. In quantum machine learning, they have been applied to classification [13] and generative modelling [11, 20] tasks. However, their optimisation can be impeded by barren plateaus [39, 26] and by poor local minima even in shallow regimes without barren plateaus [2]. Quantum data encoding. When learning from classical data, the encoding map determines the quantum representation [45]. Common schemes represent features through rotation angles, state amplitudes or computational-basis states [45]. Hamiltonian embedding constructs an evolution generator from a matrix-valued sample [49]. Block encodings provide a related but distinct matrix representation: they embed a scaled matrix as a subblock of a larger unitary used as an algorithmic primitive [7, 6]. Our block-Hamiltonian construction instead uses the matrix to define the generator of the data-encoding evolution. In reuploading models, the encoding generator constrains the frequency set available to the trainable circuit [46]. Hamiltonian learning and recognition. Hamiltonian learning infers an unknown physical generator from controlled measurements, using either structural assumptions or ansatz-free black-box protocols [18, 16, 19]. Hamiltonian recognition instead selects a generator from a finite candidate set, including through quantum-signal-processing and variational protocols [52]. Both address inverse problems for a fixed unknown generator, whereas a QSM constructs H(M)H(M) from each input matrix for forward prediction, making its candidate gap support and projector-dependent coefficients sample-conditioned. Data reuploading and quantum Fourier models. Data reuploading circuits interleave data-encoding and trainable unitaries and were introduced as universal quantum classifiers [41]. Their outputs admit truncated Fourier-series representations whose frequencies are determined by the spectrum of the encoding generator. In particular, Pauli-generated rotation gates provide a finite grid of integer-spaced coordinate frequencies, with the spacing set by the coefficient multiplying each input feature, such as α in Ry(αx)R_y(α x) [46]. When we analyse their observable outputs through this finite Fourier representation, we refer to these data-reuploading models as quantum Fourier models (QFMs). Trainable-frequency models optimise the generator spectrum [21], while recent theory quantifies the depth needed for fixed-frequency encodings to approximate tunable ones [32]. Linear-combination-of-unitaries-inspired architectures can instead select prescribed Fourier components [47]. These constructions therefore provide architecture-fixed or globally learned candidate support shared across samples. Other analyses identify coefficient constraints and vanishing expressivity in QFMs [40]. Moreover, for high-dimensional inputs, limited-qubit reuploading models can approach random-guessing performance as the encoding depth grows, motivating wider architectures rather than deeper and narrower ones [50]. QSMs instead construct the data-encoding generator from each input, so their candidate frequency support varies across samples. Spectral methods in quantum machine learning. Classical neural networks often learn low-frequency components before high-frequency ones, although the geometry of the data manifold can modify this bias [51, 43]. A recent framework derives an analogous low-frequency bias for quantum neural networks [38]. Model spectra have also been proposed as objects that quantum routines can learn, regularise or manipulate [4]. QSMs address an earlier representational choice: the input determines which candidate phase carriers are available to the model. Classical spectral architectures. Classical spectral architectures can obtain spectral structure from fixed transforms, Fourier-space maps learned across a training distribution, or domain- and graph-derived operators. Fourier neural operators parameterise integral kernels directly in Fourier space to learn maps between function spaces [30]. Spectral representations support pooling, regularisation and filter parameterisation in convolutional networks [44]. FNet uses fixed, unparameterised Fourier transforms for token mixing [28], whereas adaptive Fourier neural operators learn token mixing in the Fourier domain [17]. Graph networks derive spectral bases from graph shift operators, commonly a Laplacian; because the operator is tied to graph structure, the basis can change when the graph structure changes, while filters are learned over its spectrum [5, 12, 29, 31]. Within this broader landscape, QSMs take an input-adaptive route: each matrix sample defines a Hermitian generator whose spectral subspaces and gap-derived candidate phase carriers enter the unitary representation. 3 Preliminaries Circuit architecture and projector readout. A depth-L data-reuploading circuit alternates data-encoding unitaries with trainable mixers [41]: |ψ(;)⟩=VL(L)SL(;ϕL)⋯V1(1)S1(;ϕ1)|ψ0⟩. ψ( x; )=V_L( θ_L)S_L( x; φ_L)·s V_1( θ_1)S_1( x; φ_1) _0. (1) Here =(ℓ,ϕℓ)ℓ=1L =\( θ_ , φ_ )\_ =1^L collects the mixer and encoder parameters, with ϕℓ φ_ omitted for fixed encoders. We use |ψ0⟩=|+⟩⊗n _0= + n on n qubits. For N classes, let m=⌈log2N⌉m= _2N and associate class j with the projector [49] Pj=|j⟩⟨j|⊗(n−m),j∈0,…,N−1,P_j= j j (n-m), j∈\0,…,N-1\, (2) where |j⟩ j is an m-qubit computational-basis state. The raw projector mass qjq_j and the conditional class probability pjp_j used in the experiments are qj(;) q_j( x; ) =⟨ψ(;)|Pj|ψ(;)⟩, = ψ( x; )P_j ψ( x; ), (3) pj(;) p_j( x; ) =qj(;)∑k=0N−1qk(;),∑k=0N−1qk>0. = q_j( x; ) _k=0^N-1q_k( x; ), _k=0^N-1q_k>0. If N=2mN=2^m, the projectors exhaust the label register and pj=qjp_j=q_j. Otherwise, pjp_j conditions the measurement distribution on the named class outcomes and leaves the predicted class unchanged. The Fourier analysis below concerns the observable outputs qjq_j: the nonlinear normalisation in Equation 3 need not preserve a finite Fourier representation. Fourier representation of observable outputs. For fixed generator parameters, an observable output of a data-reuploading circuit admits a truncated Fourier expansion [46] f()=∑∈Ωc()e−i⋅,Ω⊆ℝd.f_ θ( x)= _ ω∈ c_ ω( θ)e^-i ω· x, ^d. (4) The generator eigenvalues and upload pattern determine the maximal candidate support Ω ; the corresponding spectral projectors, initial state, trainable mixers and quantum measurement observable determine the coefficients. Equal candidate support therefore need not imply equal observable outputs. Consequently, the realised Fourier support of a particular output may be smaller than Ω , because candidate frequencies with zero coefficients do not appear in the output [40]. For one upload of a single scalar input feature, let S(x)=e−ixKS(x)=e^-ixK, where K is Hermitian with distinct eigenvalues κa _a and eigenspace projectors Πa _a. Then K=∑aκaΠa,S(x)=∑ae−ixκaΠa,∑aΠa=,ΠaΠb=δabΠa.K= _a _a _a, S(x)= _ae^-ix _a _a, _a _a=I, _a _b= _ab _a. (5) Defining O=V()†OV()O_ θ=V( θ) OV( θ), the one-upload output is f1(x;) f_1(x; θ) =⟨ψ0|S(x)†OS(x)|ψ0⟩ = _0S(x) O_ θS(x) _0 (6) =∑a,be−ix(κb−κa)⟨ψ0|ΠaOΠb|ψ0⟩ = _a,be^-ix( _b- _a) _0 _aO_ θ _b _0 =∑a,bcab()e−ix(κb−κa). = _a,bc_ab( θ)e^-ix( _b- _a). The maximal one-upload support is therefore Ω1=κb−κa:a,b. _1=\ _b- _a:a,b\. (7) For frequency sets A,B⊆ℝdA,B ^d, we denote their Minkowski sum by A⊕B=+:∈A,∈BA B=\ a+ b: a∈ A, b∈ B\ and use ⨁ℓ=1LAℓ _ =1^LA_ for its repeated application. For layer-dependent generators, let Ω1,ℓ _1, denote the corresponding gap set at upload ℓ . Because the phase factors multiply across uploads, their frequencies add; hence, the maximal L-upload candidate support and the Fourier support of a particular output satisfy [46] ΩL=⨁ℓ=1LΩ1,ℓ,suppF(f)⊆ΩL. _L= _ =1^L _1, , _F(f) _L. (8) If every upload uses the same generator, this reduces to ΩL=Ω1⊕⋯⊕Ω1⏟L uploads. _L= _1 ·s _1_L uploads. (9) Frequency support of rotation encoders. We summarise the maximal one-upload supports of the three rotation encoders used in the experiments for d scalar features =(x1,…,xd) x=(x_1,…,x_d). Appendix B gives the derivations. For SRy()=⨂j=1dRy(αxj)S_R_y( x)= _j=1^dR_y(α x_j), the support is Ω1Ry=α(n1,…,nd):nj∈−1,0,1. _1^R_y= \α(n_1,…,n_d):n_j∈\-1,0,1\ \. (10) For SRyRz()=⨂j=1dRz(βxj)Ry(αxj)S_R_yR_z( x)= _j=1^dR_z(β x_j)R_y(α x_j), the one-feature support is ΩlocalRyRz=nα+mβ:n,m∈−1,0,1, _local^R_yR_z= \nα+mβ:n,m∈\-1,0,1\ \, (11) and the full vector-valued support is Ω1RyRz=(ω1,…,ωd):ωj=njα+mjβ,nj,mj∈−1,0,1. _1^R_yR_z= \( _1,…, _d): _j=n_jα+m_jβ,\ n_j,m_j∈\-1,0,1\ \. (12) For the component-wise trainable-frequency encoder, upload scale γj _j enters the rotation Ry(γjxj)R_y( _jx_j) for feature j [21]. Its one-upload support is Ω1TF()=(n1γ1,…,ndγd):nj∈−1,0,1. _1^TF( γ)= \(n_1 _1,…,n_d _d):n_j∈\-1,0,1\ \. (13) The experiments use a separate scale vector ℓ γ_ at each upload. Once α, β or the ℓ γ_ are fixed, these maximal supports are shared across samples: changing x changes the evaluation point but not the available frequency set. This sample-independent support provides the baseline for the Hamiltonian embeddings introduced in Section 4. 4 Hamiltonian Embedding Produces Sample-Conditioned Frequency Support Let M=M∈ℝp×qM=M_ x ^p× q denote the matrix representation of an input x, with p≥qp≥ q for the layouts considered here. In our previous work, we adopted the symmetric Hamiltonian embedding as the generator of a data-encoding unitary for classification [49]: Hsym(M)=M~+M~T2,H_sym(M)= M+ M^T2, (14) where M~ M is the square matrix obtained by zero-column padding when M is rectangular. Symmetrisation removes the skew-symmetric component of M~ M. We therefore also consider the block-Hamiltonian embedding Hblock(M)=(p×pMMTq×q),H_block(M)= pmatrix 0_p× p&M\\ M^T& 0_q× q pmatrix, (15) which retains M as an off-diagonal block. This construction has the same Hermitian block form used in the quantum polar decomposition algorithm, with A=MTA=M^T for real M [33]. Here it serves as a sample-derived generator of a data-encoding unitary rather than as a subroutine for implementing the polar factors. When required by the circuit dimension or projector readout, either Hamiltonian is promoted by direct-sum padding, H↦H⊕H H 0; we retain the same symbol for the padded operator. This convention preserves its nonzero spectrum and introduces additional zero modes. Within a QSM, the two global data-encoding unitaries are SHsym(t;M)=e−iHsym(M)t/2,SHblock(t;M)=e−iHblock(M)t/2.S_H_sym(t;M)=e^-iH_sym(M)t/2, S_H_block(t;M)=e^-iH_block(M)t/2. (16) We also consider a QSM with a non-overlapping patch-local block-Hamiltonian-based data-encoding unitary. This variant partitions M into R patches P1,…,PRP_1,…,P_R and applies Hblock(Pr)H_block(P_r) on disjoint registers. The experiments use four 2×22× 2 patches, defined in Appendix E. For each fixed M, the analysis below treats the upload time t, or the vector of patch times, as the Fourier variable. Accordingly, Ω1(M) _1(M) denotes sample-conditioned candidate support in the upload-time variables; the entries of M condition the generator and remain fixed throughout this expansion. Appendix C gives the full derivations, including padding, degeneracies and layer-specific upload times. Symmetric embedding produces eigengap carriers. Let sym(M)A_sym(M) denote the index set of distinct eigenvalues of Hsym(M)H_sym(M). For each a∈sym(M)a _sym(M), let λa(M) _a(M) be the corresponding eigenvalue and let Πasym(M) _a^sym(M) project onto its full eigenspace. Then Hsym(M)=∑a∈sym(M)λa(M)Πasym(M),SHsym(t;M)=∑ae−iλa(M)t/2Πasym(M).H_sym(M)= _a _sym(M) _a(M) _a^sym(M), S_H_sym(t;M)= _ae^-i _a(M)t/2 _a^sym(M). (17) For a quantum measurement observable O, with the trainable mixer absorbed into O_ θ, the resulting one-upload observable output is f1Hsym(M;t,)=∑a,bAabHsym(M,)e−i[λb(M)−λa(M)]t/2,f_1^H_sym(M;t, θ)= _a,bA_ab^H_sym(M, θ)e^-i[ _b(M)- _a(M)]t/2, (18) where the pair amplitudes and the coefficient at frequency ω are AabHsym(M,) A_ab^H_sym(M, θ) =⟨ψ0|Πasym(M)OΠbsym(M)|ψ0⟩, = _0 _a^sym(M)O_ θ _b^sym(M) _0, (19) CωHsym(M,) C_ω^H_sym(M, θ) =∑a,b:(λb(M)−λa(M))/2=ωAabHsym(M,). = _ subarrayca,b:\\ ( _b(M)- _a(M))/2=ω subarrayA_ab^H_sym(M, θ). The maximal one-upload candidate support therefore satisfies Ω1Hsym(M) _1^H_sym(M) =λb(M)−λa(M)2:a,b∈sym(M), = \ _b(M)- _a(M)2:a,b _sym(M) \, (20) suppF(f1Hsym) _F(f_1^H_sym) ⊆Ω1Hsym(M). _1^H_sym(M). Thus, for the symmetric-Hamiltonian QSM, the candidate frequencies depend on the eigengaps of the sample-derived Hamiltonian, while the realised coefficients also depend on its eigenspaces, the initial state, the trainable mixer and the readout. Block-Hamiltonian embedding produces signed singular-value gaps. If the nonzero singular values of M are σj(M) _j(M), then the nonzero eigenvalues of Hblock(M)H_block(M) are ±σj(M)± _j(M). Let block(M)A_block(M) denote the index set of distinct eigenvalues of Hblock(M)H_block(M). For each a∈block(M)a _block(M), let μa(M) _a(M) be the corresponding eigenvalue, including zero when present, and let Πablock(M) _a^block(M) project onto its full eigenspace: Hblock(M)=∑a∈block(M)μa(M)Πablock(M),μa(M)∈+σj(M),−σj(M),0.H_block(M)= _a _block(M) _a(M) _a^block(M), _a(M)∈\+ _j(M),- _j(M),0\. (21) The corresponding observable output has the expansion f1Hblock(M;t,)=∑a,bAabHblock(M,)e−i[μb(M)−μa(M)]t/2,f_1^H_block(M;t, θ)= _a,bA_ab^H_block(M, θ)e^-i[ _b(M)- _a(M)]t/2, (22) with AabHblock(M,) A_ab^H_block(M, θ) =⟨ψ0|Πablock(M)OΠbblock(M)|ψ0⟩, = _0 _a^block(M)O_ θ _b^block(M) _0, (23) CωHblock(M,) C_ω^H_block(M, θ) =∑a,b:(μb(M)−μa(M))/2=ωAabHblock(M,). = _ subarrayca,b:\\ ( _b(M)- _a(M))/2=ω subarrayA_ab^H_block(M, θ). Its maximal one-upload candidate support is Ω1Hblock(M) _1^H_block(M) =μb(M)−μa(M)2:a,b∈block(M), = \ _b(M)- _a(M)2:a,b _block(M) \, (24) suppF(f1Hblock) _F(f_1^H_block) ⊆Ω1Hblock(M). _1^H_block(M). For the block-Hamiltonian QSM, these gaps include half-differences and half-sums of singular values; when a zero eigenspace is present, they also include ±σj(M)/2± _j(M)/2. The coefficients depend on the associated singular subspaces together with the initial state, mixer and readout. Patch-local embedding factorises support across patches. For patch PrP_r, let Ω1Hblock(Pr) _1^H_block(P_r) be the candidate set in Equation 24. Because the patch uploads act on disjoint registers with independent times =(t1,…,tR) t=(t_1,…,t_R), their maximal one-upload support is Ω1patch-Hblock(M)=Ω1Hblock(P1)×⋯×Ω1Hblock(PR). _1^patch-H_block(M)= _1^H_block(P_1)×·s× _1^H_block(P_R). (25) This Cartesian product arises because the patches have independent upload-time variables. If they instead shared a single scalar time, the corresponding scalar support would be the Minkowski sum of the patch-level candidate sets. The patch-local carriers are therefore determined by the singular spectra of individual patches rather than that of the full matrix. Their amplitudes depend on the corresponding local singular subspaces, while the trainable mixer and readout can couple modes from different patches. The joint data-encoding unitary, its coefficients and the extension to layer- and patch-specific times tℓ,rt_ ,r are derived in Equations 113, 115 and 116. These three QSM data-encoding variants determine two coupled parts of the representation. Spectral values determine the maximal candidate phase carriers, whereas spectral projectors or singular subspaces participate in forming the realised coefficients through their interaction with the initial state, trainable mixer and readout. The global variants condition these quantities on the full matrix, while the patch variant conditions them separately on non-overlapping patches. The choice of data-encoding unitary therefore determines both the available sample-conditioned phase carriers and the subspaces through which the circuit forms their realised coefficients. 5 Experiments (a) (b) Figure 2: Representative training samples from the real-world and controlled synthetic benchmarks. (a) Pendigits. One example from each of the ten original classes is shown using two paired representations. The upper row shows the DYN representation, in which eight ordered two-dimensional pen coordinates are connected in their recorded sequence. The lower row shows the corresponding STA4 representation as a 4×44× 4 bitmap-like matrix. Raw feature values are displayed for visualisation, whereas the model inputs are standardised using statistics computed from the training split; the colour bar reports the raw STA4 entry values. (b) Synthetic spectral benchmarks. One noisy training input from each class is shown for the SYNTHETIC EIGENGAP task (top, 4×44× 4 observations generated from latent symmetric matrices) and the SYNTHETIC SINGULAR task (bottom, 8×28× 2 rectangular observations). For SYNTHETIC EIGENGAP, class 11 indicates that the gap between the two largest eigenvalues of the latent clean matrix exceeds τeig=0.75 _eig=0.75. For SYNTHETIC SINGULAR, class 11 indicates that the sum of the two leading latent singular values exceeds τsv=2 _sv=2; class 0 denotes the complementary condition in each task. The displayed inputs include relative Gaussian noise with ϵ=0.05ε=0.05, while the labels are computed from the corresponding clean latent spectra. Colours encode the noisy matrix entries using a separate zero-centred scale for each task. Benchmark datasets. We use Pendigits [1] as a compact real-world benchmark. Each digit is represented either by a DYN 8×28× 2 pen-trajectory matrix or by a STA4 4×44× 4 bitmap-like matrix, providing two structurally different views of the same ten-class task (Figure 2(a)). We also consider two binary synthetic tasks: SYNTHETIC EIGENGAP, based on noisy observations of latent 4×44× 4 symmetric matrices, and SYNTHETIC SINGULAR, based on noisy 8×28× 2 rectangular matrices (Figure 2(b)). Their labels are determined from the clean latent eigenvalues or singular values, making them controlled tests of sensitivity to spectral structure. Preprocessing, data splits and data-generation settings are given in Appendix Appendix D. Encoders and training protocol. All models follow the upload–mixer template in Equation 1, with |ψ0⟩=|+⟩⊗n _0= + n. We compare eight data-encoding constructions: fixed-RyR_y, fixed-RyRzR_yR_z, trainable-frequency RyR_y, fixed patch-SU(4)SU(4), trainable patch-SU(4)SU(4), patch-local block-Hamiltonian, symmetric Hamiltonian and block Hamiltonian. They are defined in Sections 3 and 4 and Appendix Appendix E. For each construction, we evaluate reuploading depths L∈1,2,4,8,16,32L∈\1,2,4,8,16,32\ over 20 random parameter-initialisation seeds. Block ℓ applies the upload Sℓ()S_ ( x) followed by the trainable mixer Vℓ(ℓ)V_ ( θ_ ). The comparison controls the dataset splits, upload–mixer template, optimiser, training budget and number of random initialisation seeds, allowing us to examine how different data-encoding constructions behave under a common experimental protocol. However, the constructions differ in both qubit and parameter count: rotation-gate data-encoding unitaries use 16 qubits, patch-local data-encoding unitaries use eight, and the global QSMs use two to four qubits before any label-register expansion, with the Pendigits global QSMs promoted to four qubits. The mixer therefore contributes 15(n−1)15(n-1) parameters per layer, while the trainable data-encoding unitaries introduce different additional parameter groups. These resource differences follow from the register layouts and parameterisations adopted by the respective data-encoding constructions. We therefore compare the predictive performance and inductive biases of the complete encoder–mixer configurations under a common experimental protocol, treating data encoding as a central architectural choice. The results compare the models as implemented, rather than isolating the encoder from all resource differences or measuring performance per qubit or trainable parameter. Appendix Appendix E gives the padding rules and exact depth-32 trainable-parameter counts. For the two global QSMs, the data-encoding unitary at layer ℓ is Ssym,ℓ(M)=exp[−i2Hsym(M)tℓ],Sblock,ℓ(M)=exp[−i2Hblock(M)tℓ],S_sym, (M)= [- i2H_sym(M)t_ ], S_block, (M)= [- i2H_block(M)t_ ], (26) where Hsym(M)H_sym(M) and Hblock(M)H_block(M) are defined in Equations 14 and 15. The trainable upload time tℓt_ scales the input-derived Hamiltonian and is initialised to 1/L1/L unless a fixed time is specified. The patch-local data-encoding constructions divide each input into four non-overlapping 2×22× 2 patches and assign each patch to a disjoint qubit pair. For fixed patch-SU(4)SU(4), a patch Pr∈ℝ2×2P_r ^2× 2 is flattened and mapped linearly to 15 coefficients of the two-qubit unitary Upatch(Pr)=exp[i∑a=115θa(Pr)Ga],U_patch(P_r)= [i _a=1^15 _a(P_r)G_a ], (27) where GaG_a are the 15 non-identity two-qubit Pauli generators. The fixed map is shared across layers and patches; the trainable variant uses a separate linear map for each layer and patch. The patch-local block-Hamiltonian QSM instead applies Equation 15 to each patch, with trainable times tℓ,rt_ ,r initialised to 1/L1/L unless fixed. After each upload, VℓV_ applies nearest-neighbour SU(4)SU(4) blocks in a brick-wall pattern. The even pairs (1,2),(3,4),…(1,2),(3,4),… precede the odd pairs (2,3),(4,5),…(2,3),(4,5),…. Each block has 15 trainable parameters, so an n-qubit mixer layer contains n−1n-1 blocks. Mixer parameters are initialised from a zero-mean Gaussian with scale 0.010.01. All trainable parameters are optimised jointly. We use Adam [22] with learning rate 0.010.01, no weight decay and 2000 optimisation steps. Pendigits runs use training and evaluation batch sizes of 32 and 128, respectively; the corresponding synthetic-task sizes are 64 and 256. 6 Results and Analysis We organise the analysis in four stages: we first compare final test accuracy across reuploading depth, then examine gradient behaviour, next analyse the final-state fidelity geometry of the trained models, and finally use value–subspace and structure-destruction ablations to identify which components of the sample-conditioned spectral representation are informative for each task. (a) (b) Figure 3: (a) Final test accuracy on the Pendigits benchmark as a function of reuploading depth. The two panels show the DYN 8×28× 2 trajectory representation and the STA4 4×44× 4 matrix representation. Each point is the mean over 20 parameter-initialisation seeds, and error bars show one standard deviation across seeds. QSMs lead the tested depth-32 comparisons. The global block-Hamiltonian QSM has the largest observed mean across the six depths on both representations, while the patch-local block-Hamiltonian QSM has the largest mean at depth 32. (b) Final test accuracy on the two synthetic spectral benchmarks as a function of reuploading depth. The SYNTHETIC EIGENGAP task uses noisy 4×44× 4 symmetric matrices with labels determined by the largest latent eigengap, while the SYNTHETIC SINGULAR task uses noisy 8×28× 2 matrices with labels determined by the sum of the two leading latent singular values. Each point is the mean over 20 seeds, and error bars show one standard deviation across seeds. The global block-Hamiltonian QSM has the largest mean at depth 32 and the largest observed mean across the six depths on both tasks. QSMs attain the largest depth-32 mean test accuracy among the tested quantum models. At depth 32, the patch-local block-Hamiltonian QSM has the largest mean accuracy on Pendigits DYN (92.75%±0.84%92.75\%± 0.84\%) and STA4 (86.14%±1.33%86.14\%± 1.33\%), while the global block-Hamiltonian QSM has the largest mean on SYNTHETIC EIGENGAP (85.15%±2.40%85.15\%± 2.40\%) and SYNTHETIC SINGULAR (92.72%±4.04%92.72\%± 4.04\%). Thus, a QSM attains the largest depth-32 mean among the tested quantum models in each of the four benchmark cases (Figure 3 and Table 1). Across the six test-accuracy-versus-depth curves, the global block-Hamiltonian QSM has the largest observed mean for all four cases, occurring at depth 16 with 93.81%±1.06%93.81\%± 1.06\%, 86.66%±2.17%86.66\%± 2.17\%, 86.65%±1.18%86.65\%± 1.18\% and 95.24%±1.53%95.24\%± 1.53\%, respectively. These values summarise the observed depth sweeps rather than performance at a validation-selected depth; the full curves and depth-32 results provide the corresponding depth-specific comparisons. Appendix Appendix F reports the validation summaries and validation-accuracy AUCs. Table 1: Final test accuracy for the main Pendigits and synthetic benchmark runs. Values are mean ± standard deviation over 20 random seeds, reported in percentages. The largest-observed-mean columns report the depth and value of the maximum mean test accuracy across the six evaluated depths for each dataset–encoder pair, providing a descriptive summary of the depth sweep. The depth-32 column reports the same metric at the largest reuploading depth used in the experiments. Boldface marks the largest mean within each dataset block and accuracy column. Dataset Encoder Depth of largest observed mean Largest observed mean test acc. (%) Depth-32 test acc. (%) Pendigits DYN fixed-RyR_y 16 75.02±2.6375.02± 2.63 72.29±2.7772.29± 2.77 fixed-RyRzR_yR_z 8 64.14±2.2664.14± 2.26 55.84±5.3055.84± 5.30 trainable-frequency(TF)-RyR_y 32 81.37±2.3681.37± 2.36 81.37±2.3681.37± 2.36 patch-SU(4)SU(4) 4 48.60±1.8748.60± 1.87 9.96±0.449.96± 0.44 trainable patch-SU(4)SU(4) 2 74.70±2.0474.70± 2.04 9.84±0.479.84± 0.47 patch HblockH_ block 32 92.75±0.8492.75± 0.84 92.75±0.84 92.75± 0.84 HsymH_ sym 32 91.89±3.7191.89± 3.71 91.89±3.7191.89± 3.71 HblockH_ block 16 93.81±1.06 93.81± 1.06 92.35±5.9592.35± 5.95 Pendigits STA4 fixed-RyR_y 8 61.10±4.3561.10± 4.35 40.22±4.4040.22± 4.40 fixed-RyRzR_yR_z 16 52.47±3.4952.47± 3.49 32.79±4.2832.79± 4.28 TF-RyR_y 8 70.11±2.3570.11± 2.35 46.85±6.3946.85± 6.39 patch-SU(4)SU(4) 4 34.51±1.0834.51± 1.08 16.41±1.8816.41± 1.88 trainable patch-SU(4)SU(4) 2 54.94±2.8854.94± 2.88 10.27±0.6210.27± 0.62 patch HblockH_ block 32 86.14±1.3386.14± 1.33 86.14±1.33 86.14± 1.33 HsymH_ sym 16 70.14±8.2970.14± 8.29 59.79±14.9059.79± 14.90 HblockH_ block 16 86.66±2.17 86.66± 2.17 76.38±16.8476.38± 16.84 Synthetic eigengap fixed-RyR_y 8 67.57±0.9867.57± 0.98 65.40±1.6965.40± 1.69 fixed-RyRzR_yR_z 8 66.47±1.1866.47± 1.18 65.55±1.2265.55± 1.22 TF-RyR_y 8 68.24±1.6268.24± 1.62 67.42±1.3367.42± 1.33 patch-SU(4)SU(4) 4 63.19±1.7263.19± 1.72 54.58±1.9754.58± 1.97 trainable patch-SU(4)SU(4) 2 71.12±1.3371.12± 1.33 50.16±2.3650.16± 2.36 patch HblockH_ block 16 74.37±2.5974.37± 2.59 71.44±2.1171.44± 2.11 HsymH_ sym 16 83.42±1.3783.42± 1.37 82.81±3.1682.81± 3.16 HblockH_ block 16 86.65±1.18 86.65± 1.18 85.15±2.40 85.15± 2.40 Synthetic singular fixed-RyR_y 16 85.00±1.1585.00± 1.15 84.34±1.2484.34± 1.24 fixed-RyRzR_yR_z 32 84.54±1.0184.54± 1.01 84.54±1.0184.54± 1.01 TF-RyR_y 32 85.40±0.6385.40± 0.63 85.40±0.6385.40± 0.63 patch-SU(4)SU(4) 4 77.50±2.0077.50± 2.00 63.41±2.0363.41± 2.03 trainable patch-SU(4)SU(4) 4 86.45±1.5286.45± 1.52 61.75±1.9161.75± 1.91 patch HblockH_ block 32 87.50±2.0687.50± 2.06 87.50±2.0687.50± 2.06 HsymH_ sym 8 90.16±1.0890.16± 1.08 88.75±2.6988.75± 2.69 HblockH_ block 16 95.24±1.53 95.24± 1.53 92.72±4.04 92.72± 4.04 The depth-32 comparisons show that the largest mean test accuracy is obtained by a QSM across structurally different inputs and spectral rules. The DYN input is an ordered 8×28× 2 trajectory, and its four non-overlapping 2×22× 2 patches correspond to consecutive trajectory segments. By contrast, STA4 is a 4×44× 4 bitmap-like representation whose patches preserve local two-dimensional neighbourhoods, while the two synthetic labels depend on full-matrix eigengap or singular-value statistics. These contrasting input structures and task definitions provide the context for the encoder-specific performance patterns discussed below. The rotation-gate baselines attain larger means on DYN than on STA4, both at depth 32 and at their highest observed means. This difference may reflect how DYN exposes coordinate values in their recorded sequence, whereas STA4 organises them within a two-dimensional bitmap neighbourhood. The strongest rotation-gate result on DYN, 81.37%±2.36%81.37\%± 2.36\% for trainable-frequency RyR_y at depth 32, remains below the depth-32 means of all three QSMs. On STA4, the explicit 2×22× 2 locality of the patch-local block-Hamiltonian QSM offers a possible interpretation of its performance, paralleling the spatial locality encoded by convolutional networks [27]. Taken together, these patterns suggest an interaction between data-encoding construction and input layout; the specific contribution of locality remains unresolved. (a) (b) Figure 4: (a) Gradient variance at initialisation as a function of reuploading depth. For each run configuration, gradients are computed on a fixed diagnostic training batch of size 32 for 20 random initialisations. The plotted quantity is the median log10 _10 variance across parameters within each parameter group, averaged over 20 run seeds; error bars show one standard deviation across run seeds. The coordinate-wise rotation-gate data-encoding unitaries exhibit strongly suppressed mixer-gradient variance at large depth, whereas the QSMs retain substantially larger initialisation variance. (b) Final-checkpoint RMS gradient as a function of reuploading depth. Gradients are computed at the final trained checkpoint using the same diagnostic-batch procedure as in the initialisation diagnostics. The RMS is computed within each parameter group, and plotted values are means over 20 run seeds with one-standard-deviation error bars. The QSMs retain finite gradients at high depth, while the rotation-gate data-encoding unitaries generally show smaller final mixer-gradient scales, especially on the synthetic spectral tasks. QSMs with Hamiltonian-based data-encoding unitaries retain stronger high-depth gradient signals, while gradient scale alone does not determine accuracy. At L=32L=32, the median log10 _10 mixer-gradient variance [39] of the fixed and trainable-frequency rotation-gate data-encoding unitaries is approximately −11-11 on Pendigits and −20-20 on the synthetic tasks. For HblockH_block, the corresponding values are −1.05±0.05-1.05± 0.05 on Pendigits DYN, 0.05±0.200.05± 0.20 on STA4, −3.72±0.05-3.72± 0.05 on SYNTHETIC EIGENGAP and −3.64±0.04-3.64± 0.04 on SYNTHETIC SINGULAR. We compute these training-loss gradients on a fixed diagnostic batch of 32 examples and separate the shared SU(4)SU(4) mixer parameters from encoder-specific parameters. Within each run seed, the initialisation statistic measures the coordinate-wise variance of each parameter gradient across 20 parameter initialisations; we take the median log variance across a parameter group and then aggregate the run-level values over 20 run seeds, with sample-standard-deviation error bars. The coordinate-wise rotation encoders therefore show much more strongly suppressed parameter-wise variation across initialisations at high depth. Appendix Appendix G defines this statistic and the remaining gradient diagnostics. The final-checkpoint root-mean-square (RMS) gradient gives a separate snapshot after optimisation (Figure 4(b)). At depth 32, the mixer-gradient RMS values for HsymH_sym and HblockH_block are 0.079±0.0340.079± 0.034 and 0.071±0.0340.071± 0.034 on Pendigits DYN, 0.164±0.0530.164± 0.053 and 0.194±0.1690.194± 0.169 on STA4, 0.108±0.0190.108± 0.019 and 0.062±0.0110.062± 0.011 on SYNTHETIC EIGENGAP, and 0.061±0.0150.061± 0.015 and 0.036±0.0090.036± 0.009 on SYNTHETIC SINGULAR. On the synthetic tasks these values exceed the rotation-gate baseline values, which are near 10−210^-2, showing that several high-performing QSMs retain larger mixer-gradient scales at the final checkpoint. The empirical-Fisher diagnostic is the Gram matrix of per-example, ground-truth-label loss gradients, following the common machine-learning convention [25]. We use its effective rank as a finite-batch summary of how the sampled gradient energy is distributed across parameter-space directions. At depth 32 on Pendigits DYN, patch-SU(4)SU(4) has empirical-Fisher effective rank 29.25±0.5229.25± 0.52 but test accuracy 9.96%±0.44%9.96\%± 0.44\%. The corresponding values for HblockH_block are 9.39±2.729.39± 2.72 and 92.35%±5.95%92.35\%± 5.95\% (Appendix Appendix G and Table 4). In this comparison, the empirical-Fisher effective-rank ordering does not match the test-accuracy ordering. The gradient results establish two complementary observations: extreme initialisation-variance suppression accompanies the lower high-depth accuracy of the rotation-gate data-encoding unitaries, while the patch-SU(4)SU(4) controls show that larger gradient variance or empirical-Fisher rank is insufficient for high accuracy. These diagnostics describe the optimisation geometry of the tested runs but do not establish a causal mechanism or a necessary gradient condition for the QSM results. Appendix Appendix G gives the complete metric definitions, while Figures 8 and 9 reports additional initialisation-RMS and empirical-Fisher results. Several high-performing quantum spectral models develop more label-aligned fidelity geometry. At high depth, several of these models show larger within-minus-between fidelity gaps and stronger kernel-target alignment than the coordinate-wise rotation and patch-SU(4)SU(4) baselines (see Figure 5). We form the pure-state fidelity kernel Kij=|⟨ψi|ψj⟩|2K_ij=| _i| _j|^2 from the final post-mixer states of a fixed 32-example validation batch. Centred kernel-target alignment [10, 9, 23] compares this kernel with the label-equality kernel, while the within-minus-between fidelity gap, motivated by class separation in quantum metric learning [34], is the mean same-label fidelity minus the mean different-label fidelity after excluding self-pairs (Equation 171). Positive gaps indicate greater average similarity within classes on the diagnostic batch. Except for the 20-seed test accuracies in Table 5, all latent diagnostics use seed 0 and have no cross-seed uncertainty estimates; the appendix expands this provenance and scope. (a) (b) Figure 5: (a) Within-minus-between fidelity gap of the final quantum states. For each trained model, we compute the pure-state fidelity kernel Kij=|⟨ψi|ψj⟩|2K_ij=| _i| _j|^2 on a fixed 32-example validation diagnostic batch and report the mean same-label off-diagonal fidelity minus the mean different-label off-diagonal fidelity. Larger values indicate stronger class separation in quantum-state geometry. At high depth, the global QSMs and the patch-local block-Hamiltonian QSM produce substantially larger gaps than the coordinate-wise rotation-gate and patch-SU(4)SU(4) baselines. (b) Kernel-target alignment of the final quantum-state fidelity kernel. The plotted metric is the centred alignment between the final fidelity kernel and the label-equality kernel on the same batch. Larger values indicate that the geometry of the trained quantum states is more aligned with class membership. Several high-performing QSMs show stronger alignment in the high-depth regime; this is a descriptive association rather than evidence that the data-encoding unitary causes the accuracy ordering. Both panels use the final checkpoint from seed 0 and contain no across-seed uncertainty estimate. At depth 32 on Pendigits DYN, HblockH_block has diagnostic-batch accuracy 1.000, target alignment 0.900 and fidelity gap 0.561; HsymH_sym gives 1.000, 0.887 and 0.538, respectively (Table 5). On STA4, patch HblockH_block has diagnostic-batch accuracy 0.938, target alignment 0.694 and gap 0.173. On SYNTHETIC SINGULAR, HsymH_sym and HblockH_block give target alignments of 0.508 and 0.461 and gaps of 0.163 and 0.159. These observations show that several high-performing QSMs form visibly class-separated fidelity geometry within the diagnostic batch. The fidelity gap and participation-ratio effective rank describe complementary aspects of the kernel geometry: the former measures label separation, while the latter measures how broadly the kernel eigenvalue mass is distributed across eigenmodes. At depth 32, the rotation-gate data-encoding unitaries have kernel effective rank near the batch-size maximum of 32 while their gaps are approximately zero (see Figure 11). The patch-SU(4)SU(4) data-encoding unitaries show the same combination of high effective rank, near-zero gap and low diagnostic-batch accuracy. Near-maximal effective rank in a fidelity kernel is consistent with an identity-like Gram matrix, a limiting geometry that can arise through exponential concentration and impair generalisation in quantum kernel methods [48]. The positive gaps of several QSMs instead indicate class-dependent off-diagonal structure within this diagnostic batch. Appendix Appendix H gives the corresponding layerwise and checkpoint trajectories, including the fidelity gap, target alignment and adjacent-layer CKA. Classical baselines probe which representations make each task accessible. The synthetic labels are defined by spectral statistics of the clean latent matrices. On SYNTHETIC EIGENGAP, the HsymH_sym-eigenvalue MLP reaches 98.01%±0.35%98.01\%± 0.35\%, compared with the largest observed QSM mean of 86.65%±1.18%86.65\%± 1.18\%. On SYNTHETIC SINGULAR, the HblockH_block-singular-value MLP reaches 98.50%±0.20%98.50\%± 0.20\%, compared with 95.24%±1.53%95.24\%± 1.53\%. These near-ceiling accuracies provide an empirical check that the noisy observations retain the intended label-defining spectral signal, supporting the use of these benchmarks as controlled tests of eigenvalue- and singular-value-dependent learning. On Pendigits, one-hidden-layer MLPs trained on raw entries reach 96.88%±0.24%96.88\%± 0.24\% on DYN and 91.32%±0.46%91.32\%± 0.46\% on STA4 (see Table 6). By contrast, MLPs trained only on HsymH_sym eigenvalues reach 59.98%±0.60%59.98\%± 0.60\% and 39.80%±0.50%39.80\%± 0.50\%, while those trained on HblockH_block singular values reach 39.67%±0.48%39.67\%± 0.48\% and 30.62%±0.53%30.62\%± 0.53\%. Under this protocol, reducing each input to its spectral values therefore removes substantial task-relevant information available in the raw matrices. The classical baselines reveal a task-dependent contrast: spectral values are almost sufficient for the controlled synthetic tasks, but not for Pendigits. The value–subspace ablations below examine how this contrast appears within the QSM representations. Ablations reveal task-dependent roles for spectral values and subspaces. The central ablation result is a reversal between the real-world and controlled tasks (see Table 2). The spectral-value controls retain each sample's values while using fixed reference bases derived from the training split. The spectral-subspace controls instead retain each sample's eigenvectors or singular vectors while replacing its values by coordinate-wise training-set medians. On Pendigits, the subspace-retaining controls preserve much more accuracy than the spectral-value controls. On the synthetic tasks, where the labels are defined by the corresponding spectral statistics, the spectral-value controls become strongest and the subspace-retaining controls are much weaker. This reversal shows that the useful component of sample-conditioned spectral geometry depends on the structure of the learning problem. To summarise performance across depth, we compare the largest mean test accuracy observed for each ablation over the six evaluated depths. Table 2: Value–subspace ablation summary. Entries give the depth and largest observed mean final test accuracy (mean ± standard deviation over 20 random seeds, in percentage) for the original encoder and the corresponding value-only and subspace-only controls. Every entry is a descriptive maximum across the tested depths rather than performance at a depth selected by independent validation; each control is retrained on its transformed dataset. Pendigits depends primarily on sample-dependent subspaces, whereas the controlled spectral tasks reverse the ordering. The complete ablation and permutation results are reported in Appendix Table 7. Task (reference encoder) Original (depth, acc.) Spectral values only (depth, acc.) Spectral subspaces only (depth, acc.) Primary observed dependence Pendigits DYN (HblockH_ block) L16, 93.81±1.0693.81± 1.06 L4, 31.65±2.4431.65± 2.44 L16, 93.56±0.8493.56± 0.84 subspaces Pendigits STA4 (HblockH_ block) L16, 86.66±2.1786.66± 2.17 L4, 25.25±2.0225.25± 2.02 L16, 84.74±4.5184.74± 4.51 subspaces Synthetic eigengap (HsymH_ sym) L16, 83.42±1.3783.42± 1.37 L8, 97.92±0.3897.92± 0.38 L1, 61.28±0.0761.28± 0.07 values Synthetic singular (HblockH_ block) L16, 95.24±1.5395.24± 1.53 L32, 98.42±0.3098.42± 0.30 L16, 49.88±1.9649.88± 1.96 values On Pendigits DYN and STA4, the HblockH_block spectral-value controls reach 31.65%±2.44%31.65\%± 2.44\% and 25.25%±2.02%25.25\%± 2.02\%, whereas the spectral-subspace controls retain 93.56%±0.84%93.56\%± 0.84\% and 84.74%±4.51%84.74\%± 4.51\%. The HsymH_sym controls show the same ordering: spectral-value accuracies are 48.87%±3.90%48.87\%± 3.90\% and 28.35%±2.11%28.35\%± 2.11\%, compared with 89.80%±4.56%89.80\%± 4.56\% and 64.30%±2.02%64.30\%± 2.02\% when sample-dependent eigenvectors are retained. Within these ablations, the Pendigits results therefore depend much more strongly on sample-dependent spectral subspaces than on spectral values alone. The synthetic controls reverse this performance ordering. The HsymH_sym spectral-value control reaches 97.92%±0.38%97.92\%± 0.38\% on SYNTHETIC EIGENGAP, whereas its spectral-subspace counterpart reaches 61.28%±0.07%61.28\%± 0.07\%. On SYNTHETIC SINGULAR, the corresponding HblockH_block values are 98.42%±0.30%98.42\%± 0.30\% and 49.88%±1.96%49.88\%± 1.96\%. This reversal agrees with the data-generating rules and demonstrates that spectral values and subspaces can play different roles under the same upload–mixer architecture. Fixed row and column permutations preserve singular values, whereas a general entry permutation does not. On SYNTHETIC SINGULAR, the largest observed mean for global HblockH_block is 95.06%±2.30%95.06\%± 2.30\% after row/column permutation, 0.18 percentage points below the original, and 90.55%±0.58%90.55\%± 0.58\% after entry permutation, 4.69 points below. For patch HblockH_block on STA4, entry and row/column permutations reduce the largest observed means by 3.04 and 2.22 points. This sensitivity is compatible with the use of fixed 2×22× 2 patches, but does not isolate locality as its cause. The complete value/subspace and permutation results are given in Appendix Table 7. 7 Discussion and Outlook Quantum Spectral Models (QSMs) operate in the spectral domain of each matrix input: we construct an input-dependent Hermitian generator from the matrix and use its time-evolution operator as the data-encoding unitary. For the symmetric- and block-Hamiltonian data-encoding unitaries, eigengaps or signed singular-value gaps determine the sample-conditioned candidate phase carriers. The associated spectral projectors represent the input-dependent subspaces that enter the corresponding Fourier coefficients. Candidate support specifies the carriers that are available in principle, whereas realised support contains only those whose coefficients remain nonzero after the initial state, trainable mixing and measurement observable are taken into account. Common coordinate-wise quantum Fourier models (QFMs) instead use architecture-defined candidate support, while trainable-frequency rotation-gate QFMs learn support that remains globally shared across samples after training [46, 21]. The key difference is that, in a QSM, both the candidate phase carriers and the spectral subspaces entering the corresponding Fourier coefficients depend on the individual input matrix. At the largest evaluated depth, a QSM variant attains the highest mean test accuracy among the tested quantum models on each of the four benchmarks. The patch-local block-Hamiltonian QSM leads on the two Pendigits representations, whereas the global block-Hamiltonian QSM leads on the two controlled spectral tasks (see Figure 3 and Table 1). Together with the differing trends across the full depth sweeps, this task-dependent ordering shows that the effect of circuit depth is specific to the encoder and dataset. One possible, non-causal interpretation of the change in the depth-32 winner concerns the scale of structure exposed by each encoder. The two synthetic labels are defined by spectral statistics of the complete clean latent matrix, whereas each Pendigits input contains an explicit local organisation. For DYN, the four non-overlapping 2×22× 2 patches are four consecutive two-point segments of the ordered pen trajectory; for STA4, they are four contiguous 2×22× 2 quadrants of the bitmap-like matrix. The patch-local QSM therefore preserves these trajectory segments or spatial neighbourhoods, while the global block-Hamiltonian QSM retains relations across the full matrix. This correspondence is consistent with a locality-based inductive-bias interpretation and motivates future controlled comparisons that isolate locality from other encoder differences and clarify its contribution to the observed accuracy ordering. The gradient and fidelity diagnostics add complementary descriptive evidence. At large depth, we observe extreme suppression of mixer-gradient variance across initialisations for the coordinate-wise rotation-gate encoders, whereas the QSMs retain greater variance and several high-performing variants retain larger final gradient scales (see Figure 4). In the fixed diagnostic batch, several QSMs also develop larger fidelity gaps and stronger kernel-target alignment (see Figure 5). The patch-SU(4)SU(4) controls further show that the gradient and fidelity diagnostics capture different aspects of the learned model geometry. Future work could investigate how these measures interact throughout training and whether their joint behaviour provides a more complete account of predictive performance. The value-versus-subspace ablations provide the most direct representational evidence. On Pendigits, value-only controls lose most of the original accuracy, while controls that preserve sample-dependent eigenvectors or singular vectors retain much more performance (Tables 2 and 7). The ordering reverses on the controlled spectral tasks, where the labels are defined by an eigengap or a leading-singular-value sum and the value-only controls become strongest. This task-dependent division of labour shows that spectral values make the synthetic rules directly accessible, whereas spectral subspaces and trainable mixing carry much of the useful information for Pendigits. These results point to a general representation-design question: should a spectral representation be shared across a dataset or adapted to each input? Coordinate-wise QFMs use architecture-defined carriers, trainable-frequency QFMs learn support shared across samples, and several classical spectral architectures use fixed or domain-derived operators or Fourier-space maps learned across a training distribution [5, 30]. A QSM takes a different approach by constructing the data-encoding generator from each input matrix, so that both the gap-derived candidate phase carriers and the spectral subspaces entering the Fourier coefficients adapt to the input. Because the time evolution of a Hermitian generator is governed by its eigenvalues and spectral projectors, using the input matrix as the generator makes spectral structure intrinsic to the data-encoding dynamics. This provides a particularly direct route to input-adaptive spectral representation learning in quantum models and motivates further investigation of when such representations are useful across quantum and classical models. Building on these results, future work can test how broadly the observed representation effects persist across model scales, datasets and resource constraints. Resource-matched comparisons could hold qubit count, parameter count and locality fixed, select depth through independent validation, retain the task-appropriate classical baselines that lead on the present benchmarks, and extend the analysis to larger matrix-structured datasets. The nonzero fidelity gaps of several QSMs at the studied sizes also motivate qubit-scaling and finite-shot analyses of whether Hamiltonian-based data encodings resist exponential kernel concentration [48]. Multi-seed latent-state diagnostics across independent batches, together with controlled global-versus-patch-local and structure-destroying ablations, could further clarify how spectral subspaces, locality and trainability interact. Together, these directions provide a path towards hardware-relevant, resource-normalised evaluations of whether the observed representation effects can support end-to-end quantum advantage. A complementary direction for future work is to develop scalable implementations of dense input-dependent matrix operations. This would involve specifying the data-access or block-encoding assumptions [35, 36], accounting for data-loading and circuit costs, and examining noise in hardware-relevant settings. At the representation level, quantum singular value transformation or eigenvalue processing could provide controlled tests of spectral values while variational operations act on the associated subspaces [15, 37]. Investigating how these approaches depend on matrix structure and access cost could connect QSM representation design with scalable quantum implementations. Taken together, our formal analysis and experiments establish QSMs as an analysable framework for input-adaptive spectral representation learning. The input-derived Hamiltonian supplies sample-conditioned candidate phase carriers through its spectral gaps, while the associated spectral subspaces interact with the initial state, trainable mixer and measurement observable to determine which carriers are realised and with what coefficients. This support–coefficient perspective provides a concrete principle for designing quantum models around matrix-valued data and offers a new lens for studying how data-dependent structure can be embedded directly into model dynamics. Data and Code Availability Code available at https://github.com/peiyong-addwater/QuantumSpectralModel. References [1] E. Alpaydin and Fevzi. Alimoglu (1996) Pen-Based Recognition of Handwritten Digits. Note: UCI Machine Learning RepositoryDOI: https://doi.org/10.24432/C5MG6K Cited by: §D.1, §1, §5. [2] E. R. Anschuetz and B. T. Kiani (2022-15 December) Quantum variational algorithms are swamped with traps. Nat. Commun. 13 (1), p. 7760. External Links: Link, Document, ISSN 2041-1723 Cited by: §2. [3] P. W. Battaglia, J. B. Hamrick, V. Bapst, A. Sanchez-Gonzalez, V. Zambaldi, M. Malinowski, A. Tacchetti, D. Raposo, A. Santoro, R. Faulkner, C. Gulcehre, F. Song, A. Ballard, J. Gilmer, G. Dahl, A. Vaswani, K. Allen, C. Nash, V. Langston, C. Dyer, N. Heess, D. Wierstra, P. Kohli, M. Botvinick, O. Vinyals, Y. Li, and R. Pascanu (2018-4 June) Relational inductive biases, deep learning, and graph networks. arXiv [cs.LG]. External Links: Link, 1806.01261, Document Cited by: §1. [4] V. Belis, J. Bowles, R. Gupta, E. Peters, and M. Schuld (2026-25 March) Spectral methods: crucial for machine learning, natural for quantum computers?. arXiv [quant-ph]. External Links: Link, 2603.24654, Document Cited by: §1, §2. [5] J. Bruna, W. Zaremba, A. Szlam, and Y. LeCun (2014) Spectral networks and locally connected networks on graphs. External Links: 1312.6203, Link Cited by: §1, §2, §7. [6] D. Camps, L. Lin, R. V. Beeumen, and C. Yang (2023) Explicit quantum circuits for block encodings of certain sparse matrices. External Links: 2203.10236, Link Cited by: §2. [7] D. Camps and R. Van Beeumen (2022-09) FABLE: fast approximate quantum circuits for block-encodings. In 2022 IEEE International Conference on Quantum Computing and Engineering (QCE), p. 104–113. External Links: Link, Document Cited by: §2. [8] M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles (2021-12 August) Variational quantum algorithms. Nat. Rev. Phys. 3 (9), p. 625–644. External Links: Link, Document, ISSN 2522-5820,2522-5820 Cited by: §2. [9] C. Cortes, M. Mohri, and A. Rostamizadeh (2012-03) Algorithms for learning kernels based on centered alignment. J. Mach. Learn. Res. 13 (1), p. 795–828. External Links: ISSN 1532-4435 Cited by: §6. [10] N. Cristianini, J. Shawe-Taylor, A. Elisseeff, and J. Kandola (2001) On kernel-target alignment. In Proceedings of the 15th International Conference on Neural Information Processing Systems: Natural and Synthetic, NIPS'01, Cambridge, MA, USA, p. 367–373. Cited by: §6. [11] P. Dallaire-Demers and N. Killoran (2018-23 July) Quantum generative adversarial networks. Phys. Rev. A 98 (1), p. 012324. External Links: Link, Document Cited by: §2. [12] M. Defferrard, X. Bresson, and P. Vandergheynst (2016) Convolutional neural networks on graphs with fast localized spectral filtering. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS'16, Red Hook, NY, USA, p. 3844–3852. External Links: ISBN 9781510838819 Cited by: §2. [13] E. Farhi and H. Neven (2018) Classification with quantum neural networks on near term processors. External Links: 1802.06002, Link Cited by: §2. [14] P. Gao, E. Trautmann, B. Yu, G. Santhanam, S. Ryu, K. Shenoy, and S. Ganguli (2017-11) A theory of multineuronal dimensionality, dynamics and measurement. Note: bioRxiv preprint External Links: Link, Document Cited by: §G.1, §H.1. [15] A. Gilyén, Y. Su, G. H. Low, and N. Wiebe (2019-23 June) Quantum singular value transformation and beyond: exponential improvements for quantum matrix arithmetics. In Proceedings of the 51st Annual ACM SIGACT Symposium on Theory of Computing, New York, NY, USA. External Links: Link, Document, ISBN 9781450367059 Cited by: §7. [16] A. Gu, L. Cincio, and P. J. Coles (2024-8 January) Practical hamiltonian learning with unitary dynamics and gibbs states. Nat. Commun. 15 (1), p. 312. External Links: Link, Document, ISSN 2041-1723,2041-1723 Cited by: §2. [17] J. Guibas, M. Mardani, Z. Li, A. Tao, A. Anandkumar, and B. Catanzaro (2022) Efficient token mixing for transformers via adaptive fourier neural operators. In International Conference on Learning Representations, External Links: Link Cited by: §1, §2. [18] J. Haah, R. Kothari, and E. Tang (2024-6 June) Learning quantum hamiltonians from high-temperature gibbs states and real-time evolutions. Nat. Phys. 20 (6), p. 1027–1031. External Links: Link, Document, ISSN 1745-2473,1745-2481 Cited by: §2. [19] H. Hu, M. Ma, W. Gong, Q. Ye, Y. Tong, S. T. Flammia, and S. F. Yelin (2025-22 October) Ansatz-free hamiltonian learning with heisenberg-limited scaling. PRX quantum 6 (4), p. 040315. External Links: Link, Document, ISSN 2691-3399 Cited by: §2. [20] H. Huang, Y. Du, M. Gong, Y. Zhao, Y. Wu, C. Wang, S. Li, F. Liang, J. Lin, Y. Xu, R. Yang, T. Liu, M. Hsieh, H. Deng, H. Rong, C. Peng, C. Lu, Y. Chen, D. Tao, X. Zhu, and J. Pan (2021-08) Experimental quantum generative adversarial networks for image generation. Physical Review Applied 16 (2). External Links: ISSN 2331-7019, Link, Document Cited by: §2. [21] B. Jaderberg, A. A. Gentile, Y. A. Berrada, E. Shishenina, and V. E. Elfving (2024-22 April) Let quantum neural networks choose their own frequencies. Phys. Rev. A 109 (4), p. 042421. External Links: Link, Document, ISSN 2469-9926,2469-9934 Cited by: §B.3, §1, §1, §2, §3, §7. [22] D. P. Kingma and J. Ba (2014-22 December) Adam: a method for stochastic optimization. arXiv [cs.LG]. External Links: Link, 1412.6980, Document Cited by: §5. [23] S. Kornblith, M. Norouzi, H. Lee, and G. Hinton (2019-24 May) Similarity of neural network representations revisited. In International Conference on Machine Learning, p. 3519–3529. External Links: Link, ISSN 2640-3498 Cited by: §H.1, §6. [24] N. Kovachki, Z. Li, B. Liu, K. Azizzadenesheli, K. Bhattacharya, A. Stuart, and A. Anandkumar (2023-01) Neural operator: learning maps between function spaces with applications to pdes. J. Mach. Learn. Res. 24 (1). External Links: ISSN 1532-4435 Cited by: §1. [25] F. Kunstner, L. Balles, and P. Hennig (2019) Limitations of the empirical fisher approximation for natural gradient descent. In Advances in Neural Information Processing Systems, Vol. 32, p. 4156–4167. External Links: Link Cited by: §G.1, §6. [26] M. Larocca, S. Thanasilp, S. Wang, K. Sharma, J. Biamonte, P. J. Coles, L. Cincio, J. R. McClean, Z. Holmes, and M. Cerezo (2025-26 March) Barren plateaus in variational quantum computing. Nat. Rev. Phys. 7 (4), p. 174–189. External Links: Link, Document, ISSN 2522-5820,2522-5820 Cited by: §2. [27] Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner (1998) Gradient-based learning applied to document recognition. Proc. IEEE Inst. Electr. Electron. Eng. 86 (11), p. 2278–2324. External Links: Link, Document, ISSN 0018-9219,1558-2256 Cited by: §1, §6. [28] J. Lee-Thorp, J. Ainslie, I. Eckstein, and S. Onta~nón (2022-07) FNet: mixing tokens with Fourier transforms. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, M. Carpuat, M. de Marneffe, and I. V. Meza Ruiz (Eds.), Seattle, United States, p. 4296–4313. External Links: Link, Document Cited by: §1, §2. [29] R. Levie, F. Monti, X. Bresson, and M. M. Bronstein (2019-01) CayleyNets: graph convolutional neural networks with complex rational spectral filters. Trans. Sig. Proc. 67 (1), p. 97–109. External Links: ISSN 1053-587X, Link, Document Cited by: §2. [30] Z. Li, N. B. Kovachki, K. Azizzadenesheli, B. liu, K. Bhattacharya, A. Stuart, and A. Anandkumar (2021) Fourier neural operator for parametric partial differential equations. In International Conference on Learning Representations, External Links: Link Cited by: §1, §2, §7. [31] Y. E. Lin, R. Talmon, and R. Levie (2024) Equivariant machine learning on graphs with nonlinear spectral filters. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2. [32] A. Y. Liu and L. Pira (2026-24 June) The cost of removing tunability in quantum data re-uploading. arXiv [quant-ph]. External Links: Link, 2606.25598, Document Cited by: §2. [33] S. Lloyd, S. Bosch, G. D. Palma, B. Kiani, Z. Liu, M. Marvian, P. Rebentrost, and D. M. Arvidsson-Shukur (2020) Quantum polar decomposition algorithm. External Links: 2006.00841, Link Cited by: §C.3, §4. [34] S. Lloyd, M. Schuld, A. Ijaz, J. Izaac, and N. Killoran (2020) Quantum embeddings for machine learning. External Links: 2001.03622, Link Cited by: §6. [35] G. H. Low and I. L. Chuang (2017-01) Optimal hamiltonian simulation by quantum signal processing. Physical Review Letters 118 (1). External Links: ISSN 1079-7114, Link, Document Cited by: §7. [36] G. H. Low and I. L. Chuang (2019-07) Hamiltonian simulation by qubitization. Quantum 3, p. 163. External Links: ISSN 2521-327X, Link, Document Cited by: §7. [37] G. H. Low and Y. Su (2026-02) Quantum eigenvalue processing. SIAM J. Comput. 55 (1), p. 135–215. External Links: Link, Document, ISSN 0097-5397,1095-7111 Cited by: §7. [38] R. Lu, R. Zhang, W. Li, Z. Wei, D. Deng, and Z. Liu (2026-6 January) A unified frequency principle for quantum and classical machine learning. arXiv [quant-ph]. External Links: Link, 2601.03169, Document Cited by: §2. [39] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven (2018-16 November) Barren plateaus in quantum neural network training landscapes. Nat. Commun. 9 (1), p. 4812. External Links: Link, Document, ISSN 2041-1723,2041-1723 Cited by: §2, §6. [40] H. Mhiri, L. Monbroussou, M. Herrero-Gonzalez, S. Thabet, E. Kashefi, and J. Landman (2025-3 September) Constrained and vanishing expressivity of quantum fourier models. Quantum 9 (1847), p. 1847. External Links: Link, 2403.09417v3, Document, ISSN 2521-327X Cited by: §2, §3. [41] A. Pérez-Salinas, A. Cervera-Lierta, E. Gil-Fuster, and J. I. Latorre (2020-6 February) Data re-uploading for a universal quantum classifier. Quantum 4, p. 226. External Links: Link, 1907.02085v2, Document, ISSN 2521-327X Cited by: §1, §2, §3. [42] A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever (2018) Improving language understanding by generative pre-training. Note: Online External Links: Link Cited by: §1. [43] N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. A. Hamprecht, Y. Bengio, and A. Courville (2018-22 June) On the spectral bias of neural networks. arXiv [stat.ML]. External Links: Link, 1806.08734, Document Cited by: §1, §2. [44] O. Rippel, J. Snoek, and R. P. Adams (2015) Spectral representations for convolutional neural networks. In Advances in Neural Information Processing Systems, Vol. 28, p. 2449–2457. External Links: Link Cited by: §1, §2. [45] M. Schuld and F. Petruccione (2021) Representing data on a quantum computer. In Machine Learning with Quantum Computers, M. Schuld and F. Petruccione (Eds.), p. 147–176. External Links: Link, Document, ISBN 9783030830984 Cited by: §2. [46] M. Schuld, R. Sweke, and J. J. Meyer (2021-24 March) Effect of data encoding on the expressive power of variational quantum-machine-learning models. Phys. Rev. A 103 (3), p. 032430. External Links: Link, Document, ISSN 1050-2947 Cited by: §1, §1, §1, §2, §2, §3, §3, §7. [47] J. Tang, J. Zhang, and X. Sun (2026-10 February) SAQNN: spectral adaptive quantum neural network as a universal approximator. arXiv [quant-ph]. External Links: Link, 2602.09718, Document Cited by: §2. [48] S. Thanasilp, S. Wang, M. Cerezo, and Z. Holmes (2024-18 June) Exponential concentration in quantum kernel methods. Nat. Commun. 15 (1), p. 5200. External Links: Link, Document, ISSN 2041-1723,2041-1723 Cited by: §6, §7. [49] P. Wang, C. R. Myers, L. C. L. Hollenberg, and U. Parampalli (2025-06) Quantum hamiltonian embedding of images for data reuploading classifiers. Quantum Mach. Intell. 7 (1), p. 35. External Links: Link, Document, ISSN 2524-4906,2524-4914 Cited by: §1, §2, §3, §4. [50] X. Wang, H. Tao, and R. Wu (2025-13–19 Jul) Predictive performance of deep quantum data re-uploading models. In Proceedings of the 42nd International Conference on Machine Learning, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research, Vol. 267, p. 64491–64524. External Links: Link Cited by: §2. [51] Z. J. Xu, Y. Zhang, T. Luo, Y. Xiao, and Z. Ma (2019-19 January) Frequency principle: fourier analysis sheds light on deep neural networks. arXiv [cs.LG]. External Links: Link, 1901.06523 Cited by: §1, §2. [52] C. Zhu, S. He, Y. Chen, L. Zhang, and X. Wang (2026-21 January) Optimal hamiltonian recognition of unknown quantum dynamics. NPJ Quantum Inf. 12 (1), p. 36. External Links: Link, Document, ISSN 2056-6387,2056-6387 Cited by: §2. Appendix Appendix A Preliminaries on Quantum Computing This appendix introduces the quantum-computing notation and basic linear-algebraic objects needed to follow the paper, assuming no prior background in quantum computing or quantum information. A.1 Quantum states and bra-ket notation A quantum state is represented by a complex vector. The standard notation for a column vector is a ket, written as |ψ⟩ ψ. Its conjugate transpose is a bra, written as ⟨ψ| ψ. If |ψ⟩=(ψ0ψ1⋮ψd−1)∈ℂd, ψ= pmatrix _0\\ _1\\ \\ _d-1 pmatrix ^d, (28) then ⟨ψ|=(ψ0⋆ψ1⋆⋯ψd−1⋆,) ψ= pmatrix _0 & _1 &·s& _d-1 , pmatrix (29) where ⋆ denotes complex conjugation. The inner product between the two states is written as ⟨ϕ|ψ⟩=∑j=0d−1ϕj⋆ψj. φ|ψ= _j=0^d-1φ _j _j. (30) This is the usual complex inner product from linear algebra. The squared norm of a state is therefore ⟨ψ|ψ⟩=∑j=0d−1|ψj|2. ψ|ψ= _j=0^d-1| _j|^2. (31) A valid pure quantum state is normalised: ⟨ψ|ψ⟩=1. ψ|ψ=1. (32) The entries of |ψ⟩ ψ are called amplitudes, not probabilities. Measurement probabilities are obtained by taking squared magnitudes. For example, if |ψ⟩=α|0⟩+β|1⟩, ψ=α 0+β 1, (33) then the probability of observing outcome 0 is |α|2|α|^2, and the probability of observing outcome 11 is |β|2|β|^2. Normalisation requires |α|2+|β|2=1.|α|^2+|β|^2=1. (34) The notation |ϕ⟩⟨ψ| φ ψ (35) denotes an outer product, hence a matrix. In particular, |ψ⟩⟨ψ| ψ ψ (36) is the rank-one projector onto the one-dimensional subspace spanned by |ψ⟩ ψ. Projectors of this type appear throughout quantum machine learning because measurements and readouts are naturally written as projections onto subspaces. A.2 Computational basis states A single qubit is a two-dimensional complex vector. The standard basis vectors are |0⟩=(10),|1⟩=(01). 0= pmatrix1\\ 0 pmatrix, 1= pmatrix0\\ 1 pmatrix. (37) These are called the computational basis states. A general one-qubit pure state is |ψ⟩=α|0⟩+β|1⟩=(αβ), ψ=α 0+β 1= pmatrixα\\ β pmatrix, (38) with |α|2+|β|2=1|α|^2+|β|^2=1. For n qubits, the state space has dimension 2n2^n. The computational basis consists of all binary strings of length n: |b1b2⋯bn⟩,bj∈0,1. b_1b_2·s b_n, b_j∈\0,1\. (39) For example, for two qubits the computational basis is |00⟩,|01⟩,|10⟩,|11⟩. 00, 01, 10, 11. (40) A general two-qubit state is |ψ⟩=α00|00⟩+α01|01⟩+α10|10⟩+α11|11⟩, ψ= _00 00+ _01 01+ _10 10+ _11 11, (41) where |α00|2+|α01|2+|α10|2+|α11|2=1.| _00|^2+| _01|^2+| _10|^2+| _11|^2=1. (42) More generally, an n-qubit state can be written as |ψ⟩=∑b∈0,1nαb|b⟩,∑b∈0,1n|αb|2=1. ψ= _b∈\0,1\^n _b b, _b∈\0,1\^n| _b|^2=1. (43) Thus an n-qubit state is a length-2n2^n complex vector indexed by binary strings. A useful single-qubit state used in the paper is |+⟩=|0⟩+|1⟩2. += 0+ 1 2. (44) The initial state in our experiment is |+⟩⊗n, + n, (45) which means that every qubit starts in the |+⟩ + state. This is the uniform superposition over all computational basis states: |+⟩⊗n=12n∑b∈0,1n|b⟩. + n= 1 2^n _b∈\0,1\^n b. (46) A.3 Tensor products The tensor product is the operation used to combine quantum systems. If |a⟩=(a0a1),|b⟩=(b0b1), a= pmatrixa_0\\ a_1 pmatrix, b= pmatrixb_0\\ b_1 pmatrix, (47) then |a⟩⊗|b⟩=(a0b0a0b1a1b0a1b1). a b= pmatrixa_0b_0\\ a_0b_1\\ a_1b_0\\ a_1b_1 pmatrix. (48) For example, |0⟩⊗|1⟩=|01⟩=(0100). 0 1= 01= pmatrix0\\ 1\\ 0\\ 0 pmatrix. (49) It is common to omit the tensor-product symbol for the basis states, so |0⟩⊗|1⟩ 0 1 (50) is written simply as |01⟩. 01. (51) Tensor products also combine operators. If A acts on the first qubit and B acts on the second qubit, then A⊗BA B (52) acts on the two-qubit system. For example, X⊗IX I means ``apply X to the first qubit and do nothing to the second,′ while I⊗XI X means ′do nothing to the first qubit and apply X to the second." For n qubits, a local single-qubit operator acting on qubit j is implicitly tensored with identities on all other qubits, such as the YjY_j operator in the main text, which means Pauli Y applied to qubit j, with identity operators on the remaining qubits. A.4 Pauli matrices The Pauli matrices are four 2×22× 2 matrices that form a basic operator basis for one qubit: I I =(1001),X=(0110), = pmatrix1&0\\ 0&1 pmatrix, X= pmatrix0&1\\ 1&0 pmatrix, (53) Y Y =(0−i0),Z=(100−1). = pmatrix0&-i\\ i&0 pmatrix, Z= pmatrix1&0\\ 0&-1 pmatrix. Here I is the identity. The other three are the Pauli X,Y,X,Y, and Z operators. The Pauli X matrix flips the computational basis states: X|0⟩=|1⟩,X|1⟩=|0⟩.X 0= 1, X 1= 0. (54) The Pauli Z matrix leaves |0⟩ 0 unchanged and changes the sign of |1⟩ 1: Z|0⟩=|0⟩,Z|1⟩=−|1⟩.Z 0= 0, Z 1=- 1. (55) The Pauli Y is similar to a bit flip but includes complex phases: Y|0⟩=i|1⟩Y|1⟩=−i|0⟩.Y 0=i 1 Y 1=-i 0. (56) All Pauli matrices are Hermitian and unitary. Hermitian means P†=P,P =P, (57) and unitary means P†P=I.P P=I. (58) Each of X,Y,ZX,Y,Z has eigenvalues +1+1 and −1-1. This simple spectrum is why Pauli rotations generate Fourier-like factors in data reuploading circuits. Multi-qubit Pauli strings are tensor products of one-qubit Pauli operators. For example, X⊗Z,Y⊗I,Z⊗X⊗YX Z, Y I, Z X Y (59) are Pauli strings. In this paper, the trainable two-qubit SU(4) blocks are parameterised using the 15 non-identity two-qubit Pauli strings. These are all tensor products of I,X,Y,ZI,X,Y,Z on two qubits except I⊗I I. A.5 Quantum gates and unitary evolution A quantum gate is a unitary matrix U. Applying a gate to a state gives |ψ′⟩=U|ψ⟩. ψ =U ψ. (60) Unitary preserves normalisation: ⟨ψ′|ψ′⟩=⟨ψ|U†U|ψ⟩=⟨ψ|ψ⟩. ψ |ψ = ψU U ψ= ψ|ψ. (61) Many gates are written as matrix exponentials of Hermitian generators. If H is Hermitian, then U(t)=exp(−iHt)U(t)= (-iHt ) (62) is unitary. In this paper, the QSM data-encoding unitaries use the convention U(t;M)=exp(−i2H(M)t),U(t;M)= (- i2H(M)t ), (63) where H(M)H(M) is a Hamiltonian constructed from the input matrix M, and t is a trainable upload-time scale. Single-qubit rotation gates are also matrix exponentials. For example, Ry(θ)=exp(−i2θY),Rz(θ)=exp(−i2θZ).R_y(θ)= (- i2θ Y ), R_z(θ)= (- i2θ Z ). (64) These gates rotate a qubit state by an angle θ around the Y or Z axis of the Bloch sphere. For the purposes of this paper, the important point is that the data value can appear inside the rotation angle. A typical coordinate-wise rotation-gate encoding has the form Ry(αxj),R_y(α x_j), (65) where xjx_j is one input feature and α is a fixed scale. A.6 Observables, measurement, and projector readout An observable is a Hermitian matrix O. Given a state |ψ⟩ ψ, the expected value of O is ⟨O⟩ψ=⟨ψ|O|ψ⟩. O_ψ= ψ|O|ψ. (66) This is a real number when O is Hermitian. For a task with N classes, let m=⌈log2N⌉m= _2N . The leading m qubits form the label register, and class j is associated with the computational-basis projector Pj=|j⟩⟨j|⊗(n−m),j∈0,…,N−1,P_j= j j (n-m), j∈\0,…,N-1\, (67) where |j⟩ j is an m-qubit basis state and the identity acts on the remaining qubits. The corresponding raw projector mass is qj=⟨ψ|Pj|ψ⟩.q_j= ψP_j ψ. (68) If N=2mN=2^m, the projectors exhaust the label register and ∑jqj=1 _jq_j=1. Otherwise, some label states are unused. The experiments condition on the named class outcomes, pj=qj∑k=0N−1qk,∑k=0N−1qk>0,p_j= q_j _k=0^N-1q_k, _k=0^N-1q_k>0, (69) and predict the class with the largest pjp_j, equivalently the largest qjq_j. As noted in Section 3, the Fourier analysis concerns the raw expectation values qjq_j rather than the normalised ratios pjp_j. Appendix B Derivation of the Frequency Support of Common Data-Encoding Unitaries This section derives the maximal one-upload frequency supports stated in Equations 10, 11, 12 and 13. Throughout, the input consists of d scalar features =(x1,…,xd) x=(x_1,…,x_d). B.1 Fixed single-qubit RyR_y encoding For fixed single-qubit RyR_y rotation-gate encoding, the encoding layer is SRy()=⨂j=1dRy(αxj),S_R_y( x)= _j=1^dR_y(α x_j), (70) where α is a constant scaling factor and Ry(αxj)=e−iαxjYj/2,R_y(α x_j)=e^-iα x_jY_j/2, (71) with YjY_j denoting the Pauli-Y operator on the j-th qubit. Its eigenvalues satisfy sj∈+1,−1s_j∈\+1,-1\, with Yj|sj⟩Y=sj|sj⟩Y_j s_j_Y=s_j s_j_Y. The corresponding joint Y-basis is |s⟩Y=|s1⟩Y⊗⋯⊗|sd⟩Y,s=(s1,…,sd)∈±1d. s_Y= s_1_Y ·s s_d_Y, s=(s_1,…,s_d)∈\± 1\^d. (72) The joint generator is K()=∑j=1dαxj2Yj.K( x)= _j=1^d α x_j2Y_j. (73) Because the YjY_j operators act on different qubits, they commute, and hence SRy()=e−iK()S_R_y( x)=e^-iK( x). In the joint Y-basis, K()|s⟩Y=(α2∑j=1dsjxj)|s⟩Y.K( x) s_Y= ( α2 _j=1^ds_jx_j ) s_Y. (74) Because the joint Y-eigenstates form a complete orthonormal basis, the spectral theorem gives the operator identity SRy()=∑s∈±1dexp(−iα2∑j=1dsjxj)ΠsY,S_R_y( x)= _s∈\± 1\^d (- iα2 _j=1^ds_jx_j ) _s^Y, (75) where ΠsY _s^Y is the rank-one projector onto |s⟩Y s_Y. In particular, when applied to a joint eigenstate, SRy()|s⟩Y=exp(−iα2∑j=1dsjxj)|s⟩Y.S_R_y( x) s_Y= (- iα2 _j=1^ds_jx_j ) s_Y. The phase-carrier vector is ηs=α2s=(αs12,…,αsd2). _s= α2s= ( α s_12,…, α s_d2 ). (76) When this encoding is used in a one-upload raw projector output qjq_j of the form in Equation 3, the phase factors from the bra and ket combine. Consequently, the candidate frequency vectors in its Fourier expansion are the pairwise differences ηs′−ηs=α2(s′−s). _s - _s= α2(s -s). (77) Since (sj′−sj)/2∈−1,0,1(s_j -s_j)/2∈\-1,0,1\, setting nj=(sj′−sj)/2n_j=(s_j -s_j)/2 yields the full vector-valued support in Equation 10. B.2 Fixed single-qubit RyRzR_yR_z encoding For fixed single-qubit RyRzR_yR_z encoding, SRyRz()=⨂j=1dRz(βxj)Ry(αxj).S_R_yR_z( x)= _j=1^dR_z(β x_j)R_y(α x_j). (78) For one feature, Sj(xj)=Rz(βxj)Ry(αxj).S_j(x_j)=R_z(β x_j)R_y(α x_j). (79) The individual spectral decompositions are Ry(αxj)=∑s=±1ΠsYe−isαxj/2,Rz(βxj)=∑r=±1ΠrZe−irβxj/2.R_y(α x_j)= _s=± 1 _s^Ye^-isα x_j/2, R_z(β x_j)= _r=± 1 _r^Ze^-irβ x_j/2. (80) Therefore, Sj(xj) S_j(x_j) =(∑r=±1ΠrZe−irβxj/2)(∑s=±1ΠsYe−isαxj/2) = ( _r=± 1 _r^Ze^-irβ x_j/2 ) ( _s=± 1 _s^Ye^-isα x_j/2 ) (81) =∑r,s=±1ΠrZΠsYe−i(rβ+sα)xj/2. = _r,s=± 1 _r^Z _s^Ye^-i(rβ+sα)x_j/2. This is a finite carrier expansion with carrier rates ηr,s=rβ+sα2. _r,s= rβ+sα2. (82) The candidate one-feature frequencies in the model output are the pairwise differences ηr′,s′−ηr,s=(r′−r)β+(s′−s)α2. _r ,s - _r,s= (r -r)β+(s -s)α2. (83) Setting m=(r′−r)/2m=(r -r)/2 and n=(s′−s)/2n=(s -s)/2, where m,n∈−1,0,1m,n∈\-1,0,1\, yields the local support in Equation 11. Taking the Cartesian product over the d independently encoded features gives the full vector-valued support in Equation 12. B.3 Component-wise trainable-frequency RyR_y encoding For component-wise trainable-frequency RyR_y encoding, generalising the scalar-value version in [21], STF(,γ)=⨂j=1dRy(γjxj)=⨂j=1de−iγjxjYj/2,S_TF( x,γ)= _j=1^dR_y( _jx_j)= _j=1^de^-i _jx_jY_j/2, (84) where γ=(γ1,…,γd)γ=( _1,…, _d) contains the trainable scaling parameters. The same joint Y-basis derivation gives the carrier vectors ηs(γ)=(s1γ12,…,sdγd2). _s(γ)= ( s_1 _12,…, s_d _d2 ). (85) For nj=(sj′−sj)/2∈−1,0,1n_j=(s_j -s_j)/2∈\-1,0,1\, the pairwise differences are ηs′(γ)−ηs(γ)=(n1γ1,…,ndγd), _s (γ)- _s(γ)=(n_1 _1,…,n_d _d), (86) which yields the full one-upload vector support in Equation 13. Appendix C Derivations for Hamiltonian-Embedding Frequency Support C.1 Conditioning variable, padding and coefficient convention The Hamiltonian-embedding expansions differ from the coordinate-wise Fourier expansions in Section 3. Here the matrix M is fixed and the upload time is the Fourier variable. Let H(0)(M)∈ℝdH×dHH^(0)(M) ^d_H× d_H denote a matrix-derived Hamiltonian before Hilbert-space padding. For a circuit dimension D=2n≥dHD=2^n≥ d_H, the implemented operator is H(M)=H(0)(M)⊕D−dH.H(M)=H^(0)(M) 0_D-d_H. (87) The main text uses the same symbol for the unpadded and padded operators. Direct-sum padding preserves every nonzero eigenvalue and enlarges the zero eigenspace. Let νa(M)a∈(M)\ _a(M)\_a (M) be the distinct eigenvalues of H(M)H(M), and let Πa(M) _a(M) be the corresponding spectral projectors. The spectral theorem gives H(M)=∑a∈(M)νa(M)Πa(M),SH(t;M)=∑a∈(M)e−iνa(M)t/2Πa(M).H(M)= _a (M) _a(M) _a(M), S_H(t;M)= _a (M)e^-i _a(M)t/2 _a(M). (88) For an initial state |ψ0⟩ _0 and an observable O_ θ that includes the trainable mixer, direct substitution gives f1H(M;t,) f_1^H(M;t, θ) =⟨ψ0|SH(t;M)†OSH(t;M)|ψ0⟩ = _0S_H(t;M) O_ θS_H(t;M) _0 (89) =∑a,bAabH(M,)e−i[νb(M)−νa(M)]t/2, = _a,bA_ab^H(M, θ)e^-i[ _b(M)- _a(M)]t/2, where AabH(M,)=⟨ψ0|Πa(M)OΠb(M)|ψ0⟩.A_ab^H(M, θ)= _0 _a(M)O_ θ _b(M) _0. (90) Several ordered pairs (a,b)(a,b) can produce the same gap. The Fourier coefficient at frequency ω is therefore CωH(M,)=∑a,b:(νb(M)−νa(M))/2=ωAabH(M,).C_ω^H(M, θ)= _ subarrayca,b:\\ ( _b(M)- _a(M))/2=ω subarrayA_ab^H(M, θ). (91) Consequently, suppF(f1H)⊆Ω1H(M)=νb(M)−νa(M)2:a,b∈(M).supp_F(f_1^H) _1^H(M)= \ _b(M)- _a(M)2:a,b (M) \. (92) The inclusion can be strict because individual amplitudes can vanish or different amplitudes at the same frequency can cancel. Using eigenspace projectors rather than chosen eigenvectors also makes Equation 89 invariant under basis changes within a degenerate eigenspace. The implemented global encoders use an independent time tℓt_ at each reuploading layer. Expanding every upload in a depth-L circuit gives a multivariate trigonometric polynomial of the form fLH(M;,)=∑,AH(M,)exp[−i2∑ℓ=1L(νbℓ(M)−νaℓ(M))tℓ],f_L^H(M; t, )= _ a, bA_ a b^H(M, ) [- i2 _ =1^L ( _b_ (M)- _a_ (M) )t_ ], (93) where the amplitudes contain the ordered products of spectral projectors and intervening mixers. Its maximal candidate support in =(t1,…,tL) t=(t_1,…,t_L) is the Cartesian product [Ω1H(M)]×L[ _1^H(M)]^× L. If all layer times are instead constrained to one scalar variable, the exponents combine and the corresponding scalar candidate set is the L-fold Minkowski sum of Ω1H(M) _1^H(M). C.2 Global symmetric Hamiltonian embedding For the symmetric encoder, M is first zero-column padded to a square matrix M~ M when p>qp>q. The natural Hamiltonian is Hsym(0)(M)=M~+M~T2,H_sym^(0)(M)= M+ M^T2, (94) followed by the direct-sum padding in Equation 87. Let Hsym(M)=∑a∈sym(M)λa(M)Πasym(M)H_sym(M)= _a _sym(M) _a(M) _a^sym(M) (95) be its decomposition into distinct eigenvalues and their eigenspace projectors. Its upload unitary is SHsym(t;M)=∑ae−iλa(M)t/2Πasym(M).S_H_sym(t;M)= _ae^-i _a(M)t/2 _a^sym(M). (96) Substitution into the observable output yields f1Hsym(M;t,) f_1^H_sym(M;t, θ) =∑a,be+iλa(M)t/2e−iλb(M)t/2 = _a,be^+i _a(M)t/2e^-i _b(M)t/2 (97) ×⟨ψ0|Πasym(M)OΠbsym(M)|ψ0⟩, × _0 _a^sym(M)O_ θ _b^sym(M) _0, which gives the pair amplitudes, aggregated coefficients and eigengap candidate set in Equations 18, 19 and 20. When the eigenspaces indexed by a and b are one-dimensional, one may choose normalised eigenvectors |va(M)⟩ v_a(M) and |vb(M)⟩ v_b(M). Writing da(M)=⟨va(M)|ψ0⟩d_a(M)= v_a(M)| _0 gives AabHsym(M,)=da(M)⋆db(M)⟨va(M)|O|vb(M)⟩.A_ab^H_sym(M, θ)=d_a(M) d_b(M) v_a(M)O_ θ v_b(M). (98) This rank-one expression illustrates the eigenvector dependence but is basis-dependent when an eigenvalue is degenerate. The projector expression in Equation 19 is the invariant statement used in the main text. Any natural or padding-induced zero eigenspaces are combined into a single projector associated with λ=0λ=0. C.3 Global block-Hamiltonian embedding For M∈ℝp×qM ^p× q, define the natural block Hamiltonian Hblock(0)(M)=(p×pMMTq×q)H_block^(0)(M)= pmatrix 0_p× p&M\\ M^T& 0_q× q pmatrix (99) on ℝp⊕ℝqR^p ^q. Lloyd et al. use the Hermitian Hamiltonian (0A†A0) ( smallmatrix0&A \\ A&0 smallmatrix ) in their quantum polar decomposition algorithm [33]. For real M, setting A=MTA=M^T gives Equation 99. Their construction uses access to this Hamiltonian to implement transformations associated with the polar decomposition; the present model instead uses its time evolution directly as a data-encoding unitary. Let the compact singular value decomposition of M be M=∑j=1rσj(M)|uj(M)⟩⟨vj(M)|,r=rank(M),M= _j=1^r _j(M) u_j(M) v_j(M), r=rank(M), (100) where σj(M)>0 _j(M)>0, uj(M)∈ℝpu_j(M) ^p and vj(M)∈ℝqv_j(M) ^q. The singular-vector equations are M|vj⟩=σj|uj⟩,MT|uj⟩=σj|vj⟩.M v_j= _j u_j, M^T u_j= _j v_j. (101) Introduce the direct-sum vectors |uj⊕0⟩=(ujq),|0⊕vj⟩=(pvj). u_j 0= pmatrixu_j\\ 0_q pmatrix, 0 v_j= pmatrix 0_p\\ v_j pmatrix. (102) Equation (99) then gives Hblock(0)|uj⊕0⟩=σj|0⊕vj⟩,Hblock(0)|0⊕vj⟩=σj|uj⊕0⟩.H_block^(0) u_j 0= _j 0 v_j, H_block^(0) 0 v_j= _j u_j 0. (103) Hence its restriction to j=span|uj⊕0⟩,|0⊕vj⟩S_j=span\ u_j 0, 0 v_j\ (104) is σjX _jX. The normalised eigenvectors on this subspace are |j,+⟩=|uj⊕0⟩+|0⊕vj⟩2,|j,−⟩=|uj⊕0⟩−|0⊕vj⟩2, j,+= u_j 0+ 0 v_j 2, j,-= u_j 0- 0 v_j 2, (105) and satisfy Hblock(0)(M)|j,+⟩=+σj(M)|j,+⟩,Hblock(0)(M)|j,−⟩=−σj(M)|j,−⟩.H_block^(0)(M) j,+=+ _j(M) j,+, H_block^(0)(M) j,-=- _j(M) j,-. (106) The natural zero eigenspace is kerHblock(0)=ker(MT)⊕ker(M). H_block^(0)= (M^T) (M). (107) Direct-sum Hilbert-space padding adds a further zero subspace of dimension D−(p+q)D-(p+q). Let Π0block(M) _0^block(M) project onto their combined span. For a chosen orthonormal singular-vector basis, Hblock(M)=∑j=1rσj(M)(|j,+⟩⟨j,+|−|j,−⟩⟨j,−|)+0Π0block(M).H_block(M)= _j=1^r _j(M) ( j,+ j,+- j,- j,- )+0\, _0^block(M). (108) If singular values are repeated, the rank-one projectors associated with each common eigenvalue are summed to obtain the invariant spectral projectors Πablock(M) _a^block(M) used in the main text. Exponentiating Equation 108 gives SHblock(t;M)=∑j=1r[e−iσj(M)t/2|j,+⟩⟨j,+|+e+iσj(M)t/2|j,−⟩⟨j,−|]+Π0block(M).S_H_block(t;M)= _j=1^r [e^-i _j(M)t/2 j,+ j,++e^+i _j(M)t/2 j,- j,- ]+ _0^block(M). (109) The generic expansion in Equation 89 now has eigenvalue labels μa(M)∈+σj(M),−σj(M),0 _a(M)∈\+ _j(M),- _j(M),0\. Pairwise differences give Ω1Hblock(M)=0∪±σj(M)−σk(M)2,±σj(M)+σk(M)2:j,k=1,…,r∪Ω0(M), _1^H_block(M)=\0\∪ \± _j(M)- _k(M)2,\ ± _j(M)+ _k(M)2:j,k=1,…,r \∪ _0(M), (110) where Ω0(M)=±σj(M)/2:j=1,…,r,Π0block(M)≠0,∅,Π0block(M)=0. _0(M)= cases \± _j(M)/2:j=1,…,r \,& _0^block(M)≠ 0,\\ ,& _0^block(M)=0. cases (111) The terms with j=kj=k in the half-sum set include ±σj(M)± _j(M). Equations (23) and (24) give the invariant coefficient and candidate-support forms; the realised support can again be smaller. C.4 Patch-local block-Hamiltonian embedding Partition M into R non-overlapping patches P1,…,PRP_1,…,P_R as defined in Appendix E. For patch r, write Hr≡Hblock(Pr)=∑ar∈rμr,ar(Pr)Πr,ar(Pr),H_r≡ H_block(P_r)= _a_r _r _r,a_r(P_r) _r,a_r(P_r), (112) where the distinct eigenvalues consist of the signed singular values of PrP_r and zero when present. Because the local uploads act on disjoint registers, their joint one-upload unitary is Spatch-Hblock(;M) S_patch-H_block( t;M) =⨂r=1Rexp[−i2Hrtr] = _r=1^R [- i2H_rt_r ] (113) =∑exp[−i2∑r=1Rμr,ar(Pr)tr]Π(M), = _ a [- i2 _r=1^R _r,a_r(P_r)t_r ] _ a(M), where =(t1,…,tR) t=(t_1,…,t_R), =(a1,…,aR) a=(a_1,…,a_R) and Π(M)=⨂r=1RΠr,ar(Pr). _ a(M)= _r=1^R _r,a_r(P_r). (114) The one-upload observable output is f1patch-Hblock(M;,) f_1^patch-H_block(M; t, θ) =∑,Apatch-Hblock(M,) = _ a, bA_ a b^patch-H_block(M, θ) (115) ×exp[−i2∑r=1R(μr,br(Pr)−μr,ar(Pr))tr], × [- i2 _r=1^R ( _r,b_r(P_r)- _r,a_r(P_r) )t_r ], with pair amplitudes Apatch-Hblock(M,)=⟨ψ0|Π(M)OΠ(M)|ψ0⟩.A_ a b^patch-H_block(M, θ)= _0 _ a(M)O_ θ _ b(M) _0. (116) The frequency vector associated with (,)( a, b) is 12(μ1,b1(P1)−μ1,a1(P1),…,μR,bR(PR)−μR,aR(PR)). 12 ( _1,b_1(P_1)- _1,a_1(P_1),…, _R,b_R(P_R)- _R,a_R(P_R) ). (117) Taking all index pairs gives the Cartesian-product candidate set in Equation 25. At layer ℓ , trt_r is replaced by the implemented parameter tℓ,rt_ ,r. For a depth-L circuit, the maximal support in the LRLR independent time variables is the Cartesian product of the R local candidate sets over all L layers; particular frequencies can still be absent through zero or cancelling amplitudes. Appendix D Additional Information on the Benchmark Datasets We evaluate the models on one real-world-data benchmark and two controlled synthetic benchmarks. The real-world data benchmark tests whether the proposed encoders are useful for a small, nontrivial handwritten-digit classification problem with matrix-valued inputs. The synthetic benchmarks are designed to isolate spectral structure: one task is controlled by eigenvalue gaps, and the other by singular values. Together, these datasets let us compare performance on realistic matrix-structured inputs and on tasks where the relevant decision is explicitly spectral. D.1 Pendigits benchmark The Pendigits dataset [1] is used as the benchmark for real-world classification tasks. We use two matrix-valued representations of each handwritten digit: the DYN representation, arranged as an 8×28× 2 rectangular trajectory matrix, and the STA4 representation, arranged as a 4×44× 4 bitmap. These two representations provide complementary views of the same underlying object. DYN retains an ordered pen-trajectory structure, whereas STA4 gives a compact square matrix representation. The two representations are useful for testing whether a QFM data-encoding unitary benefits from treating the input as a structured matrix rather than an unstructured feature vector. For all Pendigits experiments, we use the original train/test split supplied with the dataset. A validation set is formed by holding out 10% of the original training data. This validation is class-stratified so that each class remains represented in both the training and validation subsets. The final split used in the full 1010-class experiments contains 67446744 training examples, 750750 validation examples, and 34983498 test examples. Before training, Pendigits features are standardised using statistics computed only from the post-split training set. The resulting training-set mean and standard deviation are then applied to the training, validation and test sets. After standardisation, each sample is reshaped into its corresponding matrix representation for the Hamiltonian-based and patch-local data-encoding unitaries. The visual examples shown in the main text (Figure 2(a)) display raw feature values for visual interpretation; the models are trained on the standardised inputs described above. The Pendigits task is included to provide a compact, real-world benchmark whose input dimension is small enough for exact statevector simulation, especially for rotation-gate-based data-encoding unitaries, while still retaining nontrivial matrix structure. The DYN and STA4 representations also test different notions of structure: trajectory geometry in the former and spatially arranged pixel-feature geometry in the latter. This makes Pendigits a useful test case for comparing coordinate-wise rotation-gate data-encoding unitaries, patch-local data-encoding unitaries and global QSMs under the same training protocol. D.2 Synthetic spectral benchmarks In addition to Pendigits, we use two controlled binary classification tasks whose labels are defined by spectral properties of latent clean matrices. These tasks are designed to test whether the model can exploit eigenvalue and singular-value structure rather than only component-value-level features. In both tasks, the classifier observes a noisy matrix, while the label is determined by the spectrum used to construct the corresponding clean matrix. The main experiments use 40964096 samples per task, a data generation seed of 0, a validation fraction of 0.10.1, and a relative noise level of 0.050.05. For the SYNTHETIC EIGENGAP task, each clean sample is generated as a random symmetric matrix with prescribed eigenvalues. For sample index i, a Gaussian random matrix is first drawn and factorised to obtain an orthogonal matrix Q. Independently, a vector of real eigenvalues is sampled from a standard normal distribution and sorted in ascending order, λi,1≤λi,2≤⋯≤λi,d. _i,1≤ _i,2≤·s≤ _i,d. (118) The clean matrix is then Si=Qi(λi,1,λi,2,⋯,λi,d)QT.S_i=Q_idiag( _i,1, _i,2,·s, _i,d)Q^T. (119) In the main experiments, d=4d=4. The binary label is determined by the largest eigengap: yi=[|λi,d−λi,d−1|>τeig],y_i= 1 [| _i,d- _i,d-1|> _eig ], (120) with τeig=0.75 _eig=0.75. The observed input matrix is obtained by adding scaled Gaussian noise to the clean matrix. Specifically, for each clean matrix SiS_i, an independent Gaussian noise matrix ZiZ_i of the same shape is sampled and rescaled as Ei=ϵ‖Si‖F(‖Zi‖F,10−12)Zi,E_i=ε ||S_i||_Fmax(||Z_i||_F,10^-12)Z_i, (121) where ||⋅||F||·||_F denotes for Frobenius norm. Then the input data becomes Xi=Si+Ei.X_i=S_i+E_i. (122) The main experiments use ϵ=0.05ε=0.05, so the added noise has Frobenius norm equal to 5%5\% of the clean sample norm, up to the numerical denominator guard. The label is still computed from the latent clean eigenvalues, not from the noisy input data matrix. This construction is aligned with the symmetric Hamiltonian embedding. In the clean limit, the input-derived symmetric Hamiltonian has eigenvalues λj(Si) _j(S_i), and one upload has phase-difference carriers proportional to λk(Si)−λj(Si)2. _k(S_i)- _j(S_i)2. (123) Thus the label depends directly on one of the spectral gaps that can appear in the sample-conditioned Hamiltonian frequency support. The task is therefore a controlled probe of whether the model can use eigengap information rather than only raw matrix entries. For the SYNTHETIC SINGULAR value task, each clean sample is generated as a rectangular matrix with prescribed singular values. For sample index i, independent Gaussian matrices are factorised to obtain orthogonal factors UiU_i and ViV_i. A vector of singular values is sampled uniformly on [0,2][0,2], sorted in descending order, σi,1≥σi,2≥⋯≥σi,r,r=(p,q), _i,1≥ _i,2≥·s≥ _i,r, r=min(p,q), (124) and placed on the diagonal of a rectangular p×qp× q matrix Σi _i. The clean matrix is Ri=UiΣiViT.R_i=U_i _iV_i^T. (125) In the main experiments, p=8,q=2p=8,q=2, and hence r=2r=2. The label is yi=[σi,1+σi,2>τsv],y_i= 1 [ _i,1+ _i,2> _sv ], (126) where τsv=2 _sv=2. As the eigengap task, the observed SYNTHETIC SINGULAR value task input is a noisy version of the clean latent matrix. The label is determined by the latent singular values, not by the singular values of the noisy observed matrix. The main experiments again use ϵ=0.05ε=0.05. This task is aligned with the block Hamiltonian embedding. For a rectangular matrix M, the block Hamiltonian Hblock(M)=(0MMT0)H_block(M)= pmatrix0&M\\ M^T&0 pmatrix (127) has nonzero eigenvalues ±σj(M)± _j(M), where σj(M) _j(M) are the singular values of M. The one-upload phase-difference support therefore contains singular-value differences and sums, including the terms of the form ±σj(M)−σk(M)2,±σj(M)+σk(M)2.± _j(M)- _k(M)2, ± _j(M)+ _k(M)2. (128) The singular-value task uses a decision rule based on σ1+σ2 _1+ _2, which is one of the spectral combinations naturally exposed by this block construction in the clean setting. For both synthetic tasks, the train/validation/test split is generated by randomly permuting the 4096 samples with the fixed data seed. The first 20% of the permuted samples are assigned to the test set. A validation set is then formed from the 10% of the remaining samples, and the rest are used for training. This yields 819 test examples, 328 validation examples, and 2949 training examples per synthetic task. The two synthetic benchmarks therefore play complementary diagnostic roles. The eigengap task tests whether a model can exploit eigenvalue-gap information in a matrix whose clean signal is symmetric. The singular-value task tests whether a model can exploit singular-value sums in a rectangular matrix. Since the latent spectral rule is known by construction, these tasks provide controlled evidence about spectral inductive bias, while the added matrix noise prevents the benchmarks from reducing to noiseless spectral lookup. Appendix E Additional Information on the Patch-Local Encoders and Mixer Parameterisation E.1 Patch-local SU(4) and block-Hamiltonian data-encoding unitaries The patch-local data-encoding unitaries operate on four non-overlapping 2×22× 2 patches. For a 4×44× 4 matrix M=(M1:2,1:2M1:2,3:4M3:4,1:2M3:4,3:4),M= pmatrixM_1:2,1:2&M_1:2,3:4\\ M_3:4,1:2&M_3:4,3:4 pmatrix, (129) we define P1=M1:2,1:2,P2=M1:2,3:4,P3=M3:4,1:2,P4=M3:4,3:4.P_1=M_1:2,1:2, P_2=M_1:2,3:4, P_3=M_3:4,1:2, P_4=M_3:4,3:4. (130) For an 8×28× 2 matrix, the patches are Pr=M2r−1:2r,1:2,r=1,2,3,4.P_r=M_2r-1:2r,1:2, r=1,2,3,4. (131) Each patch is represented by a flattened vector r=vec(Pr)∈ℝ4 p_r=vec(P_r) ^4 (132) and is uploaded to a dedicated two-qubit system. Patch r acts on qubits 2r−12r-1 and 2r2r, so all four patch uploads together act on eight qubits. This applies to fixed patch-SU(4)SU(4), trainable patch-SU(4)SU(4), and non-overlap patch block-Hamiltonian embedding. E.2 SU(4) convention Both patch-SU(4)SU(4) data-encoding unitaries and the brick-wall mixer use the same two-qubit SU(4)SU(4) parametrisation. Let =(G1,⋯,G15)G=(G_1,·s,G_15) (133) be the ordered list of all non-identity two-qubit Pauli strings, I⊗X,I⊗Y,I⊗Z,X⊗I,X⊗X,X⊗Y,X⊗Z,Y⊗I,Y⊗X,Y⊗Y,Y⊗Z,Z⊗I,Z⊗X,Z⊗Y,Z⊗Z.I X,I Y,I Z,X I,X X,X Y,X Z,Y I,Y X,Y Y,Y Z,Z I,Z X,Z Y,Z Z. (134) For a coefficient vector θ∈ℝ15θ ^15, the corresponding two-qubit unitary is USU(4)(θ)=exp[i∑a=115θaGa].U_SU(4)(θ)= [i _a=1^15 _aG_a ]. (135) The sign convention is therefore +i+i for the SU(4)SU(4) parametrisation, different from the Hamiltonian embedding unitaries. E.3 Fixed patch-SU(4) encoder The fixed patch-SU(4) encoder maps each flattened patch r∈ℝ4 p_r ^4 to an SU(4) coefficient vector by θ(Pr)=A0r,θ(P_r)=A_0 p_r, (136) where A0∈ℝ15×4A_0 ^15× 4 is a fixed matrix. In the main experiments, the fixed map is A0=12(10000100001000011100001110100101100−101−101−100001−111−1−11−11−11−1−11)A_0= 12 pmatrix1&0&0&0\\ 0&1&0&0\\ 0&0&1&0\\ 0&0&0&1\\ 1&1&0&0\\ 0&0&1&1\\ 1&0&1&0\\ 0&1&0&1\\ 1&0&0&-1\\ 0&1&-1&0\\ 1&-1&0&0\\ 0&0&1&-1\\ 1&1&-1&-1\\ 1&-1&1&-1\\ 1&-1&-1&1 pmatrix (137) The patch upload is then Upatch(Pr)=exp(i∑a=115[A0r]aGa).U_patch(P_r)= (i _a=1^15[A_0 p_r]_aG_a ). (138) The same A0A_0 is used for every patch and every reuploading layer. E.4 Trainable patch-SU(4) encoder The trainable patch-SU(4) encoder uses the same four-patch decomposition and the same local SU(4) exponential, but the linear patch map is made trainable. For layer ℓ and patch r, let Aℓ,r∈ℝ15×4.A_ ,r ^15× 4. (139) The patch angles are θℓ,r(Pr)=Aℓ,rr, _ ,r(P_r)=A_ ,r p_r, (140) and the upload is Utrainable−patch,ℓ,r(Pr)=exp(i∑a=115[Aℓ,rr]aGa).U_trainable-patch, ,r(P_r)= (i _a=1^15[A_ ,r p_r]_aG_a ). (141) The trainable parameter tensor has shape (Aℓ,r)ℓ=1,r=1L,4∈ℝL×4×15×4. (A_ ,r )^L,4_ =1,r=1 ^L× 4× 15× 4. (142) It is initialised near the fixed patch map: Aℓ,r=A0+Δℓ,r,A_ ,r=A_0+ _ ,r, (143) where the entries of Δℓ,r _ ,r are initialised as zero-mean Gaussian noise. The main experiments use a noise scale of 0.010.01. With zero initialisation noise, the trainable patch encoder exactly reproduces the fixed patch-SU(4) angle map at initialisation. E.5 Non-overlapping patch-local block-Hamiltonian data-encoding unitary The non-overlapping patch-local block-Hamiltonian data-encoding unitary replaces the local SU(4)SU(4) patch map with structured Hamiltonian evolution. For each 2×22× 2 patch PrP_r, define Hpatch(Pr)=(0PrPrT0)∈ℝ4×4.H_patch(P_r)= pmatrix0&P_r\\ P_r^T&0 pmatrix ^4× 4. (144) At layer ℓ , the patch upload is Upatch−block,ℓ,r(Pr)=exp[−i2Hpatch(Pr)tℓ,r].U_patch-block, ,r(P_r)= [- i2H_patch(P_r)t_ ,r ]. (145) This unitary is applied to the two-qubit subsystem assigned to patch r. When trainable upload times are enabled, the parameter tensor is t∈ℝL×4.t ^L× 4. (146) Thus the model has one scalar Hamiltonian upload time for each layer and each patch. If no fixed time is supplied, all entries are initialised to 1/L1/L. The patch-local block-Hamiltonian QSM differs from its global counterpart. The global QSM constructs one block Hamiltonian from the full matrix M, so its spectral structure is determined by the singular values and singular vectors of the whole input matrix. The patch-local QSM instead computes four independent 2×22× 2 block Hamiltonians, preserving local patch singular geometry while discarding cross-patch singular-vector coupling. E.6 Brick-wall SU(4) mixer After every application of a data-encoding unitary, the model applies a trainable nearest-neighbour brick-wall mixer. For n qubits, define the ordered pair set ℬn=(1,2),(3,4),⋯⋃(2,3),(4,5),⋯,B_n= \(1,2),(3,4),·s \ \(2,3),(4,5),·s \, (147) where the even-starting pairs are applied before the odd-starting pairs. For each layer ℓ and each pair b∈ℬnb _n, the model applies Vℓ,b=exp(i∑a=115θℓ,b,aSU(4)Ga).V_ ,b= (i _a=1^15θ^SU(4)_ ,b,aG_a ). (148) The full mixer layer is the ordered product of these local blocks: Vℓ=∏b∈ℬnVℓ,b.V_ = _b _nV_ ,b. (149) For the n-qubit model used here, |ℬn|=n−1|B_n|=n-1, so the mixer contributes 15(n−1)15(n-1) trainable parameters per uploading layer. The mixer parameter tensor has shape θSU(4)∈ℝL×(n−1)×15.θ^SU(4) ^L×(n-1)× 15. (150) It is initialised from a zero-mean Gaussian distribution with scale 0.01. The number of qubits in the QFM circuit depends on the data-encoding unitary. Matrix-value-wise rotation-gate data-encoding unitaries use one qubit per scalar feature, so both 4×44× 4 and 8×28× 2 inputs use 16 qubits. The three patch-local data-encoding unitaries split the input into four 2×22× 2 patches and assign each patch to two qubits, giving 8 qubits in total. The global QSMs instead use the dimension of the input-derived Hamiltonian: HsymH_sym uses 2 qubits for 4×44× 4 inputs and 3 qubits for 8×28× 2 inputs, while HblockH_block uses 3 and 4 qubits, respectively, before any label-register expansion. For the ten-class Pendigits task, the projector readout requires label qubits, so the global QSMs are promoted to 4 qubits when their natural Hamiltonian dimension is smaller. This promotion is implemented as direct-sum padding H↦H⊕H H 0, equivalently U↦U⊕U U I, rather than by tensoring extra idle qubits. Binary synthetic tasks require only one label qubit, so no additional expansion is needed for the global QSM data-encoding unitaries. At L=32L=32, the resulting total numbers of trainable scalars, including the mixer and any trainable encoder parameters, are 7200 for fixed-RyR_y, 7200 for fixed-RyRzR_yR_z, 7712 for trainable-frequency RyR_y, 3360 for fixed patch-SU(4)SU(4), 11040 for trainable patch-SU(4)SU(4) and 3488 for the patch-local block-Hamiltonian QSM. For the four-qubit Pendigits global QSMs, both HsymH_sym and HblockH_block contain 1472 trainable scalars. On SYNTHETIC EIGENGAP, the corresponding totals are 512 for the two-qubit symmetric-Hamiltonian QSM and 992 for the three-qubit block-Hamiltonian QSM; on SYNTHETIC SINGULAR, they are 992 for the three-qubit symmetric-Hamiltonian QSM and 1472 for the four-qubit block-Hamiltonian QSM. The projector readout adds no trainable parameters. These totals make the resource mismatch explicit: the main comparison holds the data and optimisation protocol fixed, but it does not normalise circuit width or trainable-parameter count. Appendix F Additional Performance Results on the Pendigits and Synthetic Datasets The main performance archive contains 3840 complete records with finite reported metrics. Of these records, 319 are marked as having been generated from a dirty implementation worktree. However, all 20 records contributing to each of the four depth-32 winner aggregates quoted in the main text, and all 20 records contributing to each of the four global HblockH_block maxima across the tested depths, have git_dirty=False. We therefore retain those headline aggregates while disclosing the broader archive provenance. (a) (b) Figure 6: (a) Best validation accuracy on the Pendigits benchmark as a function of reuploading depth. Each point is the mean over 20 seeds, and the error bars show one standard deviation. The metric is the maximum validation accuracy observed over the 2000-step training trajectory for each run. The QSMs achieve the highest best-validation accuracies on both DYN and STA4. (b) Best validation accuracy on the two synthetic spectral benchmarks as a function of reuploading depth. Each point is the mean over 20 seeds, and the error bars show one standard deviation. The global QSMs achieve the strongest validation performance on both SYNTHETIC EIGENGAP and SYNTHETIC SINGULAR, with the block-Hamiltonian QSM giving the best overall validation accuracy. (a) (b) Figure 7: (a) Area under the validation-accuracy training curve on Pendigits. For each run, AUC is computed by trapezoidal integration of validation accuracy over recorded evaluation steps. The plotted values are means over 20 seeds with one-standard-deviation error bars. The patch-local and global block-Hamiltonian QSMs achieve the highest AUCs, indicating stronger sustained validation performance throughout training. (b) Area under the validation-accuracy training curve on SYNTHETIC EIGENGAP and SYNTHETIC SINGULAR. AUC is computed by trapezoidal integration over validation checkpoints and is reported in step units. Each point is the mean over 20 seeds, with error bars showing one standard deviation. The global block-Hamiltonian QSM gives the largest AUC on both synthetic tasks. Table 3: Additional validation-set performance summaries for the Pendigits and synthetic benchmark runs. Best validation accuracy is the largest validation accuracy reached during training for each run. Validation AUC is the unnormalised trapezoidal area under the validation-accuracy curve over recorded evaluation steps. Values are mean ± standard deviation over 20 random seeds. Dataset Encoder Best-val depth Best val. acc. (%) Best-AUC depth Val. AUC Val. AUC / 1900 Pendigits DYN fixed-RyR_y 16 78.84±1.3478.84± 1.34 8 1315.0±25.31315.0± 25.3 0.6920.692 fixed-RyRzR_yR_z 16 67.27±3.5467.27± 3.54 4 1161.2±18.21161.2± 18.2 0.6110.611 TF-RyR_y 32 84.41±2.2284.41± 2.22 8 1424.4±31.81424.4± 31.8 0.7500.750 patch-SU(4)SU(4) 4 51.61±2.1651.61± 2.16 4 859.0±32.2859.0± 32.2 0.4520.452 trainable patch-SU(4)SU(4) 2 79.97±1.8579.97± 1.85 2 1387.0±31.01387.0± 31.0 0.7300.730 patch HblockH_ block 32 96.30±0.6696.30± 0.66 32 1772.6±16.51772.6± 16.5 0.9330.933 HsymH_ sym 32 96.74±0.5796.74± 0.57 16 1740.2±26.41740.2± 26.4 0.9160.916 HblockH_ block 32 97.29±0.5397.29± 0.53 16 1780.7±15.81780.7± 15.8 0.9370.937 Pendigits STA4 fixed-RyR_y 8 65.05±3.4365.05± 3.43 8 1046.4±50.91046.4± 50.9 0.5510.551 fixed-RyRzR_yR_z 16 54.63±3.2354.63± 3.23 4 903.2±25.1903.2± 25.1 0.4750.475 TF-RyR_y 8 74.04±2.0374.04± 2.03 8 1177.8±35.21177.8± 35.2 0.6200.620 patch-SU(4)SU(4) 4 38.31±1.3438.31± 1.34 2 632.7±32.4632.7± 32.4 0.3330.333 trainable patch-SU(4)SU(4) 2 59.87±3.1659.87± 3.16 2 972.9±32.6972.9± 32.6 0.5120.512 patch HblockH_ block 32 90.78±0.6490.78± 0.64 32 1664.0±11.71664.0± 11.7 0.8760.876 HsymH_ sym 16 76.95±1.7176.95± 1.71 8 1283.8±41.71283.8± 41.7 0.6760.676 HblockH_ block 32 92.25±0.8692.25± 0.86 8 1631.4±34.91631.4± 34.9 0.8590.859 Synthetic eigengap fixed-RyR_y 8 70.64±1.0370.64± 1.03 8 1262.0±15.51262.0± 15.5 0.6640.664 fixed-RyRzR_yR_z 8 69.63±1.0269.63± 1.02 8 1248.3±20.51248.3± 20.5 0.6570.657 TF-RyR_y 8 71.51±1.1871.51± 1.18 8 1281.6±15.81281.6± 15.8 0.6750.675 patch-SU(4)SU(4) 2 66.62±1.2666.62± 1.26 2 1216.9±20.31216.9± 20.3 0.6400.640 trainable patch-SU(4)SU(4) 2 73.35±1.0973.35± 1.09 2 1338.0±19.51338.0± 19.5 0.7040.704 patch HblockH_ block 16 77.65±3.0277.65± 3.02 16 1378.2±42.61378.2± 42.6 0.7250.725 HsymH_ sym 32 85.66±2.0185.66± 2.01 32 1512.4±52.51512.4± 52.5 0.7960.796 HblockH_ block 16 88.81±1.1988.81± 1.19 16 1599.3±20.21599.3± 20.2 0.8420.842 Synthetic singular fixed-RyR_y 16 83.60±1.2883.60± 1.28 16 1536.5±18.61536.5± 18.6 0.8090.809 fixed-RyRzR_yR_z 16 82.82±1.3882.82± 1.38 16 1523.6±19.91523.6± 19.9 0.8020.802 TF-RyR_y 32 84.60±1.1584.60± 1.15 8 1556.2±14.81556.2± 14.8 0.8190.819 patch-SU(4)SU(4) 4 77.09±1.2377.09± 1.23 2 1409.7±11.01409.7± 11.0 0.7420.742 trainable patch-SU(4)SU(4) 4 86.68±1.3886.68± 1.38 4 1591.1±18.41591.1± 18.4 0.8370.837 patch HblockH_ block 32 87.74±1.9087.74± 1.90 32 1569.2±38.41569.2± 38.4 0.8260.826 HsymH_ sym 16 90.87±1.2490.87± 1.24 8 1654.2±19.01654.2± 19.0 0.8710.871 HblockH_ block 16 96.19±1.1196.19± 1.11 16 1772.0±25.51772.0± 25.5 0.9330.933 In this Appendix section, we report two additional validation-set summaries for the same training runs shown in the main text (Figure 3). The first is the best validation accuracy reached during training, see Figure 6. This metric records the largest validation accuracy observed at any evaluation checkpoint for a given run. It is useful for separating the model's best achieved validation performance from the final checkpoint used for the main held-out test evaluation. The second metric is the validation-accuracy area under the training curve (AUC), as shown in Figure 7. Let the validation checkpoints for one run be (t1,a1),(t2,a2),⋯,(tm,am),(t_1,a_1),(t_2,a_2),·s,(t_m,a_m), (151) where tit_i is the optimisation step and aia_i is the validation accuracy recorded at that step. The reported AUC is calculated by the trapezoidal rule, AUCval=∑i=1m−1ai+ai+12(ti+1−ti).AUC_val= _i=1^m-1 a_i+a_i+12(t_i+1-t_i). (152) This is the raw area in units of optimisation steps, not a normalised curve. In the main runs, validation is evaluated every 100 steps over 2000 training steps, so the maximum raw AUC is 1900 if validation accuracy is 1 at every recorded interval. A summary of these results can also be found in Table 3. The validation curves support the same qualitative conclusion as the final test results. On Pendigits DYN, the three QSMs achieve the highest best validation accuracies, with the global block-Hamiltonian QSM reaching 97.29%±0.53%97.29\%± 0.53\%, the symmetric-Hamiltonian QSM reaching 96.74%±0.57%96.74\%± 0.57\%, and the patch-local block-Hamiltonian QSM reaching 96.30%±0.66%96.30\%± 0.66\%. On Pendigits STA4, the global and patch-local block-Hamiltonian QSMs again lead, reaching 92.25%±0.86%92.25\%± 0.86\% and 90.78%±0.64%90.78\%± 0.64\%, respectively. The AUC plots provide a complementary view of training efficiency and sustained validation performance. A model can obtain a high best validation accuracy late in training but a smaller AUC if it learns slowly or is unstable. On DYN, the global block-Hamiltonian QSM achieves the highest mean validation AUC, closely followed by the patch-local block-Hamiltonian QSM. On STA4, the patch-local block-Hamiltonian QSM achieves the highest mean validation AUC, indicating more sustained validation performance throughout training. On SYNTHETIC EIGENGAP and SYNTHETIC SINGULAR, the global block-Hamiltonian QSM gives the largest validation AUC. Appendix G Additional Information on the Gradient Analysis This section defines the gradient statistics used in the main text and reports the supplementary diagnostics. G.1 Details for the Gradient Diagnostics Metrics Let B=(xi,yi)i=1mB= \(x_i,y_i) \^m_i=1 be the diagnostic training batch, with m=32m=32 in the reported experiments. The batch is selected deterministically using diagnostic seed 0 and held fixed within each run configuration. Let ℒ(θ;B)L(θ;B) be the mean cross-entropy loss on this batch. For a parameter group q, let gq(θ;B)=∇θqℒ(θ;B)∈ℝPq,g_q(θ;B)= _ _qL(θ;B) ^P_q, (153) where PqP_q is the number of parameters in group q. The parameter groups used in the plots are the shared mixer layer parameters θSU(4) _SU(4), trainable frequency parameters γ, Hamiltonian-embedding upload times tℓt_ or tℓ,rt_ ,r, and trainable patch-map parameters. For each run seed s, the initialisation diagnostic evaluates R=20R=20 parameter draws, generated with seeds s+r−1s+r-1 for r=1,…,Rr=1,…,R, on the same diagnostic batch. Let gq,r=gq(θ(r);B),r=1,⋯,R.g_q,r=g_q(θ^(r);B), r=1,·s,R. (154) For each parameter coordinate p, the initialisation gradient variance is vq,p=1R∑r=1R(gq,r,p−g¯q,p)2,g¯q,p=1R∑r=1Rgq,r,p.v_q,p= 1R _r=1^R(g_q,r,p- g_q,p)^2, g_q,p= 1R _r=1^Rg_q,r,p. (155) The plotted initialisation-variance statistic is p=1,⋯,Pqlog10(vq,p+ϵ),median_p=1,·s,P_q _10(v_q,p+ε), (156) with ϵ=10−30ε=10^-30. The RMS-gradient statistic for a single gradient vector is (gq)=1Pq∑p=1Pq|gq,p|2.RMS(g_q)= 1P_q _p=1^P_q|g_q,p|^2. (157) For initialisation, each run first averages the RMS over its R draws. For the final mode, each trained run contributes one final-checkpoint gradient vector. The reported means and error bars then aggregate these run-level statistics over 20 run seeds; the error bars are sample standard deviations across those seeds. The aggregate initialisation-diagnostic records do not store a git_dirty field. Their clean-state provenance therefore cannot be checked from the archived aggregate files, although the records used in the reported summaries are present and complete. The final-checkpoint diagnostic records do store this field, allowing the affected comparisons to be qualified below. For the empirical-Fisher diagnostic, the implementation computes one observed-label loss gradient per diagnostic-batch element: gi=∇θℓ(θ;xi,yi),i=1,⋯,m,g_i= _θ (θ;x_i,y_i), i=1,·s,m, (158) where ℓ is the per-sample cross-entropy loss and θ denotes the full trainable parameter vector. Let G=(g1T⋮gmT)∈ℝm×P.G= pmatrixg_1^T\\ \\ g_m^T pmatrix ^m× P. (159) The sample-space gradient Gram matrix is F^=1mGGT. F= 1mG^T. (160) This is the common machine-learning empirical-Fisher construction based on ground-truth labels [25]. It has the same nonzero eigenvalues as the parameter-space matrix GTG/mG^TG/m but is smaller because m=32m=32. It is not the true Fisher information, which averages over labels drawn from the model distribution, or the quantum Fisher information used in quantum natural-gradient methods. Nor should it be assumed to represent general loss curvature [25]. We use it only as a finite-batch diagnostic of the sampled gradient directions. At initialisation, each run contributes the first of its 20 parameter draws to this calculation; at the final checkpoint, each run contributes its trained parameters. Its trace is (F^)=∑jλj,tr( F)= _j _j, (161) and its participation-ratio effective rank [14] is reff=(∑jλj)2(∑jλj2,δ),r_eff= ( _j _j )^2max ( _j _j^2,δ ), (162) and the condition number is κ=λj>δλj(λj>δλj,δ),κ= max_ _j>δ _jmax (min_ _j>δ _j,δ ), (163) with damping threshold δ=10−8δ=10^-8. G.2 Additional Gradient Diagnostics Plots and Results Figure 8 reports the absolute RMS gradient scale at initialisation. It should be read together with the initialisation-variance plot in Figure 4(a). At depth 32, for example, fixed-RyR_y on Pendigits DYN has mixer-gradient RMS 0.527±0.0810.527± 0.081 but median log10 _10 mixer-gradient variance −11.04±0.08-11.04± 0.08. Nonzero gradient magnitude can therefore coexist with little coordinate-wise variation across initialisations; the two diagnostics measure different properties. Figure 8: Initialisation RMS gradient as a function of reuploading depth. For each run configuration, gradients are computed on a fixed diagnostic training batch for 20 random initialisations. The plotted value is the RMS gradient within each parameter group, averaged over initialisation draws and then over 20 run seeds; error bars show one standard deviation across seeds. This plot complements the initialisation-variance diagnostic (Figure 4(a)) by measuring absolute gradient scale rather than parameter-wise variance across initialisations. Figure 9 reports the trace, effective rank and condition number of the empirical-Fisher gradient Gram matrix at initialisation. These quantities describe total squared per-example gradient scale, the participation of sampled gradient directions and their anisotropy, respectively. At depth 32, the initial effective ranks of HsymH_sym and HblockH_block are 1.42±0.041.42± 0.04 and 1.52±0.021.52± 0.02 on SYNTHETIC EIGENGAP, and 1.10±0.021.10± 0.02 and 1.18±0.031.18± 0.03 on SYNTHETIC SINGULAR. These low ranks coexist with relatively high test accuracies, so initial empirical-Fisher rank does not track the performance ordering. Figure 9: Empirical-Fisher diagnostics at initialisation. For each diagnostic batch, per-sample gradients are collected into a gradient matrix, and the empirical Fisher is summarised by its trace, effective rank, and condition number. Each point is the mean over 20 run seeds, and error bars show one standard deviation. The plot characterises the initial tangent-space geometry introduced by each encoder but should not be interpreted as a direct predictor of final test accuracy. Figure 10 reports the same summaries at the final checkpoint. On SYNTHETIC EIGENGAP at depth 32, trainable patch-SU(4)SU(4) has effective rank 30.65±0.2330.65± 0.23 and test accuracy 50.16%±2.36%50.16\%± 2.36\%; HblockH_block has effective rank 10.86±1.7210.86± 1.72 and accuracy 85.15%±2.40%85.15\%± 2.40\%. The effective rank therefore does not reproduce the accuracy ordering. This example has a provenance limitation: all 20 trainable patch-SU(4)SU(4) final-gradient records are marked git_dirty=True, whereas all 20 HblockH_block records are marked git_dirty=False. We therefore treat it as a descriptive archived comparison rather than a clean-state reproduction. By contrast, all 20 records in each group used for the depth-32 Pendigits DYN patch-SU(4)SU(4) versus HblockH_block comparison in the main text have git_dirty=False. Figure 10: Empirical-Fisher diagnostics at the final trained checkpoint. The trace, effective rank, and condition number are computed from per-sample gradients on the fixed diagnostic batch. Each point is the mean over 20 run seeds, and error bars show one standard deviation. High final Fisher effective rank does not by itself imply high predictive performance; several patch-SU(4) models have high final Fisher rank but low depth-32 test accuracy. Table 4: Summary of the depth-32 performance and gradient analysis.. Values are mean ± standard deviation over 20 run seeds. Test accuracy is reported in percentages. Init logVar is the median log10 _10 variance of the shared SU(4)SU(4) mixer gradients across initialisation draws. Init RMS and final RMS are RMS gradients of the shared mixer group. Fisher ranks are empirical-Fisher effective ranks computed from per-sample gradients on the diagnostic batch. Dataset Encoder Test acc. (%) Init logVar Init RMS Final RMS Init Fisher rank Final Fisher rank Pendigits DYN fixed-RyR_y 72.29±2.7772.29± 2.77 −11.04±0.08-11.04± 0.08 0.527±0.0810.527± 0.081 0.054±0.0100.054± 0.010 5.71±3.115.71± 3.11 10.20±3.2310.20± 3.23 fixed-RyRzR_yR_z 55.84±5.3055.84± 5.30 −11.00±0.04-11.00± 0.04 0.333±0.0330.333± 0.033 0.050±0.0040.050± 0.004 6.63±3.466.63± 3.46 15.83±3.8815.83± 3.88 TF-RyR_y 81.37±2.3681.37± 2.36 −11.03±0.09-11.03± 0.09 0.473±0.0600.473± 0.060 0.044±0.0060.044± 0.006 5.92±2.725.92± 2.72 8.92±2.988.92± 2.98 patch-SU(4)SU(4) 9.96±0.449.96± 0.44 −1.56±0.02-1.56± 0.02 0.385±0.0060.385± 0.006 0.050±0.0020.050± 0.002 9.37±4.829.37± 4.82 29.25±0.5229.25± 0.52 trainable patch-SU(4)SU(4) 9.84±0.479.84± 0.47 −1.50±0.02-1.50± 0.02 0.359±0.0040.359± 0.004 0.050±0.0020.050± 0.002 10.71±3.3410.71± 3.34 28.46±1.0828.46± 1.08 patch HblockH_ block 92.75±0.8492.75± 0.84 −2.38±0.04-2.38± 0.04 0.201±0.0030.201± 0.003 0.048±0.0050.048± 0.005 6.42±1.216.42± 1.21 10.82±2.0010.82± 2.00 HsymH_ sym 91.89±3.7191.89± 3.71 −1.42±0.07-1.42± 0.07 0.312±0.0160.312± 0.016 0.079±0.0340.079± 0.034 5.60±0.635.60± 0.63 8.23±2.008.23± 2.00 HblockH_ block 92.35±5.9592.35± 5.95 −1.05±0.05-1.05± 0.05 0.380±0.0220.380± 0.022 0.071±0.0340.071± 0.034 5.22±1.815.22± 1.81 9.39±2.729.39± 2.72 Pendigits STA4 fixed-RyR_y 40.22±4.4040.22± 4.40 −11.33±0.03-11.33± 0.03 0.396±0.0120.396± 0.012 0.040±0.0060.040± 0.006 3.52±1.723.52± 1.72 20.51±2.9220.51± 2.92 fixed-RyRzR_yR_z 32.79±4.2832.79± 4.28 −11.19±0.07-11.19± 0.07 0.220±0.0130.220± 0.013 0.033±0.0070.033± 0.007 9.73±3.909.73± 3.90 19.98±3.4919.98± 3.49 TF-RyR_y 46.85±6.3946.85± 6.39 −11.31±0.03-11.31± 0.03 0.401±0.0280.401± 0.028 0.045±0.0070.045± 0.007 3.52±1.883.52± 1.88 18.86±2.8118.86± 2.81 patch-SU(4)SU(4) 16.41±1.8816.41± 1.88 −1.55±0.02-1.55± 0.02 0.374±0.0090.374± 0.009 0.054±0.0020.054± 0.002 9.96±4.819.96± 4.81 27.81±0.8527.81± 0.85 trainable patch-SU(4)SU(4) 10.27±0.6210.27± 0.62 −1.55±0.03-1.55± 0.03 0.384±0.0170.384± 0.017 0.050±0.0020.050± 0.002 10.50±5.0210.50± 5.02 28.02±1.0628.02± 1.06 patch HblockH_ block 86.14±1.3386.14± 1.33 −2.09±0.08-2.09± 0.08 0.244±0.0080.244± 0.008 0.079±0.0130.079± 0.013 6.36±1.786.36± 1.78 7.04±1.717.04± 1.71 HsymH_ sym 59.79±14.9059.79± 14.90 −1.37±0.04-1.37± 0.04 0.302±0.0130.302± 0.013 0.164±0.0530.164± 0.053 5.40±0.765.40± 0.76 9.27±3.079.27± 3.07 HblockH_ block 76.38±16.8476.38± 16.84 0.05±0.200.05± 0.20 0.779±0.0780.779± 0.078 0.194±0.1690.194± 0.169 4.18±1.974.18± 1.97 6.34±3.966.34± 3.96 Synthetic eigengap fixed-RyR_y 65.40±1.6965.40± 1.69 −19.73±0.06-19.73± 0.06 0.088±0.0020.088± 0.002 0.011±0.0020.011± 0.002 3.47±0.733.47± 0.73 9.22±2.039.22± 2.03 fixed-RyRzR_yR_z 65.55±1.2265.55± 1.22 −20.15±0.05-20.15± 0.05 0.046±0.0000.046± 0.000 0.011±0.0010.011± 0.001 8.97±1.348.97± 1.34 11.50±1.4711.50± 1.47 TF-RyR_y 67.42±1.3367.42± 1.33 −19.73±0.06-19.73± 0.06 0.088±0.0020.088± 0.002 0.010±0.0010.010± 0.001 3.55±0.763.55± 0.76 7.59±1.197.59± 1.19 patch-SU(4)SU(4) 54.58±1.9754.58± 1.97 −4.91±0.04-4.91± 0.04 0.074±0.0000.074± 0.000 0.017±0.0030.017± 0.003 10.17±2.3410.17± 2.34 27.02±1.5827.02± 1.58 trainable patch-SU(4)SU(4) 50.16±2.3650.16± 2.36 −4.92±0.04-4.92± 0.04 0.075±0.0010.075± 0.001 0.012±0.0000.012± 0.000 11.59±2.8111.59± 2.81 30.65±0.2330.65± 0.23 patch HblockH_ block 71.44±2.1171.44± 2.11 −9.44±0.03-9.44± 0.03 0.035±0.0040.035± 0.004 0.023±0.0040.023± 0.004 1.50±0.031.50± 0.03 7.93±1.237.93± 1.23 HsymH_ sym 82.81±3.1682.81± 3.16 −3.01±0.06-3.01± 0.06 0.094±0.0060.094± 0.006 0.108±0.0190.108± 0.019 1.42±0.041.42± 0.04 7.52±1.877.52± 1.87 HblockH_ block 85.15±2.4085.15± 2.40 −3.72±0.05-3.72± 0.05 0.052±0.0040.052± 0.004 0.062±0.0110.062± 0.011 1.52±0.021.52± 0.02 10.86±1.7210.86± 1.72 Synthetic singular fixed-RyR_y 84.34±1.2484.34± 1.24 −19.69±0.09-19.69± 0.09 0.084±0.0020.084± 0.002 0.012±0.0020.012± 0.002 4.48±1.014.48± 1.01 7.86±1.207.86± 1.20 fixed-RyRzR_yR_z 84.54±1.0184.54± 1.01 −20.04±0.05-20.04± 0.05 0.053±0.0010.053± 0.001 0.011±0.0020.011± 0.002 8.02±1.718.02± 1.71 9.13±1.289.13± 1.28 TF-RyR_y 85.40±0.6385.40± 0.63 −19.69±0.09-19.69± 0.09 0.085±0.0020.085± 0.002 0.013±0.0030.013± 0.003 4.40±1.024.40± 1.02 6.38±1.446.38± 1.44 patch-SU(4)SU(4) 63.41±2.0363.41± 2.03 −5.07±0.02-5.07± 0.02 0.061±0.0000.061± 0.000 0.026±0.0040.026± 0.004 14.15±2.6114.15± 2.61 20.22±1.6720.22± 1.67 trainable patch-SU(4)SU(4) 61.75±1.9161.75± 1.91 −5.06±0.02-5.06± 0.02 0.061±0.0010.061± 0.001 0.018±0.0010.018± 0.001 14.71±2.9514.71± 2.95 22.66±2.4722.66± 2.47 patch HblockH_ block 87.50±2.0687.50± 2.06 −9.29±0.04-9.29± 0.04 0.036±0.0050.036± 0.005 0.023±0.0090.023± 0.009 1.26±0.031.26± 0.03 5.49±1.555.49± 1.55 HsymH_ sym 88.75±2.6988.75± 2.69 −3.62±0.05-3.62± 0.05 0.069±0.0050.069± 0.005 0.061±0.0150.061± 0.015 1.10±0.021.10± 0.02 6.92±1.726.92± 1.72 HblockH_ block 92.72±4.0492.72± 4.04 −3.64±0.04-3.64± 0.04 0.070±0.0030.070± 0.003 0.036±0.0090.036± 0.009 1.18±0.031.18± 0.03 6.53±1.176.53± 1.17 Appendix H Additional Information on Latent Diagnostics H.1 Additional details on the latent diagnostic metrics Except for the 20-seed test accuracies in Table 5, the latent diagnostics use the final checkpoint from seed 0 and a fixed 32-example validation batch. They do not estimate variation across training seeds and should therefore be interpreted as illustrative within-run diagnostics rather than seed-robust estimates. The batched latent archive contains 864 complete records, of which 63 are marked git_dirty=True. Among the final-checkpoint depth-32 records, only two are dirty: Pendigits DYN with trainable patch-SU(4)SU(4) and SYNTHETIC EIGENGAP with trainable patch-SU(4)SU(4). The depth-32 QSM examples quoted in the main text—Pendigits DYN with HsymH_sym and HblockH_block, Pendigits STA4 with patch HblockH_block, and SYNTHETIC SINGULAR with HsymH_sym and HblockH_block—all have git_dirty=False. Let |ψi(s)⟩i=1m \ _i^(s) \^m_i=1 (164) denote the quantum states of the m=32m=32 diagnostic data examples at a particular stage s of the circuit. In the final-state summary plots, s is the final post-mixer state. In the layerwise plots, s indexes the post-mixer state after each reuploading layer. We then compute the pure-state fidelity kernel Kij(s)=|⟨ψi(s)|ψj(s)⟩|2.K_ij^(s)= | _i^(s)| _j^(s) |^2. (165) Let yiy_i be the class label of example i. The label-equality kernel is Yij=[yi=yj].Y_ij= 1[y_i=y_j]. (166) This is the target kernel used in the kernel-target alignment calculation. For a square matrix A∈ℝm×mA ^m× m, define the centering operator C=Im−1mT,=(11⋮1)∈ℝm,C=I_m- 1m 1 1^T, 1= pmatrix1\\ 1\\ \\ 1 pmatrix ^m, (167) so the centred Gram matrix is A~=CAC. A=CAC. (168) The centred kernel-target alignment [23] used in the plots is (K,Y)=⟨CKC,CYC⟩F‖CKC‖F‖CYC‖F,KTA(K,Y)= CKC,CYC_F||CKC||_F||CYC||_F, (169) with a small numerical guard in the denominator. This is the centred-kernel alignment (CKA) between the fidelity and label kernels. The within-minus-between fidelity gap uses only off-diagonal pairs. Let =(i,j):i<j,yi=yj,ℬ=(i,j):i<j,yi≠yj.W= \(i,j):i<j,y_i=y_j \, = \(i,j):i<j,y_i≠ y_j \. (170) Then Δfid=1||∑(i,j)∈Kij−1|ℬ|∑(i,j)∈ℬKij. _fid= 1|W| _(i,j) K_ij- 1|B| _(i,j) K_ij. (171) Positive values mean same-label examples are closer in quantum-state fidelity than different-label data examples. The participation-ratio effective rank [14] of the fidelity kernel is computed from the non-negative eigenvalues λ1,…,λm _1,…, _m of the symmetrised kernel: reff(K)=(∑jλj)2∑jλj2.r_eff(K)= ( _j _j)^2 _j _j^2. (172) This scale-invariant participation ratio summarises how broadly the kernel eigenvalue mass is spread: it equals the number of equally weighted nonzero eigenmodes in the special case where those eigenvalues are equal. It does not use the class labels and, by itself, does not measure class separation or predictive accuracy. The centred effective rank uses the same expression after centring the kernel. For adjacent-layer CKA, we apply the same centred-alignment formula to fidelity kernels from consecutive post-mixer layers: ℓ,ℓ+1=⟨CK(ℓ)C,CK(ℓ+1)C⟩F‖CK(ℓ)C‖F‖CK(ℓ+1)C‖F,CKA_ , +1= CK^( )C,CK^( +1)C_F||CK^( )C||_F||CK^( +1)C||_F, (173) This measures similarity between consecutive batch geometries; values near one indicate little layer-to-layer change. The trajectory plots evaluate the same final-state metrics at saved checkpoints rather than only at the final checkpoint. If c indexes training checkpoints, the trajectory target alignment is (Kfinal(c),Y),KTA(K^(c)_final,Y), (174) where Kfinal(c)K^(c)_final is the fidelity kernel of the final post-mixer state at checkpoint c. The diagnostic-batch projector accuracy reported in Table 5 is computed by applying the same projector readout used for the classifier to the final quantum states and comparing the argmax prediction with the diagnostic labels. H.2 Additional figures and tables The final kernel effective-rank plot in Figure 11 measures spectral spread in the final quantum-state fidelity kernel. This is useful as a control because it separates broad participation across kernel eigenmodes from class-aligned kernel geometry. In Figure 11, rotation-gate data-encoding unitaries often have an effective rank of approximately 32 (the diagnostic batch size), but their fidelity gap is approximately zero. Patch-SU(4)SU(4) data-encoding unitaries also often have high effective rank while performing poorly at depth 32. High effective rank is therefore insufficient for good classification in these runs. The label-aware evidence instead comes from the fidelity gap and kernel-target alignment, which must still be interpreted as associations rather than causes. Figure 11: Effective rank of the final quantum-state fidelity kernel. Effective rank is computed as (K)2/(K2)tr(K)^2/tr(K^2) from the eigenvalues of the final fidelity kernel. A high effective rank indicates that the kernel spectrum is spread across many eigenmodes rather than concentrated in a few; it does not, by itself, imply label alignment or high predictive performance. The layerwise fidelity-gap plot (Figure 12) shows how class separation evolves across the circuit layers at the final trained checkpoint. For each layer, the diagnostic computes the same within-minus-between fidelity gap used in the main text, but on the post-mixer state at that layer. This plot is useful because it shows whether class separation appears early or is progressively built by the reuploading architecture. The QSMs tend to produce larger layerwise gaps than the rotation-gate baselines, especially at depth 32. The patch-local block-Hamiltonian QSM is particularly relevant on STA4, where its local patch structure produces strong class separation in later layers. Figure 12: Layerwise within-minus-between fidelity gap. For each post-mixer layer, we compute the pure-state fidelity kernel on the diagnostic validation batch and report the mean same-label off-diagonal fidelity minus the mean different-label off-diagonal fidelity. The curves show where class separation emerges inside the reuploading circuit. QSMs generally build larger class-separating fidelity gaps than the coordinate-wise rotation-gate and patch-SU(4)SU(4) baselines. The layerwise target-alignment plot Figure 13 gives a complementary view of representation formation. Instead of measuring only the within-minus-between fidelity gap, it compares the full layerwise fidelity kernel with the label-equality kernel using centred kernel alignment. This plot is useful when the label structure is not captured solely by a mean gap. The main caveat is the same as for the final target-alignment plot: target alignment can be nonzero for high-rank kernels even when off-diagonal class separation is weak, so it should be interpreted together with the fidelity-gap and effective-rank plots. Figure 13: Layerwise kernel-target alignment. At each post-mixer layer, the fidelity kernel of the diagnostic validation states is centred and aligned with the label-equality kernel. The resulting curves show how class-aligned quantum-state geometry develops across depth. This diagnostic complements the layerwise fidelity-gap plot by comparing the full kernel structure to the label kernel. The layerwise effective-rank plot shown in Figure 14 tracks how the fidelity-kernel eigenvalue mass is distributed across modes as the circuit depth increases. It is most useful as a negative control. Rotation-gate data-encoding unitaries can maintain very high effective rank across layers, but this does not imply strong class separation or high test accuracy. Conversely, QSMs can have a lower effective rank while achieving larger fidelity gaps and higher test accuracy. The layerwise result therefore shows that broad kernel spectral spread alone does not reproduce the observed accuracy pattern. Figure 14: Layerwise effective rank of the fidelity kernel. Effective rank is computed from the post-mixer fidelity kernel at each layer. The plot measures spectral spread in the kernel, not label alignment. High effective rank alone is not sufficient for high test accuracy in these runs, as several rotation-gate and patch-SU(4)SU(4) models produce high-rank kernels without strong class-separated fidelity geometry. The trajectory plot in Figure 15 shows how final-state kernel-target alignment changes during training, using the stored sequence of checkpoints. This helps distinguish a model that starts with label-aligned geometry from one that learns it during optimisation. In Figure 15, the global block-Hamiltonian QSM on Pendigits DYN increases target alignment from 0.558 at initialisation to 0.900 at the final checkpoint, while fixed-RyR_y remains flat at 0.535. On SYNTHETIC SINGULAR, the global block-Hamiltonian QSM increases from 0.093 at initialisation to 0.461 at the final checkpoint, while fixed-RyR_y remains at 0.180. This supports the interpretation that the QSMs learn class-aligned quantum-state geometry over training, rather than merely starting with it. Figure 15: Training trajectory of kernel-target alignment. The curves show centred alignment between the final-state fidelity kernel and the label-equality kernel at saved checkpoints. The diagnostic is evaluated on the same validation batch used in the final latent-state summaries. QSMs tend to increase target alignment during optimisation, whereas several coordinate-wise rotation-gate baselines remain comparatively flat. The scatter plot in Figure 16 relates the fidelity-gap diagnostic to diagnostic-batch accuracy. It directly connects the latent-state geometry to prediction quality on the diagnostic batch. Points with larger within-minus-between fidelity gap generally correspond to better diagnostic-batch accuracy, especially in the Pendigits panels. The plot also makes the negative-control cases visible: some encoders may have a high effective rank, but without a positive fidelity gap, they do not form a label-separated quantum-state geometry. Figure 16: Relationship between fidelity gap and diagnostic-batch accuracy. Each point corresponds to one encoder-depth configuration on the diagnostic validation batch. The horizontal axis is the final within-minus-between fidelity gap, and the vertical axis is final projector accuracy on the same batch. Larger positive fidelity gaps generally correspond to stronger diagnostic-batch classification. The adjacent-layer CKA plot in Figure 17 measures representational shifts between consecutive post-mixer states. A high value means that the fidelity kernel changes little from one layer to the next. This diagnostic could assess whether deeper circuits are repeatedly reshaping the batch geometry or have entered a stable representation regime. In Figure 17, adjacent-layer CKA is often high at large depths, with final depth-32 values averaging about 0.986 across complete records, but this high stability does not, by itself, distinguish successful from unsuccessful encoders. Figure 17: Adjacent-layer CKA of post-mixer fidelity kernels. For each trained model, we compute centred kernel alignment between the fidelity kernels of consecutive post-mixer layers. High values indicate that adjacent layers preserve a similar batch geometry. These diagnostic measures representational stability across depth and are included as a supplementary control rather than as a direct predictor of test accuracy. Table 5: Summary of the latent diagnostic results. Test accuracy is the mean ± standard deviation over 20 random seeds, reported in percentage. The remaining columns are final-checkpoint diagnostics computed on the fixed 32-example validation diagnostic batch for seed 0. Diagnostic accuracy is the projector accuracy on this batch. KTA is centred kernel-target alignment between the final fidelity kernel and the label-equality kernel. Gap is the mean within-class off-diagonal fidelity minus the mean between-class off-diagonal fidelity. Effective rank is computed from the final fidelity kernel. Dataset Encoder Test acc. (%) Diag. acc. KTA Gap Eff. rank Pendigits DYN fixed-RyR_y 72.29±2.7772.29± 2.77 0.688 0.535 0.000 32.00 fixed-RyRzR_yR_z 55.84±5.3055.84± 5.30 0.562 0.535 0.000 32.00 TF-RyR_y 81.37±2.3681.37± 2.36 0.719 0.535 0.000 32.00 patch-SU(4)SU(4) 9.96±0.449.96± 0.44 0.062 0.535 −0.000-0.000 31.97 trainable patch-SU(4)SU(4) 9.84±0.479.84± 0.47 0.094 0.535 −0.000-0.000 31.97 patch HblockH_ block 92.75±0.8492.75± 0.84 0.938 0.761 0.260 24.87 HsymH_ sym 91.89±3.7191.89± 3.71 1.000 0.887 0.538 15.46 HblockH_ block 92.35±5.9592.35± 5.95 1.000 0.900 0.561 15.29 Pendigits STA4 fixed-RyR_y 40.22±4.4040.22± 4.40 0.438 0.535 0.000 32.00 fixed-RyRzR_yR_z 32.79±4.2832.79± 4.28 0.312 0.535 −0.000-0.000 32.00 TF-RyR_y 46.85±6.3946.85± 6.39 0.562 0.535 0.000 32.00 patch-SU(4)SU(4) 16.41±1.8816.41± 1.88 0.094 0.535 −0.000-0.000 31.97 trainable patch-SU(4)SU(4) 10.27±0.6210.27± 0.62 0.094 0.534 −0.001-0.001 31.97 patch HblockH_ block 86.14±1.3386.14± 1.33 0.938 0.694 0.173 27.03 HsymH_ sym 59.79±14.9059.79± 14.90 0.562 0.592 0.208 12.18 HblockH_ block 76.38±16.8476.38± 16.84 0.719 0.694 0.285 13.20 Synthetic eigengap fixed-RyR_y 65.40±1.6965.40± 1.69 0.594 0.180 −0.000-0.000 32.00 fixed-RyRzR_yR_z 65.55±1.2265.55± 1.22 0.625 0.180 −0.000-0.000 32.00 TF-RyR_y 67.42±1.3367.42± 1.33 0.594 0.180 −0.000-0.000 32.00 patch-SU(4)SU(4) 54.58±1.9754.58± 1.97 0.562 0.179 −0.000-0.000 31.97 trainable patch-SU(4)SU(4) 50.16±2.3650.16± 2.36 0.531 0.181 0.000 31.97 patch HblockH_ block 71.44±2.1171.44± 2.11 0.625 0.168 −0.003-0.003 29.21 HsymH_ sym 82.81±3.1682.81± 3.16 0.812 0.259 0.077 6.64 HblockH_ block 85.15±2.4085.15± 2.40 0.750 0.261 0.053 11.95 Synthetic singular fixed-RyR_y 84.34±1.2484.34± 1.24 0.812 0.180 0.000 32.00 fixed-RyRzR_yR_z 84.54±1.0184.54± 1.01 0.844 0.180 0.000 32.00 TF-RyR_y 85.40±0.6385.40± 0.63 0.812 0.180 0.000 32.00 patch-SU(4)SU(4) 63.41±2.0363.41± 2.03 0.625 0.180 0.000 31.95 trainable patch-SU(4)SU(4) 61.75±1.9161.75± 1.91 0.625 0.181 0.000 31.97 patch HblockH_ block 87.50±2.0687.50± 2.06 0.906 0.323 0.059 21.50 HsymH_ sym 88.75±2.6988.75± 2.69 0.844 0.508 0.163 6.00 HblockH_ block 92.72±4.0492.72± 4.04 0.906 0.461 0.159 11.17 Appendix I Additional Information on Classical Baseline and Ablation Study Table 6 reports the classical controls described in the main text. For every quantum ablation in Table 7, the transformation is constructed before training, applied to the training, validation and test splits, and followed by retraining. Labels are unchanged. Any reference basis or spectrum is computed from the training split only. The entry-permutation control generates one fixed permutation of the flattened matrix entries and applies it to every sample in all three splits. The row/column control similarly generates one fixed row permutation and one fixed column permutation. The permutations are therefore dataset-level transformations, not independently sampled perturbations of individual examples. For the symmetric-Hamiltonian controls, let Q0Q_0 be the eigenvector matrix of the mean training-set HsymH_sym and let 0 λ_0 be the coordinatewise median of the ordered training-set eigenvalues. The spectrum-only control retains each sample's eigenvalues and replaces its eigenvectors by Q0Q_0. The eigenvectors-only control retains each sample's eigenvectors and replaces its eigenvalues by 0 λ_0. For the block-Hamiltonian controls, let U0U_0 and V0V_0 be the left and right singular-vector matrices of the mean training matrix, and let 0 σ_0 be the coordinatewise median of the ordered training-set singular values. The singular-values-only control combines each sample's singular values with U0U_0 and V0V_0. The singular-vectors-only control retains each sample's left and right singular vectors and replaces its singular values by 0 σ_0. Because bases within degenerate spectral subspaces are not unique, the vector-only controls are implementation-dependent in those subspaces. Table 6: Classical baseline results. Each baseline is a one-hidden-layer MLP trained on either flattened raw entries or spectral-value descriptors. Values are mean ± standard deviation over 20 random seeds, reported in percentage. Δ is the test-accuracy difference, in percentage points, between the classical baseline and the largest observed QSM mean on the same dataset panel. Dataset Feature source Feature dim. Params Val. acc. (%) Test acc. (%) Δ vs. largest QSM mean Pendigits DYN raw entries 16 1468 99.31±0.2299.31± 0.22 96.88±0.2496.88± 0.24 +3.07+3.07 HsymH_ sym eig. 8 1473 66.97±1.1266.97± 1.12 59.98±0.6059.98± 0.60 −33.83-33.83 HblockH_ block sing. 2 1466 42.75±0.7042.75± 0.70 39.67±0.4839.67± 0.48 −54.14-54.14 Pendigits STA4 raw entries 16 1468 95.46±0.4795.46± 0.47 91.32±0.4691.32± 0.46 +4.66+4.66 HsymH_ sym eig. 4 1465 38.53±0.9538.53± 0.95 39.80±0.5039.80± 0.50 −46.86-46.86 HblockH_ block sing. 4 1465 31.39±1.0031.39± 1.00 30.62±0.5330.62± 0.53 −56.04-56.04 Synthetic eigengap raw entries 16 515 75.20±2.5175.20± 2.51 74.88±1.2774.88± 1.27 −11.76-11.76 HsymH_ sym eig. 4 513 98.28±0.2898.28± 0.28 98.01±0.3598.01± 0.35 +11.36+11.36 HblockH_ block sing. 4 989 67.21±1.0367.21± 1.03 67.56±0.6067.56± 0.60 −19.08-19.08 Synthetic singular raw entries 16 1465 90.37±0.8090.37± 0.80 91.32±0.6091.32± 0.60 −3.92-3.92 HsymH_ sym eig. 8 992 98.29±0.5198.29± 0.51 97.28±0.2397.28± 0.23 +2.04+2.04 HblockH_ block sing. 2 1472 98.57±0.3698.57± 0.36 98.50±0.2098.50± 0.20 +3.25+3.25 Table 7: Ablation-study summary for Hamiltonian-based QSMs. Values are mean ± standard deviation over 20 random seeds, reported in percentage. ``Original′ gives the depth and largest observed mean test accuracy of the unablated QSM variant on the same dataset panel; ``Ablated′ gives the corresponding descriptive maximum after retraining under the specified ablation. These maxima are computed across the tested depths rather than at depths selected by independent validation. Δ is the ablated-minus-original difference in percentage points. Spectrum/eigenvector controls are run for HsymH_ sym; singular-value/vector controls are run for HblockH_ block; permutation controls are shown for the global QSMs and the patch-local block-Hamiltonian QSM. Dataset Encoder Original Ablation Ablated Δ Pendigits DYN HsymH_ sym L32, 91.89±3.7191.89± 3.71 spectrum only L8, 48.87±3.9048.87± 3.90 −43.02-43.02 HsymH_ sym L32, 91.89±3.7191.89± 3.71 eigenvectors only L16, 89.80±4.5689.80± 4.56 −2.08-2.08 HsymH_ sym L32, 91.89±3.7191.89± 3.71 entry permutation L16, 89.49±10.4289.49± 10.42 −2.39-2.39 HsymH_ sym L32, 91.89±3.7191.89± 3.71 row/column permutation L32, 92.69±3.7992.69± 3.79 +0.80+0.80 HblockH_ block L16, 93.81±1.0693.81± 1.06 singular values only L4, 31.65±2.4431.65± 2.44 −62.16-62.16 HblockH_ block L16, 93.81±1.0693.81± 1.06 singular vectors only L16, 93.56±0.8493.56± 0.84 −0.24-0.24 HblockH_ block L16, 93.81±1.0693.81± 1.06 entry permutation L32, 93.94±2.3493.94± 2.34 +0.13+0.13 HblockH_ block L16, 93.81±1.0693.81± 1.06 row/column permutation L32, 94.82±1.5994.82± 1.59 +1.01+1.01 patch HblockH_ block L32, 92.75±0.8492.75± 0.84 entry permutation L32, 92.77±1.3192.77± 1.31 +0.03+0.03 patch HblockH_ block L32, 92.75±0.8492.75± 0.84 row/column permutation L32, 92.09±1.4092.09± 1.40 −0.66-0.66 Pendigits STA4 HsymH_ sym L16, 70.14±8.2970.14± 8.29 spectrum only L4, 28.35±2.1128.35± 2.11 −41.79-41.79 HsymH_ sym L16, 70.14±8.2970.14± 8.29 eigenvectors only L4, 64.30±2.0264.30± 2.02 −5.84-5.84 HsymH_ sym L16, 70.14±8.2970.14± 8.29 entry permutation L32, 72.45±6.5872.45± 6.58 +2.31+2.31 HsymH_ sym L16, 70.14±8.2970.14± 8.29 row/column permutation L16, 70.73±4.0570.73± 4.05 +0.59+0.59 HblockH_ block L16, 86.66±2.1786.66± 2.17 singular values only L4, 25.25±2.0225.25± 2.02 −61.41-61.41 HblockH_ block L16, 86.66±2.1786.66± 2.17 singular vectors only L16, 84.74±4.5184.74± 4.51 −1.91-1.91 HblockH_ block L16, 86.66±2.1786.66± 2.17 entry permutation L16, 84.59±4.0584.59± 4.05 −2.07-2.07 HblockH_ block L16, 86.66±2.1786.66± 2.17 row/column permutation L16, 86.16±2.8586.16± 2.85 −0.50-0.50 patch HblockH_ block L32, 86.14±1.3386.14± 1.33 entry permutation L16, 83.10±1.0383.10± 1.03 −3.04-3.04 patch HblockH_ block L32, 86.14±1.3386.14± 1.33 row/column permutation L16, 83.92±1.3883.92± 1.38 −2.22-2.22 Synthetic eigengap HsymH_ sym L16, 83.42±1.3783.42± 1.37 spectrum only L8, 97.92±0.3897.92± 0.38 +14.50+14.50 HsymH_ sym L16, 83.42±1.3783.42± 1.37 eigenvectors only L1, 61.28±0.0761.28± 0.07 −22.14-22.14 HsymH_ sym L16, 83.42±1.3783.42± 1.37 entry permutation L8, 67.83±1.4167.83± 1.41 −15.59-15.59 HsymH_ sym L16, 83.42±1.3783.42± 1.37 row/column permutation L4, 69.80±1.3969.80± 1.39 −13.62-13.62 HblockH_ block L16, 86.65±1.1886.65± 1.18 singular values only L16, 65.54±1.0765.54± 1.07 −21.11-21.11 HblockH_ block L16, 86.65±1.1886.65± 1.18 singular vectors only L8, 68.66±1.6568.66± 1.65 −17.99-17.99 HblockH_ block L16, 86.65±1.1886.65± 1.18 entry permutation L16, 74.46±1.3574.46± 1.35 −12.19-12.19 HblockH_ block L16, 86.65±1.1886.65± 1.18 row/column permutation L16, 83.47±1.1883.47± 1.18 −3.18-3.18 patch HblockH_ block L16, 74.37±2.5974.37± 2.59 entry permutation L16, 72.85±1.4672.85± 1.46 −1.52-1.52 patch HblockH_ block L16, 74.37±2.5974.37± 2.59 row/column permutation L16, 74.74±1.3374.74± 1.33 +0.37+0.37 Synthetic singular HsymH_ sym L8, 90.16±1.0890.16± 1.08 spectrum only L8, 97.25±0.4897.25± 0.48 +7.08+7.08 HsymH_ sym L8, 90.16±1.0890.16± 1.08 eigenvectors only L2, 51.29±1.0451.29± 1.04 −38.87-38.87 HsymH_ sym L8, 90.16±1.0890.16± 1.08 entry permutation L8, 88.35±1.0088.35± 1.00 −1.81-1.81 HsymH_ sym L8, 90.16±1.0890.16± 1.08 row/column permutation L16, 88.64±1.2788.64± 1.27 −1.53-1.53 HblockH_ block L16, 95.24±1.5395.24± 1.53 singular values only L32, 98.42±0.3098.42± 0.30 +3.17+3.17 HblockH_ block L16, 95.24±1.5395.24± 1.53 singular vectors only L16, 49.88±1.9649.88± 1.96 −45.36-45.36 HblockH_ block L16, 95.24±1.5395.24± 1.53 entry permutation L8, 90.55±0.5890.55± 0.58 −4.69-4.69 HblockH_ block L16, 95.24±1.5395.24± 1.53 row/column permutation L16, 95.06±2.3095.06± 2.30 −0.18-0.18 patch HblockH_ block L32, 87.50±2.0687.50± 2.06 entry permutation L32, 87.04±2.0487.04± 2.04 −0.46-0.46 patch HblockH_ block L32, 87.50±2.0687.50± 2.06 row/column permutation L32, 86.43±3.3786.43± 3.37 −1.06-1.06