Paper deep dive
Learning unified control of internal spin squeezing in atomic qudits for magnetometry
C. Z. Cao, J. Z. Han, M. Xiong, M. Deng, L. Wang, X. Lv, M. Xue
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/31/2026, 2:44:53 AM
Summary
The paper presents a physics-informed reinforcement learning (PIRL) framework to control internal spin squeezing in atomic qudits (specifically 161Dy) under always-on nonlinear Zeeman (NLZ) dynamics. The learned policy generates strong spin squeezing and stabilizes the measurement-relevant quadrature, achieving a magnetic sensitivity of 13.9 pT/√Hz, which is 3 dB beyond the standard quantum limit.
Entities (5)
Relation Signals (3)
161Dy → exhibits → Nonlinear Zeeman effect
confidence 95% · In a magnetic field, the multilevel Zeeman structure acquires a quadratic component, giving rise to an always-on nonlinear Zeeman (NLZ) interaction
PIRL agent → uses → Wineland parameter
confidence 95% · The agent is trained to maximize the cumulative reward defined from physically motivated squeezing metrics
Physics-informed reinforcement learning → optimizes → Nonlinear Zeeman effect
confidence 90% · physics-informed reinforcement learning can transform NLZ dynamics from a source of readout degradation into a sustained metrological resource
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Generating and preserving metrologically useful quantum states is a central challenge in quantum-enhanced atomic magnetometry. In multilevel atoms operated in the low-field regime, the nonlinear Zeeman (NLZ) effect is both a resource and a limitation. It nonlinearly redistributes internal spin fluctuations to generate spin-squeezed states within a single atomic qudit, yet under fixed readout it distorts the measurement-relevant quadrature and limits the accessible metrological gain. This challenge is compounded by the time dependence of both the squeezing axis and the effective nonlinear action. Here we show that physics-informed reinforcement learning can transform NLZ dynamics from a source of readout degradation into a sustained metrological resource. Using only experimentally accessible low-order spin moments, a trained agent identifies, in the $f=21/2$ manifold of $^{161}\mathrm{Dy}$, a unified control policy that rapidly prepares strongly squeezed internal states and stabilizes more than $4\,\mathrm{dB}$ of fixed-axis spin squeezing under always-on NLZ evolution. Including state-preparation overhead, the learned protocol yields a single-atom magnetic sensitivity of $13.9\,\mathrm{pT}/\sqrt{\mathrm{Hz}}$, corresponding to an advantage of approximately $3\,\mathrm{dB}$ beyond the standard quantum limit. Our results establish learning-based control as a practical route for converting unavoidable intrinsic nonlinear dynamics in multilevel quantum sensors into operational metrological advantage.
Tags
Links
- Source: https://arxiv.org/abs/2603.28421v1
- Canonical: https://arxiv.org/abs/2603.28421v1
Trouble viewing inline? Open PDF directly →
Full Text
60,316 characters extracted from source content.
Expand or collapse full text
†thanks: These authors contributed equally to this work.†thanks: These authors contributed equally to this work. Learning unified control of internal spin squeezing in atomic qudits for magnetometry C. Z. Cao College of Physics, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China J. Z. Han China Mobile (Suzhou) Software Technology Co., Ltd., Suzhou, 215163, China College of Physics, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China M. Xiong College of Physics, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China M. Deng Yichang Testing Technique R&D Institute, Specialized Weak Magnetic Metering Station of NDM, Yichang 443003, China L. Wang Shanghai Institute of Optics and Fine Mechanics, Chinese Academy of Sciences, Shanghai 201800, China X. Lv xudonglv@siom.ac.cn Shanghai Institute of Optics and Fine Mechanics, Chinese Academy of Sciences, Shanghai 201800, China M. Xue mxue@nuaa.edu.cn College of Physics, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China Key Laboratory of Aerospace Information Sensing and Physics (NUAA), MIIT, Nanjing 211106, China (March 30, 2026) Abstract Generating and preserving metrologically useful quantum states is a central challenge in quantum-enhanced atomic magnetometry. In multilevel atoms operated in the low-field regime, the nonlinear Zeeman (NLZ) effect is both a resource and a limitation. It nonlinearly redistributes internal spin fluctuations to generate spin-squeezed states within a single atomic qudit, yet under fixed readout it distorts the measurement-relevant quadrature and limits the accessible metrological gain. This challenge is compounded by the time dependence of both the squeezing axis and the effective nonlinear action. Here we show that physics-informed reinforcement learning can transform NLZ dynamics from a source of readout degradation into a sustained metrological resource. Using only experimentally accessible low-order spin moments, a trained agent identifies, in the f=21/2f=21/2 manifold of Dy161^161Dy, a unified control policy that rapidly prepares strongly squeezed internal states and stabilizes more than 4dB4\,dB of fixed-axis spin squeezing under always-on NLZ evolution. Including state-preparation overhead, the learned protocol yields a single-atom magnetic sensitivity of 13.9pT/Hz13.9\,pT/ Hz, corresponding to an advantage of approximately 3dB3\,dB beyond the standard quantum limit. Our results establish learning-based control as a practical route for converting unavoidable intrinsic nonlinear dynamics in multilevel quantum sensors into operational metrological advantage. I Introduction Figure 1: Physics-informed reinforcement-learning framework for control of internal spin dynamics in an atomic qudit. (a) We consider a single Dy161^161Dy atom with nuclear spin i=5/2i=5/2 and electronic angular momentum j=8j=8, focusing on the hyperfine manifold with f=21/2f=21/2. In a magnetic field, the multilevel Zeeman structure acquires a quadratic component, giving rise to an always-on nonlinear Zeeman (NLZ) interaction ∝f^z2 f_z^2. (b) Control problem. Starting from an initial coherent spin state, the qudit evolves under repeated cycles of free NLZ evolution and agent-selected transverse rotations Rμ(β)R_μ(β) with μ∈x,yμ∈\x,y\. At each step, the agent receives an observation ok∈o_k consisting of experimentally accessible low-order spin moments, chooses an action ak∈a_k from a discrete set of realizable rotations, and is trained to maximize the cumulative reward G defined from physically motivated squeezing metrics. This defines a unified control problem under always-on nonlinear dynamics. Atomic magnetometry is a leading platform for precision sensing, with applications in fundamental physics [1, 2, 3, 4, 5], biomagnetic detection [6, 7, 8, 9], and magnetic navigation [10, 11, 12]. State-of-the-art experiments have demonstrated high sensitivities in realistic operating regimes [13, 14, 15]. Quantum-enhanced magnetometry can further surpass the standard quantum limit (SQL) by exploiting nonclassical spin correlations [16, 17, 18]. In multilevel atoms, such correlations need not rely on interparticle entanglement, but can instead be generated within the internal degrees of freedom of a single spin-f system, equivalently viewed as a symmetric ensemble of 2f2f spin-1/21/2 constituents [19, 20]. Recent experiments have established the internal Hilbert space of multilevel atoms as a metrological resource. Atomic qudits can host nonclassical internal states, including non-Gaussian and cat-like states generated by controllable nonlinear dynamics such as light-induced tensor shifts, yielding enhanced magnetic sensitivity and even near-Heisenberg-limited performance under suitable control [21, 22, 23, 24]. In these settings, however, the nonlinearity is externally engineered and can be switched or shaped, so that state preparation can be cleanly separated from metrological interrogation. In realistic low-field sensing regimes, by contrast, the relevant nonlinearity is often intrinsic. The quadratic Hamiltonian ∝f^z2 f_z^2, arising from the nonlinear Zeeman (NLZ) effect, is the single-atom counterpart of the familiar one-axis-twisting (OAT) interaction in many-body spin systems [25]. It naturally generates internal spin squeezing, but simultaneously rotates the squeezing axis and drives overtwisting on experimentally relevant timescales [26]. As a result, the same dynamics that create nonclassical correlations can also degrade the metrological gain accessible under fixed readout. This difficulty is most acute when the relevant nonlinearity is intrinsic to the sensing process and therefore remains always on throughout both squeezing generation and signal interrogation. A prototypical example is Earth-field-range magnetometry, where the NLZ effect is ever present and has motivated a range of compensation, spin-locking, and staged-control approaches that either suppress its influence or recover part of the metrological gain through fixed dynamical-decoupling pulse structures and separate squeezing-generation and stabilization stages [27, 28, 29, 30, 31, 32, 33, 34, 35, 36]. Because the NLZ dynamics cannot be turned off, however, the optimal control is intrinsically nonseparable: squeezing generation and stabilization must be interleaved in time to preserve the measurement-relevant quadrature. The resulting task is therefore inherently adaptive. Useful control actions must be selected step by step during the nonlinear evolution from experimentally limited information, while their impact on fixed-readout metrological performance is not directly evident in the high-dimensional Hilbert space. This structure makes the problem a natural target for reinforcement learning and related learning-based approaches [37, 38]. In contrast to dynamical-decoupling-based strategies, which typically optimize within predefined pulse structures and staged control protocols, reinforcement learning can interact directly with the controlled dynamics and discover pulse patterns tailored to the underlying system. Here we address this challenge using physics-informed reinforcement learning (PIRL) [39]. We consider a protocol in which resonant transverse microwave rotations are interleaved with intrinsic nonlinear dynamics, and train an agent to select control operations using only experimentally accessible low-order spin moments. Guided by a reward that captures both overall squeezing and measurement-relevant fixed-axis performance, the learned unified policy rapidly generates strong internal spin squeezing while stabilizing the measurement-relevant quadrature under continuous evolution. It also reveals a simple and physically interpretable pulse structure that suppresses overtwisting while preserving useful quantum correlations. I Model and reinforcement-learning framework I.1 Intrinsic nonlinear dynamics and squeezing metrics We consider an atomic qudit encoded in the d=2f+1d=2f+1 Zeeman states |f,m⟩|f,m (m=−f,…,fm=-f,…,f) within a single hyperfine manifold in the presence of a magnetic field, which defines the quantization z axis, as illustrated in Fig. 1(a) for Dy161^161Dy with f=21/2f=21/2. The internal spin operators are denoted by f^x,f^y,f^z\ f_x, f_y, f_z\. In the weak-field regime, the hyperfine–Zeeman structure within a fixed manifold is described by the effective single-atom Hamiltonian (ℏ=1 =1) H^B≃ΩLf^z+χf^z2, H_B _L f_z+χ f_z^2, (1) where ΩL _L is the Larmor frequency and χ denotes the quadratic Zeeman effect (QZE) coefficient (see Appendix A). In the rotating frame that removes the Larmor precession, the dynamics reduce to H^QZE=χf^z2, H_ QZE=χ f_z^2, (2) which is formally equivalent to the OAT Hamiltonian, here acting on the internal (2f+1)(2f+1)-dimensional manifold of a single spin. Starting from a coherent spin state (CSS) polarized along x, |ψ0⟩=|f,mx=f⟩| _0 =|f,m_x=f , the always-on quadratic term shears the spin distribution in the plane transverse to the mean spin, thereby generating a spin-squeezed state (S). However, continued nonlinear evolution rotates the squeezing axis and induces overtwisting, so that the metrological gain under a fixed readout rapidly degrades. This competition between squeezing generation and its subsequent distortion underlies the need for active control, as schematically illustrated in Fig. 1(b). To quantify squeezing during the evolution of a single spin-f relative to a spin-coherent state, we use the Wineland parameter [40] ξ2=2fmin⟂[(Δf⟂)2]|⟨^⟩|2,ξ^2= 2f\, _ [( f_ )^2]| f |^2, (3) for which ξ2=1ξ^2=1 for a CSS. Since our sensing protocol uses a fixed readout axis, we further introduce ξy2=2f(Δfy)2|⟨f^x⟩|2, _y^2= 2f\,( f_y)^2| f_x |^2, (4) which quantifies the metrological usefulness of squeezing specifically for readout along y. I.2 Reinforcement-learning control framework The single-atom spin-f manifold is controlled by a transverse microwave drive. In the rotating frame of Eq. (2) and under the rotating-wave approximation, the dynamics are described by H^QZE+H^ctrl H_ QZE+ H_ ctrl with H^ctrl=λ(f^xcosϕ+f^ysinϕ) H_ctrl=λ( f_x φ+ f_y φ), where λ is the Rabi frequency and ϕφ sets the rotation axis in the transverse xyxy plane. For numerical optimization, we discretize the control into a sequence of piecewise-constant pulses. The total duration T is divided into NtN_t intervals of length δt=T/Ntδ t=T/N_t. In the kkth interval, the control is fully specified by the rotation angle βk=∫kδt(k+1)δtλt _k= _kδ t^(k+1)δ tλ\,dt. We denote the corresponding instantaneous transverse rotation by R^μ(βk)=e−iβkf^μ R_μ( _k)=e^-i _k f_μ, where μ labels the rotation axis determined by the phase ϕφ. Figure 2: Learned pulse protocol from PIRL. (a) Control pulse sequence selected by the PIRL agent. Colored bars denote discrete transverse rotations applied at each control step. (b) Time evolution of the Wineland squeezing parameter ξ2(t)ξ^2(t) (red solid) and the fixed-axis squeezing parameter ξy2(t) _y^2(t) (green solid) as functions of the dimensionless time χtχ t. Gray solid and dotted curves denote the QZE (OAT) evolution and the effective TACT benchmark, respectively. The vertical gray line near χt≃0.15χ t 0.15 marks the reward-defined switching time trt_r. (c) Spin Wigner distributions at representative times t1t_1–t5t_5 indicated in panel (b), covering the squeezing-generation and stabilization stages. Each control step then consists of a spin rotation followed by free evolution under the QZE Hamiltonian, |ψk+1⟩=e−iχδtf^z2R^μ(βk)|ψk⟩.| _k+1 =e^-iχδ t f_z^2 R_μ( _k)| _k . (5) This stroboscopic form makes explicit how transverse control rotations are interleaved with the intrinsic nonlinear dynamics, as illustrated in Fig. 1(b). The agent selects actions from a discrete set of experimentally natural rotations, =,Rμ(β)|μ∈x,y,β∈± π2,± π3,± π4,A= \I,\,R_μ(β)\; |\;μ∈\x,y\,\ β∈\± π2,± π3,± π4\ \, (6) where I denotes no control pulse. While this symmetric parametrization is convenient, the two rotation channels play distinct physical roles. Pairs of opposite-angle RyR_y rotations enable symmetric toggling-frame control that compensates the nonlinear rotation induced by the QZE [41, 42]. By contrast, RxR_x pulses primarily act as corrective rotations that reorient the squeezed state toward the measurement-relevant axis. For this purpose, positive rotation angles are already sufficient in practice, as confirmed by the optimized pulse sequences discussed below. To avoid an exponential dependence on the Hilbert-space dimension and to remain within experimentally feasible readout, the agent observes only low-order spin moments, =⟨f^x⟩/f,⟨f^y⟩/f,⟨f^z⟩/f,⟨f^x2⟩/f2,⟨f^y2⟩/f2,O=\ f_x /f, f_y /f, f_z /f, f_x^2 /f^2, f_y^2 /f^2\, (7) which can be obtained from measurements of spin projections and their variances, and suffice to track the mean-spin direction and transverse fluctuations relevant to squeezing and readout [rgb]0.00,0.00,1.00 [43, 44, 45, 46]. Importantly, these features are normalized by f, rendering the input representation independent of the Hilbert-space dimension and enabling transfer of the learned policy across different spin manifolds. The control objective is unified but, due to the nonlinear dynamics, the optimal strategy naturally separates into two regimes: rapid squeezing generation followed by stabilization of the measurement-relevant quadrature. In the PIRL framework, this structure is encoded directly into a scalar reward rkr_k built from the physically relevant squeezing metrics ξ2ξ^2 and ξy2 _y^2. The agent is trained to maximize the cumulative reward =∑k=0Nt−1γRLkrk,G= _k=0^N_t-1γ^k_RLr_k, (8) where γRL∈(0,1) _RL∈(0,1) is the discount factor, the cumulative reward captures long-term control performance beyond instantaneous gains. To implement this objective, we design a curriculum-style reward with a state-dependent switching time defined as the first-hitting time k⋆≡infk≥0:ξ2(k)≤ξref2,k_ ≡ \k≥ 0:\,ξ^2(k)≤ _ ref^2\, (9) namely, the first step at which the Wineland squeezing parameter reaches a prescribed threshold. The reference value ξref2 _ ref^2 is fixed throughout training and chosen as the optimal squeezing benchmark attainable under ideal two-axis counter-twisting (TACT) dynamics [25, 41]. The reward function is then defined as rk=logξ2(k−1)−logξ2(k)−c(ak),k<k⋆,ζe−κk⋆−c(ak),k=k⋆,−αlogξy2(k)−c(ak),k>k⋆,r_k= ξ^2(k-1)- ξ^2(k)-c(a_k),&k<k_ ,\\ ζ e^-κ k_ -c(a_k),&k=k_ ,\\ -α _y^2(k)-c(a_k),&k>k_ , (10) where the first term rewards rapid squeezing generation, the second provides a one-time bonus favoring earlier threshold crossing, and the third promotes post-threshold stabilization by directly optimizing the fixed-axis squeezing ξy2 _y^2, which is the quantity most relevant for metrological performance under our fixed readout. In other words, once the squeezing level reaches the target threshold, the reward switches from a state-preparation objective to a readout-aware sensing objective. The cost term c(ak)c(a_k) penalizes pulse usage, with ak∈\a_k \I\. Here ζ sets the bonus scale, κ its temporal decay with the hitting time k⋆k_ , and α the weight assigned to post-threshold stabilization (see Appendix B). We optimize the cumulative reward using Proximal Policy Optimization (PPO) [47, 48, 49], an Actor–Critic policy-gradient method [50, 51] well suited to sequential decision problems with delayed rewards. A policy network selects actions from the observation set O of low-order spin moments, while a value network estimates the expected return. This architecture enables stable model-free learning without an explicit dynamical model and is therefore well suited to discovering a single unified policy for squeezing generation and long-time stabilization [52, 53, 54, 55, 56]. I Learned control strategy, stabilization mechanism, and metrological gain I.1 Learned protocol structure We now analyze the control protocol learned by the agent. Although the reward is defined in terms of physically motivated squeezing metrics, the resulting policy exhibits a clear two-stage structure, in direct correspondence with the pulse sequence in Fig. 2(a): an initial preparation stage that rapidly generates spin squeezing, followed by a stabilization stage that preserves the measurement-relevant quadrature, as quantified in Fig. 2(b). During the preparation stage, the pulse sequence consists of three paired Ry(±π/2)R_y(±π/2) operations. These paired pulses suppress QZE-induced overtwisting and accelerate squeezing generation, enabling the system to reach ξ2(t3)=−8.19dBξ^2(t_3)=-8.19\,dB at time t3t_3, comparable to the optimal TACT benchmark of −8.07dB-8.07\,dB and below the QZE (OAT) benchmark of −7.17dB-7.17\,dB. The spin Wigner distributions in Fig. 2(c) illustrate how these pulses reshape the spin-fluctuation distribution and counteract the distortion accumulated during the intervening QZE evolution. Once the squeezing reaches the reference threshold ξref2 _ ref^2, the agent applies the isolated Rx(π/3)R_x(π/3) pulse at time t3t_3 in Fig. 2(a), rotating the state toward an orientation where the fixed-axis squeezing ξy2 _y^2 is close to its minimum. The control objective then switches from squeezing generation to stabilization. Importantly, the learned protocol does not enter the final periodic cycle immediately after this rotation. Instead, as indicated by the additional Ry(−π/4)R_y(-π/4) pulse in Fig. 2(a), the agent first rebalances the two states visited within one control period before settling into the subsequent alternating sequence of Ry(±π/2)R_y(±π/2) pulses. This intermediate adjustment is essential. Although the state at t3t_3 already has a value of ξy2 _y^2 close to the minimum of ξ2ξ^2, a direct transition into a simple alternating Ry(±π/2)R_y(±π/2) cycle would map a state oriented near the x axis close to the north or south pole, causing the other state visited in the same stabilization cycle to acquire a much larger value of ξy2 _y^2. The resulting protocol would therefore produce large oscillations of the measurement-relevant squeezing. Instead, the learned policy balances the values of ξy2 _y^2 between the two alternating states, primarily through the intermediate rotation close to Ry(−π/4)R_y(-π/4). This mechanism is visualized by the Wigner distributions at times t4t_4 and t5t_5 in Fig. 2(c), which correspond to two successive states within the stabilized cycle. The corresponding fixed-axis squeezing values, ξy2(t4)=−4.03dB _y^2(t_4)=-4.03\,dB and ξy2(t5)=−5.11dB _y^2(t_5)=-5.11\,dB, show that the learned policy maintains metrologically relevant squeezing while keeping the cycle-to-cycle oscillation bounded. As a result, the fixed-axis squeezing exhibits only weak and bounded oscillations throughout the stabilization stage, despite the continued action of the QZE term. Figure 3: Fidelity-based evidence for the toggling-frame stabilization mechanism. Fidelity evolution for several reference states. Solid and dashed curves denote evolution under f^y2 f_y^2 and f^z2 f_z^2, respectively. The reference states are the coherent spin state ,mx=f⟩ f,m_x=f , the PIRL state at t3t_3, and the OAT and TACT states at their respective minima of ξ2ξ^2. For the OAT and TACT states, an additional RxR_x rotation is applied so that ξy2=ξ2 _y^2=ξ^2. Figure 4: Metrological analysis of the PIRL strategy. (a) Phase sensitivity relative to the SQL, expressed in dB, as a function of the dimensionless interrogation time χtχ t. Encoding under the unified RL control starting from the stabilized state at time t4t_4 in Fig. 2(c) (red solid), QZE evolution from the TACT-optimal squeezed state (blue solid), and the same initial state under the RxR_x-pulse stabilization protocol of Ref. [36] (green solid). The black dashed line denotes the SQL. (b) Readout distributions Pm(ϕ)P_m(φ) for representative probe states and control protocols. Panels 1⃝ and 2⃝ show the distributions before (ϕ=0φ=0, top) and after (ϕ=0.2φ=0.2, bottom) phase encoding for the initial states used in the blue and red curves of panel (a), respectively. Panels 3⃝ and 4⃝ show the corresponding distributions at χt≈0.09χ t≈ 0.09 for the green and red curves in panel (a), respectively. (c) Single-atom magnetic-field sensitivity relative to the SQL, expressed in dB, as a function of the total protocol time TtotT_tot. The shaded region indicates the preparation time TpT_p, and the inset shows the absolute sensitivity δB(Ttot)δ B(T_tot) in nT/HznT/ Hz. The learned stabilization strategy differs qualitatively from the repeated-RxR_x correction protocol of Ref. [36]. Instead, the PIRL agent identifies a mechanism based on alternating transverse RyR_y rotations. Consider one elementary control cycle of duration dt=2δtdt=2δ t, e−iχδtf^z2Ry(−π/2)e−iχδtf^z2Ry(π/2).e^-iχδ t f_z^2\,R_y(-π/2)\,e^-iχδ t f_z^2\,R_y(π/2). (11) Under the two opposite RyR_y pulses, the second free-evolution interval is transformed as Ry(−π/2)e−iχδtf^z2Ry(π/2)=e−iχδtf^x2,R_y(-π/2)\,e^-iχδ t f_z^2\,R_y(π/2)=e^-iχδ t f_x^2, (12) because Ry(−π/2)f^z2Ry(π/2)=f^x2.R_y(-π/2)\, f_z^2\,R_y(π/2)= f_x^2. (13) One cycle therefore becomes e−iχδtf^z2e−iχδtf^x2=e−iχdt(f^z2+f^x2)+12(χδt)2[f^x2,f^z2]+O((χδt)3).e^-iχδ t f_z^2e^-iχδ t f_x^2=e^-iχ dt( f_z^2+ f_x^2)+ 12(χδ t)^2[ f_x^2, f_z^2]+O((χδ t)^3). (14) For the parameters used here, χδt∼10−3χδ t 10^-3, so the higher-order terms are negligible and the dynamics is well captured by the leading term. Using f^x2+f^z2=f(f+1)−f^y2, f_x^2+ f_z^2=f(f+1)- f_y^2, (15) we find that the effective generator is proportional to −f^y2- f_y^2, up to an irrelevant constant. The alternating Ry(±π/2)R_y(±π/2) pulses therefore suppress the direct distortion produced by f^z2 f_z^2 twisting and redirect the nonlinear evolution into a form more compatible with preserving the fixed-axis squeezing ξy2 _y^2. Since the effective dynamics is dominated by a f^y2 f_y^2-type term, its action is weak on states with small variance along f^y f_y. This picture is supported by the fidelity analysis in Fig. 3. Under f^z2 f_z^2 evolution, states with smaller ξy2 _y^2—and hence stronger anti-squeezing along f^z f_z—show faster fidelity decay, indicating greater susceptibility to QZE-induced distortion. By contrast, evolution under f^y2 f_y^2 preserves both the PIRL- and TACT-generated squeezed states much more effectively, consistent with the view that the learned protocol stabilizes the measurement-relevant squeezed quadrature. I.2 Phase sensitivity and metrological gain We now evaluate the metrological performance of the PIRL protocol under fixed projective readout of f^y f_y. In the stabilization regime, signal encoding proceeds under the learned periodic control cycle, which keeps the readout-relevant squeezing axis stable throughout the interrogation. During the sensing stage, a weak signal is encoded as a small phase ϕφ accumulated over an encoding time TeT_e. The encoded state is |ψenc⟩=∏k=1Ne[e−i(χf^z2+ϕTef^z)δtRy((−1)kπ/2)]|ψt4⟩,| _enc = _k=1^N_e [e^-i (χ f_z^2+ φT_e f_z )δ t\,R_y((-1)^kπ/2) ]| _t_4 , (16) where |ψt4⟩| _t_4 is the stabilized state at time t4t_4 in Fig. 2(c), and Ne=⌈Te/δt⌉N_e= T_e/δ t is the number of control steps during encoding. After encoding, the phase is read out from the spin projection f^y f_y. For small ϕφ, the phase sensitivity is quantified by (Δϕ)2=(Δf^y)2|∂ϕ⟨f^y⟩|2,( φ)^2= ( f_y)^2| _φ f_y |^2, which equals the inverse classical Fisher information in the linear-response regime [57]. Figure 4(a) compares the SQL-relative phase sensitivity ΔϕSQL/Δϕ _SQL/ φ as a function of the dimensionless interrogation time χtχ t. As references, we consider QZE evolution from the TACT-optimal squeezed state (blue), and the RxR_x-pulse stabilization protocol (see Appendix C) of Ref. [36] applied to the same initial state (green). The unified RL protocol (red) starts from the stabilized state at time t4t_4 in Fig. 2(c). At short interrogation times, the blue and green curves perform slightly better than the red one, because they start from the more strongly squeezed TACT-optimal probe. This initial advantage is short-lived: as the interrogation time increases and QZE-induced distortion accumulates, the freely evolving squeezed probe (blue) drops below the SQL already around χt≈0.02χ t≈ 0.02, while the optimized RxR_x protocol (green) remains above the SQL only up to about χt≈0.05χ t≈ 0.05 before degrading rapidly. By contrast, the unified RL protocol (red) keeps ΔϕSQL/Δϕ>1 _SQL/ φ>1 over a broad range of interrogation times, with only weak residual oscillations. Thus, once QZE-induced distortion becomes appreciable, the stability provided by the unified RL control outweighs the larger initial squeezing of the reference probes. Figure 4(b) shows the corresponding readout distributions Pm(ϕ)P_m(φ). Panels 1⃝ and 2⃝ compare the initial probe states used for the free-QZE reference (blue) and the unified RL protocol (red), respectively, before (ϕ=0φ=0, top) and after (ϕ=0.2φ=0.2, bottom) phase encoding. Although the RL-prepared probe (red) is slightly less squeezed than the TACT-optimal state (blue), its readout distribution remains narrow and close to Gaussian. Panels 3⃝ and 4⃝ show the corresponding distributions at χt≈0.09χ t≈ 0.09 for the RxR_x-pulse stabilization protocol (green) and the unified RL protocol (red), respectively. At this evolution time, the optimized RxR_x protocol (green) produces a strongly distorted, non-Gaussian distribution, whereas the PIRL protocol (red) preserves a narrow, approximately Gaussian profile. Accordingly, the error-propagation formula remains reliable for PIRL, but becomes inadequate once the readout statistics are strongly distorted. For a realistic Dy161^161Dy atomic qudit in a bias magnetic field at the Earth-field scale, B0=50μTB_0=50 10000\ , the relevant parameters are χ=(2π) 8.112Hzχ=(2π)\,8.112 10000\ Hz and the gyromagnetic ratio γ=(2π) 13.24GHz/Tγ\!=\!(2π)\,13.24 10000\ GHz/T. To quantify performance under a finite time budget, we convert the phase uncertainty into a magnetic-field sensitivity. Neglecting the readout time, the total protocol time is Ttot=Tp+Te,T_tot=T_p+T_e, where TpT_p is the fixed preparation time and TeT_e is the encoding time. The single-atom magnetic-field sensitivity is δB=ΔBTtot=ΔϕTtotγTe.δ B= B\, T_tot= φ\, T_totγ T_e. (17) As a benchmark, we use the coherent-spin-state (CSS) standard quantum limit (SQL), defined in the absence of QZE and with no preparation overhead, so that Ttot=TeT_tot=T_e and δBSQL(Ttot)=1γ2fTtot,δ B_SQL(T_tot)= 1γ 2f\,T_tot, (18) where ΔϕCSS=1/2f _CSS=1/ 2f. Figure 4(c) shows the SQL-relative magnetic-field sensitivity δBSQL/δBδ B_SQL/δ B as a function of TtotT_tot. Because of the finite preparation overhead, the unified RL protocol (red) is initially below the SQL. After the preparation cost is amortized, however, it surpasses the SQL at Ttot≈5.5msT_tot≈ 5.5\,ms and continues to improve at longer times. The inset shows the absolute sensitivity δB(Ttot)δ B(T_tot), which reaches 13.9pT/Hz13.9\,pT/ Hz at Ttot=30msT_tot=30\,ms, corresponding to an advantage of approximately 3dB3\,dB beyond the standard quantum limit. For fixed χ, increasing TtotT_tot corresponds to stronger accumulated QZE effects, so the continued improvement demonstrates that the learned protocol preserves metrologically useful correlations over extended interrogation times. IV summary and outlook We have shown that a unified RL control can convert the intrinsic QZE nonlinearity of an atomic qudit from a metrological liability into a useful sensing resource. Under always-on nonlinear dynamics, the learned control rapidly generates strong internal spin squeezing and then stabilizes the measurement-relevant quadrature through a simple pulse structure dominated by transverse rotations. This yields sustained phase sensitivity and, after accounting for the finite preparation overhead, a net magnetic-field sensitivity beyond the SQL. Beyond its numerical performance, the learned strategy also admits a transparent physical interpretation. The stabilization stage is dominated by alternating Ry(±π/2)R_y(±π/2) pulses, which effectively redirect the bare f^z2 f_z^2 nonlinear evolution into a form more compatible with preserving the fixed-readout squeezing. In this sense, the unified RL control does not merely discover a black-box pulse sequence, but rather identifies a physically meaningful control mechanism tailored to the always-on nonlinear dynamics. The control principle is not restricted to the specific Dy161^161Dy example studied here. Across different spin-f systems, independently trained agents consistently recover the same qualitative structure: a rapid squeezing stage followed by a stabilization stage under continued QZE evolution, although the detailed pulse sequences depend nontrivially on f. This indicates that the underlying mechanism is generic, while its optimal implementation remains system dependent. Detailed results for different spin-f systems are provided in Appendix D. More broadly, the present framework relies only on experimentally accessible low-order moments and realizable control operations, and is therefore compatible with other high-spin atoms and, in principle, with more general effective-spin models, including collective atomic ensembles. This opens a natural route toward combining internal squeezing control with interatomic entanglement generation in many-atom sensors, where additional metrological gain may become available [58, 59, 60, 61]. It is also compatible with emerging platforms offering spatially resolved readout, which may enable parallelized and site-resolved quantum sensing [62, 63, 64, 65]. Finally, the present work suggests a broader learning-based perspective on quantum sensing with intrinsic nonlinear dynamics. Here the reward is built from squeezing-based figures of merit, but the same framework can be extended to optimize sensing performance more directly. For example, when the signal has a prior distribution, one may train the controller against Bayesian posterior objectives rather than squeezing itself, thereby shifting the target from state preparation to end-to-end metrological performance [66, 67, 68]. Acknowledgements We thank Dr. Q. Liu for helpful discussions and insightful inputs. This work was supported by the National Natural Science Foundation of China (No. 12304543), the Quantum Science and Technology - National Science and Technology Major Project (No. 2021ZD0302100), the Natural Science Foundation of Jiangsu Province (Nos. BK20250404 and BG2025017), the Youth Science and Technology Talent Support Program of Jiangsu Province (Nos. JSTJ-2024-507, JSTJ-2025-600 and JSTJ-2025-829), the Frontier Technology Research Program of Suzhou (No. SYG202322), and the Postgraduate Research & Practice Innovation Program of NUAA (No. xcxjh20252107). DATA AVAILABILITY The data that support the findings of this study are publicly available at [URL]. Additional information is available from the corresponding author upon reasonable request. Appendix A Ground-State Hyperfine Structure of Dy Atoms In the presence of an external magnetic field, the internal structure of a Dy atom is governed by the combined hyperfine interaction between the electronic and nuclear angular momenta and their Zeeman couplings to the field. Denoting the nucleus and electron angular momenta by i and j, the full Hamiltonian is given by [69] alignedH^=A^⋅^+B 32^⋅^(2^⋅^+1−i(i+1)−j(j+1))2i(2i−1)j(2j−1)+gjμBj^zB+giμNi^zB. H=&A\, i\!·\! j+B 32 i\!·\! j (2 i\!·\! j+1-i(i+1)-j(j+1) )2i(2i-1)j(2j-1)\\ &+g_j _B j_zB+g_i _N i_zB. (19) where h is the Planck constant, A and B are the magnetic-dipole and electric-quadrupole hyperfine coefficients, gjg_j and gig_i are the electron and nuclear Landég factors, and μB _B and μN _N are Bohr and nuclear magnetons, respectively, BB is themagnetic field intensity. In this work, the leading magnetic field is assumed to be aligned along the laboratory z axis. We focus on the isotope Dy161^161Dy, whose ground state has j=8j=8 and i=5/2i=5/2. The hyperfine interaction splits the ground state into manifolds labeled by the total angular momentum f. In this work, we restrict attention to the ground-state hyperfine manifold with f=21/2f=21/2. The large-spin advantage of atomic spin systems has already enabled the realization of non-Gaussian and cat-like states for quantum-enhanced magnetometry with sensitivities approaching the Heisenberg limit [21, 22, 24]. In the weak-field regime, the Zeeman effect is non-negligible to but smaller than the hyperfine splitting. As a consequence, the energy spectrum within a fixed hyperfine manifold exhibits nonlinear dependence on the magnetic quantum number m. After numerically diagonalizing Eq. (19) in the |f,m⟩|f,m basis for a given magnetic-field strength BB, the eigenenergies EmE_m can be accurately fitted by a quadratic expansion, Em≃ℏ(ΩLm+χm2),E_m ( _Lm+χ m^2 ), (20) where ΩL _L is the Larmor frequency and χ is the QZE coefficient, also referred to as the NLZ coefficient or quantum-beat revival frequency. For atoms with j=1/2j=1/2, both ΩL _L and χ can be obtained analytically from the Breit–Rabi formula [70]. However, for high-spin atoms such as Dy161^161Dy with j>1/2j>1/2, no closed-form expression exists, and the numerical fitting is essential. For the leading magnetic field in Earth-field strengths around B0=50μTB_0=50 10000\ , we obtain ΩL=(2π) 661.9kHz,χ=(2π) 8.112Hz. _L=(2π) 10000\ 661.9 10000\ kHz,χ=(2π) 10000\ 8.112 10000\ Hz.The Larmor frequency satisfies ΩL=γB0 _L= _0, which corresponds to the gyromagnetic ratio γ=(2π)13.24GHz/Tγ=(2π)13.24 10000\ GHz/T [71]. The internal dynamics are therefore well described by the Hamiltonian H^=ℏ(ΩLf^z+χf^z2), H= ( _L f_z+χ f_z^2 ), (21) where f^z f_z is the z-component of the total angular-momentum operator for a single atom. The quadratic term ∝f^z2 f_z^2 gives rise to intrinsic OAT dynamics, which plays a central role in both the generation of spin squeezing and the QZE-induced degradation of metrological performance discussed in the main text. Appendix B PPO algorithm We train the PIRL agent using PPO, known as an actor-critic algorithm. The agent is parameterized by a policy network π _ θ as an actor parameterized by θ, and a value network V_ w as a critic parameterized by w. The system state at the kkth time step is the full quantum state |ψk⟩| _k of the atom, which evolves deterministically according to the controlled Schrödinger dynamics in Eq. (5). Instead of accessing the full state, the agent receives an observation oko_k defined in Eq. (7) at each discrete control step k, consisting of first moments and selected second moments of the spin operator. Based on oko_k, the policy samples an action aka_k from the discrete action set in Eq. (6), corresponding to experimentally accessible transverse rotations. The environment then evolves the underlying quantum state from |ψk⟩| _k to |ψk+1⟩| _k+1 , yielding the next observation ok+1o_k+1 and a scalar reward rkr_k in Eq. (10). Transitions (ok,ak,rk,ok+1)(o_k,a_k,r_k,o_k+1) are collected into a trajectory buffer and used for subsequent updates. The value network outputs the state value V(ok)V_ w(o_k), which estimates the expected return from observation oko_k; these value estimates are then used to compute the generalized advantage estimation (GAE), A^k=∑l=0Nt−k−1(γRLλ^)l[rk+l+γRLV(ok+l+1)−V(ok+l)], A_k= _l=0^N_t-k-1( _RL λ)^l [r_k+l+ _RLV_ w(o_k+l+1)-V_ w(o_k+l) ], (22) where γRL∈(0,1) _RL∈(0,1) is the discount factor and λ λ is a hyperparameter that controls the bias–variance trade-off. The policy network is updated by maximizing the clipped objective function Lactor()=k[min(ηk()A^k,clip(ηk(),1−ϵ,1+ϵ)A^k)],L_actor( θ)=E_k\! [ \! ( _k( θ) A_k,\,clip( _k( θ),1-ε,1+ε) A_k ) ], (23) with the probability ratio ηk()=π(ak|ok)πold(ak|ok). _k( θ)= _ θ(a_k|o_k) _ θ_ old(a_k|o_k). (24) The hyperparameter ϵε sets the clipping range for the policy ratio, which decreases the updating speed of the policy and improves the learning stability. The value network is updated by minimizing the mean-squared error, Lcritic()=k[min(V(ok)−k)2],L_critic( w)=E_k\! [ \! (V_ w(o_k)-G_k )^2 ], (25) with the discounted return k=∑l=0Nt−k−1γRLlrk+l.G_k= _l=0^N_t-k-1γ^l_RLr_k+l. (26) The total loss minimized during training is Ltotal=−Lactor()+c1Lcritic()−c2S(π),L_ total=-L_actor( θ)+c_1\,L_critic( w)-c_2\,S( _ θ), (27) where c1,c2c_1,c_2 are the hyperparameters that balance the three terms, and S(π)S( _ θ) is the policy entropy that promotes exploration. Parameters are updated over multiple epochs for each batch of collected trajectories. In our PIRL framework, the reward function is well designed to guide the agent to achieve both spin squeezing and stabilization. For this purposes,we choose ξ2(t)ξ^2(t) and ξy2(t)ξ^2_y(t) as the reward metrics, and introduce hyperparameters ζ,κ,αζ,κ,α to stabilize and accelerate the training process. To discourage the agent from taking unnecessary actions, we include a small penalty term c(ak)c(a_k) in the reward function. The parameters used for training in the Results section are listed in Table 1. Table 1: Parameters used for PPO training and environment design. Parameter Value PPO Actor learning rate 3×10−43× 10^-4 Critic learning rate 1×10−31× 10^-3 Discount factor γ 0.99990.9999 GAE parameter λ λ 0.960.96 Clip parameter ϵε 0.20.2 Value-loss weight c1c_1 0.50.5 Entropy weight c2c_2 0.010.01 Minibatch size 512 Update epochs 8 Max train steps 8×1068× 10^6 Buffer size 17920 Actor network MLP (128,128)(128,128) Critic network MLP (128,128)(128,128) Environment / reward spin f 21/2 Total steps NtN_t 70 Total time χTχ T 0.3140.314 Bonus scale ζ 5.05.0 Decay rate κ 0.050.05 Stabilization weight α 0.050.05 Action penalty c(ak)c(a_k) 0.0010.001 Appendix C Spin stabilization with D pulses As a reference strategy for comparison with the PIRL protocol, we consider a spin-stabilization scheme which incorporates a dynamic decoupling (D) sequence optimized by machine learning [36]. In this scheme, the initial state is chosen as the TACT-optimal squeezed state after an additional RxR_x rotation such that ξy2=ξ2 _y^2=ξ^2. Same as the PIRL protocol, the encoding time TeT_e is predefined and discretized with time step δt=Te/Neδ t=T_e/N_e. At each discrete time step of duration δtδ t, the system first an instantaneous rotation about the x axis and is then subjected to undergoes QZE evolution. The state update at the kkth step is |ψk+1⟩=e−iχdtf^z2e−iβkf^x|ψk⟩,| _k+1 =e^-iχ dt\, f_z^2\,e^-i _k f_x\,| _k , (28) where βk _k is the rotation angle applied at the kkth step. For a fair comparison with the discrete PIRL protocol, we use the same time step δtδ t. The rotation angle βk _k is arbitrary theoretically, while considering the running speed of the algorithm, βk _k is set to be nπ/48nπ/48, and the domain is set to [0,π/3][0,π/3]. In contrast to PIRL, this scheme is formulated as a conventional parameter optimization problem over the finite set βk\ _k\. The optimization is performed using the differential evolution (DE) algorithm [72, 73, 74]. The cost function is defined as the average value of the fixed-axis squeezing parameter during the encoding stage, =1Ne∑k=1Neξy2(k),C= 1N_e _k=1^N_e _y^2(k), (29) where NeN_e is the total number of discrete steps in the encoding stage. Appendix D Adaptability across different spin-f systems To assess the adaptability of the PIRL framework, we independently train the agent for a range of spin-f systems and compare the resulting squeezing performance and characteristic time scales. In all cases, the learned protocol exhibits the same qualitative two-stage structure: it first generates squeezing close to the effective TACT benchmark and then stabilizes the fixed-axis squeezing ξy2 _y^2 within a narrow oscillation window under continued QZE evolution, typically through alternating Ry(±π/2)R_y(±π/2) pulses. Quantitatively, the stabilized value of ξy2 _y^2 decreases overall with increasing f. For f=25/2f=25/2, corresponding to a metastable state of Dy [75], the protocol stabilizes ξy2 _y^2 at −5.14dB-5.14\,dB; for the larger spin f=16f=16, it reaches −6.31dB-6.31\,dB, approaching the regime relevant to collective-spin squeezing in larger effective-spin systems. The associated time scales show that the PIRL protocol enters the stabilized-ξy2 _y^2 regime earlier than the effective TACT model reaches its optimal squeezing. Overall, these results show that the learned strategies are not simple copies of a single pulse pattern: while the stabilization stage is generally dominated by alternating Ry(±π/2)R_y(±π/2) pulses, the squeezing stage and the onset of stabilization vary noticeably with f, highlighting the advantage of adaptive learning-based control. References Jackson Kimball et al. [2017] D. F. Jackson Kimball, J. Dudley, Y. Li, D. Patel, and J. Valdez, Constraints on long-range spin-gravity and monopole-dipole couplings of the proton, Physical Review D 96, 075004 (2017). Wang et al. [2020a] Z. Wang, X. Peng, R. Zhang, H. Luo, J. Li, Z. Xiong, S. Wang, and H. Guo, Single-species atomic comagnetometer based on rb 87 atoms, Physical Review Letters 124, 193002 (2020a). Fedderke et al. [2021a] M. A. Fedderke, P. W. Graham, D. F. Jackson Kimball, and S. Kalia, Earth as a transducer for dark-photon dark-matter detection, Physical Review D 104, 075023 (2021a). Fedderke et al. [2021b] M. A. Fedderke, P. W. Graham, D. F. Jackson Kimball, and S. Kalia, Search for dark-photon dark matter in the supermag geomagnetic field dataset, Physical Review D 104, 095032 (2021b). Arza et al. [2022] A. Arza, M. A. Fedderke, P. W. Graham, D. F. Jackson Kimball, and S. Kalia, Earth as a transducer for axion dark-matter detection, Physical Review D 105, 095007 (2022). Sander et al. [2012] T. Sander, J. Preusser, R. Mhaskar, J. Kitching, L. Trahms, and S. Knappe, Magnetoencephalography with a chip-scale atomic magnetometer, Biomedical optics express 3, 981 (2012). Kamada et al. [2015] K. Kamada, D. Sato, Y. Ito, H. Natsukawa, K. Okano, N. Mizutani, and T. Kobayashi, Human magnetoencephalogram measurements using newly developed compact module of high-sensitivity atomic magnetometer, Japanese Journal of Applied Physics 54, 026601 (2015). He et al. [2019] K. He, S. Wan, J. Sheng, D. Liu, C. Wang, D. Li, L. Qin, S. Luo, J. Qin, and J.-H. Gao, A high-performance compact magnetic shield for optically pumped magnetometer-based magnetoencephalography, Review of Scientific Instruments 90 (2019). Zhang et al. [2020] R. Zhang, W. Xiao, Y. Ding, Y. Feng, X. Peng, L. Shen, C. Sun, T. Wu, Y. Wu, Y. Yang, et al., Recording brain activities in unshielded earth’s field with optically pumped atomic magnetometers, Science Advances 6, eaba8792 (2020). Canciani and Raquet [2016] A. Canciani and J. Raquet, Absolute positioning using the earth’s magnetic anomaly field, NAVIGATION: Journal of the Institute of Navigation 63, 111 (2016). Canciani and Raquet [2017] A. Canciani and J. Raquet, Airborne magnetic anomaly navigation, IEEE Transactions on aerospace and electronic systems 53, 67 (2017). Gnadt [2022] A. Gnadt, Machine learning-enhanced magnetic calibration for airborne magnetic anomaly navigation, in AIAA SciTech 2022 forum (2022) p. 1760. Zhang et al. [2023] R. Zhang, D. Kanta, A. Wickenbrock, H. Guo, and D. Budker, Heading-error-free optical atomic magnetometry in the earth-field range, Physical Review Letters 130, 153601 (2023). Xiao et al. [2023] W. Xiao, M. Liu, T. Wu, X. Peng, and H. Guo, Femtotesla atomic magnetometer employing diffusion optical pumping to search for exotic spin-dependent interactions, Physical Review Letters 130, 143201 (2023). Lei et al. [2025] L. Lei, T. Wu, and H. Guo, Sensitivity of quantum magnetic sensing, National Science Review 12, nwaf129 (2025). Pezze et al. [2018] L. Pezze, A. Smerzi, M. K. Oberthaler, R. Schmied, and P. Treutlein, Quantum metrology with nonclassical states of atomic ensembles, Reviews of Modern Physics 90, 035005 (2018). Huang et al. [2024] J. Huang, M. Zhuang, and C. Lee, Entanglement-enhanced quantum metrology: From standard quantum limit to heisenberg limit, Appl. Phys. Rev. 11, 031302 (2024). Montenegro et al. [2025] V. Montenegro, C. Mukhopadhyay, R. Yousefjani, S. Sarkar, U. Mishra, M. G. Paris, and A. Bayat, Review: Quantum metrology and sensing with many-body systems, Phys. Rep. 1134, 1 (2025). Fernholz et al. [2008] T. Fernholz, H. Krauter, K. Jensen, J. F. Sherson, A. S. Sørensen, and E. S. Polzik, Spin Squeezing of Atomic Ensembles via Nuclear-Electronic Spin Entanglement, Phys. Rev. Lett. 101, 073601 (2008). Kurucz and Mølmer [2010] Z. Kurucz and K. Mølmer, Multilevel holstein-primakoff approximation and its application to atomic spin squeezing and ensemble quantum memories, Phys. Rev. A 81, 032314 (2010). Chalopin et al. [2018] T. Chalopin, C. Bouazza, A. Evrard, V. Makhalov, D. Dreon, J. Dalibard, L. A. Sidorenkov, and S. Nascimbene, Quantum-enhanced sensing using non-classical spin states of a highly magnetic atom, Nature communications 9, 4955 (2018). Evrard et al. [2019] A. Evrard, V. Makhalov, T. Chalopin, L. A. Sidorenkov, J. Dalibard, R. Lopes, and S. Nascimbene, Enhanced magnetic sensitivity with non-gaussian quantum fluctuations, Phys. Rev. Lett. 122, 173601 (2019). Satoor et al. [2021] T. Satoor, A. Fabre, J.-B. Bouhiron, A. Evrard, R. Lopes, and S. Nascimbene, Partitioning dysprosium’s electronic spin to reveal entanglement in nonclassical states, Phys. Rev. Res. 3, 043001 (2021). Yang et al. [2025a] Y. Yang, W.-T. Luo, J.-L. Zhang, S.-Z. Wang, C.-L. Zou, T. Xia, and Z.-T. Lu, Minute-scale schrödinger-cat state of spin-5/2 atoms, Nature Photonics 19, 89 (2025a). Kitagawa and Ueda [1993] M. Kitagawa and M. Ueda, Squeezed spin states, Physical Review A 47, 5138 (1993). Norcia and Ferlaino [2021] M. A. Norcia and F. Ferlaino, Developments in atomic control using ultracold magnetic lanthanides, Nature Physics 17, 1349 (2021). Acosta et al. [2006] V. Acosta, M. Ledbetter, S. Rochester, D. Budker, D. Jackson Kimball, D. Hovde, W. Gawlik, S. Pustelny, J. Zachorowski, and V. Yashchuk, Nonlinear magneto-optical rotation with frequency-modulated light in the geophysical field range, Physical Review A—Atomic, Molecular, and Optical Physics 73, 053404 (2006). Budker and Romalis [2007] D. Budker and M. Romalis, Optical magnetometry, Nature physics 3, 227 (2007). Seltzer et al. [2007] S. Seltzer, P. Meares, and M. Romalis, Synchronous optical pumping of quantum revival beats for atomic magnetometry, Physical Review A—Atomic, Molecular, and Optical Physics 75, 051407 (2007). Jensen et al. [2009] K. Jensen, V. Acosta, J. Higbie, M. Ledbetter, S. Rochester, and D. Budker, Cancellation of nonlinear zeeman shifts with light shifts, Physical Review A—Atomic, Molecular, and Optical Physics 79, 023406 (2009). Wasilewski et al. [2010] W. Wasilewski, K. Jensen, H. Krauter, J. J. Renema, M. Balabas, and E. S. Polzik, Quantum noise limited and entanglement-assisted magnetometry, Physical Review Letters 104, 133601 (2010). Bao et al. [2018] G. Bao, A. Wickenbrock, S. Rochester, W. Zhang, and D. Budker, Suppression of the nonlinear zeeman effect and heading error in earth-field-range alkali-vapor magnetometers, Physical review letters 120, 033202 (2018). Yang et al. [2021] P. Yang, G. Bao, L. Chen, and W. Zhang, Coherence protection of electron spin in earth-field range by all-optical dynamic decoupling, Physical Review Applied 16, 014045 (2021). Lee et al. [2021] W. Lee, V. Lucivero, M. Romalis, M. Limes, E. Foley, and T. Kornack, Heading errors in all-optical alkali-metal-vapor magnetometers in geomagnetic fields, Physical Review A 103, 063103 (2021). Bao et al. [2022] G. Bao, D. Kanta, D. Antypas, S. Rochester, K. Jensen, W. Zhang, A. Wickenbrock, and D. Budker, All-optical spin locking in alkali-metal-vapor magnetometers, Physical Review A 105, 043109 (2022). Yang et al. [2025b] P. Yang, G. Bao, J. Chen, W. Du, J. Guo, and W. Zhang, Quantum locking of intrinsic spin squeezed state in earth-field-range magnetometry, npj Quantum Information 11, 36 (2025b). Meng et al. [2023] X. Meng, Y. Zhang, X. Zhang, S. Jin, T. Wang, L. Jiang, L. Xiao, S. Jia, and Y. Xiao, Machine learning assisted vector atomic magnetometry, Nature Communications 14, 6105 (2023). Duan et al. [2025] J. Duan, Z. Hu, X. Lu, L. Xiao, S. Jia, K. Mølmer, and Y. Xiao, Concurrent spin squeezing and field tracking with machine learning, Nature Physics 21, 909 (2025). Banerjee et al. [2025] C. Banerjee, K. Nguyen, C. Fookes, and M. Raissi, A survey on physics informed reinforcement learning: Review and open problems, Expert Systems with Applications 287, 128166 (2025). Wineland et al. [1992] D. J. Wineland, J. J. Bollinger, W. M. Itano, F. Moore, and D. J. Heinzen, Spin squeezing and reduced quantum noise in spectroscopy, Physical Review A 46, R6797 (1992). Liu et al. [2011] Y. Liu, Z. Xu, G. Jin, and L. You, Spin squeezing: Transforming one-axis twisting into two-axis twisting, Physical review letters 107, 013601 (2011). Chen et al. [2019] F. Chen, J.-J. Chen, L.-N. Wu, Y.-C. Liu, and L. You, Extreme spin squeezing from deep reinforcement learning, Physical Review A 100, 041801 (2019). Cox et al. [2016] K. C. Cox, G. P. Greve, J. M. Weiner, and J. K. Thompson, Deterministic squeezed states with collective measurements and feedback, Physical review letters 116, 093602 (2016). Hosten et al. [2016] O. Hosten, N. J. Engelsen, R. Krishnakumar, and M. A. Kasevich, Measurement noise 100 times lower than the quantum-projection limit using entangled atoms, Nature 529, 505 (2016). Pedrozo-Peñafiel et al. [2020] E. Pedrozo-Peñafiel, S. Colombo, C. Shu, A. F. Adiyatullin, Z. Li, E. Mendez, B. Braverman, A. Kawasaki, D. Akamatsu, Y. Xiao, et al., Entanglement on an optical atomic-clock transition, Nature 588, 414 (2020). Robinson et al. [2024] J. M. Robinson, M. Miklos, Y. M. Tso, C. J. Kennedy, T. Bothwell, D. Kedar, J. K. Thompson, and J. Ye, Direct comparison of two spin-squeezed optical clock ensembles at the 10- 17 level, Nature Physics 20, 208 (2024). Schulman et al. [2017] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347 (2017). Wang et al. [2020b] Y. Wang, H. He, and X. Tan, Truly proximal policy optimization, in Uncertainty in artificial intelligence (PMLR, 2020) p. 113–122. Gu et al. [2021] Y. Gu, Y. Cheng, C. P. Chen, and X. Wang, Proximal policy optimization with policy feedback, IEEE Transactions on Systems, Man, and Cybernetics: Systems 52, 4600 (2021). Konda and Tsitsiklis [1999] V. Konda and J. Tsitsiklis, Actor-critic algorithms, Advances in neural information processing systems 12 (1999). Grondman et al. [2012] I. Grondman, L. Busoniu, G. A. Lopes, and R. Babuska, A survey of actor-critic reinforcement learning: Standard and natural policy gradients, IEEE Transactions on Systems, Man, and Cybernetics, part C (applications and reviews) 42, 1291 (2012). Bukov and Marquardt [2026] M. Bukov and F. Marquardt, Reinforcement learning for quantum technology, arXiv preprint arXiv:2601.18953 (2026). Yao et al. [2021] J. Yao, L. Lin, and M. Bukov, Reinforcement learning for many-body ground-state preparation inspired by counterdiabatic driving, Physical Review X 11, 031070 (2021). Li et al. [2025] S. Li, Y. Fan, X. Li, X. Ruan, Q. Zhao, Z. Peng, R.-B. Wu, J. Zhang, and P. Song, Robust quantum control using reinforcement learning from demonstration, npj Quantum Information 11, 124 (2025). Porotti et al. [2022] R. Porotti, A. Essig, B. Huard, and F. Marquardt, Deep reinforcement learning for quantum state preparation with weak nonlinear measurements, Quantum 6, 747 (2022). Porotti et al. [2019] R. Porotti, D. Tamascelli, M. Restelli, and E. Prati, Coherent transport of quantum states by deep reinforcement learning, Communications Physics 2, 61 (2019). Hayes et al. [2018] A. J. Hayes, S. Dooley, W. J. Munro, K. Nemoto, and J. Dunningham, Making the most of time in quantum metrology: concurrent state preparation and sensing, Quantum Science and Technology 3, 035007 (2018). Kurucz and Mølmer [2010] Z. Kurucz and K. Mølmer, Multilevel holstein-primakoff approximation and its application to atomic spin squeezing and ensemble quantum memories, Physical Review A—Atomic, Molecular, and Optical Physics 81, 032314 (2010). Norris et al. [2012] L. M. Norris, C. M. Trail, P. S. Jessen, and I. H. Deutsch, Enhanced squeezing of a collective spin via control of its qudit subsystems, Physical review letters 109, 173603 (2012). Zhang et al. [2025] Y. Zhang, S. Jin, J. Duan, K. Mølmer, G. Zhang, M. Wang, and Y. Xiao, Cooperative squeezing of internal and collective spins in an atomic ensemble, Phys. Rev. Lett. 135, 213604 (2025). Hu et al. [2026] Z. Hu, Y. Zhang, J. Duan, M. Wang, and Y. Xiao, Enhancing collective spin squeezing via one-axis twisting echo control of individual atoms (2026), arXiv:2602.14036 [quant-ph] . Antonijevic and Wimperis [2003] S. Antonijevic and S. Wimperis, Refocussing of chemical and paramagnetic shift anisotropies in 2h nmr using the quadrupolar-echo experiment, Journal of Magnetic Resonance 164, 343 (2003). Shaniv et al. [2019] R. Shaniv, N. Akerman, T. Manovitz, Y. Shapira, and R. Ozeri, Quadrupole shift cancellation using dynamic decoupling, Physical Review Letters 122, 223204 (2019). Na Narong et al. [2025] T. Na Narong, H. Li, J. Tong, M. Dueñas, and L. Hollberg, Quantum states imaging of magnetic field contours based on autler-townes effect in ytterbium atoms, Phys. Rev. Lett. 134, 193201 (2025). Zhang et al. [2026] Y.-W. Zhang, D.-S. Xiang, R. Liao, H.-X. Liu, B. Xu, P. Zhou, Y. Zhou, K. Zhang, and L. Li, Microwave electrometry with quantum-limited resolutions in a rydberg-atom array, Phys. Rev. Lett. 136, 110802 (2026). Kaubruegger et al. [2021] R. Kaubruegger, D. V. Vasilyev, M. Schulte, K. Hammerer, and P. Zoller, Quantum variational optimization of ramsey interferometry and atomic clocks, Phys. Rev. X 11, 041045 (2021). Liu et al. [2025] Q. Liu, M. Xue, M. Radzihovsky, X. Li, D. V. Vasilyev, L.-N. Wu, and V. Vuletić, Enhancing dynamic range of sub-standard-quantum-limit measurements via quantum deamplification, Phys. Rev. Lett. 135, 040801 (2025). Ma et al. [2025] Z. Ma, C. Han, Z. Tan, H. He, S. Shi, X. Kang, J. Wu, J. Huang, B. Lu, and C. Lee, Adaptive cold-atom magnetometry mitigating the trade-off between sensitivity and dynamic range, Science Advances 11, eadt3938 (2025). Corney [2006] A. Corney, Atomic and Laser Spectroscopy (Oxford University Press, 2006). Breit and Rabi [1931] G. Breit and I. I. Rabi, Measurement of nuclear spin, Physical Review 38, 2082 (1931). Lu et al. [2012] M. Lu, N. Q. Burdick, and B. L. Lev, Quantum degenerate dipolar fermi gas, Physical Review Letters 108, 215301 (2012). Storn and Price [1997] R. Storn and K. Price, Differential evolution–a simple and efficient heuristic for global optimization over continuous spaces, Journal of global optimization 11, 341 (1997). Lovett et al. [2013] N. B. Lovett, C. Crosnier, M. Perarnau-Llobet, and B. C. Sanders, Differential evolution for many-particle adaptive quantum metrology, arXiv preprint arXiv:1304.2246 (2013). Yang et al. [2019] X. Yang, J. Li, and X. Peng, An improved differential evolution algorithm for learning high-fidelity quantum controls, Science Bulletin 64, 1402 (2019). Budker et al. [1994] D. Budker, D. DeMille, E. D. Commins, and M. S. Zolotorev, Experimental investigation of excited states in atomic dysprosium, Phys. Rev. A 50, 132 (1994).