Paper deep dive
Free energy landscape of Dense Associative Memory
Sumedha, Abhishek Singh
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/24/2026, 1:51:57 AM
Summary
This paper derives a general expression for the free energy functional of Dense Associative Memories (DenseAMs) using large deviations theory. It analyzes polynomial interaction models and Log-Sum-Exponential (LSE) activation, revealing that memory retrieval in higher-order dense networks depends heavily on the initial state due to persistent local minima at zero overlap. The study provides exact full-retrieval thresholds for the LSE model and demonstrates that the method systematically handles complex architectures beyond traditional Hopfield networks.
Entities (7)
Relation Signals (5)
Large Deviations Theory → usedfor → Free Energy Functional
confidence 95% · Using large deviations theory, we solve and obtain a general expression for the free energy functional
Initial State → influences → memory retrieval
confidence 92% · Our analytical framework reveals how memory retrieval depends on the initial state in higher-order dense networks
Dense Associative Memory → uses → Log-Sum-Exponential
confidence 90% · dense associative memories featuring polynomial interactions and Log-Sum-Exponential (LSE) activation
Dense Associative Memory → surpasses → Hopfield Model
confidence 88% · DenseAMs are energy-based neural architectures that vastly surpass traditional Hopfield networks
Log-Sum-Exponential → provides → Exact Full-Retrieval Threshold
confidence 85% · gives the exact full-retrieval threshold for the LSE model
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative memories, including dense associative memories. We illustrate the method by reproducing classical results for the Hopfield model. For a finite number of patterns, we derive the temperature-dependent free energy functional for dense associative memories featuring polynomial interactions and Log-Sum-Exponential (LSE) activation. We also evaluate the disorder-averaged ground-state energy of these systems in the extensive limit. Our analytical framework reveals how memory retrieval depends on the initial state in higher-order dense networks, and gives the exact full-retrieval threshold for the LSE model. This method provides a systematic procedure for analyzing diverse, complex architectures in associative memory.
Tags
Links
- Source: https://arxiv.org/abs/2607.19195v2
- Canonical: https://arxiv.org/abs/2607.19195v2
Trouble viewing inline? Open PDF directly →
Full Text
26,667 characters extracted from source content.
Expand or collapse full text
Free energy landscape of Dense Associative Memory Sumedha,1,2 sumedha@niser.ac.in Abhishek Singh1,2 1School of Physical Sciences, National Institute of Science Education and Research, Bhubaneswar, P.O. Jatni, 752050, India 2Homi Bhabha National Institute, Training School Complex, Anushakti Nagar 400094, India Abstract Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative memories, including dense associative memories. We illustrate the method by reproducing classical results for the Hopfield model. For a finite number of patterns, we derive the temperature-dependent free energy functional for dense associative memories featuring polynomial interactions and Log-Sum-Exponential (LSE) activation. We also evaluate the disorder-averaged ground-state energy of these systems in the extensive limit. Our analytical framework reveals how memory retrieval depends on the initial state in higher-order dense networks, and gives the exact full-retrieval threshold for the LSE model. This method provides a systematic procedure for analyzing diverse, complex architectures in associative memory. Dense Associative Memories (DenseAMs) are energy-based neural architectures that vastly surpass traditional Hopfield networks by using higher-than-quadratic neuron interactions to expand memory storage capacity krotov1 ; gardner ; hopfield . In these networks, stored patterns represent low-energy configurations, allowing the system to perform powerful error correction by mapping partial or corrupted inputs back to the correct pattern. This process closely mirrors the biological error-correction mechanisms of the human brain, which makes it possible for it to retrieve memories from incomplete information anderson . Because of this high-capacity retrieval and error-correction, DenseAMs have become useful tools in areas like modern deep learning, generative AI, pattern recognition, and transformer architectures krotov2 ; simon ; yampolskaya ; transformer . In his seminal work, Hopfield mapped neuronal states to Ising spins and synaptic weights to the Hebb rule hebb . This established a spin-glass-like model where specific configurations serve as stored memories. As a result, the model was successfully solved using the replica method amit1 ; amit2 . In associative memory (AM) models, gradient descent dynamics drive the retrieval of stored memories by guiding the system toward the nearest local minimum on a free-energy landscape. The recent work in DenseAMs though have primarily concentrated on energy dynamics rather than the full free-energy framework. In this paper, we derive the free energy landscape for a general class of associative memories and use it to study higher DenseAMskrotov1 ; ramsauer ; demi ; lucibello . We develop a method of solution that makes use of the tilted property of the measure of exponential functions to obtain the rate function. We had earlier used similar methods to study random field quenched disorder for ferromagnets sumedha-sushant ; soheli-sumedha ; sumedha-barma . We consider models with neurons as Ising spins. Each neuron exists in two states : ±1± 1. A pattern is a configuration of N binary random variables ξ1μ,…ξNμ\ _1^μ,... _N^μ\ , where μ takes P values for P stored patterns. For a configuration C:s1,s2,…si,…,sNC:\s_1,s_2,… s_i,…,s_N\, a random variable mμ=1N∑i=1Nξiμsi m^μ= 1N _i=1^N _i^μs_i (1) captures the overlap between the μthμ^th memory and C. If the configuration C matches with the pattern μ, then mμ=1m^μ=1. We define a P dimensional vector =m1,m2…..mP m=\m^1,m^2.....m^P\ to represent the overlap between a configuration and the patterns. A general DenseAM krotov1 is defined by an energy function/Hamiltonian of the form: H()=−∑μ=1PF(∑i=1Nξiμsi)=−Nf() H( m)=- _μ=1^PF ( _i=1^N _i^μs_i )=-Nf( m) (2) where F(x)F(x) is a smooth function. The gradient descent dynamics then works in the direction of lowering the energy and stops when energy can no longer be decreased via single spin flipkrotov2 . Hence H can be considered the cost function of the memory retrieval which is to be minimised. In this work we develop a formalism using large deviations to get the free energy for any function f. For P patterns represented by the vectors ξ^μ, the joint probability distribution of the order parameter ()( m), i.e., Q()Q( m) in the absence of cost function is given by Q() Q( m) =∑sip(si)∏μ=1Pδ(∑i=1Nξiμsi=Nmμ) = _\s_i\p(\s_i\) _μ=1^Pδ ( _i=1^N _i^μs_i=Nm^μ ) (3) here p(si)p(\s_i\) is the probability of having a configuration C=siC=\s_i\. Using the integral representation of the delta function by introducing a set of variables =λ1,…λμ…,λP λ=\ _1,... _μ..., _P\ we get Q() Q( m) =∫dP(2π)Pexp(−N⋅) = d^P λ(2π)^P (-N 1.0muλ -1.0mu· m)\, (4) (∑sip(si)exp(∑μλμ∑iξiμsi)) ( _\s_i\p(\s_i\) ( _μ _μ _i _i^μs_i) ) The sum over spins sis_i can be done easily to give Q() Q( m) =∫dP(2π)Pexp(−N⋅)∏i=1Ncosh(∑μλμξiμ) = d^P λ(2π)^P (-N 1.0muλ -1.0mu· m) _i=1^N ( _μ _μ _i^μ ) (5) =∫dP(2π)Pexp[−N(⋅−Φ())] = d^P λ(2π)^P [-N ( λ· m- ( λ) ) ] where Φ(,)=1N∑i=1Nlog(cosh(∑μλμξiμ)). ( λ, ξ)= 1N _i=1^N ( ( _μ _μ _i^μ)). In the limit N→∞N→∞, via the law of large numbers we get Φ()=[logcosh(∑μλμξμ)], ( λ)=E_ ξ\! [ ( _μ _μξ^μ) ], (6) where, [⋯]E_ ξ\! [·s ] is the average over the distribution from which the patterns ξ have been drawn. The contours of integration in Eq. 5 are understood to be analytically deformed, so that they pass the saddle point. The probability Q()Q( m) satisfies large deviation principle with rate function I0()I_0( m) , i.e Q()≍exp[−NI0()]Q( m) [-NI_0( m)]. The saddle point approximation then gives I0() I_0( m) =sup∈ℝP[⋅−Φ()] = _ λ ^P [ λ· m- ( λ) ] (7) Let ∗ λ^* be the supremum of the function on the right. It is a solution of a set of P equations of the kind m∗μ=∂ϕ()∂λμ|∗ m^μ_*= . ∂φ( λ)∂ _μ _ m_* (8) In general, in order to find ∗ λ^*, we need to find the solution of the above P dimensional array of equations. For P=1P=1 it is easy and I0()I_0( m) is the rate function of a set of N non-interacting Ising spins. For higher P we circumvent the step where the hard inversion need to be performed by developing a procedure similar to the one used by us for random field ferro-magnets soheli-sumedha ; sumedha-barma . The probability of a configuration C of neurons under Gibbs measure is proportional to exp(−βH()) (-β H( m)) with β as the inverse temperature that controls the noise. Since m is a random variable drawn from a distribution Q()Q( m) that is defined on −1,1P\-1,1\^P, the probability QH,β()Q_H,β( m), in the presence of the cost function f()f( m) is given by QH,β()=∫AQ()exp(Nβf()) Q_H,β( m)= _AQ( m) (Nβ f( m)) (9) where A is the subset of all possible configurations compatible with the value m for P patterns. For QH,β()∼exp(−NI)Q_H,β( m) (-NI), the rate function I can be calculated using the tilted large deviations principal hollander ; sumedha-nabin that connects I and I0I_0 through the relation Iβ() I_β( m) =I0()−βf() =I_0( m)-β f( m) (10) where I0()=∗⋅−Φ(∗)I_0( m)= λ^*· m- ( λ^*). The function Iβ()I_β( m) is the free energy landscape for the system whose global minima gives the value of the free energy at a given β. Let ∗ m_* be a fixed point of IβI_β, then at ∗ m_*, since ∂Iβ∂mμ|∗=0 . ∂ I_β∂ m^μ _ m_ =0, ∂I0∂mμ|∗=β∂f∂mμ|∗ . ∂ I_0∂ m^μ _ m_ =β . ∂ f∂ m^μ _ m_ (11) Hence, we have ∂I0∂mμ I_0m^μ =∂[∗()⋅−Φ(∗())]∕∂mμ = * [ λ ( m)· m- ( λ ( m)) ]m^μ (12) =λμ∗+∑νmν∂λν∂mμ|∗−∑ν∂λν∂mμ|∗∂Φ()∂λν|∗ =λ _μ+ _νm^ν . ∂ _ν∂ m^μ _ λ^*- _ν . ∂ _ν∂ m^μ _ λ^* . ∂ ( λ)∂λ^ν _ λ^* Substituting back in Eq. 11, we get: λμ∗=β∂f∂mμ|∗ λ _μ=β . ∂ f∂ m^μ _ m_ (13) For our case of binary spins we use Eq. 6 to arrive at the P self-consistency equations for fixed point m∗μ\m^μ_*\. They are m∗μ=[ξμtanh(∑νβ∂f∂mν|∗ξν)] m_ ^μ=E_ ξ\! [ξ^μ ( _νβ . ∂ f∂ m^ν _ m_ ξ^ν) ] (14) The rate function, which is the generalised free energy functional of the DenseAMs comes out to be IH,β() I_H,β( m) =βℱ()=β∑νm∗ν∂f∂mν|∗ = ( m)=β _νm^ν_ . ∂ f∂ m^ν _ m_ (15) −[logcosh(∑νβ∂f∂mν|∗ξν)]−βf(∗). -E_ ξ\! [ ( _νβ . ∂ f∂ m^ν _ m_ ξ^ν) ]-β f( m_ ). We have defined ℱ()=βIH,βF( m)=β I_H,β , as the free energy functional of the system and the free energy is 1βinfIH,β 1βinf_ mI_H,β. Dense Associative memory Hopfield Model with polynomial interaction: The Hamiltonian is HN=−Nk!∑μ=1P(mμ)k H_N=- Nk! _μ=1^P(m^μ)^k (16) The f()=1k!∑μ=1P(mμ)kf( m)= 1k! _μ=1^P(m^μ)^k and λμ∗=β(k−1)!(mμ)k−1λ _μ= β(k-1)!(m^μ)^k-1. Substituting in Eq. 14 and 15 we get mμ m^μ =[ξμtanh(∑νβ(k−1)!(mμ)k−1ξν)] =E_ ξ\! [ξ^μ ( _ν β(k-1)!(m^μ)^k-1ξ^ν) ] (17) ℱ() ( m) =k−1k!∑μ=1P(mμ)k = k-1k! _μ=1^P(m^μ)^k (18) −1β[logcosh(∑νβ(k−1)!(mν)k−1ξν)] - 1βE_ ξ\! [ ( _ν β(k-1)!(m^ν)^k-1ξ^ν) ] Note that while for even k, the above equations hold for −1<mμ<1-1<m^μ<1, for odd k they are valid only for positive values 0<mμ<10<m^μ<1. This is because for odd k, negative values of mμm^μ results in a higher energy than the mirror state with positive mμm^μ. The Eq. 17 matches with the stochastic dynamics fixed point equation for DenseAMs obtained recentlyrooke . For k=2k=2, the expression of ℱ()F( m) derived above matches with the classic results in amit1 . The ℱ()F( m) for k>2k>2 however was not known from the earlier studies. Our method gives the exact expression of the free energy functional for any k. The large deviations approach allowed us to handle higher order interactions with ease, which is not possible with the standard Hubbard-Stratonovich transformation. Figure 1: Plot of disorder averaged ground state energy of retrieval(R(m)R(m)) of a random memory when N,P→∞N,P→∞. We have taken ek=1e_k=1 and hence γ=αγ=α. a) for Hopfield model,b) k=4k=4 DenseAM. In both cases the function R(m)R(m) is plotted near, above and below the transition, In (c) k=4k=4 basin of m=0m=0 state is shown to illustrtate that it is always present. In (d) R(m)R(m) for LSE model below the transition is shown. We can study any finite P using Eqs. 17 and 18. For P=1P=1, the ℱF and mμm_μ equation are similar to that of a k-spin Ising model. Hence single pattern DenseAMs has continuous transition for k=2k=2 and first order transition for k≥3k≥ 3. The transition temperatures for example for k=2,3,4k=2,3,4 are Tc=1/βc=1,0.23,0.04T_c=1/ _c=1,0.23,0.04 respectively. The stability analysis can also be performed straightforwardly by finding the eigenvalues of the Hessian for P=2P=2. We would though focus on the extensive limit in the rest of this section. The capacity of a network is defined as the maximum number of patterns that can be stored such that a randomly chosen pattern is retrieved with small error as N→∞N→∞, P→∞P→∞ at zero temperature. We study the ℱ()F( m) by taking N→∞N→∞, P→∞P→∞ along with β→∞β→∞ for the retrieval of a randomly chosen pattern. Each of (mμ)k(m^μ)^k corresponding to other patterns is taken as a random variable distributed with mean 0 and standard deviation ekN(k−1)/2 e_kN^(k-1)/2 rooke ; krotov1 , where eke_k is a constant dependent on k. Then by central limit theorem, the z=∑μ=2P(mμ)k−1ξμz= _μ=2^P(m^μ)^k-1 _μ can be taken as a Gaussian random variable with mean 0 and variance Pek2/Nk−1Pe_k^2/N^k-1. For P=αNk−1P=α N^k-1, γ=ek2αγ=e_k^2α is the variance and p(z)=12πγe−z2/2γp(z)= 1 2πγe^-z^2/2γ. Separating m1=m^1=m from other patterns and performing the average over ξ1ξ^1 and z in the limit of β→∞β→∞, we get the disorder averaged ground state energy of the system as R() R( m) =k−1k!mk−1(k−1)!2γπexp(−m2(k−1)2γ) = k-1k!m^k- 1(k-1)! 2γπexp ( -m^2(k-1)2γ ) (19) +mk−1(k−1)!erf(mk−12γ) + m^k-1(k-1)!erf ( m^k-1 2γ ) We dropped the term ∑μ=2P(mμ)k _μ=2^P(m^μ)^k while writing the above expression as its mean is 0 and variance decays as 1/N1/N. The Eq. 19 gives the landscape of the gradient dynamics sumedha . The fixed points of Eq. 19 are given by m=erf(mk−12γ) m=erf ( m^k-1 2γ ) (20) For k=2k=2 this equation is the same as the fixed point equation for m obtained in amit2 , with γ=rαγ=rα, where r is mean square random overlap obtained via replica calculation. Let us study the consequence of R(m)R(m) for DenseAMs in a bit more detail: for k=2k=2, the function R(m)R(m) has a minima at m=0m=0 for high γ, which continuously changes into a double well as γ is lowered as shown in Fig. 1. The system falls into one of the two minima spontaneously with decreasing γ at γc=2/π _c=2/π. The local and global minima are the same and gradient descent always reaches the global minima. For k>2k>2 the system undergoes a first order transition from m=0m=0 to m≠0m≠ 0 state. The m=0m=0 stops being a global minima at γg _g but continues to be a local minima as γ is decreased. This can be seen by evaluating the second derivative of R(m)R(m) at m=0m=0. χ(m)=∂2R(m)/∂m2=1−(k−1)2/πγmk−2em−2(k−1)/2γχ(m)=∂^2R(m)/∂ m^2=1-(k-1) 2/πγm^k-2e^m^-2(k-1)/2γ. For k>2k>2, χ(0)=1χ(0)=1 and hence m=0m=0 is a minima for all values of γ, though it stops being a global minima at a certain γg _g (see Fig. 1). As a result, the steady state of gradient descent depends on the initial starting state of the system. If the initial state is in the basin of attraction of m=0m=0, the pattern is not retrieved for any γ. At a threshold γl>γg _l> _g, the m≠0m≠ 0 minima shows up for the first time. As k increases the basin of attraction of this non zero fixed point shrinks, though the minima gets closer to m=1m=1. As a result there is less error in retrieval, provided the starting state is in the basin of the rerieval state. Since basin reduces, choices of initial state for successful memory retrieval gets more restricted. The threshold on αk _k for reliable retrieval of memory depends on the percentage of allowed error (given by (1−m)/2(1-m)/2). This for 1.5%1.5\% error for Hopfield model is α≈0.14α≈ 0.14 hopfield ; amit2 . A bound for 0.5%0.5\% was obtained in krotov1 . This threshold is different than the transition threshold discussed above as the m at the transition though non zero approaches 11 only in the limit of k→∞k→∞ (see Table 1). We show next for LSE the two thresholds match completely due to full pattern retrieval at and below the transition. k γg _g γl _l m∗m_* 2 0.66 0.66 0 3 0.18 0.26 0.85 4 0.1 0.2 0.92 5 0.063 0.17 0.95 10 0.015 0.13 0.98 Table 1: For γ<γlγ< _l a local minima at m∗m_* in R(m)R(m) appears which becomes a global minima at γg _g. Log-Sum-Exponential (LSE) model: The Energy function is HN(|)=12‖2−1λlog(∑μ=1Pexp(λ⋅)) H_N( σ| ξ)= 12|| σ||^2- 1λ ( _μ=1^P (λ σ· ξ^μ)) (21) where σ is state of the neuron and λ is the interaction strength. The second term is the LSE term which is equal to 1λlog(∑μ=1Pexp(Nλmμ)) 1λ ( _μ=1^P (Nλ m^μ)). We define ϕ=1λNln∑μ=2Pexp(λNmμ)φ= 1λ Nln _μ=2^P (λ Nm^μ) as was done in lucibello . Then we get fLSE=−1Nλlog(eλNm1+eλNϕ) f_LSE=- 1Nλlog (e^λ Nm^1+e^λ Nφ ) (22) here we have taken f=HN/Nf=H_N/N to maintain the concavity of the LSE function. For small N, substituting the f above in Eqs. 14 and 15 gives the exact ℱ()F( m) as: ℱN() _N( m) =−∑ν(mνeλNmν∑μeNλmμ)+1λNlog(∑μeλNmμ) =- _ν ( m^νe^λ Nm^ν _μe^Nλ m^μ )+ 1λ N ( _μe^λ Nm^μ) −1β[logcoshβ∑ν(ξνeλNmν∑μeNλmμ)] - 1βE_ ξ\! [ β _ν ( ξ^νe^λ Nm^ν _μe^Nλ m^μ ) ] (23) We now focus on the large N limit here. Using the law of large numbers we approximate ϕ∼P⟨exp(λNmμ)⟩mμφ P (λ Nm^μ) _m^μ. The mμm^μ for μ≠1μ≠ 1 are random variables drawn from a Gaussian distribution with mean 0 and standard deviation σ/Nσ/ N. Hence p(mμ)∼exp(−NI0(mμ))p(m^μ) (-NI_0(m^μ)) with I0=(mμ)2/(2σ2)I_0=(m^μ)^2/(2σ^2). Using the tilted LDP, we then get, ⟨exp(λNmμ)⟩p(mμ)∼exp(−Nλ2σ2/2) (λ Nm^μ) _p(m^μ) (-Nλ^2σ^2/2). Assuming P∼exp(αN)P (α N), we get ϕ=1λ(α+λ2σ22) φ= 1λ (α+ λ^2σ^22 ) (24) The extrema of ϕφ occurs at λ∗=2α/σλ^*= 2α/σ. For λ<λ∗λ<λ^*, ϕφ increases rapidly with increasing λ and is a slow increasing function for λ>λ∗λ>λ^*. The fLSE=−m1f_LSE=-m^1 for ϕ<m1φ<m^1 and fLSE=−ϕf_LSE=-φ when ϕ>m1φ>m^1. For ϕ>m1φ>m^1 there is no retrieval. So let us consider the ϕ<m1φ<m^1 scenario. For binary variables the norm is 11. But to consider the more general case, we take ‖2=N(m1)2|| σ||^2=N(m^1)^2 (assuming that m1m^1 depends on σ through its rescaled norm), the f is fϕ<m1=12m12−m1 f_φ<m^1= 12m^1^2-m^1 (25) This gives ℱϕ<m1() _φ<m^1( m) =12m12−1βξ[logcosh(β(m1−1)ξ1)] = 12m^1^2- 1βE_ξ [ (β(m^1-1)ξ^1) ] (26) For β→∞β→∞, we get Rϕ<m1(m1)=12(m1)2−m1+1 R_φ<m^1(m^1)= 12(m^1)^2-m^1+1 (27) with a fixed point at m∗1=1m^1_*=1. This is the free energy landscape of the retrieval phase of LSE. As shown in Fig. 1, the landscape has tilted towards the m=1m=1 state and the retrieval of the memory is error free. Equating ϕ=1φ=1 then gives us the threshold αc(λ) _c(λ) below which the pattern is completely retrieved with no error. We fix σ=1σ=1. Then αc(λ)≤1/2 _c(λ)≤ 1/2 for real λ. Since λ=1λ=1 at αc=1/2 _c=1/2 αc=λ(1−λ2) _c=λ (1- λ2 ) (28) for λ≤λ∗=1λ≤λ^*=1. We have hence reproduced the threshold for retrieval derived recently using the random energy modellucibello . The treatment of LSE here is done assuming exponential number of stored patterns, but the method can be applied to any number of patterns. Unlike polynomial interaction DenseAMs, for LSE the αc _c gives the threshold of retrieval as m∗=1m_*=1 for α<αcα< _c. Discussion: Our method can evaluate any cost function defined by Eq. (1). Although currently demonstrated using binary variables, the framework generalizes effortlessly to other discrete or continuous variables. We validated our approach by analyzing two prominent AM models. We showed that the disorder-averaged ground state energy, R(m)R(m), is useful for decoding gradient dynamics. In DenseAMs, we showed a strong initial-state dependence during memory retrieval, as the basin at m=0m=0 persists for all α. One would require a mechanism like a stochastic noise rooke to come out of the basin of m=0m=0 state. References (1) D. Krotov and J. J. Hopfield, Dense Associative Memory for Pattern Recognition Advances in Neural Information Processing Systems, vol. 29, (2016). (2) E. Gardner, Multiconnected neural network models, J. Phys. A: Math. Gen. 20 3453 (1987). (3) J. J. Hopfield, Neural networks and physical systems with emergent collective computational abilities, Proceedings of the National Academy of Sciences 79, 2554–2558 (1982). (4) J. R. Anderson and G. H. Bower, Human Associative Memory (Psychology Press, 1974). (5) D. Krotov, B. Hoover, P. Ram and B. Pham, Modern methods in Associative memory, arXiv:2507.06211 (2025). (6) J. Simon, D. Kunin, A. Atanasov, E. Boix-Adserà, B. Bordelon, J. Cohen, N. Ghosh, F. Guth, A. Jacot, M. Kamb, D. Karkada, E. J. Michaud, B. Ottlik, J. Turnbull, There will be a scientific theory of Deep Learning arXiv:2604.21691 (7) B. Hoover, Y. Liang, B. Pham, R. Panda, H. Strobelt, D. H. Chau, M. Zaki, and D. Krotov. Energy transformer, Advances in Neural Information Processing Systems, 36, (2024). (8) M. Yampolskaya and P. Mehta, Hopfield Networks as Models of Emergent Function in Biology, Annual Review of Biophysics 55 323-342(2026). (9) Donald O. Hebb, The organization of behavior, new york: Wiley,The first stage of perception: growth of the assembly, In Neurocomputing, Volume 1: Foundations of Research. The MIT Press, (1949). (10) D. J. Amit, H. Gutfreund, and H. Sompolinsky, Physical Review A 32, 1007–1018 (1985). (11) D. J. Amit, H. Gutfreund, and H. Sompolinsky, Physical Review Letters 55, 1530–1533 (1985). (12) C. Lucibello and M. M´ezard, Exponential Capacity of Dense Associative Memories, Physical Review Letters 132 077301 (2024). (13) M. Demircigil, J. Heusel, M. Löwe, S. Upgang and F. Vermet, On a Model of Associative Memory with Huge Storage Capacity, Journal of Statistical Physics 168 288(2017). (14) H. Ramsauer, B. Sch¨afl, J. Lehner, P. Seidl, M. Widrich, T. Adler, L. Gruber, M. Holzleitner, M. Pavlovi´c, G. K. Sandve, V. Greiff, D. Kreil, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter, Hopfield networks is all you need, International Conference on Learning Representations (2021); arXiv:2008.02217. (15) Sumedha, and S. K. Singh, Effect of random field disorder on the first order transition in p-spin interaction model Physica A: Statistical Mechanics and its Applications 442, 276–283 (2016). (16) S. Mukherjee and Sumedha, Phase Transitions in the Blume–Capel Model with Trimodal and Gaussian Random Fields, Journal of Statistical Physics, 188 22(2022). (17) Sumedha, and M. Barma Phase transitions in XYXY models with randomly oriented crystal fields, Phys. Rev. E, 105 105, 024111(2022). (18) Frank den Hollander, Large Deviations, Fields Institute Monographs, AMS (2000) Theorem I.17. (19) Sumedha, and Nabin K Jana, Absence of first order transition in the random crystal field Blume–Capel model on a fully connected graph (see Appendix), J. Phys. A: Math. Theor. 50 015003 (2017). (20) S. Rooke, D. Krotov, V. Balasubramanian, and D. Wolpert, Stochastic thermodynamics of associative memory, arXiv:2601.01253 (2026). (21) Gradient dynamics is same as the zero temperature Glauber dynamics. For zero temperature Glauber dynamics we have recently shown this for a different model, namely the random field Blume Capel model in Sumedha and Aldrin B E, Glauber dynamics phase transitions in athermal random field Blume-Capel and Blume-Emery-Grifitths models, arXiv:2607:16561.