Paper deep dive
Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout
Wenjie Xi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 7/5/2026, 1:57:52 AM
Summary
The paper introduces an 'action-operator' framework for molecular diffusion models to provide a mathematically consistent way to interpret their learned representations and extract thermodynamic quantities. By defining a base action (S0) for a fixed environment and an alchemical perturbation operator (O), the authors demonstrate that diffusion models can be used to audit and estimate free-energy differences (ΔF). The framework utilizes a 'noisy operator bridge' that leverages the fact that normalized forward noising preserves the partition function, allowing for free-energy readout from endpoint ensembles even with negligible phase-space overlap. The method was validated through experiments on alanine dipeptide (ALA2) systems and a challenging C6-H to C6-F ligand-pocket perturbation in the 185L/IND dataset, showing high accuracy compared to MBAR references.
Entities (8)
Relation Signals (6)
Action-Operator Framework → defines → Base Action (S0)
confidence 100% · By defining a fixed molecular environment as a base action S0(x)
Action-Operator Framework → defines → Alchemical Perturbation Operator (O)
confidence 100% · and an alchemical perturbation as an operator O(x)
Noisy Operator Bridge → estimates → Free-Energy Difference (ΔF)
confidence 100% · enables a 'noisy operator bridge' capable of reading out free-energy differences (ΔF)
Diffusion Models → isanalyzedby → Action-Operator Framework
confidence 100% · we introduce a mathematically consistent action-operator framework natively compatible with diffusion models.
Noisy Operator Bridge → validatedon → Alanine Dipeptide (ALA2)
confidence 100% · In controlled experiments on alanine dipeptide systems, we show...
Noisy Operator Bridge → validatedon → 185L/IND
confidence 100% · When applied to a challenging C6-H to C6-F ligand-pocket nonbonded perturbation (185L/IND)...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Diffusion models are increasingly utilized for modeling molecular structures and conformational ensembles, yet the thermodynamic meaning of their learned representations and scores remains elusive. To resolve this ambiguity, we introduce a mathematically consistent action-operator framework natively compatible with diffusion models. By defining a fixed molecular environment as a base action $S_0(x)$ and an alchemical perturbation as an operator $O(x)$, standard diffusion noising induces effective noised actions and operators whose gradients and alchemical derivatives are directly represented by the model's learned fields. This rigorous self-consistency enables a ``noisy operator bridge'' capable of reading out free-energy differences ($\Delta F$) from endpoint ensembles and per-frame evaluations. In controlled experiments on alanine dipeptide systems, we show that incorporating physical inductive biases enables partial recovery of the base action and perturbation operator. When applied to a challenging C6-H to C6-F ligand-pocket nonbonded perturbation (185L/IND) with negligible phase-space overlap, our supervised bridge estimates the alchemical $\Delta F$ within approximately $1\ k_\mathrm{B}T$ of a stable 19-state MBAR reference. Finally, we demonstrate that endpoint coordinates and binary labels alone are sufficient to partially recover the operator shape and a centered free-energy scale without any force or action supervision. This work provides a rigorous path toward transforming generative molecular diffusion models from black-box coordinate samplers into auditable thermodynamic estimators.
Tags
Links
- Source: https://arxiv.org/abs/2606.30687v1
- Canonical: https://arxiv.org/abs/2606.30687v1
Trouble viewing inline? Open PDF directly →
Full Text
37,845 characters extracted from source content.
Expand or collapse full text
Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout Wenjie Xi1 1Department of Physics and HK Institute of Quantum Science & Technology, The University of Hong Kong, Pokfulam Road, Hong Kong, China E-mail: wjxi@connect.hku.hk Abstract Diffusion models are increasingly utilized for modeling molecular structures and conformational ensembles, yet the thermodynamic meaning of their learned representations and scores remains elusive. To resolve this ambiguity, we introduce a mathematically consistent action-operator framework natively compatible with diffusion models. By defining a fixed molecular environment as a base action S0(x)S_0(x) and an alchemical perturbation as an operator O(x)O(x), standard diffusion noising induces effective noised actions and operators whose gradients and alchemical derivatives are directly represented by the model’s learned fields. This rigorous self-consistency enables a “noisy operator bridge” capable of reading out free-energy differences (ΔF F) from endpoint ensembles and per-frame evaluations. In controlled experiments on alanine dipeptide systems, we show that incorporating physical inductive biases enables partial recovery of the base action and perturbation operator. When applied to a challenging C6-H to C6-F ligand-pocket nonbonded perturbation (185L/IND) with negligible phase-space overlap, our supervised bridge estimates the alchemical ΔF F within approximately 1kBT1\ k_BT of a stable 19-state MBAR reference. Finally, we demonstrate that endpoint coordinates and binary labels alone are sufficient to partially recover the operator shape and a centered free-energy scale without any force or action supervision. This work provides a rigorous path toward transforming generative molecular diffusion models from black-box coordinate samplers into auditable thermodynamic estimators. 1 Introduction Diffusion models have emerged as a powerful paradigm for representation and simulation of molecular structures and conformational ensembles [1, 2, 3, 4]. While these models are typically evaluated on the physical plausibility of their generated coordinates, such spatial metrics leave their underlying thermodynamic ensembles underdetermined. Though recent studies have begun to investigate how diffusion scores correspond to physical quantities [5, 6, 7, 8, 9], establishing a mathematically consistent framework that connects local scores directly to auditable thermodynamic ensembles remains highly desirable. To bridge this gap, we propose an action-operator interpretation for molecular diffusion models that is mathematically and structurally consistent with standard diffusion frameworks. The core physical intuition is that a declared molecular environment defines a reduced base action S0(x)S_0(x), while a chemical alteration—such as a ligand condition or an alchemical mutation—defines a perturbation operator O(x)O(x) that shifts the system’s action to Sκ(x)=S0(x)+κO(x)S_κ(x)=S_0(x)+κ O(x) along an auxiliary path coordinate κ. The normalized forward diffusion kernel naturally maps these clean actions to noise-dependent effective actions, such that the model’s learned score fields correspond exactly to the gradients of these noised actions when equipped with appropriate inductive biases. Simultaneously, the alchemical derivatives of these effective actions yield the conditional expectations of the operators. Under this semantics, molecular configurations generated by the model during generalization are rendered thermodynamically auditable, enabling us to assess whether the underlying S0S_0 and O representations are physically sound. Crucially, a key thermodynamic consequence of this formulation is that normalized noising preserves the partition function (Zκ,σ=ZκZ_κ,σ=Z_κ). This invariant allows us to construct a “noisy operator bridge” to read out absolute free-energy differences (ΔF F) through a path integral of the effective operator in noise space. In this work, we carry out four controlled experiments to evaluate this action-operator semantics sequentially, establishing a clear path from representation learning to thermodynamic diagnostics: First, we investigate whether diffusion-style neural networks can accurately capture and represent the base-action S0S_0 of a fixed molecular ensemble (alanine dipeptide), and demonstrate how this representation scales with model capacity and data size. Second, we fix S0S_0 and evaluate whether the condition-dependent perturbation operator O can be recovered from biased ensembles using cross-entropy classification and clean score matching. While complete recovery remains challenging, the recovery accuracy scales positively with larger dataset sizes and model capacities. Third, we apply the fully integrated noisy operator bridge to a challenging C6-HC6-H to C6-FC6-F ligand-pocket nonbonded perturbation in 185L/IND characterized by negligible endpoint phase-space overlap. Given accurate evaluations of the base action and operator, we show that the bridge estimates the alchemical free-energy difference as −24.74±0.10kBT-24.74± 0.10~k_BT, which is within 1.08±0.10kBT1.08± 0.10~k_BT of a stable 19-state MBAR reference (−25.82kBT-25.82~k_BT). Finally, in an endpoint-only gauge audit (Experiment 4A), we demonstrate that combining a frozen base representation with endpoint thermodynamic consistency can identify the nonconstant operator shape and a gauged physical free-energy scale purely from coordinate distributions and binary labels, completely bypassing physical action or force supervision. This approach bypasses the fundamental statistical bottleneck of traditional free-energy perturbation (FEP) and thermodynamic integration (TI) methods, which typically require simulating a large number of intermediate alchemical states to overcome poor phase-space overlap [10, 11, 12, 13, 14, 15, 16]. While advanced transport-based estimators, such as LBAR, Boltzmann Generators, Neural TI, and flow-matching methods, improve endpoint overlap by learning invertible molecular mappings or interpolating intermediate densities [17, 18, 19, 20, 21, 22, 6, 23, 24, 25], our noisy bridge represents a conceptually distinct paradigm. It does not compute change-of-variables Jacobians or generate nonequilibrium work trajectories. We make the methodological boundaries explicit: while the supervised bridge utilizes endpoint ensembles combined with declared per-frame reduced-action evaluations, the endpoint-only audit demonstrates that coordinate distributions alone can identify a relative free-energy gauge without online energy supervision. The key contribution of our Action-Operator semantics is demonstrating how a generative diffusion architecture can be systematically white-boxed to deliver auditable thermodynamic readouts. 2 Fixed Environment as a Base Action S0S_0 2.1 Physical Picture At a fixed temperature and molecular environment, an equilibrium molecular ensemble in coordinate space x residing on a Riemannian manifold ℳM is characterized by the probability measure: p0(x)=Z0−1e−S0(x).p_0(x)=Z_0^-1e^-S_0(x). (1) For an all-atom classical force field with potential energy U0(x)U_0(x), this action corresponds to the dimensionless, reduced potential S0(x)=βU0(x)S_0(x)=β U_0(x). A diffusion model trained on a fixed environment does not merely fit an untyped distribution of coordinates, but rather learns noised information about S0S_0 through a normalized noising kernel Kσ(z∣x)K_σ(z x). This process induces a noised density p0,σ(z)=∫Kσ(z∣x)p0(x)dxp_0,σ(z)= K_σ(z x)p_0(x)\,dx and a corresponding effective noised action S0,σ(z)=−logp0,σ(z)+constS_0,σ(z)=- p_0,σ(z)+const. For constrained or mass-metric coordinates, this effective action defines a natural score vector field: v0,σ(z)=−1∇zlogp0,σ(z)=−1∇zS0,σ(z),v_0,σ(z)= M^-1 _z p_0,σ(z)=- M^-1 _zS_0,σ(z), (2) where M is the symmetric, positive-definite mass matrix (or the Riemannian metric tensor of the constraint manifold), ∇z _z is the Euclidean gradient, and the natural score v0,σ(z)v_0,σ(z) is transformed from the Euclidean score ∇zlogp0,σ(z) _z p_0,σ(z) via the metric inverse −1 M^-1. This formulation establishes the noised action S0,σS_0,σ as a native, integrable diffusion quantity. 2.2 Model architecture For semantic readout, we enforce score integrability by parameterizing a componentized scalar-action network Aθ(z,σ)A_θ(z,σ). This network receives molecular features (atom identities, distances, local internal coordinates, and noise embeddings) and is assembled from topology-conditioned component heads—specifically bond, angle, torsion, 1–4 exception, and direct nonbonded terms in our ALA2 implementation. Rather than employing an independent vector head, the score is obtained via automatic differentiation of this assembled action: vθ(z,σ)=−1∇zAθ(z,σ),v_θ(z,σ)=- M^-1 _zA_θ(z,σ), (3) followed by tangent projection when molecular constraints or mass-metric coordinates are used. This architecture ensures that Aθ(z,σ)A_θ(z,σ) directly represents the effective noised action up to a z-independent additive constant, establishing the necessary mathematical gauge for downstream auditing. For clean samples x∼p0x p_0, we define noised positions on the constraint manifold via tangent noise and retraction: z=Retrx(σξ),ξ∼(0,−1).z=Retr_x(σξ), ξ (0, M^-1). (4) For nonlinear retractions, this procedure induces a normalized pushforward Markov kernel rather than a flat Euclidean Gaussian density. Dropping the retraction Jacobian in local coordinates would necessitate a state-dependent reweighting term to ensure normalization. To bypass this explicit Jacobian computation, we avoid closed-form density evaluation in Experiments 1 and 2, relying solely on sampling and score matching. In the low-noise local-coordinate regime used here, the sampling-defined denoising target and the base-action loss are: tσ(z,x)=Πzx−zσ2,t_σ(z,x)= _z x-zσ^2, (5) ℒS0(θ)=x,σ,ξ[σ2d‖vθ(z,σ)−tσ(z,x)‖2],L_S_0(θ)=E_x,σ,ξ [ σ^2d \|v_θ(z,σ)-t_σ(z,x) \|_ M^2 ], (6) where Πz _z is the tangent projection, d is the effective dimension, and ‖v‖2=v⊤v\|v\|_ M^2=v Mv. 2.3 Audit: Experiment 1 In Experiment 1, we use a low-dimensional ALA2 ensemble to test whether our diffusion architecture can represent nontrivial base-action and operator structures. Trajectories are obtained from a public dataset [26], with OpenMM providing the reference reduced-actions and forces [27]. We first train a configuration-only, low-noise DSM model for S0S_0. As detailed in Table 1, we evaluate the scaling performance across frame counts and pretraining settings, followed by a component-wise audit of the 256-frame checkpoint. These capacity-sensitive trends demonstrate successful (albeit uncalibrated) S0S_0 recovery, even though the individual physical components are not equally well-recovered due to representation degeneracies. Having established the base-action audit, we next turn to ligand or condition perturbations as operator fields. Table 1: Experiment 1A: Base action S0S_0 recovery under scaling gates and component-head audit. Config. / Component Score Action Action NRMSE Pearson Aff. NRMSE S0S_0 scaling gate 16 frames 0.6392 0.3452 0.9359 64 frames 0.5814 0.6600 0.7474 256 frames 0.5802 0.7721 0.6354 16 frames, 4096 pretrain 0.5532 0.4668 0.8833 16 frames, larger model 0.6367 0.4570 0.8866 Component-head audit, 256-frame checkpoint Assembled total – 0.780 0.621 Bond (1–2) – 0.759 0.629 Angle (1–3) – 0.920 0.392 Torsion – 0.612 0.790 Direct – −0.175-0.175 0.983 Exception (1–4) – 0.313 0.907 Direct + exception – −0.053-0.053 0.998 3 Ligand perturbations as operator fields 3.1 Physical picture Attempting to learn the base action S0(x)S_0(x) and the perturbation operator O(x)O(x) simultaneously via joint conditional score matching is ill-conditioned, as the unconstrained base action and perturbation operator readily absorb each other’s representation errors. To resolve this degeneracy, we must declare or fix S0(x)S_0(x) first. Once the base environment is fixed, a ligand perturbation can be represented as a systematic action shift along an alchemical path coordinate κ: Sκ(x)=S0(x)+κO(x).S_κ(x)=S_0(x)+κ O(x). (7) For noised variables, the unnormalized effective action Sκ,σ(z)S_κ,σ(z) satisfies: e−Sκ,σ(z)=∫Kσ(z∣x)e−S0(x)−κO(x)dx.e^-S_κ,σ(z)= K_σ(z x)e^-S_0(x)-κ O(x)\,dx. (8) Differentiating Equation (8) with respect to κ under the integral yields the effective operator field: Oσ(κ,z)=∂κSκ,σ(z)=κ,σ[O(x)∣z].O_σ(κ,z)= _κS_κ,σ(z)=E_κ,σ[O(x) z]. (9) Assuming a κ-independent noising kernel, Equation (9) demonstrates that the effective operator field is diffusion-native, corresponding directly to the alchemical derivative of the noised effective action. 3.2 Cross-entropy operator selection Condition labels provide a density-ratio route to scalar operator learning. For samples from a set of parameterized ensembles pκj(x)∝e−S0(x)−κjO(x)p_ _j(x) e^-S_0(x)- _jO(x), a Bayesian classifier yields logits of the form: ℓj(x)=bj−κjOθ(x), _j(x)=b_j- _jO_θ(x), (10) where bjb_j absorbs the class priors and free-energy constants. The cross-entropy loss is formulated as: ℒCE(θ,b)=−(x,y)logexp[by−κyOθ(x)]∑jexp[bj−κjOθ(x)].L_CE(θ,b)=-E_(x,y) [b_y- _yO_θ(x)] _j [b_j- _jO_θ(x)]. (11) For a binary endpoint pair (κ∈0,1κ∈\0,1\), this recovers a logistic density-ratio model: logp1(x)p0(x)=−O(x)+ΔF. p_1(x)p_0(x)=-O(x)+ F. (12) Because cross-entropy is a classification objective providing only scalar-level density-ratio supervision, the resulting gradients of OθO_θ do not reliably represent fine-grained microscopic forces. We therefore use cross-entropy primarily to locate a restricted operator basin, leaving the resolution of local forces to subsequent score matching. 3.3 Clean score matching for operator fields With S0S_0 fixed, the score of the perturbed action is: vκ(x)=−1∇Sκ(x)=v0(x)−κ−1∇O(x).v_κ(x)=- M^-1∇ S_κ(x)=v_0(x)-κ M^-1∇ O(x). (13) Defining the operator score via automatic differentiation of the parameterized scalar operator, gθ(x)=−1∇Oθ(x)g_θ(x)=- M^-1∇ O_θ(x), we train OθO_θ on clean samples from pκp_κ. Dropping θ-independent terms, the Hyvarinen score-matching objective for v0+κgθv_0+κ g_θ is: ℒcleanSM(θ)=pκ[κ⟨gθ,v0⟩+12κ2‖gθ‖2+κdivgθ],L_cleanSM(θ)=E_p_κ [κ g_θ,v_0 _ M+ 12κ^2\|g_θ\|_ M^2+κ\,div_ Mg_θ ], (14) where the divergence is estimated via sliced finite differences in tangent directions. Crucially, fixing the base action S0S_0 prevents mutual error absorption between S0S_0 and O, rendering the score-matching objective in Equation (14) mathematically well-posed. To maintain global thermodynamic consistency, this local score-matching refinement is initialized or regularized within the restricted operator basin previously located by cross-entropy selection, preventing unconstrained global drift during force alignment. 3.4 Audit: Experiment 2 In Experiment 2, we evaluate whether the operator field can be recovered and audited once S0S_0 is fixed. Experiment 2A uses a theoretical OpenMM S0S_0 and applies cross-entropy to locate the scalar operator basin, followed by clean on-manifold sliced score matching to refine the local operator score. As shown in Table 2, CE already recovers a strong held-out scalar operator signal, while CE+CSM preserves this scalar recovery and improves both operator-score and finite-difference gradient audits. Shuffled-κ CSM does not reproduce the force-level gain, and shuffled-CE collapses both scalar and local-gradient recovery. These results indicate that once S0S_0 is anchored, operator-directed supervision can recover an auditable scalar O together with partial tangent-gradient information. Table 2: Experiment 2A: fixed-S0S_0 operator recovery with cross-entropy (CE) and clean score matching (CSM). Values are mean ± standard deviation over five seeds in the high-budget common-random audit. Audit metric CE CE+CSM CE+shuf.-CSM shuf.-CE+CSM Scalar O Pearson 0.855±0.0310.855± 0.031 0.830±0.0450.830± 0.045 0.872±0.0430.872± 0.043 0.219±0.1180.219± 0.118 Operator-score cosine 0.360±0.0790.360± 0.079 0.487±0.0440.487± 0.044 0.358±0.0820.358± 0.082 0.153±0.0610.153± 0.061 Operator-score NRMSE 1.015±0.0641.015± 0.064 0.905±0.0250.905± 0.025 0.984±0.0660.984± 0.066 1.051±0.0531.051± 0.053 Finite-difference Pearson 0.405±0.0940.405± 0.094 0.531±0.0530.531± 0.053 0.409±0.0980.409± 0.098 0.129±0.0970.129± 0.097 Finite-difference NRMSE 0.909±0.0420.909± 0.042 0.845±0.0320.845± 0.032 0.906±0.0480.906± 0.048 0.987±0.0150.987± 0.015 In this ALA2 audit, the operator O(x)O(x) is a smooth periodic function of selected backbone dihedral coordinates. The operator models are trained on biased ALA2 ensembles at κ∈−2,−1,0,1,2κ∈\-2,-1,0,1,2\, using 320 configurations per training κ. Evaluation is performed on held-out biased ensembles at κ∈−1.5,−0.75,0.75,1.5κ∈\-1.5,-0.75,0.75,1.5\. 4 A noisy reduced-potential bridge for free-energy readout 4.1 Why a noisy bridge can preserve free energy The preceding sections define the requisite semantics of S0S_0 and O. In this section, we demonstrate how to extract free energy from a diffusion model whether S is declared or learned. For the protein–ligand endpoint studies below, we denote the two endpoint alchemical states by q=0q=0 and q=1q=1; this q is the binary endpoint specialization of the path coordinate κ. Free energy requires the partition function: Fκ=−logZκ,Zκ=∫e−Sκ(x)dx.F_κ=- Z_κ, Z_κ= e^-S_κ(x)\,dx. (15) Because learned actions are identified only up to arbitrary additive constants, we must manually fix an action gauge to resolve the partition-function ambiguity. Even when the operator is known, endpoint coordinate ensembles often suffer from negligible phase-space overlap. The key thermodynamic invariant of our formulation is that normalized forward noising preserves the partition function ZκZ_κ: Zκ,σ=∫e−Sκ,σ(z)dν(z)=∬Kσ(z∣x)e−Sκ(x)dxdν(z)=∫e−Sκ(x)dx=Zκ,Z_κ,σ= e^-S_κ,σ(z)\,dν(z)= K_σ(z x)e^-S_κ(x)\,dx\,\,dν(z)= e^-S_κ(x)\,dx=Z_κ, (16) which holds under the condition that the noising kernel is κ-independent and normalized over the noisy space, such that ∫Kσ(z∣x)dν(z)=1 K_σ(z x)\,dν(z)=1 for every clean configuration x without state-dependent truncation or reweighting. When the molecular coordinate space is reduced to a lower-dimensional feature space, we define the forward noising kernel KσK_σ and the reference measure directly on this reduced feature space. By using the empirical endpoint ensembles to define the pushforward measure, we completely bypass the need for a Cartesian-to-feature change-of-variables Jacobian. This feature-space formulation ensures that the partition function invariance Zκ,σ=ZκZ_κ,σ=Z_κ holds exactly, provided a common normalized pushforward kernel is applied consistently to both endpoints. Consequently, we can write a density-ratio identity in noisy space that preserves the clean free energy: logp1,σ(z)p0,σ(z)=−[S1,σ(z)−S0,σ(z)]+ΔF. p_1,σ(z)p_0,σ(z)=- [S_1,σ(z)-S_0,σ(z) ]+ F. (17) Representing the action difference as a path integral of the effective operator: dσ(z)=∫01Oσ(κ,z)dκ=S1,σ(z)−S0,σ(z),d^σ(z)= _0^1O_σ(κ,z)\,dκ=S_1,σ(z)-S_0,σ(z), (18) we obtain the fundamental noisy density-ratio identity: logp1,σ(z)p0,σ(z)=−dσ(z)+ΔF. p_1,σ(z)p_0,σ(z)=-d^σ(z)+ F. (19) 4.2 Shared-operator bridge objective To compute the quadrature in Equation (19), we train a sigma-conditioned operator network Oθ(κ,z,σ)≈Oσ(κ,z)O_θ(κ,z,σ)≈ O_σ(κ,z) alongside a shared scalar ΔFθ F_θ. The predicted action difference is computed via quadrature: dθσ(z)=∫01Oθ(κ,z,σ)dκ.d^σ_θ(z)= _0^1O_θ(κ,z,σ)\,dκ. (20) The unified training objective minimizes four terms: ℒbridge=wOℒO+wAℒA+wRℒR+wSℒS.L_bridge=w_OL_O+w_AL_A+w_RL_R+w_SL_S. (21) The first term, ℒOL_O, regresses the endpoint operator evaluations when physical targets O^σ O_σ are available (as in supervised settings): ℒO=[(Oθ(κ,z,σ)−O^σ)2].L_O=E [ (O_θ(κ,z,σ)- O_σ )^2 ]. (22) To fix the additive gauge of dθσd^σ_θ at the zero-noise limit (σ→0+σ→ 0^+), we introduce the clean action-integral anchor ℒAL_A: ℒA=x[(∫01Oθ(κ,x,0+)dκ−ΔS(x))2],ΔS(x)=S1(x)−S0(x),L_A=E_x [ ( _0^1O_θ(κ,x,0^+)\,dκ- S(x) )^2 ], S(x)=S_1(x)-S_0(x), (23) which evaluates the direct-path action difference ΔS(x) S(x) on clean endpoint features with a near-zero sigma label for numerical stability. In unsupervised coordinate-only learning where analytical potentials are absent, this shift degeneracy is alternatively resolved by fixing an explicit action convention (such as centering OθO_θ on the source training pool). Without such an anchor or centering convention, dθσd^σ_θ and ΔFθ F_θ retain a shift degeneracy even if the endpoint operator and density-ratio consistency are otherwise accurate. The third term, ℒRL_R, enforces density-ratio consistency using the noisy endpoint identity: ℒR=z,y[BCE(ΔFθ−dθσ(z),y)],L_R=E_z,y [BCE ( F_θ-d^σ_θ(z),y ) ], (24) where y=0y=0 for noisy source endpoint samples, y=1y=1 for noisy target endpoint samples, and training uses balanced batches to prevent the logistic intercept from absorbing unbalanced class priors. Finally, ℒSL_S is a mild smoothness regularizer in κ. The scalar ΔFθ F_θ is learned directly by the bridge without relying on intermediate-κ ensembles, nonequilibrium trajectories, or accept/reject steps. 4.3 Controlled study: Experiment 3 Experiment 3 transfers our action-operator semantics to a controlled 185L/IND protein–ligand bound-state system based on PDB entry 185L, utilizing OpenMM and OpenFF-family small-molecule parameterizations [28, 29, 27, 30, 31]. We first introduce a physical C6-H to C6-F ligand-atom mutation. Importantly, this operator is not a full force-field replacement or a complete ligand mutation free energy. It is a restricted single-topology ligand-pocket nonbonded perturbation: O(x)O(x) is defined as the change in nonbonded interaction energy between the IND ligand atoms and a fixed set of pocket heavy atoms. Due to negligible phase-space overlap between the endpoints, direct endpoint FEP fails, yielding a bidirectional forward-reverse discrepancy (hysteresis) of 28 to 33 kBTk_BT. An independent 19-state split-path MBAR calculation provides a stable, converged validation reference of −25.8219kBT-25.8219\ k_BT for subsequent comparison. We then implement the main noisy operator bridge, utilizing only the endpoint states (q=0,1q=0,1) along with per-frame evaluations of S0(x)S_0(x) and the direct-path action difference ΔS(x) S(x) to supervise the clean action-integral anchor ℒAL_A. In this setup, ΔS S is withheld from the input features and acts solely as the zero-noise target to fix the additive gauge. Table 3: Experiment 3: Free-energy estimates and high-noise diagnostics for the 185L/IND C6-H to C6-F mutation. All values are in units of kBTk_BT. Errors are absolute errors relative to the 19-state split-path MBAR reference. Method / Configuration Estimate Abs. Error 19-state split-path MBAR reference −25.8219-25.8219 – Noisy operator bridge, 6 seeds −24.7405±0.0952-24.7405± 0.0952 1.0814±0.09521.0814± 0.0952 high-noise BAR layer, σ=0.03σ=0.03 −24.7041±0.1685-24.7041± 0.1685 1.1178±0.16851.1178± 0.1685 high-noise BAR layer, σ=0.04σ=0.04 −24.6877±0.1856-24.6877± 0.1856 1.1341±0.18561.1341± 0.1856 As summarized in Table 3, the main bridge yields a shared free-energy estimate ΔFθ=−24.7405±0.0952kBT F_θ=-24.7405± 0.0952\ k_BT, matching the stable 19-state reference within 1.0814±0.0952kBT1.0814± 0.0952\ k_BT. 4.4 Endpoint-only gauge audit: Experiment 4 Experiment 4 addresses a more difficult question: can we identify free energies from endpoint coordinates and labels alone? To address this, we apply our endpoint-only audit to the 185L/IND protein–ligand bound state, where the base action S0S_0 corresponds to the unmutated (q=0q=0) pocket environment, and the operator O represents either the C6-H3 to F or the IND C4-H7 to F ligand-atom mutation. The protocol first trains and freezes a q=0q=0 base representation S0,θS_0,θ via feature-space denoising score matching. It then learns a scalar operator OθO_θ purely from balanced q=0/1q=0/1 endpoint ensembles. Because the learned operator has an arbitrary additive gauge, we fix its constant offset by centering the predictions on the q=0q=0 training pool: Oθ(x)=O~θ(x)−x∈train,q=0[O~θ(x)].O_θ(x)= O_θ(x)-E_x ,q=0[ O_θ(x)]. (25) The corresponding audit target is thus the q=0q=0-centered physical scale, verifying whether coordinate distributions alone can recover the nonconstant operator shape. We evaluate this unsupervised audit on both the original C6-H3 to F mutation and a second single-topology mutation (IND C4-H7 to C4-F7) to rule out single-site bias. As summarized in Table 4, the main model successfully recovers the q=0q=0-centered operator scale for both sites, achieving endpoint-pool audit errors of 0.8149±0.3496kBT0.8149± 0.3496\ k_BT for C6 and 0.7324±0.0747kBT0.7324± 0.0747\ k_BT for C4. In contrast, all control variants fail: shuffling labels collapses the BAR estimates to near-zero, while removing the q=0q=0 centering or the endpoint thermodynamic consistency constraints severely inflates the audit errors. To establish a rigorous thermodynamic baseline, we generate an independent 19-state split-path MBAR reference for the C4 mutation, yielding a converged absolute free energy of −19.0432kBT-19.0432\ k_BT. Translated to the q=0q=0-centered gauge, this corresponds to a reference value of −3.5935±0.1651kBT-3.5935± 0.1651\ k_BT. Against this independent split-MBAR-centered reference, our coordinate-only estimator predicts the scale with an error of 1.3676±0.0747kBT1.3676± 0.0747\ k_BT. Table 4: Experiment 4: endpoint-only gauge audit for C6 and C4 mutations. Values are in units of kBTk_BT and reported as mean ± standard deviation over five seeds. Errors are evaluated against the endpoint-pool q=0q=0-centered audit reference. Method / Control C6-H3 → F C4-H7 → F BAR Err. BAR Err. CE + TC −2.2419±0.1125-2.2419± 0.1125 0.8149±0.34960.8149± 0.3496 −2.2259±0.1513-2.2259± 0.1513 0.7324±0.07470.7324± 0.0747 Shuffled labels −0.0002±0.0013-0.0002± 0.0013 3.0565±0.44163.0565± 0.4416 0.0011±0.00440.0011± 0.0044 1.4946±0.16521.4946± 0.1652 No q=0q=0 centering −0.1182±0.5476-0.1182± 0.5476 2.9386±0.38902.9386± 0.3890 0.0620±0.43130.0620± 0.4313 1.5555±0.32281.5555± 0.3228 No endpoint TC −8.1991±3.4096-8.1991± 3.4096 5.1423±3.19225.1423± 3.1922 −5.8005±2.4880-5.8005± 2.4880 4.3070±2.33134.3070± 2.3313 These unsupervised results strengthen the semantic narrative of this work. They confirm that endpoint coordinates can support a useful gauge-fixed free-energy readout when thermodynamic consistency is enforced. 5 Discussion and conclusion This work establishes a physically rigorous action-operator semantics for molecular diffusion models, enabling the direct readout and audit of internalized thermodynamic quantities such as base actions (S0S_0), perturbation operators (O), and free energy differences (ΔF F). This framework successfully extracts absolute free energy differences in a supervised protein–ligand mutation (Experiment 3) and uncovers relative free energy scales purely from endpoint coordinates and labels in an unsupervised audit (Experiment 4). While this work provides a foundational language to interpret and audit diffusion models as thermodynamic estimators, significant challenges remain for real-world pharmaceutical applications. First, fully robust free energy calculation in drug discovery often requires evaluating relative free energy differences (ΔΔF F) between multiple ligands. This necessitates solving the more complex problem of a universally fixed thermodynamic gauge across diverse chemical transformations, moving beyond our current system-specific or endpoint-centered conventions. Second, while our controlled studies confirm the identifiability of S0S_0, O, and F within a specific framework, it remains crucial to validate whether these physically interpretable quantities can be accurately and robustly extracted in truly large-scale generative models trained on vast and diverse biomolecular datasets. The ultimate goal is to evolve current black-box diffusion models into auditable thermodynamic machines capable of predicting drug-like properties. Achieving this requires scaling the joint recovery of S0S_0, condition-dependent O, and Oσ(κ,z)O_σ(κ,z) across numerous protein–ligand environments, guided by the robust semantic framework laid out in this work. This represents a rigorous path toward transforming generative molecular models into powerful, interpretable tools for chemical and biological discovery. Acknowledgments This work was supported by Research Grants Council of Hong Kong (GRF 17311322 and CRF C7012-21GF) and National Natural Science Foundation of China (Grant No. 12222416). References [1] Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N. & Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of ICML (2015). URL https://proceedings.mlr.press/v37/sohl-dickstein15.html. [2] Ho, J., Jain, A. & Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (2020). URL https://papers.nips.c/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html. [3] Song, Y. et al. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (2021). URL https://openreview.net/forum?id=PxTIG12RRHS. [4] Arts, M. et al. Two for one: Diffusion models and force fields for coarse-grained molecular dynamics. Journal of Chemical Theory and Computation 19, 6151–6159 (2023). URL https://doi.org/10.1021/acs.jctc.3c00702. [5] Plainer, M., Wu, H., Klein, L., Günnemann, S. & Noé, F. Consistent sampling and simulation: Molecular dynamics with energy-based diffusion models. In Advances in Neural Information Processing Systems (2025). URL https://arxiv.org/abs/2506.17139. [6] Mate, B., Fleuret, F. & Bereau, T. Neural thermodynamic integration: Free energies from energy-based diffusion models. Journal of Physical Chemistry Letters 15, 11395–11404 (2024). URL https://doi.org/10.1021/acs.jpclett.4c01958. [7] Sarma, S. et al. Can we extract physics-like energies from generative protein diffusion models? bioRxiv (2025). URL https://doi.org/10.1101/2025.11.28.690021. [8] Xie, Y. et al. Enhanced diffusion sampling: Efficient rare event sampling and free energy calculation with diffusion models. arXiv preprint (2026). URL https://arxiv.org/abs/2602.16634. [9] Xi, W. & Chen, W.-Q. Autonomous emergence of hamiltonian in deep generative models. arXiv preprint (2026). URL https://arxiv.org/abs/2604.20821. [10] Zwanzig, R. W. High-temperature equation of state by a perturbation method. i. nonpolar gases. Journal of Chemical Physics 22, 1420–1426 (1954). URL https://doi.org/10.1063/1.1740409. [11] Kirkwood, J. G. Statistical mechanics of fluid mixtures. Journal of Chemical Physics 3, 300–313 (1935). URL https://doi.org/10.1063/1.1749657. [12] Bennett, C. H. Efficient estimation of free energy differences from monte carlo data. Journal of Computational Physics 22, 245–268 (1976). URL https://doi.org/10.1016/0021-9991(76)90078-4. [13] Shirts, M. R. & Chodera, J. D. Statistically optimal analysis of samples from multiple equilibrium states. Journal of Chemical Physics 129, 124105 (2008). URL https://doi.org/10.1063/1.2978177. [14] Meng, X.-L. & Wong, W. H. Simulating ratios of normalizing constants via a simple identity: A theoretical exploration. Statistica Sinica 6, 831–860 (1996). URL https://w3.stat.sinica.edu.tw/statistica/j6n4/j6n43/j6n43.htm. [15] Jarzynski, C. Nonequilibrium equality for free energy differences. Physical Review Letters 78, 2690–2693 (1997). URL https://doi.org/10.1103/PhysRevLett.78.2690. [16] Crooks, G. E. Entropy production fluctuation theorem and the nonequilibrium work relation for free energy differences. Physical Review E 60, 2721–2726 (1999). URL https://doi.org/10.1103/PhysRevE.60.2721. [17] Jarzynski, C. Targeted free energy perturbation. Physical Review E 65, 046122 (2002). URL https://doi.org/10.1103/PhysRevE.65.046122. [18] Wirnsberger, P. et al. Targeted free energy estimation via learned mappings. Journal of Chemical Physics 153, 144112 (2020). URL https://doi.org/10.1063/5.0018903. [19] Yoo, S., Kang, L. & Minh, D. D. L. Learned mappings for targeted free energy perturbation between peptide conformations. Journal of Chemical Physics 159, 124104 (2023). URL https://doi.org/10.1063/5.0164662. [20] Noe, F., Olsson, S., Kohler, J. & Wu, H. Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning. Science 365, eaaw1147 (2019). URL https://doi.org/10.1126/science.aaw1147. [21] Ding, X. & Zhang, B. DeepBAR: A fast and exact method for binding free energy computation. Journal of Physical Chemistry Letters 12, 2509–2515 (2021). URL https://doi.org/10.1021/acs.jpclett.1c00189. [22] Mate, B. & Fleuret, F. Learning interpolations between boltzmann densities. Transactions on Machine Learning Research (2023). URL https://openreview.net/forum?id=TH6YrEcbth. [23] Mate, B., Fleuret, F. & Bereau, T. Solvation free energies from neural thermodynamic integration. Journal of Chemical Physics 162, 124107 (2025). URL https://doi.org/10.1063/5.0251736. [24] Erdogan, E. et al. FreeFlow: Latent flow matching for free energy difference estimation. OpenReview preprint (2025). URL https://openreview.net/forum?id=D2EdWRWEQo. [25] Du, Y. et al. FEAT: Free energy estimators with adaptive transport. In Advances in Neural Information Processing Systems (2025). URL https://openreview.net/forum?id=GQXeLGYMda. [26] McGibbon, R. T. MD trajectories of ALA2 (fileset). figshare Dataset (2014). URL https://doi.org/10.6084/m9.figshare.1026131.v8. [27] Eastman, P. et al. OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLoS Computational Biology 13, e1005659 (2017). URL https://doi.org/10.1371/journal.pcbi.1005659. [28] RCSB Protein Data Bank. PDB ID 185L: Specificity of ligand binding in a buried non-polar cavity of T4 lysozyme: linkage of dynamics and structural plasticity (1995). URL https://doi.org/10.2210/pdb185L/pdb. [29] Morton, A. & Matthews, B. W. Specificity of ligand binding in a buried nonpolar cavity of T4 lysozyme: linkage of dynamics and structural plasticity. Biochemistry 34, 8576–8588 (1995). URL https://doi.org/10.1021/bi00027a007. [30] Mobley, D. L. et al. Escaping atom types in force fields using direct chemical perception. Journal of Chemical Theory and Computation 14, 6076–6092 (2018). URL https://doi.org/10.1021/acs.jctc.8b00640. [31] Open Force Field Initiative. OpenFF force field file openff-2.2.1.offxml (2024). URL https://github.com/openforcefield/openff-forcefields/blob/main/openforcefields/offxml/openff-2.2.1.offxml.