Paper deep dive
A Hyperfinite Framework for Score-Based Generative Modeling
Sunder Ram Krishnan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/5/2026, 4:30:14 AM
Summary
The paper develops a hyperfinite formulation of score-based generative modeling using Nonstandard Analysis. It derives the infinitesimal generator and Fokker-Planck equation on a hyperfinite grid, establishes a backward-mean identity for reverse-time drift, and proves that minimizing an internal score-matching objective recovers the required score function. The work also derives a hyperfinite Girsanov formula linking likelihood optimization to Fisher-divergence objectives and analyzes second-order consistency, showing dependence on the fourth moment of increment distributions.
Entities (10)
Relation Signals (8)
Nonstandard Analysis → providesframeworkfor → Hyperfinite Grid
confidence 95% · develop a hyperfinite formulation of score-based generative modeling within the framework of Nonstandard Analysis
Score Matching → recovers → Score Function
confidence 94% · minimization of an internal score-matching objective recovers the score function required by the reverse-time dynamics
Hyperfinite Grid → supports → Internal Diffusion Process
confidence 93% · Starting from an internal diffusion process on a hyperfinite grid
Internal Diffusion Process → correspondsto → Fokker-Planck equation
confidence 92% · derive the associated infinitesimal generator and establish its correspondence with the classical Fokker--Planck equation
Second-order consistency → dependson → Fourth moment of increment distribution
confidence 91% · analyze the second-order consistency of the hyperfinite dynamics and show that the leading correction term depends explicitly on the fourth moment of the increment distribution
Backward-mean identity → yields → Reverse-time SDE
confidence 90% · obtain a hyperfinite backward-mean identity that yields the reverse-time drift and provides a constructive derivation of the reverse-time SDE
Hyperfinite Girsanov Formula → establishesrelationshipbetween → Likelihood Optimization
confidence 88% · derive a hyperfinite Girsanov formula and establish a relationship between likelihood optimization and Fisher-divergence objectives
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Score-based diffusion models are typically formulated using continuous-time stochastic differential equations and measure-theoretic stochastic calculus. In this paper, we develop a hyperfinite formulation of score-based generative modeling within the framework of Nonstandard Analysis. Starting from an internal diffusion process on a hyperfinite grid, we derive the associated infinitesimal generator and establish its correspondence with the classical Fokker--Planck equation. We then obtain a hyperfinite backward-mean identity that yields the reverse-time drift and provides a constructive derivation of the reverse-time SDE. Building on these results, we show that minimization of an internal score-matching objective recovers the score function required by the reverse-time dynamics, thereby connecting score estimation with generative sampling directly at the hyperfinite level. Under suitable assumptions, we further derive a hyperfinite Girsanov formula and establish a relationship between likelihood optimization and Fisher-divergence objectives. Finally, we analyze the second-order consistency of the hyperfinite dynamics and show that the leading correction term depends explicitly on the fourth moment of the increment distribution, with the Gaussian value $\kappa=3$ eliminating the leading dispersion contribution. Taken together, these results provide a unified hyperfinite framework for diffusion-based generative modeling--while laying foundations for further extensions--that links discrete grid dynamics, reverse-time diffusion, score matching, and likelihood-based formulations within a common nonstandard setting.
Tags
Links
- Source: https://arxiv.org/abs/2608.02799v1
- Canonical: https://arxiv.org/abs/2608.02799v1
Trouble viewing inline? Open PDF directly →
Full Text
187,159 characters extracted from source content.
Expand or collapse full text
A Hyperfinite Framework for Score-Based Generative Modeling Ram Krishnan CONTACT S. R. Krishnan. Email: eeksunderram@gmail.com Abstract Score-based diffusion models are typically formulated using continuous-time stochastic differential equations and measure-theoretic stochastic calculus. In this paper, we develop a hyperfinite formulation of score-based generative modeling within the framework of Nonstandard Analysis. Starting from an internal diffusion process on a hyperfinite grid, we derive the associated infinitesimal generator and establish its correspondence with the classical Fokker–Planck equation. We then obtain a hyperfinite backward-mean identity that yields the reverse-time drift and provides a constructive derivation of the reverse-time SDE. Building on these results, we show that minimization of an internal score-matching objective recovers the score function required by the reverse-time dynamics, thereby connecting score estimation with generative sampling directly at the hyperfinite level. Under suitable assumptions, we further derive a hyperfinite Girsanov formula and establish a relationship between likelihood optimization and Fisher-divergence objectives. Finally, we analyze the second-order consistency of the hyperfinite dynamics and show that the leading correction term depends explicitly on the fourth moment of the increment distribution, with the Gaussian value κ=3κ=3 eliminating the leading dispersion contribution. Taken together, these results provide a unified hyperfinite framework for diffusion-based generative modeling–while laying foundations for further extensions–that links discrete grid dynamics, reverse-time diffusion, score matching, and likelihood-based formulations within a common nonstandard setting. keywords: Score-based diffusion models; Nonstandard analysis; Hyperfinite stochastic processes; Reverse-time stochastic differential equations; Fokker–Planck equation; Score matching; Girsanov theorem; Generative modeling 1 Introduction The mathematical architecture of modern stochastic calculus is traditionally built upon the sophisticated machinery of measure theory and the Itô-isometry. Since the seminal work of [26], Nonstandard Analysis (NSA) has offered a powerful alternative to this continuum-based approach by providing a rigorous logical framework for the use of infinitesimals. While the ‘standard’ viewpoint provides rigorous foundations, its inherent complexity often obscures the underlying pathwise intuition, a limitation that can be bridged by leveraging the infinitesimal techniques of NSA. In the realm of probability, this paradigm reached a state of maturity through the development of the Loeb measure [21] and the “radically elementary” approach pioneered by [23]. The central object of interest in the internal framework of NSA is the hyperfinite random walk, a discrete-time process on a lattice with a positive infinitesimal time step. Unlike standard discrete approximations, hyperfinite walks live on an internal lattice with infinitesimal time increments; under suitable moment and regularity assumptions, their standard parts recover the corresponding continuous-time diffusions through transfer and Loeb-space limit constructions [3, 7, 16]. The development of this perspective spans the 1970s and 1980s. Anderson [3] gave a nonstandard representation of Brownian motion, Lindstrøm [18, 19] developed a theory of hyperfinite stochastic integration and martingale representation, and Keisler [16] showed how stochastic differential equations (SDEs) can be formulated and solved on hyperfinite grids. Cutland [7, 8] subsequently established foundational existence results for SDEs and Loeb-space probability theory that underpin these constructions. While these foundational works established the existence and regularity of solutions to SDEs, the technical barrier to entry remained high. Much of the literature from this period—including the influential contributions of [12] and [18, 19]—relied on substantial machinery from Loeb measure theory and model-theoretic NSA. While mathematically rigorous, this machinery can make the underlying hyperfinite constructions less transparent to readers coming from applied probability, statistics, or machine learning. The power of NSA in stochastic analysis lies in its ability to convert many analytic difficulties of the continuum—such as the non-differentiability and infinite variation of Brownian paths or the convergence of stochastic integrals and martingales—into the algebra of hyperfinite processes. On a hyperfinite grid, finite-difference and summation-by-parts identities are exact. Also, stochastic change-of-variables formulas arise as exact hyperfinite telescoping identities. Taking standard parts via the standard-part map, denoted by st(⋅)st(·) or (⋅)∘ (·), then recovers the classical Itô formula and related results as consequences of the underlying hyperfinite structure. In recent years, the landscape of generative modeling has been transformed by the advent of score-based diffusion models [11, 27]. These frameworks treat data generation as the time-reversal of a diffusion process, where a complex data distribution is progressively transformed into Gaussian noise via a forward SDE. A neural network is then trained to reverse this process by estimating the score function, ∇logpt(x)∇ p_t(x), where pt(x)p_t(x) denotes the marginal density of the forward process at time t. Although the statistical foundations of score matching were established by [15] and linked to denoising by [31], the success of this paradigm in high-dimensional settings rests upon deep results in stochastic analysis—most notably Anderson’s reverse-time SDE formula [2] and the Girsanov theorem. These provide the rigorous link between path measures and the variational objectives used for optimization. However, the standard theoretical treatment of diffusion-based generative models relies heavily on continuum arguments. Establishing the connection between discrete-time training objectives and continuous-time reverse dynamics typically requires a combination of limiting procedures, regularity assumptions on the underlying diffusion, and path-space measure-theoretic techniques. In particular, the analysis often involves reverse-time diffusion theory, change-of-measure arguments, and the control of singular behavior of scores and Fisher-information-type quantities as t→0t→ 0. While these tools are mathematically well understood, they can obscure the underlying microscopic structure of the dynamics and the relationship between discrete stochastic evolutions and their continuum limits. Despite its success in stochastic analysis, NSA has received little attention in the modern generative-modeling literature. Classical works such as [3], [16], and [13, 14] established nonstandard foundations for Brownian motion, stochastic integration, and SDEs, while [4] developed an elementary hyperfinite framework based on grid functions and infinitesimal difference equations. However, these methods were not developed in the context of score matching, reverse diffusion, or likelihood-based generative modeling. To the best of our knowledge, a systematic hyperfinite treatment connecting internal diffusion dynamics, score-based objectives, reverse-time SDEs, and variational learning principles has not previously been developed. The present work aims to build such a connection. Rather than treating hyperfinite processes merely as approximations of standard diffusions, we formulate score-based diffusion modeling directly on a hyperfinite probability space and then recover the standard theory through standard-part constructions. This viewpoint yields transparent pathwise derivations of several classical diffusion identities while also providing variational connections between denoising score matching (DSM), internal score fields, and surrogate likelihood objectives. 1.1 Summary of results To provide a precise overview of our results, we first introduce the requisite notation, which will be expanded upon in Appendix A where the fundamentals are discussed. Let ℝd∗ R^d be the nonstandard extension of Euclidean space and ⊂ℝd∗G⊂ R^d be a hyperfinite grid with mesh size δx=1/N′∈ℝ+∗δ x=1/N ∈ R^+, for a hypernatural N′∈ℕ∗∖ℕN ∈ N . Define the hyperfinite timeline δt=kδt:k∈0,1,…,NT_δ t=\kδ t:k∈\0,1,…,N\\ for an infinitesimal δt∈ℝ+∗δ t∈ R^+, where N=1/δtN=1/δ t is a hypernatural. It is assumed that N,N′N,N are such that δx2/δtδ x^2/δ t is a finite hyperreal that is not infinitesimal. For a,b∈ℝd∗a,b∈ R^d, we use the notation a≈ba≈ b to indicate that a−ba-b is an infinitesimal (componentwise). Define an internal forward walk Xtt∈δt\X_t\_t _δ t on G governed by the transition: Xt+δt=Xt+b∗(t,Xt)δt+σ∗(t,Xt)ΔWt,X_t+δ t=X_t+^*b(t,X_t)δ t+^*σ(t,X_t) W_t, where ΔWt W_t is a hyperfinite increment satisfying the internal filtration constraints: int[ΔWt∣ℱt]=0E_int[ W_t _t]=0 and int[ΔWtΔWt⊤∣ℱt]=Iδt+O(δt3/2)E_int[ W_t W_t _t]=Iδ t+O(δ t^3/2). We assume the drift and diffusion coefficients, b and σ, satisfy minimal regularity conditions, summarized as Assumption 2.2. To facilitate a transparent exposition, we employ a modular approach to these regularity requirements: while baseline results utilize Assumption 2.2, specific propositions may impose additional constraints–such as stronger differentiablity in x and t, smoothness of the initial distribution p0p_0, or uniform ellipticity of the diffusion matrix a=σσ⊤a=σ –to ensure the existence of densities or the stability of the reverse-time process. The requisite assumptions are stated explicitly alongside each result to maintain clarity. To maintain mathematical transparency while retaining full foundational rigor, our framework selectively bridges two distinct paradigms of nonstandard stochastic analysis. In treating optimization and generative objectives (such as score matching and discrete updates), we adopt a pathwise, internal grid-function perspective inspired by Nelson’s radically elementary approach [23], allowing algebraic identities to be preserved without immediate recourse to outer measure spaces. Conversely, when mapping these internal trajectories back to classical continuous-time processes, we utilize the full model-theoretic machinery of Loeb measures [21] and S-lifting (cf. [3, 13, 16]), ensuring that our hyperfinite paths project rigorously onto standard Itô diffusion laws. A grid function (cf. [5]), which may be viewed as an infinitesimal discretization of a standard function on ℝdR^d, is an internal function u:→ℝ∗,u:G→ R, and the space of all grid functions on G is denoted by ()G(G). Let 1,…,de_1,…,e_d denote the standard basis vectors in ℝdR^d and ⊂′⊂ℝd∗G ⊂ R^d be hyperfinite grids, where ′G is assumed to be large enough to contain all stencils required by the differential operators acting on functions in (′)G(G ). For a grid function u∈(′)u (G ), the forward grid derivative is the internal function in ()G(G) defined by Δi+u(x)=u(x+δxi)−u(x)δx, _i^+u(x)= u(x+δ xe_i)-u(x)δ x, for all x∈x . With α=(α1,…,αd)∈ℕdα=( _1,…, _d) ^d a multi-index, the α-th order forward grid derivative is defined recursively by Δαu=(Δ1+)α1⋯(Δd+)αdu, ^αu=( _1^+) _1·s( _d^+) _du, provided ′G contains the required k-step neighborhood of G for |α|=k|α|=k, where the order of the derivative is |α|=∑i=1dαi|α|= _i=1^d _i. We outline our contributions below. The first two results synthesize established insights from the literature on nonstandard stochastic analysis; our objective in presenting them is to provide a unified, accessible framework that bridges classical theory with contemporary generative modeling while offering an elementary and detailed nonstandard proof; some proofs—most prominently those concerning the reverse SDE—are, to our knowledge, novel in the NSA literature in the form presented here. Our initial goal is to rigorously derive the infinitesimal generator ℒtL_t and its dual relationship to the time-reversed process. 1. Internal Generator and Fokker-Planck Duality: For internal grid time t∈δt _δ t, any φ∈Cc3(ℝd) ∈ C_c^3(R^d) and x∈x , the internal grid dynamics (formally given in Definition 2.3) satisfy: int[φ∗(Xt+δt)−φ∗(x)∣Xt=x]=δtℒt∗φ∗(x)+Rδt(x),E_int[^* (X_t+δ t)-^* (x) X_t=x]=δ t\,\,^*L_t^* (x)+R_δ t(x), where the grid generator ℒt∗^*L_t is defined by: ℒt∗φ∗(x)=∑i=1dbi∗(x,t)Δi+φ∗(x)+12∑i,j=1daij∗(x,t)Δi+Δj+φ∗(x). L_t^* (x)= _i=1^d^*b_i(x,t) _i^+^* (x)+ 12 _i,j=1^d^*a_ij(x,t) _i^+ _j^+^* (x). Furthermore, the remainder satisfies |Rδt(x)|≤K(δt)3/2|R_δ t(x)|≤ K(δ t)^3/2 for some finite hyperreal K (see Lemma 3.1). Using the hyperfinite walk framework and internal adjoint identities from nonstandard stochastic calculus (cf. [4]), we recover the standard Fokker–Planck equation for the marginal density pt(x)p_t(x): ∂p∂t=−∑i=1d∂xi[bi(t,x)p(t,x)]+12∑i,j=1d∂2∂xi∂xj[aij(t,x)p(t,x)]=ℒt∗p, ∂ p∂ t=- _i=1^d ∂ x_i[b_i(t,x)p(t,x)]+ 12 _i,j=1^d ∂^2∂ x_i∂ x_j[a_ij(t,x)p(t,x)]=L_t^*p, for standard t∈[ϵ,1]t∈[ε,1] with any real ϵ>0ε>0, where ℒt∗L_t^* is the adjoint of the standard generator ℒtL_t (cf. Theorem 3.2). This result is classical in both standard and nonstandard stochastic analysis; our purpose is to derive it within a unified hyperfinite framework that will later support the reverse-time and score-based constructions. 2. Backward Mean Identity and Score-Lifting: By analyzing the time-reversed internal transition, we derive in Theorem 3.3 that for y∈ℝd∗y∈^*R^d such that the marginal pt+δt∗(y)^*p_t+δ t(y) is appreciable (i.e., limited and not infinitesimal), the backward mean satisfies: int[Xti−Xt+δti∣Xt+δt=y]=βi(t,y)δt+O(δt3/2)E_int[X_t^i-X_t+δ t^i X_t+δ t=y]= _i(t,y)δ t+O(δ t^3/2) for t∘>ϵ t>ε with any real ϵ>0ε>0, where βi _i denotes the internal reverse-drift, defined as: βi(t,y)=−bi∗(t,y)+∑j=1daij∗(t,y)Δj+logpt∗(y)+∑j=1dΔj+aij∗(t,y). _i(t,y)=-^*b_i(t,y)+ _j=1^d^*a_ij(t,y) _j^+ ^*p_t(y)+ _j=1^d _j^+^*a_ij(t,y). Drawing on ideas from the hyperfinite SDE framework of Keisler and Benci [16, 4], and the infinitesimal stochastic analysis of Stroyan and Bayod [29], we derive the above expression for β through an explicit internal Jacobian calculation and a perturbative analysis of the hyperfinite increments. Within this framework, the score term ∇logpt(⋅)∇ p_t(·), represented by the grid difference vector Δ+logpt∗:=(Δj+logpt∗)j ^+ ^*p_t:=( _j^+ ^*p_t)_j, emerges naturally as the correction required by time reversal. While the initial derivation of the vector field β rests on purely internal, elementary grid-function algebra in the style of Nelson [23], the formal pathwise convergence of Ys=X1−sY_s=X_1-s to a standard continuous Itô process crucially leverages the model-theoretic lifting theorems and Loeb-space regularities established by the classical NSA school. To be precise, Theorem 3.4 shows that the standard part of the hyperfinite reverse process solves the classical reverse-time Itô SDE. Specifically, for Loeb-almost every path, Y¯s=st(Ys) Y_s=st(Y_s) is the unique solution in law to the standard Itô SDE dY¯s=b^(1−s,Y¯s)ds+σ(1−s,Y¯s)dW¯sd Y_s= b(1-s, Y_s)ds+σ(1-s, Y_s)d W_s for any s: 1−s∈[ϵ,1]s:\,1-s∈[ε,1] with real ϵ>0ε>0, where b^(t,x)=st(β(t,x)) b(t,x)=st(β(t,x)). In addition, a second nonstandard proof is provided using the hyperfinite walk dynamics that provides a distributional perspective. This provides a pathwise hyperfinite derivation of the reverse-drift formula associated with the classical reverse-time diffusion results of [24, 2], replacing many of the usual limiting arguments by finite-dimensional internal calculations. 3. Internal Score Equivalence: Our first genuinely generative-modeling contribution is to show that the internal minimization of a hyperfinite DSM objective recovers precisely the score term appearing in the reverse-time dynamics of Theorem 3.3. Specifically, we define an internal score matching objective ℐ(θ)I(θ) based on the hyperfinite transition probabilities and show that, if ℐ(θ)≈0I(θ)≈ 0, for Loeb-almost all (x,t)∈×(δt∩[τ,1])(x,t) ×(T_δ t∩[τ,1]) (τ appreciable so that pt∗(x)^*p_t(x) is appreciable), the learned score sθs_θ satisfies the internal identity sθ(x,t)≈Δ+logpt∗(x)s_θ(x,t)≈ ^+ ^*p_t(x) (Theorem 4.1). This confirms that the learned score function is the unique algebraic lifting required to ensure the consistency of the reverse-time dynamics, effectively bridging the gap between empirical loss minimization and the theoretical conditions for time-reversal. 4. Generative Convergence: We establish a constructive procedure for sampling from the target distribution q, defined as an internal probability mass function on the grid G, through a hyperfinite reverse recursion. By defining the sampling process for appreciable forward times (i.e. 1−s1-s) as Ys+δt=Ys+β^(s,Ys)δt+σ∗(1−s,Ys)ΔW~s,Y_s+δ t=Y_s+ β(s,Y_s)δ t+^*σ(1-s,Y_s) W_s, where W~s W_s are internal independent increments of a hyperfinite Brownian motion and the reverse drift is given by β^i(s,y)=−bi∗(1−s,y)+∑j=1daij∗(1−s,y)sθ,j(y,1−s)+∑j=1dΔj+aij∗(1−s,y), β_i(s,y)=-^*b_i(1-s,y)+ _j=1^d^*a_ij(1-s,y)s_θ,j(y,1-s)+ _j=1^d _j^+^*a_ij(1-s,y), we prove that the standard-part projection Y¯s=st(Ys) Y_s=st(Y_s) forms a diffusion process for s∈[0,1−ϵ]s∈[0,1-ε] for any real ϵ>0ε>0. Consequently, under the assumptions of Theorem 4.3, the terminal law of the standard-part process coincides with the target distribution q∘ q. With this generative convergence result formulated as a mapping between internal grid distributions and their standard-part projections, we retain an elementary, discrete validation layout that maps naturally to deep learning frameworks, bypassing some of the measure-theoretic complexities typical of traditional continuum proofs. That is, this provides a hyperfinite, pathwise justification for the generative correctness of score-based diffusion models while avoiding many of the limiting arguments that typically appear in continuum formulations. 5. Hyperfinite Girsanov and Likelihood-Score Equivalence: Restricting to a state-independent diffusion matrix a, we employ a local exponential tilt to derive a hyperfinite Girsanov identity (Lemma 5.1). Subsequently, it is shown that minimizing the Fisher score-matching objective maximizes a lower bound on the surrogate likelihood, while replacing the usual Novikov-type integrability arguments by an internal hyperfinite calculation (Theorem 5.2 and Lemma 5.3). 6. Second-Order Consistency: We examine the approximation fidelity of the hyperfinite walk by fixing a finite point (t,x)(t,x) and setting P=p∗P= p. Consider the internal Euler increment h=b∗(t,x)δt+σ∗(t,x)δtξh= b(t,x)δ t+ σ(t,x) δ tξ, where ξ is an internal random variable satisfying int[ξ]=0E_int[ξ]=0, int[ξ2]=1E_int[ξ^2]=1, int[ξ3]=0E_int[ξ^3]=0, and int[ξ4]=κ<∞E_int[ξ^4]=κ<∞. The source-based master equation (cf. [23]) is given by: P(t+δt,x)=ξ[P(t,x−h)].P(t+δ t,x)=E_ξ [P(t,x-h) ]. Defining the internal frozen operator =−b∗∂+a∗2∂2A=- b∂+ a2∂^2, which uses internal spatial derivatives and serves as the nonstandard representative of the adjoint generator ℒt∗L_t^*, one gets st(P)=ℒt∗pst(AP)=L_t^*p for locally frozen coefficients. For d=1d=1, Theorem 36 establishes that for any finite (t,x)(t,x) with t>0t>0, the master equation update satisfies: P(t+δt,x)−(P+δtP+δt222P)δt2≈κ−324a(t∘,x∘)2∂4p∂x4(t∘,x∘). P(t+δ t,x)- (P+δ tAP+ δ t^22A^2P )δ t^2≈ κ-324a( t, x)^2 ∂^4p∂ x^4( t, x). This demonstrates that while first-order consistency requires only matching the first two moments, second-order weak consistency fundamentally depends on the fourth-order statistics of the hyperfinite increment. Since diffusion models are trained against quantities derived from the forward density evolution, the result highlights a connection between the fourth-order statistics of the driving noise and the accuracy of score-based targets generated by hyperfinite discretizations. 1.2 Broader outlook Bridging the Discrete-Continuous Divide: A White-Box Perspective. The generative modeling community experiences a persistent tension between discrete-time diffusion implementations and continuous-time SDE theoretical justifications, which typically rely on the Stroock-Varadhan martingale problem [28]. We argue that this gap stems from the traditional continuum formalism and its limit-based tools. Drawing inspiration from the radically elementary perspective of Nelson [23] and grid-based stochastic calculus (for example, [4, 16]), we position the hyperfinite grid as the fundamental probabilistic domain. This yields a “white-box” framework where standard diffusion is the shadow—the standard part—of a hyperfinite walk. This nonstandard approach bridges discrete algorithmic steps and continuous models, recasting continuum diffusion via exact algebraic identities on a hyperfinite grid to provide a rigorous yet intuitive discrete calculus toolset. From Model Theory to Machine Learning: Simplifying the NSA Canon. Some of the foundational work in nonstandard stochastic analysis includes [3, 13, 16, 18]. These developments introduced hyperfinite representations of stochastic processes, nonstandard stochastic integration, and infinitesimal approaches to SDEs. Most of the classical literature relies heavily on Loeb measure spaces, saturation, and model-theoretic machinery, which can obscure basic algebraic intuition. Furthermore, modern generative concepts like forward-reverse diffusion dynamics and score-based modeling were outside the scope of foundational NSA. We revisit these classical NSA constructions through the lens of score-based generative modeling. Even though we remain within the ‘Robinsonian’ scheme of NSA, we prioritize direct infinitesimal-calculus arguments (motivated by [23] in spirit) while retaining Loeb-space rigor. Rather than replacing classical theory, we synthesize hyperfinite walk constructions, reverse-time drift identities, and internal densities into a unified framework. This clarifies the role of score functions, likelihood objectives, and time reversal, hopefully making NSA accessible to researchers in stochastic modeling, numerical analysis, and generative artificial intelligence (AI). Some preliminary material is presented in the appendices; readers unfamiliar with the topic may find it helpful to consult these before proceeding to the main sections. In the substantial but foundational Appendix A, we establish the mathematical underpinnings of the hyperfinite framework, including the superstructure construction, internal objects, hyperfinite sets, transfer principles, properties of the Loeb measure, and rigorous definitions of S-regularity and grid calculus. Even readers already acquainted with NSA are advised to consult Appendix A for a concise overview. With this background, hyperfinite grid SDEs, their standard parts, and solutions are discussed in Section 2, with some classical proofs deferred to Appendices B and C. Appendix D develops the discrete Itô calculus on hyperfinite grids, providing pathwise control over stochastic increments and deriving the nonstandard Itô formula. This machinery is then leveraged in Section 3 to analyze the forward generator and prove the Fokker–Planck duality, culminating in the hyperfinite derivation of the reverse‑time SDE. These insights are applied in Section 4 to score‑based generative modeling, establishing the internal consistency of score estimation and the convergence of hyperfinite sampling dynamics. The hyperfinite Girsanov theorem is derived in Section 5 to link path‑space likelihoods with the Fisher divergence. A local expansion of the source‑based hyperfinite master equation is provided in Section 6 to analyze higher‑order operator consistency. Finally, Section 7 offers concluding remarks and outlines future research directions. 2 Hyperfinite grid SDEs and their standard part Building on the foundations laid in Section A.6, where hyperfinite Brownian motion is introduced using the Rademacher increment array (Definition A.32), we proceed to define SDEs on the grid and show their convergence to classical Itô SDEs. Definition 2.1 (Hyperfinite Grid SDE). Let b:[0,1]×ℝd→ℝdb:[0,1]×R^d ^d and σ:[0,1]×ℝd→ℝd×mσ:[0,1]×R^d ^d× m be standard functions with natural extensions b∗^*b and σ∗^*σ. For a hyperfinite time grid δtT_δ t and an m-dimensional Brownian increment array ΔWt W_t, a grid diffusion process is an internal ℝd∗^*R^d-valued process satisfying the recursion [16]: Xt+δt=Xt+b∗(t,Xt)δt+σ∗(t,Xt)ΔWt.X_t+δ t=X_t+^*b(t,X_t)\,δ t+^*σ(t,X_t)\, W_t. (1) Assumption 2.2 (Regularity of Coefficients). The standard coefficients b and σ satisfy global Lipschitz continuity and linear growth: there exists L∈ℝ+L ^+ such that for all t∈[0,1]t∈[0,1] and x,y∈ℝdx,y ^d: ‖b(t,x)−b(t,y)‖+‖σ(t,x)−σ(t,y)‖≤L‖x−y‖, \|b(t,x)-b(t,y)\|+\|σ(t,x)-σ(t,y)\|≤ L\|x-y\|, ‖b(t,x)‖2+‖σ(t,x)‖2≤L2(1+‖x‖2). \|b(t,x)\|^2+\|σ(t,x)\|^2≤ L^2(1+\|x\|^2). By the transfer principle, these properties hold internally for b∗^*b and σ∗^*σ across all x,y∈ℝd∗x,y∈^*R^d. Note on Generalization. While we utilize the specific Rademacher construction for the driving noise in the introductory material (Appendices A to D) for pedagogical clarity, our core developments starting from Section 3 are established within a more general framework. We characterize the dynamics through the increment ΔXt:=Xt+δt−Xt X_t:=X_t+δ t-X_t, relying solely on its infinitesimal moment conditions. Specifically, assuming the diffusion tensor a=σσ⊤a=σ is uniformly elliptic, one may define normalized increments ΔW~t:=a∗(t,Xt)−1/2(ΔXt−b∗(t,Xt)δt) W_t:=^*a(t,X_t)^-1/2( X_t-^*b(t,X_t)δ t) that recover d-dimensional Brownian motion in the Loeb limit. This approach ensures our results are (driving noise) distribution-agnostic. Definition 2.3 (Hyperfinite Grid Dynamics). Let G be a hyperfinite spatial grid with uniform spacing δx=Cδtδ x=C δ t. For Xt=x∈X_t=x , let the increment ΔXt=Xt+δt−Xt X_t=X_t+δ t-X_t take values in a finite local neighborhood x⊂V_x . The transition probabilities p(x,x+v)p(x,x+v) (or pv(x)p_v(x)) for v∈xv _x are internal hyperreals satisfying ∑vpv=1 _vp_v=1, pv≥0p_v≥ 0, and the following infinitesimal moment conditions: 1. First Moment (Drift): ∑v∈xvp(x,x+v)=b∗(t,x)δt _v _xvp(x,x+v)=^*b(t,x)δ t, 2. Second Moment (Diffusion): ∑v∈xvv⊤p(x,x+v)=a∗(t,x)δt+O(δt3/2) _v _xv p(x,x+v)=^*a(t,x)δ t+O(δ t^3/2), where a=σσ⊤a=σ is a uniformly elliptic diffusion tensor i.e. inft∈[0,1],x∈ℝd,‖v‖=1v⊤a(t,x)v>0, _t∈[0,1],x ^d,\,\|v\|=1v a(t,x)v>0, and b,σb,σ are jointly continuous in addition to the basic Assumption 2.2. A valid probability measure exists on G if the grid constant C satisfies the nonstandard Courant-Friedrichs-Lewy (CFL) condition C2I⪰a∗(t,x)C^2I ^*a(t,x) for all (x,t)(x,t) in the support of the process [6]. This ensures that the local covariance of the target SDE remains contained within the reachable neighborhood of the hyperfinite grid at each infinitesimal step, maintaining the non-negativity of internal probabilities [1]. The above construction ensures that Xt∈X_t for all t∈δt _δ t while maintaining infinitesimal consistency with the target SDE coefficients. Proposition 2.4 (Existence and S-continuity of Grid Solutions). Under Assumption 2.2, for any internal initial datum X0(ω)X_0(ω) such that ‖X0(ω)‖\|X_0(ω)\| is limited for all ω∈Ωω∈ , the grid dynamics in Definition 2.3 (or recursion (1)) admit a unique internal solution Xt(ω)X_t(ω) for all t∈δt _δ t. Furthermore, XtX_t is μL _L-almost surely S-continuous. Proof. Existence and uniqueness follow by internal induction on the hyperfinite set δtT_δ t. For a fixed ω, the value Xt+δtX_t+δ t is uniquely determined by the internal transition law. To establish that XtX_t is limited Loeb-a.s., we analyze the second moment. Squaring the recursion and applying the linear growth property from Assumption 2.2, we obtain the internal inequality: int[‖Xt+δt‖2]≤int[‖Xt‖2]+K(1+int[‖Xt‖2])δt,E_int[\|X_t+δ t\|^2] _int[\|X_t\|^2]+K(1+E_int[\|X_t\|^2])δ t, where K is a limited constant depending on the linear growth coefficient L. By the internal discrete Grönwall inequality, int[‖Xt‖2]≤(int[‖X0‖2]+Kt)eKtE_int[\|X_t\|^2]≤(E_int[\|X_0\|^2]+Kt)e^Kt, which is limited for all t∈δt _δ t. The details regarding S-continuity are provided in Appendix B. ∎ Theorem 2.5 (Convergence to the Standard Itô Solution). Let b:[0,1]×ℝd→ℝdb:[0,1]×R^d ^d and σ:[0,1]×ℝd→ℝd×mσ:[0,1]×R^d ^d× m be standard functions that are jointly continuous and satisfy the global Lipschitz and linear growth conditions of Assumption 2.2. Let X:δt×Ω→ℝd∗X:T_δ t× →^*R^d be the unique internal solution to the hyperfinite grid recursion (1) where st(X0)st(X_0) is limited μL _L-almost surely. Define the standard-part process X¯t X_t for each standard t∈[0,1]t∈[0,1] by X¯t(ω)≔st(Xτ(ω)) X_t(ω) (X_τ(ω)) for any τ∈δtτ _δ t such that τ≈tτ≈ t. Then: 1. X¯t X_t is well-defined μL _L-almost surely and has continuous paths. 2. X¯t X_t is the unique strong solution to the standard Itô SDE: X¯t=X¯0+∫0tb(r,X¯r)r+∫0tσ(r,X¯r)W¯r, X_t= X_0+ _0^tb(r, X_r)\,dr+ _0^tσ(r, X_r)\,d W_r, (2) where W¯ W is the standard Brownian motion on the Loeb space (Ω,ℒ,μL)( ,L, _L). Proof. The proof may be found in Appendix C. ∎ The next section carries out a derivation of the forward and reverse evolutions constructively using the grid calculus framework developed in Section A. 3 Hyperfinite derivation of forward and reverse diffusion This section establishes the infinitesimal generator of the grid dynamics. While Rademacher increments were used previously for clarity, the following results hold for any walk dynamics satisfying Definition 2.3. The proof of Lemma 3.1 utilizes internal moment conditions within the grid geometry. Even though the abstract convergence of hyperfinite internal expectations to standard Markovian generators is well established via Loeb measures and hyperfinite semigroups (for broad contours, see Albeverio et al. [1], Keisler [16], Lindstrøm [20]), Lemma 3.1 provides a coordinate-wise expansion formulated entirely within the internal grid algebra using the native forward differences. 3.1 Forward generator analysis Lemma 3.1 (Grid Moment Expansion). Let G be a hyperfinite grid with spacing δx=Cδtδ x=C δ t. Under the assumptions on coefficients b,σb,σ as in Definition 2.3 and internal grid time t∈δt _δ t, for any φ∈Cc3(ℝd) ∈ C_c^3(R^d) and x∈x , the internal grid dynamics satisfy: int[φ∗(Xt+δt)−φ∗(x)∣Xt=x]=δtℒt∗φ∗(x)+Rδt(x),E_int[^* (X_t+δ t)-^* (x) X_t=x]=δ t\,\,^*L_t^* (x)+R_δ t(x), (3) where the grid generator ℒt∗^*L_t is defined by: ℒt∗φ∗(x)=∑i=1dbi∗(t,x)Δi+φ∗(x)+12∑i,j=1daij∗(t,x)Δi+Δj+φ∗(x),^*L_t^* (x)= _i=1^d^*b_i(t,x) _i^+^* (x)+ 12 _i,j=1^d^*a_ij(t,x) _i^+ _j^+^* (x), (4) and the forward difference operators Δi+ _i^+ are the native grid derivatives. The remainder satisfies |Rδt(x)|≤K(δt)3/2|R_δ t(x)|≤ K(δ t)^3/2 for some finite K∈ℝfin∗K∈^*R_fin depending only on the C3C^3 norm of φ , the ∥⋅∥∞\|·\|_∞ norm of b,σb,σ restricted to supp(φ)( ), and the grid constant C. In order to avoid cluttered notation, we occasionally forgo differentiating a function f (or an operator) and its extension f∗^*f, prioritizing clarity and comprehension for the reader. Proof. Let ΔX=Xt+δt−x X=X_t+δ t-x be the jump increment. By the definition of the grid dynamics, ΔX X takes values in a finite local neighborhood x⊂V_x with probability p(x,x+v)p(x,x+v) for v∈xv _x. For any v∈x⊂v _x , since φ is S-C3C^3, we have the internal identity: φ∗(x+v)−φ∗(x)=∑ivi∂i(φ∗)(x)+12∑i,jvivj∂ij(φ∗)(x)+ℛ3(x,v),^* (x+v)-^* (x)= _iv_i _i(^* )(x)+ 12 _i,jv_iv_j _ij(^* )(x)+R_3(x,v), where the third-order remainder is bounded by |ℛ3(x,v)|≤M‖v‖3|R_3(x,v)|≤ M\|v\|^3 for M=16sup|D3φ|M= 16 |D^3 |. By the S-smoothness of φ and the diffusion scaling of the grid, we substitute ∂iφ=Δi+φ+O(δt) _i = _i^+ +O( δ t) and ∂ijφ=Δi+Δj+φ+O(δt) _ij = _i^+ _j^+ +O( δ t) into the expansion. Next, we take the internal expectation int[⋅]=∑v∈xp(x,x+v)(⋅)E_int[·]= _v _xp(x,x+v)(·). 1. Substitution of Moments: By the construction of the grid dynamics, we have ∑vvipv=bi∗δt _vv_ip_v=^*b_iδ t and ∑vvivjpv=aij∗δt+O(δt3/2) _vv_iv_jp_v=^*a_ijδ t+O(δ t^3/2). Substituting these into the expected Taylor expansion: int[Δφ∣Xt=x]=(∑ibi∗Δi+φ)δt+12(∑i,jaij∗Δi+Δj+φ)δt+O(δt3/2)+int[ℛ3].E_int[ X_t=x]= ( _i^*b_i _i^+ )δ t+ 12 ( _i,j^*a_ij _i^+ _j^+ )δ t+O(δ t^3/2)+E_int[R_3]. 2. Remainder Scaling: Since ‖v‖≤C′δt\|v\|≤ C δ t for all v∈xv _x with a limited C′C proportional to C, the third-order term is bounded as |int[ℛ3]|≤∑v∈xpvM‖v‖3≤M(C′δt)3=MC′3(δt)3/2.|E_int[R_3]|≤ _v _xp_vM\|v\|^3≤ M(C δ t)^3=MC 3(δ t)^3/2. Combining the terms, we obtain int[Δφ∣Xt=x]=δtℒtφ+O(δt3/2)E_int[ X_t=x]=δ tL_t +O(δ t^3/2). The finiteness of K follows from the fact that φ∈Cc3 ∈ C_c^3 and the S-boundedness of the coefficients b and a. ∎ 3.2 Derivation of the Fokker-Planck equation While the transition from hyperfinite walks to the Fokker-Planck equation is addressed (quite explicitly) in the discrete SDE framework of Benci et al. [4], our derivation offers an explicit, entry-by-entry correspondence between the stochastic recursion and the distributional evolution. Specifically, we utilize the internal grid adjoint identities to show that the discrete update operator acts as the exact internal grid analogue of the continuous Fokker-Planck operator ℒt∗L_t^*. By keeping the duality entirely algebraic on the hyperfinite lattice prior to taking standard parts, we bypass the need for continuous measure-theoretic limits, allowing the classical equation to emerge directly from the discrete grid conservation law. Theorem 3.2 (Forward Density Evolution). In addition to the conditions on b,σb,σ from Definition 2.3, assume that that b and σ are Cα2,1+αC α2,1+α and Cα2,2+αC α2,2+α, respectively, for some α∈(0,1)α∈(0,1). Then the standard marginal density pt(x)p_t(x) satisfies the Fokker-Planck equation: ∂p∂t=−∑i=1d∂xi[bi(t,x)p(t,x)]+12∑i,j=1d∂2∂xi∂xj[aij(t,x)p(t,x)]=ℒt∗p, ∂ p∂ t=- _i=1^d ∂ x_i[b_i(t,x)p(t,x)]+ 12 _i,j=1^d ∂^2∂ x_i∂ x_j[a_ij(t,x)p(t,x)]=L_t^*p, for standard t∈[ϵ,1]t∈[ε,1] with any real ϵ>0ε>0. Here ℒt∗L_t^* is the adjoint of the standard generator ℒtL_t. Extending Theorem 3.2 to t=0t=0 would require additional smoothness assumptions on the initial distribution p0p_0 and the forward diffusion. Proof. Given a real ϵ>0ε>0, consider ϵ≈τ∈δtε≈τ _δ t. Henceforth, we work on the (sub)grid δt∩[τ,1]T_δ t∩[τ,1] and do not distinguish between standard functions and their hyperreal extensions, unless necessary. Consider the discrete time evolution of the density on this grid. Let ⟨f,g⟩=∑x∈f(x)g(x)(δx)d f,g _G= _x f(x)g(x)(δ x)^d, where G is a uniform, infinite, hyperfinite grid with the usual diffusion scaling such that δx/δtδ x/ δ t is limited and noninfinitesimal. For any test function φ∈Cc∞(ℝd) ∈ C_c^∞(R^d), the internal probability mass conservation implies: ⟨pt+δt,φ⟩=⟨pt,Pδtφ⟩ p_t+δ t, _G= p_t,P_δ t _G where Pδtφ(x)=int[φ(Xt+δt)∣Xt=x]P_δ t (x)=E_int[ (X_t+δ t) X_t=x]. Applying Lemma 3.1 and noting uniformity of error: ⟨pt+δt,φ⟩=⟨pt,φ+δtℒtφ+O(δt3/2)⟩. p_t+δ t, _G= p_t, +δ tL_t +O(δ t^3/2) _G. Rearranging and forming the internal difference quotient: ⟨pt+δt−ptδt,φ⟩=⟨pt,ℒtφ⟩+O(δt1/2). p_t+δ t-p_tδ t, _G= p_t,L_t _G+O(δ t^1/2). Using Propositions A.29 and A.30, we have the following internal identity for nonstandard extensions of all φ∈Cc∞ ∈ C_c^∞ (that have grid-compact support): ⟨pt+δt−ptδt,φ⟩=⟨−∑iΔi−(bip)+12∑ijΔj−Δi−(aijp),φ⟩+O(δt1/2). p_t+δ t-p_tδ t, _G= - _i _i^-(b_ip)+ 12 _ij _j^- _i^-(a_ijp), _G+O(δ t^1/2). Since the coefficients satisfy the stated regularity assumptions and a is uniformly elliptic, standard parabolic regularity theory implies that p∈C1,2p∈ C^1,2; see, for example, Schauder estimates in [10]. Passing to standard parts in the weak identity and using the infinitesimal consistency of grid derivatives with standard derivatives, we obtain for every test function φ∈Cc∞(ℝd) ∈ C_c^∞(R^d): ⟨∂p∂t,φ⟩=⟨−∑i∂xi(bip)+12∑i,j∂2∂xi∂xj(aijp),φ⟩. ∂ p∂ t, = - _i ∂ x_i(b_ip)+ 12 _i,j ∂^2∂ x_i∂ x_j(a_ijp), . Thus the standard density p satisfies the Fokker–Planck equation in the weak sense on [ϵ,1][ε,1]. Since the Schauder estimates ensure that the first argument in the inner products on both sides are continuous, via the classical du Bois-Reymond lemma, ∂p∂t=−∑i=1d∂xi(bip)+12∑i,j=1d∂2∂xi∂xj(aijp)=ℒt∗p ∂ p∂ t=- _i=1^d ∂ x_i(b_ip)+ 12 _i,j=1^d ∂^2∂ x_i∂ x_j(a_ijp)=L_t^*p pointwise for all standard t∈[ϵ,1]t∈[ε,1]. ∎ Remark 1 (Alternative Purely Nonstandard Regularity Pathway). An alternative to invoking standard parabolic PDE regularity theory is to operate entirely within the internal nonstandard domain by exploiting the intrinsic smoothing properties of the grid. Under the uniform ellipticity of a(t,x)a(t,x) and the hyperfinite grid scaling, the discrete fundamental solution (the grid transition kernel) and its finite differences satisfy internal Aronson-type Gaussian bounds for all appreciable times t≥ϵ>0t≥ε>0. By the linearity of the internal grid convolution and a nonstandard induction argument, the hyperfinite random walk undergoes localized combinatorial dispersion. This dispersion induces S-smoothness on pt∗^*p_t up to the required order in both space and time, thereby ensuring that the standard part [pt∗(x)]∘ [^*p_t(x)] natively inherits classical C1,2C^1,2 regularity under the conditions of Theorem 3.2. This pathway validates the corresponding arguments in the proof of the theorem from first principles inside the hyperfinite world, bypassing external continuum machinery. 3.3 Backward mean analysis Theorem 3.3 (Backward Mean Identity). Suppose b,σb,σ from Definition 2.3 belong, in addition, to Cα2,1+αC α2,1+α and Cα2,2+αC α2,2+α respectively for some α∈(0,1)α∈(0,1), implying p∈C1,2([ϵ,1]×ℝd)p∈ C^1,2([ε,1]×R^d) for any standard ϵ>0ε>0. Then, for any t such that st(t)>ϵst(t)>ε, the backward conditional expectation satisfies int[Xti−Xt+δti∣Xt+δt=y]=βi(t,y)δt+O(δt3/2),E_int[X_t^i-X_t+δ t^i X_t+δ t=y]= _i(t,y)δ t+O(δ t^3/2), (5) for any limited y∈y such that pt+δt∗(y)^*p_t+δ t(y) is appreciable (non-infinitesimal), where the backward drift β is given by: βi(t,y)=−bi∗(t,y)+∑j=1daij∗(t,y)Δj+logpt∗(y)+∑j=1dΔj+aij∗(t,y). _i(t,y)=-^*b_i(t,y)+ _j=1^d^*a_ij(t,y) ^+_j ^*p_t(y)+ _j=1^d ^+_j^*a_ij(t,y). (6) We provide two distinct proofs for this identity. Despite being long and arduous, the first, utilizing a Rademacher-based perturbation analysis, offers significant insight into the Jacobian transformations and local density perturbations at the microscopic level. The second proof is a bit more concise, leveraging the general hyperfinite walk dynamics without relying on a specific noise distribution. Notably, the first proof generalizes to any increment ΔWk=ξkδt W_k= _k δ t with appropriate moment conditions (cf. step 3 of the first proof), demonstrating that the Rademacher assumption is a convenience rather than a constraint. The derivation of the backward drift β presented here situates itself at the intersection of classical nonstandard stochastic analysis and modern generative modeling. While the foundational identity for β is a cornerstone of Nelson’s stochastic mechanics [24] (where he actually used standard calculus) and was also established in the standard setting by Anderson [2], its treatment within the hyperfinite framework offers unique structural transparency. Although previous NSA literature–such as Keisler [16] and Albeverio et al. [1]–addresses the theory of internal processes, to the best of our knowledge, the time-reversal identity has not been explicitly shown via hyperfinite algebra and our specific expansion—utilizing the internal Jacobian transformation and pathwise perturbation of the hyperfinite increment—provides a granular ‘microscopic’ derivation. This approach, mirroring the infinitesimal bridge constructions of Stroyan and Bayod [29], rigorously recovers the score term ∇logpt(⋅)∇ p_t(·) from the hyperfinite conditional expectation expansion. By explicitly capturing how spatial variations in the diffusion tensor a and the local density p interact via the Jacobian inversion, the following proof bridges the gap between the discrete hyperfinite world and the score-based modeling objectives of [27]. Proof. To simplify notation, we drop the ∗ from internal functions and work on the hyperfinite grid. 1. The Implicit Jump Equation: Let h=Xt−Xt+δth=X_t-X_t+δ t conditioned on Xt+δt=yX_t+δ t=y, so that Xt=y+hX_t=y+h. From the forward recursion: −hi=bi(y+h)δt+∑k=1mσik(y+h)ΔWk.-h^i=b_i(y+h)δ t+ _k=1^m _ik(y+h) W_k. Expanding b and σ around the destination point y∈y : bi(y+h) b_i(y+h) =bi+∑j∂jbihj+O(‖h‖2), =b_i+ _j _jb_ih^j+O(\|h\|^2), σik(y+h) _ik(y+h) =σik+∑j∂jσikhj+O(‖h‖2). = _ik+ _j _j _ikh^j+O(\|h\|^2). 2. Perturbation Expansion of h: Since the forward increment satisfies hi=O(δt),h^i=O( δ t), repeated substitution of the implicit jump equation yields an asymptotic expansion of the form hi=δtH0i+δtH1i+O(δt3/2).h^i= δ t\,H_0^i+δ t\,H_1^i+O(δ t^3/2). Substituting ΔWk=ξkδt W_k= _k δ t and equating powers of δtδ t: 1. Order δt1/2δ t^1/2: −H0i=∑kσikξk⟹H0i=−∑kσikξk-H_0^i= _k _ik _k H_0^i=- _k _ik _k. 2. Order δtδ t: −H1i=bi+∑k,j∂jσikH0jξk-H_1^i=b_i+ _k,j _j _ikH_0^j _k. Substituting H0jH_0^j: H1i=−bi+∑k,j,l(∂jσik)σjlξlξk.H_1^i=-b_i+ _k,j,l( _j _ik) _jl _l _k. 3. Expectations of the Terms: By the symmetry of the noise int[ξk]=0E_int[ _k]=0 and independence, we obtain: (i) int[H0i]=0E_int[H_0^i]=0 and (i) int[H1i]=−bi+∑j,k(∂jσik)σjkE_int[H_1^i]=-b_i+ _j,k( _j _ik) _jk. Thus, the unconditional mean is int[hi]=int[H1i]δt+O(δt3/2)E_int[h^i]=E_int[H_1^i]δ t+O(δ t^3/2). 4. Density Expansion via Bayes’ Rule: The probability of a particle landing at y depends on the local concentration of paths. Let J be the internal Jacobian matrix Jij=∂yi∂xj=δij+∂jbiδt+∑k∂jσikξkδtJ_ij= ∂ y^i∂ x^j= _ij+ _jb_iδ t+ _k _j _ik _k δ t. The determinant expansion is: det(J)=1+δt∑j,k∂jσjk(t,y)ξk+O(δt). (J)=1+ δ t _j,k _j _jk(t,y) _k+O(δ t). By the transfer of the change of variables formula, the density weighting is given by the inverse Jacobian (detJ)−1=1−δt∑j,k∂jσjkξk+O(δt)( J)^-1=1- δ t _j,k _j _jk _k+O(δ t). The conditional expectation under consideration is: int[hi∣y]=ξ[hipt(y+h)(detJ)−1]pt+δt(y).E_int[h^i y]= E_ξ[h^ip_t(y+h)( J)^-1]p_t+δ t(y). Expanding pt(y+h)p_t(y+h) and collecting terms of order δtδ t in the numerator NiN^i: Niδt=ptξ[H1i]+∑j∂jptξ[H0iH0j]−ptξ[H0i∑j,k∂jσjkξk], N^iδ t=p_tE_ξ[H_1^i]+ _j _jp_tE_ξ[H_0^iH_0^j]-p_tE_ξ [H_0^i _j,k _j _jk _k ], where, from left to right, Term A (Drift): pt(−bi+∑j,k(∂jσik)σjk) p_t (-b_i+ _j,k( _j _ik) _jk ) Term B (Score): ∑jaij∂jpt _ja_ij _jp_t Term C (Volumic): −ptξ[H0i∑j,k∂jσjkξk]=pt∑j,kσik∂jσjk. -p_tE_ξ [H_0^i _j,k _j _jk _k ]=p_t _j,k _ik _j _jk. 5. Synthesis: Summing A and C via the product rule ∂jaij=∑k(σik∂jσjk+σjk∂jσik) _ja_ij= _k( _ik _j _jk+ _jk _j _ik): Ni=δt[pt(−bi+∑j∂jaij)+∑jaij∂jpt]+O(δt3/2).N^i=δ t [p_t(-b_i+ _j _ja_ij)+ _ja_ij _jp_t ]+O(δ t^3/2). Dividing by pt+δt(y)=pt(y)+O(δt)p_t+δ t(y)=p_t(y)+O(δ t) and using ∂jptpt=∂jlogpt _jp_tp_t= _j p_t: int[hi∣y]=(−bi+∑j∂jaij+∑jaij∂jlogpt)δt+O(δt3/2).E_int[h^i y]= (-b_i+ _j _ja_ij+ _ja_ij _j p_t )δ t+O(δ t^3/2). Replacing internal partial derivatives with grid difference operators Δj+ _j^+ incurs an error of O(δx)=O(δt)O(δ x)=O( δ t), which is absorbed into the O(δt3/2)O(δ t^3/2) factor. ∎ Aliter. Let h=Xt−Xt+δth=X_t-X_t+δ t be the backward jump conditioned on Xt+δt=yX_t+δ t=y. By the internal discrete Bayes’ rule [23], the conditional expectation is: int[hi∣y]=1pt+δt(y)∑x∈(x−y)ipt(x)P(y∣x).E_int[h^i y]= 1p_t+δ t(y) _x (x-y)^ip_t(x)P(y x). (7) Let v=y−xv=y-x be the forward increment from x. Then h=−vh=-v and we rewrite the sum over the local neighborhood xV_x that can reach y. For a fixed v, let Φv(x)=pt(x)p(x,x+v) _v(x)=p_t(x)p(x,x+v). We perform a Taylor expansion of this ‘flux density’ around the destination point y: Φv(y−v)=Φv(y)−∑jvj∂jΦv(y)+ℛ2(y,v), _v(y-v)= _v(y)- _jv^j _j _v(y)+R_2(y,v), where ℛ2=O(‖v‖2)R_2=O(\|v\|^2). Substituting this into the numerator NiN^i of (7): Ni=∑v(−vi)[pt(y)p(y,y+v)−∑jvj∂j(pt(y)p(y,y+v))+ℛ2].N^i= _v(-v^i) [p_t(y)p(y,y+v)- _jv^j _j (p_t(y)p(y,y+v) )+R_2 ]. 1. Moment Substitution: Distributing the sum over v and applying the forward grid moments ∑vipv=biδtΣ v^ip_v=b_iδ t and ∑vivjpv=aijδt+O(δt3/2)Σ v^iv^jp_v=a_ijδ t+O(δ t^3/2): Ni=−pt(y)biδt+∑j∂j(pt(y)aijδt)−∑vviℛ2+O(δt3/2).N^i=-p_t(y)b_iδ t+ _j _j (p_t(y)a_ijδ t )- _vv^iR_2+O(δ t^3/2). Since viℛ2=O(‖v‖3)=O(δt3/2)v^iR_2=O(\|v\|^3)=O(δ t^3/2), and the sum involves a finite number of neighbors, the total remainder is O(δt3/2)O(δ t^3/2). 2. Discrete Product Rule and Density Shift: We replace the internal partial derivative ∂j _j with the grid operator Δj+ _j^+, incurring an O(δt)O( δ t) error which, when multiplied by δtδ t, is absorbed into the O(δt3/2)O(δ t^3/2) remainder. Applying the discrete product rule (Proposition A.27): Δj+(aijpt)=aij(y)Δj+pt(y)+pt(y+δxj)Δj+aij(y). _j^+(a_ijp_t)=a_ij(y) _j^+p_t(y)+p_t(y+δ xe_j) _j^+a_ij(y). By the S-differentiability of ptp_t in space, pt(y+δxj)=pt(y)+O(δt)p_t(y+δ xe_j)=p_t(y)+O( δ t). Substituting this back: Ni=δt[−ptbi+∑j(aijΔj+pt+ptΔj+aij)]+O(δt3/2).N^i=δ t [-p_tb_i+ _j(a_ij _j^+p_t+p_t _j^+a_ij) ]+O(δ t^3/2). 3. Synthesis: Dividing by pt+δt(y)=pt(y)+O(δt)p_t+δ t(y)=p_t(y)+O(δ t) and noting that since pt(y)p_t(y) is appreciable by hypothesis and ptp_t is S-C2C^2, Δj+logpt(y)=Δj+pt(y)pt(y)+O(δt), _j^+ p_t(y)= _j^+p_t(y)p_t(y)+O( δ t), int[hi∣y] _int[h^i y] =(−bi+∑jΔj+aij+∑jaijΔj+logpt)δt+O(δt3/2). = (-b_i+ _j _j^+a_ij+ _ja_ij _j^+ p_t )δ t+O(δ t^3/2). This confirms the backward drift βi _i from Theorem 3.3. This alternative proof provides a distributional (Eulerian) perspective that complements the preceding pathwise (Lagrangian) derivation. By applying the internal discrete Bayes’ rule directly to the hyperfinite grid G, we characterize the backward mean as a local redirection of probability flux rather than a perturbation of individual noise increments. This approach—leveraging a discrete Taylor expansion of the flux density Φv(x) _v(x)—highlights how the backward drift β naturally emerges from the interaction between forward grid moments and the local geometry of the density ptp_t. This “macro-grid” methodology, inspired by the nonstandard frameworks of [23] and [4], offers a robust algebraic verification that the discrete score-like term Δj+logpt _j^+ p_t arises naturally as the exact correction required when expressing backward transport via forward transition statistics on a hyperfinite lattice. Consequently, the identity (5) follows directly from the defining moment structure of the hyperfinite random walk, establishing that the result is invariant to the specific choice of internal increment distribution. 3.4 Hyperfinite derivation of reverse SDE Time-Reversal Setup. Define the time-reversed hyperfinite process Ys=X1−sY_s=X_1-s for s∈δts _δ t. Let ℱsY\ F_s^Y\ be the internal filtration generated by Y, i.e., ℱsY=σ(Yr:r≤s,r∈δt) F_s^Y=σ(Y_r:r≤ s,r _δ t). Equivalently, this is the internal filtration generated by the forward process looked at in reverse time. The preservation of the adapted structure between the hyperfinite and standard settings is rigorously justified by the theory of adapted distributions [12]. Theorem 3.4 (Time-Reversal). Let XtX_t satisfy the hyperfinite walk dynamics (Definition 2.3) under the assumptions of Theorem 3.3. Let Ys=X1−sY_s=X_1-s be the time-reversed process. For μL _L-almost every ω∈Ωω∈ , the standard-part process Y¯s=st(Ys) Y_s=st(Y_s) is the unique solution in law to the standard Itô SDE: dY¯s=b^(1−s,Y¯s)ds+σ(1−s,Y¯s)dW¯sd Y_s= b(1-s, Y_s)ds+σ(1-s, Y_s)d W_s (8) for any 1−s∈[ϵ,1]1-s∈[ε,1] with real ϵ>0ε>0, where the standard reverse drift b b is the standard part of the internal backward drift β from Theorem 3.3. Proof. 1. Internal Moment Identification: Let s∈δts _δ t and t=1−st=1-s appreciable. The backward increment is ΔYs=Ys+δt−Ys=Xt−δt−Xt Y_s=Y_s+δ t-Y_s=X_t-δ t-X_t. By Theorem 3.3, the conditional expectation satisfies: int[Xt−δti−Xti∣Xt=y]=βi(t−δt,y)δt+O(δt3/2).E_int[X_t-δ t^i-X_t^i X_t=y]= _i(t-δ t,y)δ t+O(δ t^3/2). (9) Since β is S-continuous in t, βi(t−δt,y)=βi(t,y)+O(δtα/2) _i(t-δ t,y)= _i(t,y)+O(δ t^α/2), so (9) becomes βi(t,y)δt+O(δt1+α2) _i(t,y)δ t+O(δ t^1+ α2). For the second moment: int[(Xt−δti−Xti)(Xt−δtj−Xtj)∣Xt=y]=aij(t,y)δt+O(δt1+α2).E_int[(X_t-δ t^i-X_t^i)(X_t-δ t^j-X_t^j) X_t=y]=a_ij(t,y)δ t+O(δ t^1+ α2). (10) 2. Martingale Construction: For φ∈Cc3(ℝd) ∈ C_c^3(R^d), consider the internal Doob decomposition of φ∗(Ys)^* (Y_s): Ms=φ∗(Ys)−φ∗(Y0)−∑r<sint[φ∗(Yr+δt)−φ∗(Yr)∣ℱrY].M_s=^* (Y_s)-^* (Y_0)- _r<sE_int[^* (Y_r+δ t)-^* (Y_r) F^Y_r]. By Lemma 3.1 and the backward moments (9)-(10), the internal compensator above is: ∑r<s(∑iβiΔi+φ∗+12∑i,jaijΔi+Δj+φ∗)δt+∑r<sO(δt1+α2). _r<s ( _i _i _i^+^* + 12 _i,ja_ij _i^+ _j^+^* )δ t+ _r<sO(δ t^1+ α2). The remainder ∑r<sO(δt1+α2)≈0 _r<sO(δ t^1+ α2)≈ 0. Taking the standard part: M¯s=φ(Y¯s)−φ(Y¯0)−∫0sℒ1−urevφ(Y¯u)u, M_s= ( Y_s)- ( Y_0)- _0^sL^rev_1-u ( Y_u)du, where ℒtrevφ=b^⋅∇φ+12tr(aD2φ)L^rev_t = b·∇ + 12tr(aD^2 ) (standard reverse Kolmogorov operator). 3. Quadratic Variation and Characterization: The internal quadratic variation [M]s=∑r<s(ΔMr)2[M]_s= _r<s( M_r)^2 satisfies: st([M]s)=st(∑r<s∑i,j(Δi+φ∗Δj+φ∗)aijδt)=∫0s(∇φ)⊤a(1−u,Y¯u)∇φdu.st([M]_s)=st ( _r<s _i,j( _i^+^* _j^+^* )a_ijδ t )= _0^s(∇ ) a(1-u, Y_u)∇ \,du. Since MsM_s is an S-continuous internal martingale with limited quadratic variation, its standard part M¯s M_s is an a.s. continuous Loeb-space martingale. This identifies Y¯s Y_s as a solution to the martingale problem for (b^,a)( b,a). By the Stroock–Varadhan uniqueness theorem, under the assumed regularity and uniform ellipticity, the martingale problem associated with (b^,a)( b,a) is well posed. Consequently, Y¯ Y has the same law as the unique diffusion solving the reverse Itô SDE. Alternative (Brownian Reconstruction): Alternatively, define the normalized internal increment ΔW~s=a∗(t,Ys)−1/2(ΔYs−βδt) W_s=^*a(t,Y_s)^-1/2( Y_s-βδ t). One easily verifies int[ΔW~s]=O(δt1+α/2)E_int[ W_s]=O(δ t^1+α/2) and int[ΔW~sΔW~s⊤]=Iδt+O(δt1+α/2)E_int[ W_s W_s ]=Iδ t+O(δ t^1+α/2). By Anderson’s theorem [3], W¯s=st(∑r<sΔW~r) W_s=st( _r<s W_r) is a standard Brownian motion, and the SDE follows. ∎ The rigorous derivation of the reverse SDE in Theorem 3.4 follows the structural blueprint of classical nonstandard stochastic analysis, yet it shifts the analytical focus toward pathwise consistency. While we build upon the foundational framework of Keisler [16] regarding hyperfinite SDEs and the martingale limit theorems of Lindstrøm [18, 19], our proof explicitly bridges these techniques with the classical time-reversal identities of Anderson [2]. By constructing the reverse-time process directly on the hyperfinite grid, we demonstrate that the score correction does not emerge abstractly after a passage to standard parts; rather, it appears concretely at the internal level via the discrete Bayes’ rule on the hyperfinite lattice. This perspective provides a “white-box” justification for the reverse-time process, showing how the discrete backward drift β—derived via our local perturbation and flux density expansions—serves as the exact internal lifting of the standard reverse drift b b. 4 Hyperfinite approach to generative modeling To implement the reverse-time diffusion process derived in Theorem 3.4, popular in the generative models literature [27], one requires the score function ∇logpt(x)∇ p_t(x), which is generally unknown. In the standard literature on diffusion models, this quantity is estimated via score matching [15, 31]. Classical derivations typically require the interchange of derivatives and integrals through the Leibniz integral rule together with decay and integrability assumptions on the underlying densities. Within the hyperfinite framework, these analytic difficulties are replaced by finite-dimensional internal algebra. Since the data distribution is represented as an internal probability mass function on the hyperfinite grid and the forward dynamics are governed by an internal Markov recursion, the DSM objective admits an exact internal conditional-expectation representation. In particular, the optimal score field is defined by an orthogonal projection which holds exactly at the internal level. For appreciable times t, parabolic regularity and S-smoothness of the transition densities imply that this projected score is infinitesimally close to the marginal grid score Δ+logpt(x) ^+ p_t(x). Consequently, the learned score provides a mathematically consistent internal representation of the backward drift appearing in the reverse diffusion dynamics. 4.1 Hyperfinite score estimation and internal consistency In the hyperfinite framework, the data distribution q(x0)q(x_0) is represented by an internal probability mass function q on the grid G. The “learned” score function sθ(x,t)s_θ(x,t) is an S-continuous internal function s:×δt→ℝd∗s:G×T_δ t→^*R^d in the limited part of the domain. We now justify that the internal minimization of the DSM objective yields the precise backward drift required for Theorem 3.4. Theorem 4.1 (Internal Score Equivalence). Assume regularity as in Theorem 3.3. Let G be a hyperfinite grid and δt∈δtδ t _δ t be the infinitesimal time step with diffusion scaling. Let Pt|0(xt|x0)P_t|0(x_t|x_0) be the internal transition probability of the forward grid recursion. For an arbitrary real ϵ>0ε>0, consider a representative in δtT_δ t given by τ≈ϵτ≈ε. Define the internal objective: ℐ(θ)=12∑t∈δt∩[τ,1]∑x0∈∑xt∈q(x0)Pt|0(xt|x0)∥sθ(xt,t)−Δ+logPt|0(xt|x0)∥2δt.I(θ)= 12 _t _δ t∩[τ,1] _x_0 _x_t q(x_0)P_t|0(x_t|x_0) \|s_θ(x_t,t)- ^+ P_t|0(x_t|x_0) \|_G^2δ t. (11) If there exists θ∈ℝk∗θ∈^*R^k such that sθs_θ is limited and ℐ(θ)≈0I(θ)≈ 0, then for Loeb-almost all (x,t)∈×(δt∩[τ,1])(x,t) ×(T_δ t∩[τ,1]) (where pt(x)p_t(x) is appreciable), sθ(x,t)≈Δ+logpt(x).s_θ(x,t)≈ ^+ p_t(x). Consequently, the standard part st(sθ)st(s_θ) coincides with the standard score ∇logpt∘∇ p_ t almost everywhere on the support of the process in every appreciable time interval [τ,1][τ,1]. Proof. 1. Linearity of the Marginal Difference. By the definition of the internal marginal density pt(x)=∑x0q(x0)Pt|0(x|x0)p_t(x)= _x_0q(x_0)P_t|0(x|x_0) and the linearity of the grid derivative Δ+ ^+, we have the exact identity: Δ+pt(x)pt(x)=∑x0∈[q(x0)Pt|0(x|x0)pt(x)]Δ+Pt|0(x|x0)Pt|0(x|x0). ^+p_t(x)p_t(x)= _x_0 [ q(x_0)P_t|0(x|x_0)p_t(x) ] ^+P_t|0(x|x_0)P_t|0(x|x_0). (12) 2. Optimal Score as Projection. The internal objective ℐ(θ)I(θ) is minimized when sθs_θ is the orthogonal projection of the noisy score Δ+logPt|0 ^+ P_t|0 onto the subspace of functions independent of x0x_0. Solving the first-order condition yields the unique internal minimizer: sopt(x,t)=∑x0∈[q(x0)Pt|0(x|x0)pt(x)]Δ+logPt|0(x|x0).s_opt(x,t)= _x_0 [ q(x_0)P_t|0(x|x_0)p_t(x) ] ^+ P_t|0(x|x_0). (13) 3. Equivalence for Appreciable t. Let =(Δ+Pt|0)/Pt|0D=( ^+P_t|0)/P_t|0. For appreciable t, uniform ellipticity together with the regularity assumptions on b and σ imply standard parabolic smoothing estimates for the forward transition kernel. Consequently, Pt|0(xt|x0)P_t|0(x_t|x_0) is strictly positive and S-C2C^2 on limited regions, and the discrete derivative D is limited. Hence the internal Taylor expansion for log(1+ε) (1+ ) applies to ε=δx =δ x\,D. By the definition of the grid derivative and the internal Taylor theorem: Δ+logPt|0=log(1+δx)δx=−12δx2+(δx23). ^+ P_t|0= (1+δ xD)δ x=D- 12δ xD^2+O(δ x^2D^3). Substituting this into (13) and comparing with (12): sopt(x,t)=Δ+pt(x)pt(x)−∑x0q(x0|x)[12δx2+(δx23)].s_opt(x,t)= ^+p_t(x)p_t(x)- _x_0q(x_0|x) [ 12δ xD^2+O(δ x^2D^3) ]. (14) Since D is limited for appreciable t and δx≈0δ x≈ 0, the second term in (14) is strictly infinitesimal. Furthermore, with appreciable pt(x)p_t(x), Δ+pt(x)pt(x)≈Δ+logpt(x) ^+p_t(x)p_t(x)≈ ^+ p_t(x) since ptp_t is (at least) S-C2C^2 for appreciable t. Therefore, we have the pointwise proximity: sopt(x,t)≈Δ+logpt(x).s_opt(x,t)≈ ^+ p_t(x). 4. Orthogonal Projection and Convergence. We define δt′:=δt∩[τ,1]T _δ t:=T_δ t∩[τ,1] and the hyperfinite Hilbert space ℋ=L2(2×δt′,ν)H=L^2(G^2×T _δ t,ν) equipped with the internal measure ν(x0,x,t)=q(x0)Pt|0(x|x0)δtν(x_0,x,t)=q(x_0)P_t|0(x|x_0)δ t. Let ⊂ℋU be the subspace of internal functions independent of the initial state x0x_0. By the internal projection theorem, the function sopt(x,t)s_opt(x,t) defined in (13) is the unique orthogonal projection of the noisy score Δ+logPt|0 ^+ P_t|0 onto U. Since sθ∈s_θ , the internal Pythagorean identity holds exactly: ℐ(θ)=‖sθ−sopt‖ℋ2+‖sopt−Δ+logPt|0‖ℋ2.I(θ)=\|s_θ-s_opt\|_H^2+\|s_opt- ^+ P_t|0\|_H^2. (15) Given the hypothesis ℐ(θ)≈0I(θ)≈ 0, it follows from the non-negativity of the internal norms that ‖sθ−sopt‖ℋ2≈0\|s_θ-s_opt\|_H^2≈ 0. By summing over x0x_0 and using the internal marginal density pt(x)=∑x0q(x0)Pt|0(x|x0)p_t(x)= _x_0q(x_0)P_t|0(x|x_0), we define the marginal internal measure μ(x,t)=pt(x)δtμ(x,t)=p_t(x)δ t and obtain: ‖sθ−sopt‖L2(μ)2=∑t,xpt(x)|sθ(x,t)−sopt(x,t)|2δt≈0.\|s_θ-s_opt\|_L^2(μ)^2= _t,xp_t(x)|s_θ(x,t)-s_opt(x,t)|^2δ t≈ 0. (16) Let ℒ(μ)L(μ) be the Loeb measure constructed from μ. By standard properties of hyperfinite LpL^p spaces (cf. [22]), an internal function with an infinitesimal L2L^2 norm has a standard part that vanishes ℒ(μ)L(μ)-almost everywhere. Thus, sθ(x,t)≈sopt(x,t)s_θ(x,t)≈ s_opt(x,t) for ℒ(μ)L(μ)-almost all (x,t)(x,t). For all appreciable t∈δt _δ t, Step 3 establishes the infinitesimal estimate sopt(x,t)≈Δ+logpt(x)s_opt(x,t)≈ ^+ p_t(x). By the properties of grid derivatives of liftings (Proposition A.24), the standard part of the grid derivative coincides with the standard gradient: st(Δ+logpt(x))=∇logpt∘(x∘)st( ^+ p_t(x))=∇ p_ t( x). Combining these results, we have the Loeb-almost everywhere equality: st(sθ(x,t))=∇logpt∘(x∘),st(s_θ(x,t))=∇ p_ t( x), which completes the proof that the standard part of the learned score recovers the score of the continuous marginal distribution. ∎ 4.2 Hyperfinite sampling and generative dynamics With the internal score sθs_θ established as an infinitesimal approximation of the standard score on the appreciable time domain, we define the constructive procedure for generating a sample from the target distribution q. In the hyperfinite framework, this corresponds to an internal recursion starting from a state of maximal entropy (the terminal forward measure) and “crystallizing” into a data point. Let XtX_t satisfy the hyperfinite walk dynamics (Definition 2.3) under the assumptions of Theorem 3.3. Definition 4.2 (Hyperfinite Reverse Process on the Appreciable Domain). Let G be the hyperfinite grid and δtT_δ t be the infinitesimal time discretization of [0,1][0,1] with diffusion scaling. For any standard ϵ∈ℝ+ε ^+, let δtϵ=δt∩[0,1−ϵ]T_δ t^ε=T_δ t∩[0,1-ε]. The internal reverse sampling process Yss∈δtϵ\Y_s\_s _δ t^ε is defined by the recursion: Ys+δt=Ys+β^(s,Ys)δt+σ∗(1−s,Ys)ΔW~s,Y_s+δ t=Y_s+ β(s,Y_s)δ t+^*σ(1-s,Y_s) W_s, where the initial state Y0Y_0 is an internal random variable on G governed by the internal grid density p1p_1 inherited from the terminal state of the forward process at t=1t=1. The terms ΔW~s W_s denote internal independent increments of a hyperfinite Brownian motion satisfying int[ΔW~s]=O(δt1+α/2)E_int[ W_s]=O(δ t^1+α/2) and int[ΔW~sΔW~s⊤]=Iδt+O(δt1+α/2)E_int[ W_s W_s ]=Iδ t+O(δ t^1+α/2). The internal reverse drift field β β is evaluated at the reflected forward time t=1−s≥ϵt=1-s≥ε and defined component-wise via the learned score field sθs_θ as: β^i(s,y)=−bi∗(1−s,y)+∑j=1daij∗(1−s,y)sθ,j(y,1−s)+∑j=1dΔj+aij∗(1−s,y). β_i(s,y)=-^*b_i(1-s,y)+ _j=1^d^*a_ij(1-s,y)s_θ,j(y,1-s)+ _j=1^d _j^+^*a_ij(1-s,y). Theorem 4.3 (Generative Convergence). Assume the regularity conditions of Theorem 4.1. If sθs_θ is S-continuous on the support of the process and ℐ(θ)≈0I(θ)≈ 0, then for any standard ϵ∈ℝ+ε ^+, the standard-part process Y¯s=st(Ys) Y_s=st(Y_s) is a well-defined continuous diffusion for s∈[0,1−ϵ]s∈[0,1-ε]. Moreover, the standard terminal state extension defined by the standard limit Y¯1:=limϵ→0+Y¯1−ϵ Y_1:= _ε→ 0^+ Y_1-ε exists almost surely, and its law is identical to the initial standard data distribution q∘ q. Proof. 1. Appreciable Pathwise Identification. Let ϵ∈ℝ+ε ^+ be any strictly positive standard real number, and restrict the reverse process to the compact interval s∈δtϵs _δ t^ε as specified in Definition 4.2. For all such s, the forward time parameter t=1−s≥ϵt=1-s≥ε is strictly appreciable. By the assumption ℐ(θ)≈0I(θ)≈ 0, Theorem 4.1 applies pointwise ℒ(μ)L(μ)-almost everywhere on this domain, yielding sθ(y,1−s)≈Δ+logp1−s(y)s_θ(y,1-s)≈ ^+ p_1-s(y). Substituting this identity into the expression for β β shows that β^(s,y)≈β(1−s,y) β(s,y)≈β(1-s,y) holds for ℒ(μ)L(μ)-almost every (y,s)∈ℝfind∗×δtϵ(y,s)∈^*R_fin^d×T_δ t^ε, where β is the exact theoretical backward drift from Theorem 3.3. 2. Martingale Approximation on Compact Subsets. From the definitions and regularity assumptions, it is easy to check that the paths of YsY_s are S-continuous almost surely on s∈δtϵs _δ t^ε, making the standard-part process Y¯s=st(Ys) Y_s=st(Y_s) a well-defined, sample-continuous process on this compact interval. Since β^(s,y) β(s,y) is infinitesimally close to the true drift β(1−s,y)β(1-s,y) ℒ(μ)L(μ)-almost everywhere, the Loeb law induced by Y¯s Y_s solves the standard martingale problem dictated by the time-reflected coefficients (b^,a)( b,a). For any test function f∈Cc3(ℝd)f∈ C_c^3(R^d), the process M¯s=f(Y¯s)−f(Y¯0)−∫0sℒ1−urevf(Y¯u)u M_s=f( Y_s)-f( Y_0)- _0^sL^rev_1-uf( Y_u)du is a Loeb-a.s. continuous standard martingale over the interval s∈[0,1−ϵ]s∈[0,1-ε]. By the pathwise uniqueness of the solution to the standard martingale problem [28], the law of Y¯s Y_s on the space C([0,1−ϵ],ℝd)C([0,1-ε],R^d) coincides exactly with the law of the time-reversed forward process st(X1−s)st(X_1-s). 3. Weak Convergence at the Boundary. To extend the distributional identity to the terminal boundary s=1s=1, let ϕ∈Cb(ℝd)φ∈ C_b(R^d) be any standard bounded continuous test function. By the formal definition of the terminal state extension, Y¯1=limϵ→0+Y¯1−ϵ Y_1= _ε→ 0^+ Y_1-ε holds pathwise almost surely. Because ϕφ is a continuous mapping, this implies the almost sure convergence of random variables: limϵ→0+ϕ(Y¯1−ϵ)=ϕ(Y¯1) _ε→ 0^+φ( Y_1-ε)=φ( Y_1). By the Lebesgue dominated convergence theorem, this yields the weak convergence: [ϕ(Y¯1)]=limϵ→0+[ϕ(Y¯1−ϵ)].E[φ( Y_1)]= _ε→ 0^+E[φ( Y_1-ε)]. From Step 2, for any fixed standard ϵ>0ε>0, the law of Y¯1−ϵ Y_1-ε matches the law of st(Xϵ)st(X_ε), meaning [ϕ(Y¯1−ϵ)]=[ϕ(st(Xϵ))]E[φ( Y_1-ε)]=E[φ(st(X_ε))]. Finally, the standard forward diffusion st(Xt)st(X_t) is sample-continuous at its initial boundary t=0t=0 with st(X0)∼q∘st(X_0) q. Applying the standard dominated convergence theorem to the forward process yields: [ϕ(Y¯1)]=limϵ→0+[ϕ(st(Xϵ))]=[ϕ(st(X0))]=∫ℝdϕ(y)d(q∘)(y).E[φ( Y_1)]= _ε→ 0^+E[φ(st(X_ε))]=E[φ(st(X_0))]= _R^dφ(y)\,d( q)(y). Since this equality holds for all standard ϕ∈Cb(ℝd)φ∈ C_b(R^d), the uniqueness of weak limits guarantees that the law of Y¯1 Y_1 is exactly the initial standard data distribution q∘ q. ∎ Remark 2. It is standard in the diffusion models literature to suppose that the forward process XtX_t is constructed such that the terminal density p1p_1 is S-equivalent (see Appendix E) to the standard Gaussian density φ(⋅) (·) in the L1()L^1(G) sense. This assumption, denoted by p1≈φ∗p_1≈^* , is satisfied by standard noise schedules (cf. [27]), where the forward drift induces sufficient contraction to reach the stationary measure. In the hyperfinite setting, this ensures that the forgetting of the initial state data X0X_0 is achieved up to an infinitesimal error at t=1t=1, a result consistent with the convergence of internal heat kernels on hyperfinite grids [1]. Furthermore, because ℐ(θ)≈0I(θ)≈ 0, the L2L^2 error of the score estimator is infinitesimal on the appreciable time domain (see (16)), ensuring the generated paths do not deviate macroscopically from the theoretical time-reversed paths. The generative convergence established in Theorem 4.3 provides a rigorous hyperfinite justification for the sampling procedures central to modern score-based generative modeling [27]. While the standard-analysis approach typically relies on the stability theory of SDEs and the Stroock-Varadhan martingale problem [28] to prove that a learned score-drift yields the correct terminal distribution, our proof utilizes the discrete pathwise structure of the hyperfinite grid to simplify the transition. By leveraging the internal score equivalence (Theorem 4.1) on the appreciable time domain alongside the S-continuity of the hyperfinite walk, we demonstrate that an infinitesimal optimization error ℐ(θ)≈0I(θ)≈ 0 implies that the standard-part trajectories of the sampling recursion coincide for Loeb-almost all sample paths with the theoretical time-reversal derived in Section 3. This methodology follows the lifting principles for hyperfinite martingales established by Lindstrøm [18, 19] and the adapted distribution theory of Hoover and Keisler [12], confirming that the “crystallization” of noise into data is a structural consequence of the hyperfinite walk’s internal algebra. 5 Hyperfinite Girsanov theorem and maximum likelihood In the hyperfinite framework, the relationship between score-matching and log-likelihood is established by analyzing the Radon-Nikodym derivative of internal measures on the grid G. Unlike the continuum case, we do not require the noise increments to be strictly Gaussian, provided they satisfy the standard hyperfinite grid moment conditions. We specialize here to a state-independent, constant diffusion matrix σ such that a=σσ⊤≻0a=σ 0. Lemma 5.1 (Internal Girsanov Identity for Hyperfinite Walks). Let G be a hyperfinite grid with spacing δx=Cδtδ x=C δ t. Let ℙP be the internal path measure of a hyperfinite walk (cf. Definition 2.3) on the space Ω=δt =G^T_δ t corresponding to the internal drift coefficient b(t,x)b(t,x). Let ℚQ be the internal path measure constructed via the local exponential tilt qv(ωt,t)=pv(ωt,t)exp(sθ(ωt,t)⊤(v−b(t,Xt)δt))Zt(ωt),q_v( _t,t)=p_v( _t,t) (s_θ( _t,t) (v-b(t,X_t)δ t) )Z_t( _t), (17) where v=ΔXt=Xt+δt−Xt∈xv= X_t=X_t+δ t-X_t _x given Xt=xX_t=x. Letting ΔWt=v−b(t,Xt)δt W_t=v-b(t,X_t)δ t, assume that sθs_θ is a limited internal function on the support of the process. Then the internal log-density of the path measures satisfies: logdℚdℙ(ω)=∑t∈δt∖1[sθ(Xt,t)⊤ΔWt−12sθ(Xt,t)⊤asθ(Xt,t)δt]+ℰ(ω), dQdP(ω)= _t _δ t \1\ [s_θ(X_t,t) W_t- 12s_θ(X_t,t) as_θ(X_t,t)δ t ]+E(ω), (18) where ℰ(ω)≈0E(ω)≈ 0 uniformly for all paths ω whose states remain limited. Furthermore, if the local transition ratios take the exponential form in (17), the induced local drift under ℚQ satisfies ℚ[ΔXt∣Xt]δt=b+asθ+O(δt1/2). E_Q[ X_t X_t]δ t=b+as_θ+O(δ t^1/2). Conversely, among all internal transition laws satisfying this drift constraint, the unique minimizer of the local internal relative entropy DKL(q∥p)=∑v∈xqvlog(qvpv)D_KL(q p)= _v _xq_v \! ( q_vp_v ) has transition ratios of the exponential form (17) up to an S-equivalent factor. Proof. By the internal multiplicative structure of hyperfinite path measures, the global Radon-Nikodym derivative for an internal sample path ω∈Ωω∈ is given by the hyperfinite product of the local transition ratios: dℚdℙ(ω)=∏t∈δt∖1qΔXt(ωt,t)pΔXt(ωt,t). dQdP(ω)= _t _δ t \1\ q_ X_t( _t,t)p_ X_t( _t,t). Taking the internal logarithm converts the product into a hyperfinite sum: logdℚdℙ(ω)=∑t∈δt∖1log(qΔXt(ωt,t)pΔXt(ωt,t)). dQdP(ω)= _t _δ t \1\ \! ( q_ X_t( _t,t)p_ X_t( _t,t) ). (19) 1. Expansion of the Local Partition Function: From the definition of the exponential tilt in (17), the local normalization factor is Zt=∑v∈xpvexp(sθ⊤(v−bδt))=ℙ[exp(sθ⊤ΔWt)∣Xt].Z_t= _v _xp_v \! (s_θ (v-bδ t) )=E_P [ (s_θ W_t) X_t ]. Since ΔWt=O(δt) W_t=O( δ t) uniformly on limited states and sθs_θ is limited, sθ⊤ΔWt=O(δt).s_θ W_t=O( δ t). Applying the internal Taylor expansion with remainder gives exp(sθ⊤ΔWt)=1+sθ⊤ΔWt+12(sθ⊤ΔWt)2+16(sθ⊤ΔWt)3eξt, (s_θ W_t)=1+s_θ W_t+ 12(s_θ W_t)^2+ 16(s_θ W_t)^3e _t, where ξt _t lies between 0 and sθ⊤ΔWts_θ W_t. Taking conditional expectation under ℙP: Zt Z_t =1+sθ⊤ℙ[ΔWt∣Xt]+12sθ⊤ℙ[ΔWtΔWt⊤∣Xt]sθ+ℛt,3. =1+s_θ E_P[ W_t X_t]+ 12s_θ E_P[ W_t W_t X_t]s_θ+R_t,3. The remainder term satisfies ℛt,3=16ℙ[(sθ⊤ΔWt)3eξt∣Xt].R_t,3= 16E_P [(s_θ W_t)^3e _t X_t ]. Because sθs_θ is limited and the third moments are uniformly bounded by O(δt3/2)O(δ t^3/2), uniformly over limited states, ℛt,3=O(δt3/2).R_t,3=O(δ t^3/2). Using the moment assumptions, ℙ[ΔWt∣Xt]=0,E_P[ W_t X_t]=0, and ℙ[ΔWtΔWt⊤∣Xt]=aδt+O(δt3/2),E_P[ W_t W_t X_t]=aδ t+O(δ t^3/2), we obtain Zt=1+12sθ(Xt,t)⊤asθ(Xt,t)δt+O(δt3/2). Z_t=1+ 12s_θ(X_t,t) as_θ(X_t,t)δ t+O(δ t^3/2). Applying the logarithmic expansion log(1+x)=x−12x2+O(x3), (1+x)=x- 12x^2+O(x^3), with x=12sθ⊤asθδt+O(δt3/2),x= 12s_θ as_θδ t+O(δ t^3/2), gives logZt=12sθ(Xt,t)⊤asθ(Xt,t)δt+O(δt3/2). Z_t= 12s_θ(X_t,t) as_θ(X_t,t)δ t+O(δ t^3/2). 2. Pathwise Summation: Substituting (17) into (19) yields log(qΔXt(ωt,t)pΔXt(ωt,t))=sθ(Xt,t)⊤ΔWt−logZt. ( q_ X_t( _t,t)p_ X_t( _t,t) )=s_θ(X_t,t) W_t- Z_t. Using the expansion for logZt Z_t, log(qΔXt(ωt,t)pΔXt(ωt,t))=sθ(Xt,t)⊤ΔWt−12sθ(Xt,t)⊤asθ(Xt,t)δt+ηt, ( q_ X_t( _t,t)p_ X_t( _t,t) )=s_θ(X_t,t) W_t- 12s_θ(X_t,t) as_θ(X_t,t)δ t+ _t, where ηt=O(δt3/2), _t=O(δ t^3/2), uniformly over limited states. Summing over the hyperfinite timeline: logdℚdℙ(ω)=∑t∈δt∖1[sθ(Xt,t)⊤ΔWt−12sθ(Xt,t)⊤asθ(Xt,t)δt]+ℰ(ω), dQdP(ω)= _t _δ t \1\ [s_θ(X_t,t) W_t- 12s_θ(X_t,t) as_θ(X_t,t)δ t ]+E(ω), where ℰ(ω)=∑t∈δtηt.E(ω)= _t _δ t _t. Also, ℰ(ω)=O(δt1/2)≈0.E(ω)=O(δ t^1/2)≈ 0. This proves (18). 3. Verification of the Induced Local Drift (Forward Direction): By definition, ΔXt=b(t,Xt)δt+ΔWt. X_t=b(t,X_t)δ t+ W_t. Taking conditional expectation under ℚQ: ℚ[ΔXt∣Xt]=b(t,Xt)δt+ℚ[ΔWt∣Xt].E_Q[ X_t X_t]=b(t,X_t)δ t+E_Q[ W_t X_t]. (20) Using the Radon-Nikodym ratio, ℚ[ΔWt∣Xt]=1Ztℙ[ΔWtexp(sθ⊤ΔWt)∣Xt].E_Q[ W_t X_t]= 1Z_tE_P [ W_t (s_θ W_t) X_t ]. (21) Expanding the exponential: ΔWtesθ⊤ΔWt W_te^s_θ W_t =ΔWt(1+sθ⊤ΔWt+12(sθ⊤ΔWt)2+O(δt3/2)) = W_t (1+s_θ W_t+ 12(s_θ W_t)^2+O(δ t^3/2) ) =ΔWt+ΔWtΔWt⊤sθ+O(δt3/2). = W_t+ W_t W_t s_θ+O(δ t^3/2). Taking conditional expectation: ℙ[ΔWtesθ⊤ΔWt∣Xt] _P [ W_te^s_θ W_t X_t ] =ℙ[ΔWt∣Xt]+ℙ[ΔWtΔWt⊤∣Xt]sθ+O(δt3/2) =E_P[ W_t X_t]+E_P[ W_t W_t X_t]s_θ+O(δ t^3/2) =asθδt+O(δt3/2). =as_θδ t+O(δ t^3/2). Substituting into (21), ℚ[ΔWt∣Xt]=asθδt+O(δt3/2)1+12sθ⊤asθδt+O(δt3/2).E_Q[ W_t X_t]= as_θδ t+O(δ t^3/2)1+ 12s_θ as_θδ t+O(δ t^3/2). Using (1+x)−1=1−x+O(x2),(1+x)^-1=1-x+O(x^2), with x=O(δt)x=O(δ t): ℚ[ΔWt∣Xt] _Q[ W_t X_t] =(asθδt+O(δt3/2))(1−12sθ⊤asθδt+O(δt3/2)) = (as_θδ t+O(δ t^3/2) ) (1- 12s_θ as_θδ t+O(δ t^3/2) ) =asθδt+O(δt3/2). =as_θδ t+O(δ t^3/2). Substituting back into (20), ℚ[ΔXt∣Xt]=(b+asθ)δt+O(δt3/2).E_Q[ X_t X_t]=(b+as_θ)δ t+O(δ t^3/2). Dividing by δtδ t: ℚ[ΔXt∣Xt]δt=b+asθ+O(δt1/2)≈b+asθ. E_Q[ X_t X_t]δ t=b+as_θ+O(δ t^1/2)≈ b+as_θ. 4. Converse Direction and Uniqueness: Suppose an internal transition law q satisfies ℚ[ΔXt∣Xt]δt=b+asθ+O(δt1/2). E_Q[ X_t X_t]δ t=b+as_θ+O(δ t^1/2). Consider the constrained optimization problem minqDKL(q∥p)=minq∑v∈xqvlog(qvpv), _qD_KL(q p)= _q _v _xq_v \! ( q_vp_v ), (22) subject to ∑vqv=1 and ∑vqv=(b+asθ)δt+O(δt3/2). _vq_v=1 and _vvq_v=(b+as_θ)δ t+O(δ t^3/2). The functional q↦∑vqvlog(qv/pv)q _vq_v (q_v/p_v) is strictly convex on the probability simplex, while the constraints are affine. Hence the constrained minimizer is unique. Introduce Lagrange multipliers λ0∈ℝ∗ _0∈^*R and λ∈ℝd∗λ∈^*R^d: ℒ(q,λ0,λ) (q, _0,λ) =∑vqvlog(qvpv)−λ0(∑vqv−1) = _vq_v \! ( q_vp_v )- _0 ( _vq_v-1 ) −λ⊤(∑vqv−ℚ[ΔXt∣Xt]). -λ ( _vvq_v-E_Q[ X_t X_t] ). Setting ∂ℒ/∂qv=0 /∂ q_v=0 gives log(qvpv)+1−λ0−λ⊤v=0, \! ( q_vp_v )+1- _0-λ v=0, hence qv=pvexp(λ⊤v)Z~t,q_v=p_v (λ v) Z_t, where Z~t=exp(1−λ0). Z_t= (1- _0). Writing v=bδt+ΔWt,v=bδ t+ W_t, and cancelling the deterministic factor exp(λ⊤bδt) (λ bδ t), qv=pvexp(λ⊤ΔWt)Zt,q_v=p_v (λ W_t)Z_t, where Zt=ℙ[exp(λ⊤ΔWt)∣Xt].Z_t=E_P[ (λ W_t) X_t]. Repeating the same Taylor expansion argument as above yields ℚ[ΔWt∣Xt]=aλδt+O(δt3/2).E_Q[ W_t X_t]=aλδ t+O(δ t^3/2). Since ℚ[ΔXt∣Xt]=bδt+ℚ[ΔWt∣Xt],E_Q[ X_t X_t]=bδ t+E_Q[ W_t X_t], matching the target drift gives asθδt+O(δt3/2)=aλδt+O(δt3/2).as_θδ t+O(δ t^3/2)=aλδ t+O(δ t^3/2). Because a is uniformly nondegenerate, λ=sθ+O(δt1/2).λ=s_θ+O(δ t^1/2). Substituting this back into the transition ratio: qvpv q_vp_v =exp((sθ+O(δt1/2))⊤ΔWt)Zt = \! ((s_θ+O(δ t^1/2)) W_t )Z_t =exp(sθ⊤ΔWt)Ztexp(O(δt)). = (s_θ W_t)Z_t (O(δ t)). Since exp(O(δt))=1+O(δt)≈1, (O(δ t))=1+O(δ t)≈ 1, we conclude qvpv≈exp(sθ⊤ΔWt)Zt. q_vp_v≈ (s_θ W_t)Z_t. This establishes the converse characterization and completes the proof. ∎ Remark 3. To identify a canonical change of measure on the hyperfinite grid, we restrict attention to internal transition kernels that preserve the local second-moment structure of the reference dynamics. More precisely, if the reference process ℙP satisfies the infinitesimal moment condition ℙ[ΔXtΔXt⊤∣Xt]=aδt+O(δt3/2),E_P[ X_t X_t X_t]=a\,δ t+O(δ t^3/2), then any admissible perturbed transition kernel defining an internal path measure ℚQ must satisfy this same scaling. This invariance ensures that the quadratic variation of the induced continuous-time Loeb diffusion remains unaltered under the change of measure. The first-order drift constraint alone does not uniquely determine the local transition probabilities; indefinitely many internal perturbations can yield the target drift while differing at higher orders. Such arbitrary variations alter the higher conditional moments of the grid increments, potentially destabilizing the fine-scale sample path behavior or violating the nonstandard CFL condition. To isolate a distinguished, minimal perturbation, we impose the internal minimum relative entropy condition minqDKL(q∥p)=minq∑v∈xqvlogqvpv, _qD_KL(q p)= _q _v _xq_v q_vp_v, subject to the normalization and target drift constraints. The corresponding Euler–Lagrange equations uniquely yield the exponential tilt qv=pvexp(sθ(Xt,t)⊤(v−b(t,Xt)δt))Zt,q_v=p_v (s_θ(X_t,t) (v-b(t,X_t)δ t) )Z_t, which serves as the unique information-theoretic minimizer preserving positivity and internal absolute continuity with respect to the reference walk. This optimization framework establishes that the exponential tilt is not an arbitrary modeling heuristic, but rather the unique configuration required for a rigorous hyperfinite Girsanov theorem. Because the logarithm of the resulting partition function admits the pathwise expansion logZt=12sθ⊤asθδt+O(δt3/2), Z_t= 12s_θ as_θ\,δ t+O(δ t^3/2), it naturally generates the quadratic compensation term in the global Radon–Nikodym derivative upon hyperfinite summation. Consequently, the exponential tilt is uniquely compatible with the prescribed drift and the underlying geometry of the hyperfinite grid. The hyperfinite Girsanov identity established in Lemma 5.1 shows that the Radon–Nikodym derivative between internal path measures arises directly from the transition algebra of the hyperfinite grid. Unlike the classical continuous-time proof of Girsanov’s theorem, which is formulated through stochastic integration and martingale criteria such as the Novikov condition, the hyperfinite construction proceeds entirely through finite-step transition ratios and hyperfinite summation. As has been the case throughout, our internal approach aligns with the foundational frameworks of hyperfinite diffusions (cf. [16]). More precisely, the change of measure is implemented locally through the exponential tilt of the reference transition probabilities, and the logarithm of the resulting density ratio is expanded using the internal Taylor expansion of the partition function ZtZ_t. The accumulated pathwise exponent is obtained by summing over the hyperfinite timeline, and the remainder is infinitesimal under our assumptions. This structural decomposition mirrors applications of nonstandard measure theory to stochastic control [1]. An internal perspective has also been leveraged by Cutland [9] and Osswald [25] to extend Girsanov transformations to anticipative settings. In our case, the internal logarithmic density ratio provides a rigorous hyperfinite representation of the continuous-time drift-energy functional associated with score-based diffusion dynamics. Theorem 5.2 (Hyperfinite Surrogate Likelihood Decomposition). Let ⊂ℝd∗G⊂^*R^d be a hyperfinite spatial lattice and consider the grid dynamics in Definition 2.3 with constant symmetric positive-definite diffusion tensor a=σσ⊤≻0.a=σ 0. Let ℙ∗P^* denote the internal path measure of the true backward diffusion process on the hyperfinite timeline δtT_δ t, initialized at time t=1t=1 with terminal forward distribution p1p_1, and governed by backward drift b¯fwd(t,x)=−b(t,x)+asopt(x,t). b_fwd(t,x)=-b(t,x)+a\,s_opt(x,t). Let ℚQ denote the parameterized backward generative path measure initialized with the same terminal distribution p1p_1, with drift b¯θ(t,x)=−b(t,x)+asθ(x,t) b_θ(t,x)=-b(t,x)+a\,s_θ(x,t) inducing the distribution qθq_θ at t=0t=0, where both transition laws minimise the corresponding relative entropy as in Lemma 5.1, and b¯θ,b¯fwd b_θ, b_fwd denote the specializations of the drift from Definition 4.2 and (6), respectively.. With limited sθs_θ, assume regularity of b,σb,σ implying limitedness of the score field sopts_opt on its support so that the internal Fisher score-matching functional ℐa(θ)I_a(θ) is limited: ℐa(θ)=ℙ∗[∑t∈δt∖012‖sθ(Xt,t)−sopt(Xt,t)‖a2δt],I_a(θ)=E_P^* [ _t _δ t \0\ 12\|s_θ(X_t,t)-s_opt(X_t,t)\|_a^2\,δ t ], and define the surrogate likelihood functional ℒsurr(θ)=ℙ∗[logqθ(X0)].L_surr(θ)=E_P^*[ q_θ(X_0)]. Then: ℒsurr(θ)=H0−ℐa(θ)+ℛbridge(θ)+O(δt1/2), _surr(θ)=H_0-I_a(θ)+R_bridge(θ)+O(δ t^1/2), (23) where H0=ℙ∗[logℙ∗(X0)]H_0=E_P^*[ ^*(X_0)] is independent of θ, and ℛbridge(θ)=X0∼ℙ∗[KL(ℙ∗(⋅|X0)∥ℚ(⋅|X0))]≥0.R_bridge(θ)=E_X_0 ^* [D_KL (P^*(·|X_0)\|Q(·|X_0) ) ]≥ 0. Consequently, ℒsurr(θ)≥H0−ℐa(θ)+O(δt1/2).L_surr(θ)≥ H_0-I_a(θ)+O(δ t^1/2). (24) Applying the standard part map yields the exact standard inequality st(ℒsurr(θ))≥st(H0−ℐa(θ)).st(L_surr(θ)) (H_0-I_a(θ)). Proof. Step 1: Global Hyperfinite Radon–Nikodym Expansion Let Ω=δt =G^T_δ t denote the internal hyperfinite trajectory space. The unconditional path measures factorize as dℙ∗(ω) dP^*(ω) =p1(X1)∏t∈δt∖0pΔXt∗(Xt,t), =p_1(X_1) _t _δ t \0\p^*_ X_t(X_t,t), dℚ(ω) dQ(ω) =p1(X1)∏t∈δt∖0qΔXt(Xt,t), =p_1(X_1) _t _δ t \0\q_ X_t(X_t,t), where ΔXt=Xt−δt−Xt. X_t=X_t-δ t-X_t. Since both path measures share the identical initialization p1(X1)p_1(X_1), cancellation yields logdℙ∗dℚ(ω)=∑tlog(pΔXt∗(Xt,t)qΔXt(Xt,t)). dP^*dQ(ω)= _t ( p^*_ X_t(X_t,t)q_ X_t(X_t,t) ). (25) Define the background backward increment ΔWt∘=ΔXt+b(t,Xt)δt. W_t = X_t+b(t,X_t)δ t. By Lemma 5.1, the local transition kernels admit the exponential tilt representations pΔXt∗(Xt,t) p^*_ X_t(X_t,t) =pvexp(sopt(Xt,t)⊤ΔWt∘)Zt∗, =p_v (s_opt(X_t,t) W_t )Z_t^*, qΔXt(Xt,t) q_ X_t(X_t,t) =pvexp(sθ(Xt,t)⊤ΔWt∘)Zt. =p_v (s_θ(X_t,t) W_t )Z_t. Substituting these formulas into (25) cancels the baseline walk kernel pvp_v: logdℙ∗dℚ(ω) dP^*dQ(ω) =∑t[(sopt−sθ)⊤ΔWt∘+logZt−logZt∗]. = _t [(s_opt-s_θ) W_t + Z_t- Z_t^* ]. (26) Define the true backward martingale increment under ℙ∗P^*: ΔWt∗=ΔXt−b¯fwd(t,Xt)δt=ΔWt∘−asopt(Xt,t)δt. W_t^*= X_t- b_fwd(t,X_t)δ t= W_t -a\,s_opt(X_t,t)δ t. Equivalently, ΔWt∘=ΔWt∗+asopt(Xt,t)δt. W_t = W_t^*+a\,s_opt(X_t,t)δ t. Substituting this identity into (26) gives logdℙ∗dℚ(ω) dP^*dQ(ω) =∑t(sopt−sθ)⊤ΔWt∗+∑t(sopt−sθ)⊤asoptδt = _t(s_opt-s_θ) W_t^*+ _t(s_opt-s_θ) a\,s_opt\,δ t +∑t(logZt−logZt∗). + _t( Z_t- Z_t^*). (27) Using the partition-function expansion from Lemma 5.1, logZt=12‖sθ(Xt,t)‖a2δt+O(δt3/2), Z_t= 12\|s_θ(X_t,t)\|_a^2δ t+O(δ t^3/2), and similarly, logZt∗=12‖sopt(Xt,t)‖a2δt+O(δt3/2). Z_t^*= 12\|s_opt(X_t,t)\|_a^2δ t+O(δ t^3/2). Therefore, logZt−logZt∗=12(‖sθ‖a2−‖sopt‖a2)δt+O(δt3/2). Z_t- Z_t^*= 12 (\|s_θ\|_a^2-\|s_opt\|_a^2 )δ t+O(δ t^3/2). Combining the deterministic terms yields (sopt−sθ)⊤asopt+12‖sθ‖a2−12‖sopt‖a2=12‖sopt−sθ‖a2.(s_opt-s_θ) as_opt+ 12\|s_θ\|_a^2- 12\|s_opt\|_a^2= 12\|s_opt-s_θ\|_a^2. Hence, logdℙ∗dℚ(ω)=∑t(sopt−sθ)⊤ΔWt∗+∑t12‖sopt−sθ‖a2δt+ℰ(ω), dP^*dQ(ω)= _t(s_opt-s_θ) W_t^*+ _t 12\|s_opt-s_θ\|_a^2δ t+E(ω), (28) where ℰ(ω)=∑tO(δt3/2),E(ω)= _tO(δ t^3/2), which satisfies ℰ(ω)=O(δt1/2).E(ω)=O(δ t^1/2). Step 2: Evaluation of the Global Relative Entropy Taking expectation under ℙ∗P^* in (28) gives KL(ℙ∗∥ℚ) _KL(P^*\|Q) =∑tℙ∗[(sopt−sθ)⊤ΔWt∗] = _tE_P^* [(s_opt-s_θ) W_t^* ] +ℙ∗[∑t12‖sopt−sθ‖a2δt]+O(δt1/2). +E_P^* [ _t 12\|s_opt-s_θ\|_a^2δ t ]+O(δ t^1/2). (29) Let ℱt∗=σ(Xs:s≥t)F_t^*=σ(X_s:s≥ t) denote the backward filtration. Since ΔWt∗ W_t^* is the backward martingale increment under ℙ∗P^*, ℙ∗[ΔWt∗∣ℱt∗]=0.E_P^*[ W_t^* _t^*]=0. Because sopt(Xt,t)s_opt(X_t,t) and sθ(Xt,t)s_θ(X_t,t) are ℱt∗F_t^*-measurable, the tower property gives ℙ∗[(sopt−sθ)⊤ΔWt∗] _P^* [(s_opt-s_θ) W_t^* ] =0. =0. Therefore, KL(ℙ∗∥ℚ)=ℐa(θ)+O(δt1/2).D_KL(P^*\|Q)=I_a(θ)+O(δ t^1/2). (30) Step 3: Boundary Disintegration Identity Disintegrating the unconditional path measures with respect to the terminal state X0X_0 gives dℙ∗(ω) dP^*(ω) =ℙ∗(X0)dℙ∗(ω|X0), =P^*(X_0)\,dP^*(ω|X_0), dℚ(ω) dQ(ω) =qθ(X0)dℚ(ω|X0). =q_θ(X_0)\,dQ(ω|X_0). Consequently, logdℙ∗dℚ=logℙ∗(X0)−logqθ(X0)+logdℙ∗(⋅|X0)dℚ(⋅|X0). dP^*dQ= ^*(X_0)- q_θ(X_0)+ dP^*(·|X_0)dQ(·|X_0). (31) Taking expectation under ℙ∗P^* yields KL(ℙ∗∥ℚ) _KL(P^*\|Q) =ℙ∗[logℙ∗(X0)]−ℙ∗[logqθ(X0)] =E_P^*[ ^*(X_0)]-E_P^*[ q_θ(X_0)] +X0∼ℙ∗[KL(ℙ∗(⋅|X0)∥ℚ(⋅|X0))]. +E_X_0 ^* [D_KL (P^*(·|X_0)\|Q(·|X_0) ) ]. (32) Using the definitions H0=ℙ∗[logℙ∗(X0)],H_0=E_P^*[ ^*(X_0)], ℒsurr(θ)=ℙ∗[logqθ(X0)],L_surr(θ)=E_P^*[ q_θ(X_0)], and ℛbridge(θ)=X0∼ℙ∗[KL(ℙ∗(⋅|X0)∥ℚ(⋅|X0))],R_bridge(θ)=E_X_0 ^* [D_KL (P^*(·|X_0)\|Q(·|X_0) ) ], equation (32) becomes KL(ℙ∗∥ℚ)=H0−ℒsurr(θ)+ℛbridge(θ).D_KL(P^*\|Q)=H_0-L_surr(θ)+R_bridge(θ). (33) Combining (30) and (33) yields H0−ℒsurr(θ)+ℛbridge(θ)=ℐa(θ)+O(δt1/2).H_0-L_surr(θ)+R_bridge(θ)=I_a(θ)+O(δ t^1/2). Rearranging gives ℒsurr(θ)=H0−ℐa(θ)+ℛbridge(θ)+O(δt1/2).L_surr(θ)=H_0-I_a(θ)+R_bridge(θ)+O(δ t^1/2). Since relative entropy is nonnegative, ℛbridge(θ)≥0.R_bridge(θ)≥ 0. Therefore, ℒsurr(θ)≥H0−ℐa(θ)+O(δt1/2).L_surr(θ)≥ H_0-I_a(θ)+O(δ t^1/2). Finally, applying the standard part operator and using O(δt1/2)≈0O(δ t^1/2)≈ 0 gives st(ℒsurr(θ))≥st(H0−ℐa(θ)).st(L_surr(θ)) (H_0-I_a(θ)). This completes the proof. ∎ The equivalence established in Theorem 5.2 provides a rigorous structural explanation for the effectiveness of score-based training and shows that it is not merely an approximation to maximum likelihood, but a pathwise realization of it within the hyperfinite probabilistic framework. Standard derivations in the diffusion-model literature [11, 27] typically relate score matching to likelihood optimization through variational lower bounds (ELBOs), Gaussian perturbation identities, or continuum integration-by-parts arguments. In contrast, the hyperfinite framework shows that the connection emerges directly from the algebra of the grid transition structure itself, where calculations are performed through finite-step transition ratios and internal summation, with the continuum limit recovered through the Loeb measure construction. More precisely, the internal Girsanov identity of Lemma 5.1 expresses the logarithmic density ratio between path measures as a hyperfinite sum of local transition contributions. After expanding the local exponential tilt and summing over the hyperfinite timeline, the stochastic first-order terms appear as martingale increments with respect to the backward filtration. Their contribution vanishes exactly under the internal expectation by the tower property and the martingale property of hyperfinite increments [18, 19]. The remaining deterministic contribution is precisely the quadratic score discrepancy functional. Consequently, the Fisher divergence does not arise as a heuristic surrogate for likelihood, but as the quadratic compensation term in the hyperfinite change-of-measure formula. In particular, after taking standard parts, minimizing the score-matching objective maximizes a lower bound on the surrogate likelihood. Since a is a constant positive-definite matrix, there exist constants 0<λmin≤λmax0< _ ≤ _ such that λmin‖v‖2≤‖v‖a2≤λmax‖v‖2, _ \|v\|_G^2≤\|v\|_a^2≤ _ \|v\|_G^2, implying norm equivalence. As in the proof of Theorem 4.1, for an arbitrary real ϵ>0ε>0, consider a representative in δtT_δ t given by τ≈ϵτ≈ε and define δt′:=δt∩[τ,1]T _δ t:=T_δ t∩[τ,1]. Lemma 5.3 (Internal Score Decomposition). Let q be the initial distribution and let Pt|0P_t|0 denote the internal forward transition kernel on the hyperfinite grid G. Assume the regularity hypotheses of Theorem 4.1. For appreciable times t in δt′T _δ t, standard parabolic regularity and uniform ellipticity imply that the logarithmic grid derivatives Δ+logPt|0 ^+ P_t|0 are limited and internally square-integrable with respect to the internal measure q(x0)Pt|0(xt|x0)δt.q(x_0)P_t|0(x_t|x_0)δ t. Let sopt(xt,t)=∑x0∈[q(x0)Pt|0(xt|x0)pt(xt)]Δ+logPt|0(xt|x0)s_opt(x_t,t)= _x_0 [ q(x_0)P_t|0(x_t|x_0)p_t(x_t) ] ^+ P_t|0(x_t|x_0) denote the optimal internal score projection from Theorem 4.1. Then the internal DSM objective admits the orthogonal decomposition ℐ(θ)=12∑t∈δt′pt[‖sθ−sopt‖2]δt+C′,I(θ)= 12 _t _δ tE_p_t [\|s_θ-s_opt\|_G^2 ]δ t+C , where C′=12∑t∈δt′q,Pt|0[‖sopt−Δ+logPt|0‖2]δtC = 12 _t _δ tE_q,P_t|0 [\|s_opt- ^+ P_t|0\|_G^2 ]δ t is a limited hyperreal independent of θ. Proof. For each fixed t∈δt′t _δ t, expand the square appearing in the DSM objective: ‖sθ−Δ+logPt|0‖2 \|s_θ- ^+ P_t|0\|_G^2 =‖sθ−sopt‖2+‖sopt−Δ+logPt|0‖2 =\|s_θ-s_opt\|_G^2+\|s_opt- ^+ P_t|0\|_G^2 +2⟨sθ−sopt,sopt−Δ+logPt|0⟩. +2 s_θ-s_opt,\,\,s_opt- ^+ P_t|0 _G. We now take the expectation with respect to the internal measure q(x0)Pt|0(xt|x0).q(x_0)P_t|0(x_t|x_0). Applying the internal tower property to the cross-term, we condition on XtX_t. Since sθ(Xt,t)s_θ(X_t,t) and sopt(Xt,t)s_opt(X_t,t) are measurable with respect to the σ-algebra generated by XtX_t, they factor out of the inner conditional expectation: q,Pt|0[⟨sθ−sopt,sopt−Δ+logPt|0⟩] _q,P_t|0 [ s_θ-s_opt,\,\,s_opt- ^+ P_t|0 _G ] =pt[⟨sθ−sopt,q,Pt|0[sopt−Δ+logPt|0|Xt]⟩]. =E_p_t [ s_θ-s_opt,\,\,E_q,P_t|0 [s_opt- ^+ P_t|0\, |\,X_t ] _G ]. By the definition of the optimal score projection, q,Pt|0[Δ+logPt|0∣Xt]=sopt(Xt,t)E_q,P_t|0[ ^+ P_t|0 X_t]=s_opt(X_t,t). Because sopt(Xt,t)s_opt(X_t,t) is constant given XtX_t, the inner conditional expectation vanishes identically: q,Pt|0[sopt−Δ+logPt|0|Xt]=sopt(Xt,t)−sopt(Xt,t)=0.E_q,P_t|0 [s_opt- ^+ P_t|0\, |\,X_t ]=s_opt(X_t,t)-s_opt(X_t,t)=0. Hence, q,Pt|0[‖sθ−Δ+logPt|0‖2] _q,P_t|0 [\|s_θ- ^+ P_t|0\|_G^2 ] =pt[‖sθ−sopt‖2]+q,Pt|0[‖sopt−Δ+logPt|0‖2]. =E_p_t [\|s_θ-s_opt\|_G^2 ]+E_q,P_t|0 [\|s_opt- ^+ P_t|0\|_G^2 ]. Multiplying by 12δt 12δ t and summing over t∈δt′t _δ t yields ℐ(θ) (θ) =12∑t∈δt′pt[‖sθ−sopt‖2]δt = 12 _t _δ tE_p_t [\|s_θ-s_opt\|_G^2 ]δ t +12∑t∈δt′q,Pt|0[‖sopt−Δ+logPt|0‖2]δt. + 12 _t _δ tE_q,P_t|0 [\|s_opt- ^+ P_t|0\|_G^2 ]δ t. The second term depends only on the forward process and the optimal projection sopts_opt, and is therefore independent of θ. Defining this term to be C′C completes the proof. ∎ Remark 4. The restriction to the truncated timeline δt′=δt∩[τ,1],τ≈ϵ>0,T _δ t=T_δ t∩[τ,1], τ≈ε>0, excludes the singular initial layer near t=0t=0, where the logarithmic grid derivatives of the transition kernel need not remain uniformly limited. This restriction does not affect the standard-part conclusions of the theory under some conditions: Indeed, if sθs_θ and sopts_opt are limited internal functions, since the interval [0,τ][0,τ] has arbitrarily small Loeb measure, its contribution to the integrated quadratic objective is (ϵ)O(ε). Consequently, the truncated objective converges to the full-timeline counterpart as ϵ↓0ε 0, and both have the same standard-part limit. The decomposition established in Lemma 5.3 provides a rigorous justification for the use of DSM as a surrogate for likelihood optimization. By the internal Pythagorean theorem and the orthogonality of the conditional projection sopts_opt, the DSM objective decomposes into the projection error 12∑t∈δt′pt[‖sθ−sopt‖2]δt 12 _t _δ tE_p_t\! [\|s_θ-s_opt\|_G^2 ]δ t plus a θ-independent residual term C′C . Since the diffusion tensor a is uniformly positive definite, the norms ∥⋅∥\|·\|_G and ∥⋅∥a\|·\|_a are equivalent, so ℐ(θ)I(θ) and the corresponding pathwise score-matching functional induce equivalent minimization problems. This recovers, within the hyperfinite framework, the structural relationship between DSM and Fisher divergence established by Vincent [31]. Moreover, on appreciable time intervals, uniform ellipticity and parabolic regularity imply that the transition densities are S-smooth and possess limited logarithmic grid derivatives on the finite part of the Loeb space, ensuring that the residual term C′C is a limited hyperreal. Taken together, Theorem 4.1, Lemma 5.3, and Theorem 5.2 establish a variational connection between DSM and surrogate likelihood maximization on the hyperfinite grid. On every appreciable time interval [τ,1][τ,1], Theorem 4.1 identifies the DSM minimizer with the optimal internal projection field sopts_opt, while Lemma 5.3 shows that the DSM objective differs from the corresponding pathwise score-matching functional only by a θ-independent constant. Theorem 5.2 then relates this pathwise score-matching functional to the surrogate likelihood through a non-negative bridge divergence term. Consequently, DSM training admits a likelihood-based interpretation on appreciable time intervals. Extending this identification to the full timeline requires additional control of the infinitesimal initial layer [0,τ)[0,τ) as τ→0τ→ 0, which may be obtained under further regularity or integrability assumptions on the forward diffusion and its associated densities (also see Remark 4). 6 Local expansion of a source–based hyperfinite master equation To keep the exposition and notation at an elementary level, we focus on the one-dimensional case here, even though the result may be extended to the general d-dimensional case we have been discussing. Fix a finite point (t,x)(t,x) and set P=p∗P= p. The internal Euler increment is h=b∗(t,x)δt+σ∗(t,x)δtξh= b(t,x)δ t+ σ(t,x) δ tξ, where ξ is an internal random variable satisfying int[ξ]=0,int[ξ2]=1,int[ξ3]=0,int[ξ4]=κ<∞.E_int[ξ]=0, _int[ξ^2]=1, _int[ξ^3]=0, _int[ξ^4]=κ<∞. The source-based master equation is (cf. [23]): P(t+δt,x)=ξ[P(t,x−h)].P(t+δ t,x)=E_ξ [P(t,x-h) ]. (34) 6.1 Operator consistency Define the internal frozen operator =−b∗∂+a∗2∂2A=- b∂+ a2∂^2, where ∂≡∂x∂≡ _x, the internal spatial derivative. We have the exact internal identity: 2P=b2∗∂2P−b∗a∗∂3P+a2∗4∂4P.A^2P= b^2∂^2P- b a∂^3P+ a^24∂^4P. (35) The operator =−b∗(t,x)∂+a∗(t,x)2∂2A=- b(t,x)∂+ a(t,x)2∂^2 is the frozen-coefficient representation of the adjoint generator at the spacetime point (x,t)(x,t). For S-regular liftings P of a standard density p, one has st(P)=ℒt,fr∗p,st(AP)=L_t,fr^*p, where ℒt,fr∗=−b(t,x)∂x+a(t,x)2∂x2L_t,fr^*=-b(t,x) _x+ a(t,x)2 _x^2 denotes the Fokker–Planck generator with coefficients frozen at (t,x)(t,x). Theorem 6.1 (Local Second-Order Consistency). In addition to the conditions on b,σb,σ from Definition 2.3, assume that b and σ are C1+α2,3+αC^1+ α2,3+α and C1+α2,4+αC^1+ α2,4+α, respectively, for some α∈(0,1)α∈(0,1). For any finite (t,x)(t,x) with t>0t>0, the master equation update satisfies: st(P(t+δt,x)−(P+δtP+δt222P)δt2)=κ−324a(t∘,x∘)2∂4p∂x4(t∘,x∘).st ( P(t+δ t,x)- (P+δ tAP+ δ t^22A^2P )δ t^2 )= κ-324a( t, x)^2 ∂^4p∂ x^4( t, x). (36) Proof. By the internal Taylor theorem and (34): P(t+δt,x) P(t+δ t,x) =P+δtP+δt2(b2∗2∂2P−b∗a∗2∂3P+κa2∗24∂4P)+o(δt2), =P+δ tAP+δ t^2 ( b^22∂^2P- b a2∂^3P+ κ a^224∂^4P )+o(δ t^2), where we used ξ[h] _ξ[h] =b∗δt, = bδ t, ξ[h2] _ξ[h^2] =a∗δt+b2∗δt2, = aδ t+ b^2δ t^2, ξ[h3] _ξ[h^3] =3b∗a∗δt2+O(δt3), =3 b aδ t^2+O(δ t^3), ξ[h4] _ξ[h^4] =κa2∗δt2+O(δt3). =κ a^2δ t^2+O(δ t^3). Using (35) and subtracting the semigroup expansion: P(t+δt,x) P(t+δ t,x) −(P+δtP+δt222P) - (P+δ tAP+ δ t^22A^2P ) =δt2[(b2∗2∂2P−b∗a∗2∂3P+κa2∗24∂4P) =δ t^2 [ ( b^22∂^2P- b a2∂^3P+ κ a^224∂^4P ) −(b2∗2∂2P−b∗a∗2∂3P+3a2∗24∂4P)]+o(δt2) - ( b^22∂^2P- b a2∂^3P+ 3 a^224∂^4P ) ]+o(δ t^2) =δt2(κ−324a2∗∂4P)+o(δt2). =δ t^2 ( κ-324 a^2∂^4P )+o(δ t^2). Dividing by δt2δ t^2 and taking the standard part, noting that st(∂4P)=∂4pst(∂^4P)=∂^4p due to S-smoothness resulting from parabolic regularity, yields the result. ∎ Remark 5 (Significance of Second-Order Consistency). Theorem 36 shows that, for the source–based hyperfinite master equation considered here, local second–order consistency with the diffusion semigroup is achieved precisely when the driving increment has kurtosis κ=3κ=3. Any mismatch in the fourth moment produces a non–vanishing second–order correction proportional to the fourth spatial derivative of the density. The resulting term κ−324a2∂x4p κ-324a^2 _x^4p may be interpreted as a numerical dispersion error induced solely by the fourth–order statistics of the increment. Consequently, while first–order (Euler–scale) consistency requires only matching the first two moments, second–order weak consistency additionally depends on the kurtosis of the underlying noise source. Rademacher increments do not satisfy this condition, whereas the following non–Gaussian distribution does: ΔWt=δtξt, W_t= δ t\, _t, where the ξt _t are independent and satisfy ℙ(ξt=±3)=16,ℙ(ξt=0)=23.P( _t=± 3)= 16, ( _t=0)= 23. While Theorem 36 assumes locally frozen coefficients, it identifies the leading contribution arising from fourth–moment mismatch in the driving noise. In the general case where b and a vary spatially, additional O(δt2)O(δ t^2) terms appear through derivatives of the coefficients and their interaction with the density P. Those contributions depend on the specific discretization and coefficient structure, whereas the term isolated in Theorem 36 depends only on the kurtosis of the increment. Consequently, any scheme based on the source–equation representation must match the fourth moment of the target diffusion in order to eliminate this particular second–order error. The significance of second–order consistency for generative modeling lies in the fidelity of the score target. Since the optimal estimator sopts_opt is determined by the density generated by the hyperfinite forward process, a second–order error in the master equation propagates into the corresponding score field. Theorem 36 shows that a mismatch in kurtosis produces a density error of order O(δt2)O(δ t^2), and hence a corresponding second–order perturbation of the learned score. Thus, within the source–based hyperfinite framework, matching the fourth moment of the increment is necessary to remove this particular source of second–order bias. This provides a precise mathematical explanation for the advantage of Gaussian noise, and more generally of kurtosis–matched noise sources, in approximating diffusion semigroups at second order [17]. 7 Conclusion and Outlook In this paper, we developed a hyperfinite framework for score-based diffusion models using the tools of nonstandard stochastic analysis. Rather than treating discrete-time dynamics as approximations to a continuum limit, we studied stochastic evolution directly on a hyperfinite grid and showed how the familiar objects of diffusion theory arise through the standard-part map. Our analysis demonstrates that several key structures underlying score-based generative modeling admit natural hyperfinite representations. In particular: • Generator and Fokker–Planck Structure. We derived the internal infinitesimal generator associated with hyperfinite walk dynamics and established its correspondence with the classical Fokker–Planck equation after taking standard parts. • Backward Mean and Reverse-Time Dynamics. We obtained an explicit hyperfinite backward-mean identity and showed how the score term arises from the interaction between local density variation and the geometry of the internal grid transition. This yields a constructive derivation of the reverse drift and leads to a hyperfinite derivation of the reverse-time SDE. • Score Matching and Generative Consistency. We established that minimization of an internal score-matching objective recovers the score required by the reverse-time dynamics and used this connection to justify the convergence of the corresponding hyperfinite sampling procedure. • Likelihood and Higher-Order Structure. Under appropriate assumptions, we derived a hyperfinite Girsanov formula and related likelihood optimization to Fisher-divergence objectives. We also quantified the leading second-order consistency error of the hyperfinite dynamics, identifying the role played by the fourth moment of the increment distribution. From a broader perspective, the results suggest that many constructions appearing in modern diffusion modeling can be formulated directly at the hyperfinite level prior to standard-part projection. This viewpoint provides a complementary perspective on the relationship between discrete algorithms and continuous-time diffusion theory, while retaining compatibility with the classical framework of Loeb spaces, stochastic integration, and martingale methods. The theoretical framework established in this paper opens several fertile avenues for future research: 1. Heavy-Tailed SDEs and Lévy-Flight Models: While the present work focuses on Gaussian-like noise, the hyperfinite framework is naturally suited to the study of heavy-tailed and non-local processes. Building on the hyperfinite jump models of [20] and the nonstandard stochastic calculus of [1], one could derive reverse-time identities for stable-distribution diffusion processes. This would provide a rigorous foundation for Lévy-flight generative models in regimes where classical martingale machinery and standard density regularity often become intractable. 2. Numerical Analysis of Higher-Order Sampling: Our identification of the numerical dispersion term κ−324a2∂4p κ-324a^2∂^4p suggests a new frontier in the design of sampling kernels. By constructing internal noise increments that explicitly minimize higher-order Taylor residuals on the hyperfinite grid, it may be possible to develop sampling algorithms with order-of-convergence properties that surpass standard Euler-Maruyama benchmarks [17]. This shifts the focus from purely empirical step-size tuning to the analytical design of the hyperfinite increment’s spectral properties. 3. Path-Space Inference and Uncertainty Quantification: The hyperfinite Girsanov identity offers a rigorous, white-box tool for analyzing inference in generative AI. The ability to calculate exact internal likelihood ratios between path measures provides a more granular framework for studying mode collapse, distributional shift, and out-of-distribution detection—phenomena that part of the deep learning community currently addresses via empirical heuristics rather than rigorous path-space analysis. We hope that the framework developed here helps clarify the connections between nonstandard stochastic analysis and modern generative modeling, and that it encourages further interaction between these two areas. Declaration No potential conflict of interest was reported by the author(s). Additional information Funding No funding was received for conducting this research. References [1] S. Albeverio, J. E. Fenstad, R. Høegh-Krohn, and T. Lindstrøm (2009) Nonstandard methods in stochastic analysis and mathematical physics. Courier Dover Publications. Cited by: §A.6, Corollary A.16, Definition A.33, Definition 2.3, §3.3, §3, §5, item 1, Remark 2. [2] B. D. Anderson (1982) Reverse-time diffusion equation models. Stochastic Processes and their Applications 12 (3), p. 313–326. Cited by: item 2, §1, §3.3, §3.4. [3] R. M. Anderson (1976) A non-standard representation for brownian motion and itô integration. Israel Journal of Mathematics 25 (1), p. 15–46. Cited by: §A.5.5, Corollary A.16, Definition A.32, Definition A.33, Proposition A.35, Appendix C, Appendix C, §D.1, §D.2, §1.1, §1.2, §1, §1, §3.4. [4] V. Benci, S. Galatolo, and M. Ghimenti (2008) An elementary approach to stochastic differential equations using the infinitesimals. arXiv preprint arXiv:0807.3477. Cited by: §A.5, item 1, item 2, §1.2, §1, §3.2, §3.3. [5] E. Bottazzi (2019) Grid functions of nonstandard analysis in the theory of distributions and in partial differential equations. Advances in Mathematics 345, p. 429–482. Cited by: §A.5, §1.1. [6] R. Courant, K. Friedrichs, and H. Lewy (1928) Über die partiellen differenzengleichungen der mathematischen physik. Mathematische annalen 100 (1), p. 32–74. Cited by: Definition 2.3. [7] N. J. Cutland (1982) On the existence of solutions to stochastic differential equations on loeb spaces. Zeitschrift für Wahrscheinlichkeitstheorie und Verwandte Gebiete 60 (3), p. 335–357. Cited by: §1. [8] N. J. Cutland (1983) Nonstandard measure theory and its applications. Bulletin of the London Mathematical Society 15 (6), p. 529–589. Cited by: §1. [9] N. J. Cutland (1987) Infinitesimals in action. Journal of the London Mathematical Society 2 (2), p. 202–216. Cited by: §5. [10] A. Friedman (2008) Partial differential equations of parabolic type. Courier Dover Publications. Cited by: §3.2. [11] J. Ho, A. Jain, and P. Abbeel (2020) Denoising diffusion probabilistic models. Advances in neural information processing systems 33, p. 6840–6851. Cited by: §1, §5. [12] D. N. Hoover and H. J. Keisler (1984) Adapted probability distributions. Transactions of the American Mathematical Society 286 (1), p. 159–201. Cited by: §1, §3.4, §4.2. [13] D. N. Hoover and E. Perkins (1983) Nonstandard construction of the stochastic integral and applications to stochastic differential equations. i. Transactions of the American Mathematical Society 275 (1), p. 1–36. Cited by: item 3, §1.1, §1.2, §1. [14] D. N. Hoover and E. Perkins (1983) Nonstandard construction of the stochastic integral and applications to stochastic differential equations. i. Transactions of the American Mathematical Society 275 (1), p. 37–58. Cited by: §1. [15] A. Hyvärinen and P. Dayan (2005) Estimation of non-normalized statistical models by score matching.. Journal of Machine Learning Research 6 (4). Cited by: §1, §4. [16] H. J. Keisler (1984) An infinitesimal approach to stochastic analysis. American Mathematical Society. Cited by: §A.5, Definition A.10, Definition A.17, §B.1, §B.2, Appendix C, item 2, §1.1, §1.2, §1, §1, Definition 2.1, §3.3, §3.4, §3, §5. [17] P. E. Kloeden, E. Platen, and H. Schurz (2012) Numerical solution of sde through computer experiments. Springer Science & Business Media. Cited by: §6.1, item 2. [18] T. L. Lindstrøm (1980) Hyperfinite stochastic integration i: the nonstandard theory. Mathematica Scandinavica 46 (2), p. 265–292. Cited by: §B.1, §1.2, §1, §3.4, §4.2, §5. [19] T. L. Lindström (1980) Hyperfinite stochastic integration i: hyperfinite representations of standard martingales. Mathematica Scandinavica 46 (2), p. 315–331. Cited by: item 1, §1, §3.4, §4.2, §5. [20] T. Lindstrøm (2004) Hyperfinite lévy processes. Stochastics: An International Journal of Probability and Stochastic Processes 76 (6), p. 517–548. Cited by: §3, item 1. [21] P. A. Loeb (1975) Conversion from nonstandard to standard measure spaces and applications in probability theory. Transactions of the American Mathematical society 211, p. 113–122. Cited by: §A.3, §A.5.5, Theorem A.15, §D.2, §1.1, §1. [22] P. A. Loeb (1979) An introduction to nonstandard analysis and hyperfinite probability theory. In Probabilistic analysis and related topics, p. 105–142. Cited by: §A.3, Theorem A.12, §4.1, Remark 9. [23] E. Nelson (1987) Radically elementary probability theory. Princeton University Press. Cited by: §A.6, 2nd item, item 2, item 6, §1.1, §1.2, §1, §3.3, §3.3, §6. [24] E. Nelson (2020) Dynamical theories of brownian motion. Cited by: item 2, §3.3. [25] H. Osswald (2009) On anticipative girsanov transformations for lévy processes. Journal of Theoretical Probability 22 (2), p. 474–481. Cited by: §5. [26] A. Robinson (1974) Non-standard analysis. Princeton University Press. Cited by: Theorem A.12, Appendix A, §1, Remark 10. [27] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole (2021) Score-based generative modeling through stochastic differential equations. External Links: 2011.13456, Link Cited by: §1, §3.3, §4.2, §4, §5, Remark 2. [28] D. W. Stroock and S. S. Varadhan (2007) Multidimensional diffusion processes. Springer. Cited by: §1.2, §4.2, §4.2. [29] K. D. Stroyan and J. M. Bayod (2011) Foundations of infinitesimal stochastic analysis. Elsevier. Cited by: §A.5.4, §D.1, §D.2, item 2, §3.3. [30] K. D. Stroyan and W. A. J. Luxemburg (1977) Introduction to the theory of infiniteseimals. Vol. 72, Academic Press. Cited by: §A.1, Definition A.17, Appendix A, Remark 10. [31] P. Vincent (2011) A connection between score matching and denoising autoencoders. Neural computation 23 (7), p. 1661–1674. Cited by: §1, §4, §5. Appendix A Mathematical preliminaries This section collects the requisite nonstandard-analytic material. Our presentation follows the superstructure framework for rigorous handling of probability spaces and higher-order objects; for much of this material, the reader may refer to [26, 30]. Readers familiar with the basic constructions in NSA may skip directly to Section A.5, reviewing only the remainder of this section and Appendix D as needed. A.1 Hyperreal field and ultrapower construction We start with the definition of hyperreal numbers and some classical concepts of fundamental importance to this work. Definition A.1 (Free Ultrafilter). Let J be a nonempty set. A filter U on J is a subset of (J)P(J), the power set of J, satisfying the following properties: 1. Properness: ∅∉ . 2. Finite intersection property: If A,B∈A,B , then A∩B∈A∩ B . 3. Upward closure: If A∈A and A⊆B⊆JA B J, then B∈B . The filter U is called an ultrafilter if it also satisfies: 4. Maximality: For every A⊆JA J, either A∈A or J∖A∈J A . An ultrafilter U is called a free ultrafilter if it also satisfies: 5. Freeness/Nonprincipality: U contains no finite subsets of J. The existence of a free ultrafilter may be proved using Zorn’s lemma [30]. Definition A.2 (Ultrapower/Hyperreal Numbers). Let U be a fixed free ultrafilter on ℕN. Let ℝℕR^N denote the set of real sequences. Define an equivalence relation ∼ on ℝℕR^N by declaring (xn)∼(yn)if and only ifn∈ℕ:xn=yn∈.(x_n) (y_n) and only if \n :x_n=y_n\ . The ultrapower of ℝR is the quotient ℝ∗:=ℝℕ/∼. R:=R^N/\! . Denote the equivalence class of a sequence (xn)(x_n) by [xn][x_n]. Remark 6. Arithmetic operations and the order on ℝR lift coordinatewise to ℝℕR^N and descend to well-defined operations on ℝ∗ R via [xn]+[yn]:=[xn+yn],[xn]⋅[yn]:=[xnyn],[x_n]+[y_n]:=[x_n+y_n], [x_n]·[y_n]:=[x_ny_n], and [xn]<[yn]if and only ifn∈ℕ:xn<yn∈.[x_n]<[y_n] and only if \n :x_n<y_n\ . The resulting structure ℝ∗ R is a totally ordered field extending ℝR. We identify each real number r∈ℝr with the equivalence class of the constant sequence (r,r,r,…)(r,r,r,…). Thus, we consider |x|<M|x|<M, for a hyperreal x=[xn]x=[x_n] and positive real M, as equivalent to n:|xn|<M∈\n:|x_n|<M\ . Definition A.3 (Finite, Infinite, Infinitesimal). An element x∈ℝ∗x∈ R is finite (or limited) if there exists M∈ℝ>0M _>0 such that |x|<M|x|<M; x is infinitesimal if |x|<r|x|<r for every real r>0r>0; x is infinite if |x|>M|x|>M for every real M>0M>0. We write x≈yx≈ y when x−yx-y is infinitesimal, and say x is infinitely close to y. For x,y∈ℝd∗x,y∈^*R^d, we define x≈yx≈ y componentwise. Proposition A.4 (Standard Part Map). Every finite x∈ℝ∗x∈ R has a unique nearest real number st(x)∈ℝst(x) , called its standard part, satisfying x≈st(x)x (x). The map st:x∈ℝ∗:x finite→ℝst:\x∈ R:x finite\ is a ring homomorphism onto ℝR with kernel the set of infinitesimals. Proof. Let x=[xn]x=[x_n] be represented by a real sequence (xn)(x_n). Finiteness implies (xn)(x_n) is bounded in ℝR. Let L=supr∈ℝ:n:xn<r∉L= \r :\n:x_n<r\ \ and U=infr∈ℝ:n:xn>r∉U= \r :\n:x_n>r\ \. One can show L=U≕ℓL=U and that n:|xn−ℓ|<ϵ∈\n:|x_n- |<ε\ for every standard ϵ>0ε>0. Define st(x)=ℓst(x)= . Uniqueness and the homomorphism properties follow from the properties of U. ∎ Remark 7. In certain contexts, it is less cumbersome, and also common in the NSA literature, to use the notation x∘ x for st(x)st(x); we will use both. A.2 Superstructure framework, internal objects, hyperfinite sets, and transfer Constructing hyperreal numbers is a good start, but that does not take us very far in terms of analysis. For rigorous handling of relations, functions, probability spaces, and higher-order objects in general, we work within the superstructure framework. Definition A.5 (Superstructure). Let V0(ℝ)=ℝV_0(R)=R. Define recursively: Vn+1(ℝ)=Vn(ℝ)∪(Vn(ℝ)).V_n+1(R)=V_n(R) (V_n(R)). The superstructure over ℝR is V(ℝ)=⋃n=0∞Vn(ℝ)V(R)= _n=0^∞V_n(R). Elements of V(ℝ)V(R) are called standard objects. Definition A.6 (Bounded Ultrapower). Let U be a nonprincipal ultrafilter on ℕN. The bounded ultrapower V(ℝ∗)V( R) is constructed as follows. For each n, define: Vn(ℝ∗)=[(Am)m∈ℕ]:Am∈Vn(ℝ),V_n( R)= \[(A_m)_m ]:A_m∈ V_n(R) \, where [An]=[Bn][A_n]=[B_n] if n:An=Bn∈\n:A_n=B_n\ . The bounded ultrapower is V(ℝ∗)=⋃n=0∞Vn(ℝ∗)V( R)= _n=0^∞V_n( R). The hyperreal field ℝ∗ R is the set V1(ℝ∗)V_1( R) in our superstructure framework identified with the image of sequences of real numbers. Definition A.7 (Internal Objects). An object A∈V(ℝ∗)A∈ V( R) is internal if A∈Vn(ℝ∗)A∈ V_n( R) for some n. Objects in V(ℝ∗)V( R) that are not internal are called external. To be more specific, we have: Definition A.8 (Internal Sets). A subset A⊆ℝ∗A R is internal if A∈V2(ℝ∗)A∈ V_2( R), i.e., there exists a sequence of subsets (An)n∈ℕ(A_n)_n with An⊆ℝA_n such that A=[An]:=[xn]∈ℝ∗:n:xn∈An∈A=[A_n]:=\[x_n]∈ R:\n:x_n∈ A_n\ \. A subset of ℝ∗ R that is not internal is called external. Definition A.9 (Internal Functions). A function f:A→Bf:A→ B between internal sets is internal if its graph (x,f(x)):x∈A\(x,f(x)):x∈ A\ is an internal subset of A×BA× B. Equivalently, with A=[An]A=[A_n] and B=[Bn]B=[B_n], f=[fn]f=[f_n] where each fn:An→Bnf_n:A_n→ B_n is a function between standard sets. Before proceeding, it is necessary to emphasize a subtle, yet critical, nuance. Remark 8 (Notation for internal extensions). For the sake of readability, we frequently omit the asterisk prefix ∗ for standard functions and operations when their use on non-standard domains is clear from context. For instance, we write ‖x‖\|x\| for the internal extension of the Euclidean norm ∥∗⋅∥ \|·\|, and for a standard test function φ , we may write φ(x) (x) instead of φ∗(x) (x) for x in a nonstandard domain. Furthermore, in much of our presentation, derivatives would also need to be understood in a proper internal sense even if we use the “standard” notation. All of these should be reasonably clear from context. Definition A.10 (Hyperfinite Sets). An internal set H⊂ℝ∗H⊂ R is hyperfinite if there exists a hypernatural number N∈ℕ∗N∈ N and an internal bijection f:1,2,…,N→Hf:\1,2,…,N\→ H. Equivalently, it is of the form H=[an,1,an,2,…,an,Nn]H=[\a_n,1,a_n,2,…,a_n,N_n\], where Nn∈ℕN_n , N=[Nn]∈ℕ∗N=[N_n]∈ N, and the an,ia_n,i are standard reals [16]. Example A.11 (Hyperfinite Time Grid). Let N∈ℕ∗N∈ N be an infinite hypernatural. Define the time step δt=1/N∈ℝ∗δ t=1/N∈ R (a positive infinitesimal) and the hyperfinite time grid δt=0,δt,2δt,…,1.T_δ t=\0,δ t,2δ t,…,1\. This is a hyperfinite internal subset of ℝ∗ R with internal cardinality N+1N+1. It can be represented as δt=[0,1/Nn,2/Nn,…,1]T_δ t=[\0,1/N_n,2/N_n,…,1\] for any sequence (Nn)(N_n) such that [Nn]=N[N_n]=N. The following theorem highlights the power of NSA: it permits carrying out relatively simple manipulations in the extended nonstandard world and transferring them directly to their standard counterparts. Theorem A.12 (Full Transfer Principle [22, 26]). If φ(x1,…,xk) (x_1,…,x_k) is a bounded-quantifier formula in the superstructure language of set theory, and A1,…,AkA_1,…,A_k are internal objects with Ai=[Ai,n]A_i=[A_i,n], then φ(A1,…,Ak) (A_1,…,A_k) is true in V(ℝ∗)V( R) if and only if n∈ℕ:φ(A1,n,…,Ak,n) is true in V(ℝ)∈.\n : (A_1,n,…,A_k,n) is true in V(R)\ . Remark 9. Hyperfinite sets behave in many ways like finite sets: one can sum over them, define internal functions on them, and apply the transfer principle to finite combinatorial facts. This allows us to perform discrete algebra on an infinitesimally fine grid, which is the cornerstone of our approach. It is important to note, however, that it does not automatically transfer statements involving quantification over all subsets (which is a second-order concept), hence the critical distinction between internal and external sets [22]. A.3 Hyperfinite counting measure and Loeb measure To perform probability theory on hyperfinite sample spaces, we start with the internal counting measure and convert it into a standard probability measure via the Loeb construction. Definition A.13 (Internal Counting Measure). Let H be a hyperfinite set. The internal counting measure νint _int is the finitely-additive measure defined on the internal algebra ℐ(H)I(H) of all internal subsets of H by νint(A)=|A| _int(A)=|A| (the internal cardinality). Definition A.14 (Internal Probability Measure). If H is a nonempty hyperfinite set, the normalized internal counting measure μint _int is defined on ℐ(H)I(H) by μint(A)=|A||H|,for A∈ℐ(H). _int(A)= |A||H|, A (H). This is an internal, finitely additive probability measure. Theorem A.15 (Loeb Measure [21]). Let (H,ℐ(H),μint)(H,I(H), _int) be an internal hyperfinite probability space. There exists a unique standard σ-additive probability measure μL _L on the σ-algebra σ(ℐ(H))σ(I(H)) generated by ℐ(H)I(H) such that for every internal set A∈ℐ(H)A (H), μL(A)=st(μint(A)). _L(A)=st ( _int(A) ). The measure μL _L is called the Loeb measure associated to μint _int. Construction Sketch. For each internal A∈ℐ(H)A (H), define μL−(A)=st(μint(A)) _L^-(A)=st( _int(A)). For arbitrary B⊆HB H, define the outer measure: μL∗(B)=inf∑n=1∞μL−(An):B⊆⋃n=1∞An,An∈ℐ(H). _L (B)= \ _n=1^∞ _L^-(A_n):B _n=1^∞A_n,\ A_n (H) \. A set B is Loeb measurable if for every ϵ>0ε>0, there exists A∈ℐ(H)A (H) such that μL∗(B△A)<ϵ _L (B A)<ε. The collection of Loeb measurable sets forms a sigma-algebra ℒL, and μL=μL∗|ℒ _L= _L |_L is the Loeb measure. This description is equivalent to the standard Carathéodory construction of Loeb measure [21, 22]. ∎ Note that (H,ℒ,μL)(H,L, _L) is referred to as the Loeb space. Corollary A.16 (Pushforward via Standard Part). Let X:H→ℝ∗X:H→ R be an internal function on a hyperfinite probability space (H,μint)(H, _int). If X is finite except on a Loeb-null set, then the pushforward of the Loeb measure μL _L by the map ω↦st(X(ω))ω (X(ω)) defines a standard probability measure on ℝR, which is the distribution of the standard part of X [1, 3]. A.4 Nonstandard regularity classes We now define differentiability and smoothness directly in the nonstandard setting. Definition A.17 (S-continuity). An internal function f:ℝd∗→ℝ∗f: R^d→ R is said to be S-continuous if x≈y⇒f(x)≈f(y),x≈ y f(x)≈ f(y), for all finite x,y∈ℝd∗x,y∈ R^d [16, 30]. Definition A.18 (S-differentiability). An internal function f:ℝd∗→ℝ∗f: R^d→ R is S-differentiable if there exists an internal function Df:ℝd∗→ℝd∗Df: R^d→ R^d (or ∇f∇ f) such that for all finite x∈ℝd∗x∈ R^d and all infinitesimal h, f(x+h)−f(x)=Df(x)⋅h+ϵ‖h‖where ϵ≈0.f(x+h)-f(x)=Df(x)· h+ε\|h\| ε≈ 0. Definition A.19 (S-CkC^k functions). An internal function f:ℝd∗→ℝ∗f: R^d→ R is said to be S-CkC^k if for all multi-indices |α|≤k|α|≤ k, the internal partial derivatives ∂αf∂^αf exist, take finite values, and are S-continuous for all finite x∈ℝd∗x∈ R^d. A.4.1 Relation to standard smoothness Proposition A.20 (Nonstandard characterization of CkC^k regularity). Let f:ℝd→ℝf:R^d be a standard function and k≥1k≥ 1. 1. If f∈Ck(ℝd)f∈ C^k(R^d), then for every multi-index α with |α|≤k|α|≤ k, the internal derivative ∂α(f∗)∂^α( f) exists and is S-continuous at every finite point x∈ℝd∗x∈ R^d. 2. Conversely, if for every |α|≤k|α|≤ k the internal derivative ∂α(f∗)∂^α( f) exists and is S-continuous at every finite point x∈ℝd∗x∈ R^d, then the standard function f is in Ck(ℝd)C^k(R^d). Moreover, for all multi-indices |α|≤k|α|≤ k and all finite x∈ℝd∗x∈ R^d, the following shadow relation holds: ∂α(f∗)(x)≈(∂αf)(st(x)).∂^α( f)(x)≈(∂^αf)(st(x)). Remark 10. The distinction between continuity and uniform continuity is elegantly captured here: continuity of a standard function on ℝdR^d corresponds to S-continuity of its extension at all finite points, whereas uniform continuity on ℝdR^d corresponds to S-continuity at all points (including infinite ones) of ℝd∗ R^d [26, 30]. A.5 Grid functions and grid calculus For later use, we introduce hyperfinite spatial grids, grid functions, and the associated discrete calculus. These objects provide an infinitesimal discretization of ℝdR^d on which differential operators are realized as finite difference operators. The framework is entirely internal and amenable to transfer [16], following the approach of treating SDEs via infinitesimal discretizations as developed in [4]; most of these concepts may also be found in [5]. A.5.1 Hyperfinite spatial grids We saw an example of a hyperfinite time grid above, but now proceed to define a slightly more elaborate (spatial) grid. Let N′∈ℕ∗N ∈ N be an infinite hypernatural and define the spatial mesh size δx=1N′∈ℝ∗.δ x= 1N ∈ R. Fix d∈ℕd . The associated d-dimensional hyperfinite grid is δx=(δxℤ∗)d=(k1δx,…,kdδx):ki∈ℤ∗.G_δ x=(δ x\, Z)^d=\(k_1δ x,…,k_dδ x):k_i∈ Z\. This is an internal subset of ℝd∗ R^d. In applications, we could restrict attention to an internal hyperfinite subset ⊂δx.G _δ x. We assume N′N and N are chosen such that the ratio (δx)2/δt(δ x)^2/δ t is a finite (limited), noninfinitesimal hyperreal. (In our later developments, this parabolic grid scaling ensures that the discrete finite difference diffusion operator maps cleanly to its standard continuous counterpart upon taking the standard part.) A.5.2 Grid functions Definition A.21 (Grid functions). A grid function is an internal function u:→ℝ∗,u:G→ R, where ⊂ℝd∗G⊂ R^d is a hyperfinite grid. The space of all grid functions on G is denoted by ()G(G). Grid functions are purely internal objects and may be viewed as infinitesimal discretizations of standard functions on ℝdR^d. A.5.3 Grid difference operators Let 1,…,de_1,…,e_d denote the standard basis vectors in ℝdR^d, and ⊂′⊂ℝd∗G ⊂ R^d be hyperfinite grids. We typically assume ′G is large enough to contain all stencils required by the differential operators acting on functions in (′)G(G ). Definition A.22 (First-order grid derivatives). For a grid function u∈(′)u (G ), the forward and backward grid derivatives are the internal functions in ()G(G) defined by Δi+u(x)=u(x+δxi)−u(x)δx,Δi−u(x)=u(x)−u(x−δxi)δx _i^+u(x)= u(x+δ xe_i)-u(x)δ x, _i^-u(x)= u(x)-u(x-δ xe_i)δ x for all x∈x . Definition A.23 (Higher-order grid derivatives). Let α=(α1,…,αd)∈ℕdα=( _1,…, _d) ^d be a multi-index. The α-th order forward grid derivative is defined recursively by Δαu=(Δ1+)α1⋯(Δd+)αdu, ^αu=( _1^+) _1·s( _d^+) _du, provided ′G contains the required k-step neighborhood of G for |α|=k|α|=k, where the order of the derivative is |α|=∑i=1dαi|α|= _i=1^d _i. All grid derivatives are internal operators on (′)G(G ). A.5.4 Grid derivatives and standard derivatives Proposition A.24 (Consistency of grid derivatives). Let f∈Ck+1(ℝd)f∈ C^k+1(R^d) and let u=f∗|′u= f|_G be the restriction to a hyperfinite grid ′G with spacing δxδ x as above. Then for every multi-index α with |α|≤k|α|≤ k and every finite x∈x , Δαu(x)≈∂α(f∗)(x)≈(∂αf)(st(x)). ^αu(x)≈∂^α( f)(x)≈(∂^αf)(st(x)). Proof. We proceed by induction on |α||α|. For the base case |α|=1|α|=1, fix i∈1,…,di∈\1,…,d\ and a finite x∈x . By definition: Δi+u(x)=f∗(x+δxi)−f∗(x)δx. _i^+u(x)= f(x+δ xe_i)- f(x)δ x. Since f∈C2(ℝd)f∈ C^2(R^d), we apply the internal Taylor theorem [29] (the transfer of the standard Taylor theorem). There exists some ξ between x and x+δxix+δ xe_i such that: f∗(x+δxi)=f∗(x)+δx∂i(f∗)(x)+(δx)22∂i2(f∗)(ξ). f(x+δ xe_i)= f(x)+δ x\, _i( f)(x)+ (δ x)^22∂^2_i( f)(ξ). Since x is finite and δxδ x is infinitesimal, ξ is also finite. Because ∂i2f∂^2_if is continuous, its extension ∂i2(f∗) _i^2( f) is S-continuous at finite points and thus takes only finite values. Therefore, the second-order term is O(δx2)O(δ x^2). Dividing by δxδ x gives: Δi+u(x)=∂i(f∗)(x)+O(δx). _i^+u(x)= _i( f)(x)+O(δ x). Since O(δx)≈0O(δ x)≈ 0, we have Δi+u(x)≈∂i(f∗)(x) _i^+u(x)≈ _i( f)(x). By the S-continuity of ∂i(f∗) _i( f) at finite points, this is also ≈(∂if)(st(x))≈( _if)(st(x)). The higher-order case |α|≤k|α|≤ k follows by induction, noting that the Ck+1C^k+1 assumption ensures that the remainder term after k applications of the difference operator remains infinitesimal. ∎ Obviously, if we are content with stating the remainder to be o(1)o(1), the smoothness condition can be relaxed to f∈Ckf∈ C^k. A.5.5 Grid integrals Definition A.25 (Grid integral). Let ⊂ℝd∗G⊂ R^d be a hyperfinite grid and let u:→ℝ∗u:G→ R be a grid function. The grid integral of u over G is defined by ∫u(x)x:=∑x∈u(x)(δx)d, _Gu(x)\,dx:= _x u(x)(δ x)^d, where the sum is the internal hyperfinite sum. If the grid integral is finite, its standard part coincides with the Lebesgue integral of the associated standard function, when such a function exists (otherwise, the grid integral represents the Loeb integral) [3, 21]. Remark 11 (Transfer of Calculus Identities). Any first-order discrete identity valid for all finite N (e.g., discrete product rule, summation-by-parts) transfers to the hyperfinite grid. This makes many algebraic manipulations on the grid immediate and rigorous, as illustrated below. A.5.6 Some grid calculus identities To ensure that some of the following identities hold without boundary terms, we define the notion of support on a hyperfinite grid. These are, in particular, of use to us in rigorously deriving the Fokker-Planck equation, and we state them in a convenient form towards realizing this goal. Definition A.26 (Grid-Compact Support). An internal grid function u∈()u (G) has grid-compact support if there exists a finite R∈ℝ+R ^+ such that u(x)=0u(x)=0 for all x∈x with ‖x‖≥R\|x\|≥ R. Proposition A.27 (Discrete Product Rule). Let u,v∈(′)u,v (G ) be any internal grid functions. Then for each i∈1,…,di∈\1,…,d\ and x∈x : u(x)Δi+v(x)=Δi+(u(x)v(x))−v(x+δxi)Δi+u(x).u(x) _i^+v(x)= _i^+(u(x)v(x))-v(x+δ xe_i) _i^+u(x). Proof. This is a purely algebraic identity. Expanding the right-hand side using the definition of Δi+ _i^+: Δi+(u(x)v(x))−v(x+δxi)Δi+u(x) _i^+(u(x)v(x))-v(x+δ xe_i) _i^+u(x) =u(x+δxi)v(x+δxi)−u(x)v(x)δx−v(x+δxi)u(x+δxi)−u(x)δx = u(x+δ xe_i)v(x+δ xe_i)-u(x)v(x)δ x-v(x+δ xe_i) u(x+δ xe_i)-u(x)δ x =u(x+δxi)v(x+δxi)−u(x)v(x)−v(x+δxi)u(x+δxi)+v(x+δxi)u(x)δx = u(x+δ xe_i)v(x+δ xe_i)-u(x)v(x)-v(x+δ xe_i)u(x+δ xe_i)+v(x+δ xe_i)u(x)δ x =u(x)v(x+δxi)−u(x)v(x)δx=u(x)Δi+v(x). = u(x)v(x+δ xe_i)-u(x)v(x)δ x=u(x) _i^+v(x). ∎ Proposition A.28 (Forward-Backward Adjoint Identity). Let u,v∈()u,v (G). If v has grid-compact support, then: ∑x∈u(x)Δi+v(x)δxd=−∑x∈v(x)Δi−u(x)δxd. _x u(x) _i^+v(x)δ x^d=- _x v(x) _i^-u(x)δ x^d. Proof. Proposition A.27 yields: ∑x∈u(x)Δi+v(x)δxd=∑x∈Δi+(u(x)v(x))δxd−∑x∈v(x+δxi)Δi+u(x)δxd. _x u(x) _i^+v(x)δ x^d= _x _i^+(u(x)v(x))δ x^d- _x v(x+δ xe_i) _i^+u(x)δ x^d. Now, the hyperfinite grid G is defined for indices up to some infinite M∈ℕ∗M∈ N such that MδxMδ x is infinite. The first sum on the right-hand side is a telescoping sum of an internal function. Specifically, the internal sum in the i-th direction, with w=uvw=uv: ∑k=−MMw((k+1)δxi)−w(kδxi)δxδx=w((M+1)δxi)−w(−Mδxi). _k=-M^M w((k+1)δ xe_i)-w(kδ xe_i)δ xδ x=w((M+1)δ xe_i)-w(-Mδ xe_i). Since v has grid-compact support R, there exists a finite R>0R>0 such that v(x)=0v(x)=0 for all ‖x‖≥R\|x\|≥ R. Because MδxMδ x is infinite, we have ‖(M+1)δxi‖>R\|(M+1)δ xe_i\|>R and ‖(−M)δxi‖>R\|(-M)δ xe_i\|>R. Thus: w((M+1)δxi)=0andw(−Mδxi)=0,w((M+1)δ xe_i)=0 w(-Mδ xe_i)=0, which ensures the telescoping sum vanishes. For the second sum, let y=x+δxiy=x+δ xe_i. As x ranges over G, y ranges over the shifted grid +δxiG+δ xe_i. However, since v(y)=0v(y)=0 for all y∈(△(+δxi))y∈(G (G+δ xe_i)), we can maintain the summation over the original domain G: ∑x∈v(x+δxi)u(x+δxi)−u(x)δxδxd _x v(x+δ xe_i) u(x+δ xe_i)-u(x)δ xδ x^d =∑y∈v(y)u(y)−u(y−δxi)δxδxd = _y v(y) u(y)-u(y-δ xe_i)δ xδ x^d =∑y∈v(y)Δi−u(y)δxd. = _y v(y) _i^-u(y)δ x^d. ∎ Proposition A.29 (First-Order Adjoint). Let p,b∈()p,b (G) and let φ∈() (G) have grid-compact support. Then: ∑x∈p(x)b(x)Δi+φ(x)δxd=−∑x∈φ(x)Δi−(b(x)p(x))δxd. _x p(x)b(x) _i^+ (x)δ x^d=- _x (x) _i^-(b(x)p(x))δ x^d. Proof. Define u(x)=p(x)b(x)u(x)=p(x)b(x). Since u∈()u (G) and φ has grid-compact support, the result follows immediately from Proposition A.28. ∎ Proposition A.30 (Second-Order Adjoint). Let p,a∈()p,a (G) and let φ∈() (G) have grid-compact support. Then: ∑x∈p(x)a(x)Δi+Δj+φ(x)δxd=∑x∈φ(x)Δj−Δi−(a(x)p(x))δxd. _x p(x)a(x) _i^+ _j^+ (x)δ x^d= _x (x) _j^- _i^-(a(x)p(x))δ x^d. Proof. Let w(x)=Δj+φ(x)w(x)= _j^+ (x). Since φ has grid-compact support, its grid derivatives also have grid-compact support. Applying Proposition A.28 to u=pau=pa and v=wv=w: ∑x∈(pa)Δi+(Δj+φ)δxd=−∑x∈(Δj+φ)Δi−(pa)δxd. _x (pa) _i^+( _j^+ )δ x^d=- _x ( _j^+ ) _i^-(pa)δ x^d. Applying Proposition A.28 a second time with u=Δi−(pa)u= _i^-(pa) and v=φv= , the display above is equal to: −(−∑x∈φ(x)Δj−[Δi−(pa)]δxd)=∑x∈φ(x)Δj−Δi−(ap)δxd.- (- _x (x) _j^-[ _i^-(pa)]δ x^d )= _x (x) _j^- _i^-(ap)δ x^d. ∎ Starting with the next subsection, we look at nonstandard stochastic calculus. A.6 Hyperfinite stochastic increments and hyperfinite Brownian motion To model Brownian motion on the grid, we construct an internal array of random increments. This construction mirrors the standard Donsker invariance principle in the internal setting and also forms the basis for the approach in [23]. Hyperfinite Probability Space. Let δt=0,δt,2δt,…,1T_δ t=\0,δ t,2δ t,…,1\ be the hyperfinite time grid. Consider the internal sample space Ω=−1,+1δt∖1 =\-1,+1\^T_δ t \1\, representing the set of all possible paths of binary increments. We endow Ω with the uniform internal counting measure μint(A)=|A||Ω| _int(A)= |A|| | for all internal A⊆ΩA . The Loeb space (Ω,ℒ,μL)( ,L, _L) is the completion of the measure space (Ω,σ(),st∘μint)( ,σ(A),st _int), where A is the internal algebra of all internal subsets of Ω . Definition A.31 (Internal Independence). A family Xτ∈δt\X_τ\_τ _δ t of internal random variables is internally independent if for any internal subset of indices τ1,…,τn\ _1,…, _n\ and any internal Borel sets B1,…,BnB_1,…,B_n: ℙint(⋂i=1nXτi∈Bi)=∏i=1nℙint(Xτi∈Bi).P_int ( _i=1^n\X_ _i∈ B_i\ )= _i=1^nP_int(X_ _i∈ B_i). Definition A.32 (Hyperfinite Brownian Motion [3]). Let ξ(t,ω)ξ(t,ω) be an internal Rademacher array such that for each t, ξ(t,⋅)ξ(t,·) is internally independent and ℙint(ξ=±1)=1/2P_int(ξ=± 1)=1/2. We define the Brownian increments as ΔWt(ω):=ξ(t,ω)δt W_t(ω):=ξ(t,ω) δ t. The hyperfinite Brownian motion is the internal process defined by the summation: Wt(ω):=∑s∈δt∩[0,t)ΔWs(ω).W_t(ω):= _s _δ t∩[0,t) W_s(ω). It is important to clarify the relationship between the specific noise model employed in the current development and the broader theoretical framework of this paper. For pedagogical clarity and to illustrate the varied hyperfinite constructions available, the introductory sections utilize a specific Rademacher noise increment to define the discrete SDE. This allows for a more intuitive introduction to the ‘Calculus-first’ perspective. However, the majority of our results are established under the more general framework of Definition 2.3. In this broader setting, we characterize the dynamics through a general hyperfinite walk, which endogenously specifies the moments of the driving noise required to recover hyperfinite Brownian motion. We emphasize that all derivations presented in these sections remain valid under the generalized definition, as the latter encompasses the Rademacher case while remaining agnostic to the specific incremental distribution. Definition A.33 (S-integrability [1, 3]). An internal random variable X:Ω→ℝ∗X: →^*R is said to be S-integrable if: 1. int[|X|]E_int[|X|] is a limited hyperreal. 2. For every null internal set A⊆ΩA (i.e., μint(A)≈0 _int(A)≈ 0), the internal expectation over A is infinitesimal: int[|X|A]≈0E_int[|X| 1_A]≈ 0. Proposition A.34 (Moment Identities and S-integrability). For each t∈δt _δ t, the internal expectations satisfy int[Wt]=0E_int[W_t]=0 and int[Wt2]=tE_int[W_t^2]=t. Furthermore, the process Wt2W_t^2 is S-integrable, ensuring that L[st(Wt2)]=st(int[Wt2])=st(t)E_L[st(W_t^2)]=st(E_int[W_t^2])=st(t). Proof. By construction, any t∈δt _δ t is of the form kδtkδ t for some k∈0,1,…,N⊆ℕ∗k∈\0,1,…,N\ ^*N. The identities int[Wt]=0E_int[W_t]=0 and int[Wt2]=tE_int[W_t^2]=t follow directly from the internal independence of the increments ΔWs W_s and the fact that int[(ΔWs)2]=δtE_int[( W_s)^2]=δ t, which yields int[(∑i=0k−1ΔWiδt)2]=∑i=0k−1δt=kδt=tE_int[( _i=0^k-1 W_iδ t)^2]= _i=0^k-1δ t=kδ t=t. To establish S-integrability, we demonstrate that int[Wt4]E_int[W_t^4] is limited for all limited t. Using the internal multinomial expansion for the hyperfinite sum: int[Wt4]=int[(∑i=0k−1ΔWiδt)4]=∑i,j,l,m=0k−1int[ΔWiδtΔWjδtΔWlδtΔWmδt].E_int[W_t^4]=E_int [ ( _i=0^k-1 W_iδ t )^4 ]= _i,j,l,m=0^k-1E_int[ W_iδ t W_jδ t W_lδ t W_mδ t]. By internal independence and the vanishing of the first moments, the only non-vanishing terms in the internal sum are those where indices coincide in pairs. Specifically, we have k terms where i=j=l=mi=j=l=m, and 3k(k−1)3k(k-1) terms where the indices form two distinct pairs (e.g., i=j≠l=mi=j≠ l=m, including permutations): int[Wt4] _int[W_t^4] =∑i=0k−1int[ΔWiδt4]+3∑0≤i<j≤k−1int[ΔWiδt2]int[ΔWjδt2] = _i=0^k-1E_int[ W_iδ t^4]+3 _0≤ i<j≤ k-1E_int[ W_iδ t^2]E_int[ W_jδ t^2] =k(δt)2+3k(k−1)(δt)2 =k(δ t)^2+3k(k-1)(δ t)^2 =tδt+3t(t−δt)=3t2−2tδt. =tδ t+3t(t-δ t)=3t^2-2tδ t. For any limited t, int[Wt4]≈3t2E_int[W_t^4]≈ 3t^2, which is a limited hyperreal. Since the p-th internal moment is limited for p=2p=2, the random variable Wt2W_t^2 satisfies the de la Vallée Poussin theorem transferred to the Loeb setting [1] implying S-integrability. Consequently, the standard part and internal expectation commute, recovering the standard second moment on the Loeb space. ∎ Proposition A.35 (Loeb Process is a Wiener Process [3]). Let WtW_t be the hyperfinite Brownian motion constructed on (Ω,ℒ,μL)( ,L, _L). For standard t∈[0,1]t∈[0,1], define the standard-part process W¯t(ω)≔st(Wτ(ω)) W_t(ω) (W_τ(ω)) by choosing any τ∈δtτ _δ t with τ≈tτ≈ t. The process W¯t W_t is a standard Brownian motion; it is well-defined μL _L-almost surely, has continuous paths, and independent Gaussian increments. Detailed Proof Sketch. 1. Finite-Dimensional Distributions: Fix standard times 0=t0<t1<⋯<tk≤10=t_0<t_1<·s<t_k≤ 1 and choose τj∈δt _j _δ t such that τj≈tj _j≈ t_j. Let V be the vector of increments Vj=Wτj−Wτj−1V_j=W_ _j-W_ _j-1. The internal characteristic function ϕV(u) _V(u) for u∈ℝku ^k is given by: ϕV(u)=int[exp(i∑j=1kujVj)]=∏j=1k∏s∈δt∩[τj−1,τj)cos(ujδt). _V(u)=E_int [ (i _j=1^ku_jV_j ) ]= _j=1^k _s _δ t∩[ _j-1, _j) (u_j δ t). Using the internal Taylor expansion cos(z)=1−z2/2+O(z4) (z)=1-z^2/2+O(z^4) and the fact that log(1−ϵ)=−ϵ+O(ϵ2) (1-ε)=-ε+O(ε^2) for infinitesimal ϵε, we obtain: st(logϕV(u))=st(∑j=1kτj−τj−1δt(−uj22δt+O(δt2)))=−12∑j=1kuj2(tj−tj−1).st( _V(u))=st ( _j=1^k _j- _j-1δ t (- u_j^22δ t+O(δ t^2) ) )=- 12 _j=1^ku_j^2(t_j-t_j-1). Since the standard part of the internal characteristic function is the characteristic function of the Loeb distribution, the increments are independent Gaussians with variances tj−tj−1t_j-t_j-1. 2. S-continuity and Path Continuity: We demonstrate that WtW_t is μL _L-almost surely S-continuous. For any s,t∈δts,t _δ t with s<ts<t, the fourth-moment identity from Proposition A.34 yields: int[|Wt−Ws|4]=3(t−s)2−2(t−s)δt<3(t−s)2.E_int[|W_t-W_s|^4]=3(t-s)^2-2(t-s)δ t<3(t-s)^2. By the transferred Kolmogorov-Chentsov theorem, for μL _L-almost all ω, the map t↦Wt(ω)t W_t(ω) is S-continuous on δtT_δ t. This ensures that t≈s⟹Wt≈Wst≈ s W_t≈ W_s, rendering W¯t W_t independent of the choice of τ≈tτ≈ t and continuous in the standard topology. 3. Covariance: The covariance Cov(W¯ti,W¯tj)=min(ti,tj)Cov( W_t_i, W_t_j)= (t_i,t_j) follows directly from the independence and variance of the increments established in Step 1. ∎ Appendix B Proof of S-continuity of grid solutions To show that XtX_t is μL _L-a.s. S-continuous, we establish a fourth-moment bound on the increments over any internal interval [s,t)⊆δt[s,t) _δ t. B.1 Fourth Moment Bound Let s,t∈δts,t _δ t with s<ts<t. The internal increment Xt−XsX_t-X_s decomposes into a drift term D and an internal martingale term M [16]: Xt−Xs=∑r∈[s,t)b∗(r,Xr)δt⏟D+∑r∈[s,t)σ∗(r,Xr)ΔWr⏟M.X_t-X_s= _r∈[s,t)^*b(r,X_r)δ t_D+ _r∈[s,t)^*σ(r,X_r) W_r_M. Using the elementary inequality ‖a+b‖4≤8(‖a‖4+‖b‖4)\|a+b\|^4≤ 8(\|a\|^4+\|b\|^4), we bound the moments of D and M separately. 1. The Drift Term (D). By Jensen’s inequality for internal sums, we have: ‖D‖4=‖∑r∈[s,t)b∗(r,Xr)δt‖4≤(t−s)3∑r∈[s,t)‖b∗(r,Xr)‖4δt.\|D\|^4= \| _r∈[s,t)^*b(r,X_r)δ t \|^4≤(t-s)^3 _r∈[s,t)\|^*b(r,X_r)\|^4δ t. Taking internal expectations and applying the linear growth condition from Assumption 2.2: int[‖D‖4]≤(t−s)3∑r∈[s,t)K(1+int[‖Xr‖4])δt≤K1(t−s)4,E_int[\|D\|^4]≤(t-s)^3 _r∈[s,t)K(1+E_int[\|X_r\|^4])δ t≤ K_1(t-s)^4, where we utilize the fact that int[‖Xr‖4]E_int[\|X_r\|^4] is limited for all r∈δtr _δ t via the discrete internal Grönwall inequality. 2. The Diffusion Term (M). Since M is an internal martingale, we apply the internal Burkholder-Davis-Gundy (BDG) inequality [18] for p=4p=4: int[‖M‖4]≤K⋅int[(∑r∈[s,t)‖σ∗(r,Xr)‖2δt)2].E_int[\|M\|^4]≤ K·E_int [ ( _r∈[s,t)\|^*σ(r,X_r)\|^2δ t )^2 ]. Applying Jensen’s inequality to the internal sum within the expectation: int[‖M‖4]≤K(t−s)∑r∈[s,t)int[‖σ∗(r,Xr)‖4]δt≤K2(t−s)2.E_int[\|M\|^4]≤ K(t-s) _r∈[s,t)E_int[\|^*σ(r,X_r)\|^4]δ t≤ K_2(t-s)^2. 3. Combining Terms. Given t−s≤1t-s≤ 1, the bound (t−s)4≤(t−s)2(t-s)^4≤(t-s)^2 holds. Thus, there exists a limited constant C′C such that: int[‖Xt−Xs‖4]≤C′(t−s)2.E_int[\|X_t-X_s\|^4]≤ C (t-s)^2. This internal inequality is the nonstandard analogue of the Kolmogorov-Chentsov condition. B.2 Loeb Measure Argument For standard n,k∈ℕn,k , define the internal event Bn,kB_n,k as the set of ω∈Ωω∈ for which there exist s,t∈δts,t _δ t with |t−s|<1/n|t-s|<1/n and ‖Xt(ω)−Xs(ω)‖>1/k\|X_t(ω)-X_s(ω)\|>1/k. By the transferred Kolmogorov inequality [16], we have: ℙint(Bn,k)≤C′(1/n)2(1/k)4⋅n=C′k4n.P_int(B_n,k)≤ C (1/n)^2(1/k)^4· n= C k^4n. The set of non-S-continuous paths is given by A=⋃k∈ℕ⋂n∈ℕBn,kA= _k _n B_n,k. For each standard k, μL(⋂n∈ℕBn,k)≤infn∈ℕst(C′k4/n)=0 _L( _n B_n,k)≤ _n st(C k^4/n)=0. By the countable subadditivity of the Loeb measure μL _L, it follows that μL(A)=0 _L(A)=0. Thus, XtX_t is μL _L-almost surely S-continuous. Appendix C Convergence to the standard Itô solution The proof identifies the standard part of the hyperfinite recursion as the unique strong solution to the Itô SDE by establishing S-continuity and matching the limits of internal quadratic forms. Step 1: S-continuity and Finiteness. By the internal Grönwall inequality and the fourth-moment bound established in Appendix B, XtX_t is S-bounded and S-continuous μL _L-a.s. [3, 16]. Thus, for almost all ω, the standard part X¯t=st(Xτ) X_t=st(X_τ) with continuous paths exists and is well-defined for all t∈[0,1]t∈[0,1]. Since b and σ are continuous in t and Lipschitz in x, their extensions b∗^*b and σ∗^*σ are S-continuous. The composition of S-continuous functions is S-continuous; thus, the internal processes Bs≔b∗(s,Xs)B_s ^*b(s,X_s) and Σs≔σ∗(s,Xs) _s ^*σ(s,X_s) are S-continuous and S-bounded on δtT_δ t μL _L-a.s. Step 2: Convergence of the Drift. The drift term Dτ=∑s<τBsδtD_τ= _s<τB_sδ t is an internal Riemann sum. For S-continuous and S-bounded integrands, the standard part of the sum is the Riemann integral of the standard part [3]: st(Dτ)=∫0st(τ)st(Bs)s=∫0tb(r,X¯r)r.st(D_τ)= _0^st(τ)st(B_s)\,ds= _0^tb(r, X_r)\,dr. Step 3: Convergence of the Diffusion. Let Mτ=∑s<τΣsΔWsM_τ= _s<τ _s W_s. We identify its standard part M¯t≔st(Mτ) M_t (M_τ) via its martingale properties: 1. Martingale Property: MτM_τ is an S-continuous internal martingale. Its standard part M¯t M_t is therefore a continuous standard martingale on the Loeb space [19]. 2. Quadratic Variation: The internal quadratic variation [M]τ=∑s<τΣsΣs⊤δt[M]_τ= _s<τ _s _s δ t is an internal Riemann sum. Its standard part is: st([M]τ)=∫0tσ(r,X¯r)σ(r,X¯r)⊤r.st([M]_τ)= _0^tσ(r, X_r)σ(r, X_r) \,dr. 3. Identification: By the internal covariation [M,W]τ=∑s<τΣsδt[M,W]_τ= _s<τ _sδ t, we have st([M,W]τ)=∫0tσ(r,X¯r)rst([M,W]_τ)= _0^tσ(r, X_r)\,dr. By the Lévy’s characterization/martingale representation theorem, M¯t M_t is uniquely identified as the Itô integral ∫0tσ(r,X¯r)W¯r _0^tσ(r, X_r)\,d W_r [13]. Step 4: Conclusion. Taking the standard part of the internal identity Xτ=X0+Dτ+MτX_τ=X_0+D_τ+M_τ for τ≈tτ≈ t yields the standard integral equation: X¯t=X¯0+∫0tb(r,X¯r)r+∫0tσ(r,X¯r)W¯r. X_t= X_0+ _0^tb(r, X_r)\,dr+ _0^tσ(r, X_r)\,d W_r. By the global Lipschitz condition, this SDE admits a unique strong solution, which must coincide with X¯t X_t μL _L-almost surely. Appendix D Discrete Itô calculus on a hyperfinite grid This section provides rigorous proofs for the discrete Itô formula within the framework of hyperfinite grid diffusions and its consequences in the standard world. D.1 Discrete Itô formula with pathwise control We utilize the fact that for μL _L-almost all paths, the process remains within the finite galaxy of ℝd∗^*R^d [3], allowing us to control the Taylor remainder pathwise (see also [29]). Proposition D.1 (Discrete Itô Formula). Let f∈C3(ℝd;ℝ)f∈ C^3(R^d;R) and let (Xt)t∈δt(X_t)_t _δ t satisfy the grid recursion (1) with Rademacher noise increments. Suppose the coefficients b,σb,σ satisfy the conditions of Theorem 2.5. Then for μL _L-almost every ω∈Ωω∈ : f∗(Xt)−f∗(X0)=∑s<t[∇(f∗)(Xs)⋅ΔXs+12ΔXs⊤D2(f∗)(Xs)ΔXs]+ϵt(ω), f(X_t)- f(X_0)= _s<t [∇( f)(X_s)· X_s+ 12 X_s D^2( f)(X_s) X_s ]+ _t(ω), (37) where ΔXs=Xs+δt−Xs X_s=X_s+δ t-X_s and the error satisfies |ϵt(ω)|≤K(ω)δt≈0| _t(ω)|≤ K(ω) δ t≈ 0 for all t∈δt _δ t. Proof. By the linear growth of coefficients and the path properties established in Theorem 2.5, there exists a set Ω0 _0 with μL(Ω0)=1 _L( _0)=1 such that for every ω∈Ω0ω∈ _0, the path is S-bounded: maxs∈δt‖Xs(ω)‖ _s _δ t\|X_s(ω)\| is limited. Consequently, the internal coefficients along the path satisfy maxs∈δt(‖b∗(s,Xs)‖+‖σ∗(s,Xs)‖)≤C′(ω) _s _δ t(\|^*b(s,X_s)\|+\|^*σ(s,X_s)\|)≤ C (ω) for some limited C′(ω)∈ℝ∗C (ω)∈^*R. Fix ω∈Ω0ω∈ _0. Applying the internal Taylor theorem to f∗^*f: f∗(Xs+δt)=f∗(Xs)+∇(f∗)(Xs)⊤ΔXs+12ΔXs⊤D2(f∗)(Xs)ΔXs+Rs, f(X_s+δ t)= f(X_s)+∇( f)(X_s) X_s+ 12 X_s D^2( f)(X_s) X_s+R_s, where the remainder term in Lagrange form is: Rs=16∑|α|=3∂α(f∗)(ξs)(ΔXs)α,R_s= 16 _|α|=3∂^α(^*f)( _s)( X_s)^α, for some ξs _s on the internal line segment [Xs,Xs+δt][X_s,X_s+δ t]. 1. Local Finite Bound: As argued above, the set of reachable states is contained in an internal ball B(0,H(ω))B(0,H(ω)) of limited radius H(ω)H(ω). Because f is a standard C3C^3 function, its third-order internal derivatives are S-continuous at all finite points and thus S-bounded on B(0,H(ω))B(0,H(ω)). Hence, there exists a limited M(ω)M(ω) such that: max|α|=3sups∈δt|∂α(f∗)(ξs)|≤M(ω). _|α|=3 _s _δ t|∂^α(^*f)( _s)|≤ M(ω). 2. Remainder Estimate: From the recursion (1), ‖ΔXs‖≤‖b∗‖δt+‖σ∗‖mδt\| X_s\|≤\|^*b\|δ t+\|^*σ\| mδ t. Given the boundedness of the coefficients, there exists a limited L(ω)L(ω) such that ‖ΔXs‖≤L(ω)δt\| X_s\|≤ L(ω) δ t. Substituting this into the remainder: |Rs|≤d36M(ω)L(ω)3(δt)3/2=K(ω)(δt)3/2,|R_s|≤ d^36M(ω)L(ω)^3(δ t)^3/2=K(ω)(δ t)^3/2, where K(ω)K(ω) is a limited hyperreal. 3. Summation of Errors: The total error ϵt(ω) _t(ω) is the internal sum of these remainders over t/δt/δ t steps: |ϵt(ω)|=|∑s<tRs|≤∑s<tK(ω)(δt)3/2≤1δtK(ω)(δt)3/2=K(ω)δt.| _t(ω)|= | _s<tR_s |≤ _s<tK(ω)(δ t)^3/2≤ 1δ tK(ω)(δ t)^3/2=K(ω) δ t. Since K(ω)K(ω) is limited and δt≈0 δ t≈ 0, it follows that ϵt(ω) _t(ω) is infinitesimal uniformly for all t∈δt _δ t. ∎ D.2 Itô formula in the limit Of course, we have the following [29]: Corollary D.2 (Itô Formula in the Limit). Let X¯t X_t be the standard part of the grid diffusion XtX_t satisfying the assumptions of Theorem 2.5. Then for any f∈C3(ℝd;ℝ)f∈ C^3(R^d;R), the process X¯t X_t satisfies the classical Itô formula: f(X¯t) f( X_t) =f(X¯0)+∫0t∇f(X¯s)⊤b(s,X¯s)s+∫0t∇f(X¯s)⊤σ(s,X¯s)W¯s =f( X_0)+ _0^t∇ f( X_s) b(s, X_s)\,ds+ _0^t∇ f( X_s) σ(s, X_s)\,d W_s +12∫0ttr(σ(s,X¯s)σ(s,X¯s)⊤D2f(X¯s))s. + 12 _0^ttr (σ(s, X_s)σ(s, X_s) D^2f( X_s) )ds. Proof. For brevity, we denote both the standard function and its natural extension by f. 1. Pathwise Regularity and Lifting. There exists a Loeb-measurable set Ω0 _0 with μL(Ω0)=1 _L( _0)=1 such that for all ω∈Ω0ω∈ _0: 1. The process Xt(ω)X_t(ω) is S-bounded (|Xt|<H(ω)|X_t|<H(ω)) and S-continuous on δtT_δ t. 2. The standard-part process X¯r=st(Xτ) X_r=st(X_τ) (for τ≈rτ≈ r) is well-defined and continuous. 3. The functions b,σ,∇f,b,σ,∇ f, and D2fD^2f evaluated along the path are S-continuous and S-bounded. Consequently, the internal integrands are liftings of their standard counterparts. For instance, ∇f(Xs)≈∇f(X¯st(s))∇ f(X_s)≈∇ f( X_st(s)) holds for all s∈δts _δ t [21]. 2. The Drift and Martingale Components. Substituting the increment ΔXs=b∗δt+σ∗ΔWs X_s=^*bδ t+^*σ W_s into the first-order sum of the Taylor expansion: ∑s<t∇f(Xs)⊤ΔXs=∑s<t∇f(Xs)⊤b(s,Xs)δt⏟I1+∑s<t∇f(Xs)⊤σ(s,Xs)ΔWs⏟I2. _s<t∇ f(X_s) X_s= _s<t∇ f(X_s) b(s,X_s)δ t_I_1+ _s<t∇ f(X_s) σ(s,X_s) W_s_I_2. Drift Term I1I_1: As an internal Riemann sum with an S-continuous integrand, its standard part satisfies: st(I1)=∫0st(t)st(∇f(Xr)⊤b(r,Xr))r=∫0st(t)∇f(X¯r)⊤b(r,X¯r)r.st(I_1)= _0^st(t)st(∇ f(X_r) b(r,X_r))dr= _0^st(t)∇ f( X_r) b(r, X_r)dr. Martingale Term I2I_2: Let Mτ=∑s<τΦsΔWsM_τ= _s<τ _s W_s with Φs=∇f(Xs)⊤σ(s,Xs) _s=∇ f(X_s) σ(s,X_s). MτM_τ is an S-continuous internal martingale. Its standard part M¯t M_t is a continuous martingale on the Loeb space with quadratic variation ⟨M¯⟩t=st([M]t)=∫0t‖Φr‖2r M _t=st([M]_t)= _0^t\| _r\|^2dr. This, paired with the internal covariation [M,W]t[M,W]_t, uniquely identifies st(I2)st(I_2) as the Itô integral ∫0st(t)(∇f⊤σ)W¯r _0^st(t)(∇ f σ)d W_r [3]. 3. The Itô Term and Quadratic Collapse. We expand the second-order sum 12∑s<tΔXs⊤D2f(Xs)ΔXs 12 _s<t X_s D^2f(X_s) X_s. Note that: ΔXsΔXs⊤=(σΔWs)(σΔWs)⊤+Oω(δt3/2). X_s X_s =(σ W_s)(σ W_s) +O_ω(δ t^3/2). The Oω(δt3/2)O_ω(δ t^3/2) terms vanish upon summation. For the principal term, let As=σ(s,Xs)⊤D2f(Xs)σ(s,Xs)A_s=σ(s,X_s) D^2f(X_s)σ(s,X_s). We evaluate the internal sum: J=∑s<t∑i,j=1m(As)ijΔWs(i)ΔWs(j).J= _s<t _i,j=1^m(A_s)_ij W_s^(i) W_s^(j). For Rademacher increments where (ΔWs(i))2=δt( W_s^(i))^2=δ t: • Diagonal terms (i=ji=j): The sum is ∑s<tTr(As)δt _s<tTr(A_s)δ t. By the cyclic property, Tr(As)=Tr(σσ⊤D2f)Tr(A_s)=Tr(σ D^2f). • Off-diagonal terms (i≠ji≠ j): The terms Kτ=∑s<τ∑i≠j(As)ijΔWs(i)ΔWs(j)K_τ= _s<τ _i≠ j(A_s)_ij W_s^(i) W_s^(j) form an internal martingale. Its quadratic variation [K]t[K]_t is O(∑δt2)=O(δt)O(Σδ t^2)=O(δ t). Thus, by the internal martingale property, Kt≈0K_t≈ 0 Loeb-a.s. [23]. Thus, st(J)=∫0st(t)Tr(σ(r,X¯r)σ(r,X¯r)⊤D2f(X¯r))rst(J)= _0^st(t)Tr(σ(r, X_r)σ(r, X_r) D^2f( X_r))dr. 4. Conclusion. Summing the components and noting the remainder ϵt≈0 _t≈ 0 from Proposition D.1, the standard part of the discrete expansion recovers the standard Itô Formula: f(X¯t)−f(X¯0)=∫0t(∇f⊤b+12Tr(σσ⊤D2f))r+∫0t∇f⊤σdW¯r.f( X_t)-f( X_0)= _0^t (∇ f b+ 12Tr(σ D^2f) )dr+ _0^t∇ f σ\,d W_r. ∎ Appendix E S-equivalence of measures Definition E.1 (S-Equivalence of Internal Grid Densities). Let G be a hyperfinite grid in ℝd∗^*R^d with cell volume (δx)d(δ x)^d. Let p1p_1 and p2p_2 be internal, non-negative probability densities on G (i.e. ∑x∈p1(x)(δx)d=∑x∈p2(x)(δx)d=1 _x p_1(x)(δ x)^d= _x p_2(x)(δ x)^d=1). We say that p1p_1 and p2p_2 are S-equivalent, denoted p1≈p2p_1≈ p_2, if for every internal set A⊆A : ∑x∈Ap1(x)(δx)d≈∑x∈Ap2(x)(δx)d. _x∈ Ap_1(x)(δ x)^d≈ _x∈ Ap_2(x)(δ x)^d. If p1,p2p_1,p_2 are S-integrable, then p1≈p2p_1≈ p_2 if and only if their induced nonstandard measures define the exact same Loeb measure on the σ-algebra of Loeb-measurable subsets of G. Furthermore, for these S-integrable probability densities, a hyperfinite analogue of Scheffé’s theorem guarantees that the above condition is equivalent to convergence in the internal L1()L^1(G) norm: ‖p1−p2‖L1()=∑x∈|p1(x)−p2(x)|(δx)d≈0.\|p_1-p_2\|_L^1(G)= _x |p_1(x)-p_2(x)|(δ x)^d≈ 0.