Paper deep dive
Gradient Flow Drifting: Generative Modeling via Wasserstein Gradient Flows of KDE-Approximated Divergences
Jiarui Cao, Zixuan Wei, Yuxin Liu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/22/2026, 6:16:07 AM
Summary
The paper introduces 'Gradient Flow Drifting', a mathematical framework that unifies various generative models (including Drifting Models and MMD-based generators) as Wasserstein gradient flows of KDE-approximated divergence functionals. It proves that the drifting field of the Drifting Model is equivalent to the Wasserstein-2 gradient flow of the forward KL divergence under KDE approximation, provides a unified identifiability proof, and proposes a mixed-divergence strategy to balance mode collapse and mode blurring.
Entities (5)
Relation Signals (3)
Gradient Flow Drifting → includes → MMD-based generators
confidence 95% · Besides that, this broad family of generative models can also include MMD-based generators
Drifting Model → isequivalentto → Wasserstein-2 gradient flow
confidence 95% · the drifting field of drifting model equals... the particle velocity field of the Wasserstein-2 gradient flow of KL(q||p)
Gradient Flow Drifting → subsumes → Drifting Model
confidence 95% · we prove an equivalence between the recently proposed Drifting Model and the Wasserstein gradient flow of the forward KL divergence under kernel density estimation (KDE) approximation.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We reveal a precise mathematical framework about a new family of generative models which we call Gradient Flow Drifting. With this framework, we prove an equivalence between the recently proposed Drifting Model and the Wasserstein gradient flow of the forward KL divergence under kernel density estimation (KDE) approximation. Specifically, we prove that the drifting field of drifting model (arXiv:2602.04770) equals, up to a bandwidth-squared scaling factor, the difference of KDE log-density gradients $\nabla \log p_{\mathrm{kde}} - \nabla \log q_{\mathrm{kde}}$, which is exactly the particle velocity field of the Wasserstein-2 gradient flow of $KL(q\|p)$ with KDE-approximated densities. Besides that, this broad family of generative models can also include MMD-based generators, which arises as special cases of Wasserstein gradient flows of different divergences under KDE approximation. We provide a concise identifiability proof, and a theoretically grounded mixed-divergence strategy. We combine reverse KL and $\chi^2$ divergence gradient flows to simultaneously avoid mode collapse and mode blurring, and extend this method onto Riemannian manifold which loosens the constraints on the kernel function, and makes this method more suitable for the semantic space. Preliminary experiments on synthetic benchmarks validate the framework.
Tags
Links
- Source: https://arxiv.org/abs/2603.10592v1
- Canonical: https://arxiv.org/abs/2603.10592v1
Trouble viewing inline? Open PDF directly →
Full Text
67,675 characters extracted from source content.
Expand or collapse full text
Gradient Flow Drifting: Generative Modeling via Wasserstein Gradient Flows of KDE-Approximated Divergences Jiarui Cao, Zixuan Wei The Chinese University of Hong Kong Hong Kong 1155244613, 1155245852@link.cuhk.edu.hk &Yuxin Liu Civil Aviation University of China Tianjin yx_liu2025061016@163.com Abstract We reveal a precise mathematical framework about a new family of generative models which we call Gradient Flow Drifting. With this framework, we prove an equivalence between the recently proposed Drifting Model and the Wasserstein gradient flow of the forward KL divergence under kernel density estimation (KDE) approximation. Specifically, we prove that the drifting field of drifting model Deng et al. (2026) equals, up to a bandwidth-squared scaling factor, the difference of KDE log-density gradients ∇logpkde−∇logqkde∇ p_kde-∇ q_kde, which is exactly the particle velocity field of the Wasserstein-2 gradient flow of KL(q∥p)KL(q\|p) with KDE-approximated densities. Besides that, this broad family of generative models can also include MMD-based generators, which arises as special cases of Wasserstein gradient flows of different divergences under KDE approximation. We provide a concise identifiability proof, and a theoretically grounded mixed-divergence strategy. We combine reverse KL and χ2χ^2 divergence gradient flows to simultaneously avoid mode collapse and mode blurring, and extend this method onto Riemannian manifold which loosens the constraints on the kernel function, and makes this method more suitable for the semantic space. Preliminary experiments on synthetic benchmarks validate the framework. 1 Introduction Generative modeling seeks to learn a mapping f such that the pushforward f#pϵf_\#p_ ε of a simple prior pϵp_ ε approximates a data distribution pdatap_data. The recently proposed Drifting Model Deng et al. (2026) introduces a new paradigm: rather than relying on iterative inference-time dynamics (as in diffusion or flow-based models), it evolves the pushforward distribution during training time via a drifting field p,qV_p,q, and naturally admits one-step generation. Drifting Models achieve state-of-the-art one-step FID on ImageNet 256×256256× 256 (1.54 in latent space and 1.61 in pixel space). Despite their empirical success, theoretical foundations of Drifting Models remain underdeveloped. The original paper’s analysis is somewhat heuristic and the identifiability proof (Appendix C.1 therein) requires additional smoothness assumptions. We argue that this complexity stems from a failure to recognize a fundamental connection. Our key observation. The drifting field of Deng et al. (2026), when instantiated with a Gaussian kernel kh(,)=exp(−‖−‖22h2)k_h(x,y)= (- \|x-y\|^22h^2), satisfies the exact identity: p,q()=h2(∇logpkde()−∇logqkde()),V_p,q(x)=h^2 (∇ p_kde(x)-∇ q_kde(x) ), (1) where pkde()=p[kh(,)]p_kde(x)=E_p[k_h(x,y)] is the Kernel Density Estimation (KDE) of p. The right-hand side is precisely the particle velocity field of the Wasserstein-2 gradient flow of the KL divergence KL(q∥p)KL(q\|p), with true densities replaced by their KDE approximations with the same kernel. This identification has several consequences: 1. Unified framework: By varying the divergence functional, we obtain a family of gradient flow drifting models. MMD-based generators correspond to the ℒ2L^2 distribution distance, and drifting models to the KL divergence. We can construct new models from any f-divergence and any other divergence that can prove the distribution convergence. 2. Mixed gradient flows: Convex combinations of divergences yield legitimate mixed gradient flows (Theorem 4.12), enabling strategies that combine the complementary strengths of different divergences—e.g., MMD for global mode coverage and reverse KL for local sharpness, and reverse KL divergence and χ2χ^2 provide more specific precise forcing. 3. Simplified identifiability: The equilibrium condition p,q=⇒p=qV_p,q=0 p=q follows in lines from the injectivity of the kernel mean embedding under characteristic kernels. 4. Drifting Model as a special case: The standard energy dissipation inequality for Wasserstein gradient flows immediately yields dtKL(qtkde∥pkde)≤0 ddtKL(q_t^kde\|p^kde)≤ 0. Concurrent work by Li and Zhu (2026) reinterprets Drifting Models through a flow-map semigroup decomposition, but does not identify the KDE–gradient flow connection. Belhadji et al. (2025) unifies MMD gradient flows with mean shift but does not extend to f-divergences. Our framework subsumes both perspectives. 2 Related Work Drifting Models. Deng et al. (2026) propose learning a one-step pushforward map by evolving the generated distribution during training via a kernel-based drifting field. They achieve strong empirical results but provide limited theoretical analysis. Li and Zhu (2026) reinterpret Drifting Models via long-short flow-map factorization, connecting them to closed-form flow matching based on semigroup consistency. Our work gives a new perspective to view it and includes it into a big family of generative models. Wasserstein gradient flows in generative modeling. Wasserstein gradient flows Jordan et al. (1998); Ambrosio et al. (2005); Santambrogio (2015) provide a variational framework for the evolution of probability measures. Several works leverage this framework for generative modeling: Arbel et al. (2019) study MMD gradient flows for sampling; Yi et al. (2023) use Wasserstein gradient flows to unify divergence GANs, introducing MonoFlow with a monotone rescaling of the log density ratio; Choi et al. (2024) propose scalable Wasserstein gradient descent. Our work makes optimizing the gradient work directly possible through kernel density estimation. Kernel density estimation and score estimation. The connection between mean shift and KDE gradients is classical Cheng (1995); Comaniciu and Meer (2002). Belhadji et al. (2025) recently unified mean shift, MMD-optimal quantization, and gradient flows. Our work extends this connection to arbitrary f-divergences. MMD and kernel methods for generation. MMD-based generative models Dziugaite et al. (2015); Li et al. (2015) minimize the MMD between generated and data distributions. Zhou et al. (2025) extend moment matching to one-/few-step diffusion. Chizat et al. (2026) provide quantitative convergence rates for MMD Wasserstein gradient flows. Our framework reveals MMD generators as one member of a broader family. f-divergence minimization. f-divergence variational estimation has been widely studied Nguyen et al. (2010); Nowozin et al. (2016). Yi et al. (2023) connects f-divergence GANs to Wasserstein gradient flows but requires a discriminator to estimate density ratios. Our KDE-based approach avoids adversarial training entirely, but aligns with an approximation based on particulars. 3 Preliminaries 3.1 Notation Let (ℝd)P(R^d) denote the set of Borel probability measures on ℝdR^d, and 2(ℝd)P_2(R^d) the subset with finite second moments. For μ∈(ℝd)μ (R^d), we write μ also for its density with respect to Lebesgue measure when it exists. We use ⟨⋅,⋅⟩ ·,· for inner products and ∥⋅∥\|·\| for norms, with subscripts indicating the space when ambiguous. 3.2 Kernel Density Estimation Definition 3.1 (KDE operator). Given a kernel k:ℝd×ℝd→ℝk:R^d×R^d and μ∈(ℝd)μ (R^d), the KDE operator is k[μ]():=∫ℝdk(,)dμ().T_k[μ](x):= _R^dk(x,y)dμ(y). (2) For the Gaussian kernel kh(,)=exp(−‖−‖2/(2h2))k_h(x,y)= (-\|x-y\|^2/(2h^2)) with bandwidth h>0h>0, we write μkde():=kh[μ]() _kde(x):=T_k_h[μ](x). 3.3 Reproducing Kernel Hilbert Spaces Definition 3.2 (RKHS and kernel mean embedding). A symmetric positive definite kernel k induces a unique reproducing kernel Hilbert space ℋkH_k with inner product ⟨⋅,⋅⟩ℋk ·,· _H_k satisfying the reproducing property: f()=⟨f,k(,⋅)⟩ℋkf(x)= f,k(x,·) _H_k for all f∈ℋkf _k. The kernel mean embedding of μ∈(ℝd)μ (R^d) is mkμ:=∫k(⋅,)dμ()∈ℋkm_k^μ:= k(·,y)dμ(y) _k. Definition 3.3 (Characteristic kernel). A kernel k is characteristic if the kernel mean embedding map μ↦mkμ m_k^μ is injective on (ℝd)P(R^d). 3.4 Wasserstein Gradient Flows Definition 3.4 (Wasserstein-2 gradient flow). Given a functional ℱ:2(ℝd)→ℝF:P_2(R^d) with first variation δℱδq δ q, its Wasserstein-2 gradient flow is the curve qtt≥0\q_t\_t≥ 0 satisfying the continuity equation: ∂tqt=∇⋅(qt∇δℱδq|qt). _tq_t=∇· (q_t∇ δ q |_q_t ). (3) Equivalently, particles t∼qtx_t q_t evolve as dtdt=(t) dx_tdt=v(x_t) where ()=−∇δℱδq()v(x)=-∇ δ q(x). 3.5 The Drifting Model We recall the core formulation of Deng et al. (2026). Given a data distribution p and a generated distribution q=f#pϵq=f_\#p_ ε, the drifting field is: p,q()=p[k(,+)(+−)]p[k(,+)]⏟p+()−q[k(,−)(−)]q[k(,−)]⏟q−(),V_p,q(x)= E_p[k(x,y^+)(y^+-x)]E_p[k(x,y^+)]_V_p^+(x)- E_q[k(x,y^-)(y^--x)]E_q[k(x,y^-)]_V_q^-(x), (4) with training loss ℒ=ϵ[‖fθ(ϵ)−stopgrad(fθ(ϵ)+p,qθ(fθ(ϵ)))‖2]L=E_ ε[\|f_θ( ε)-stopgrad(f_θ( ε)+V_p,q_θ(f_θ( ε)))\|^2]. 4 Method: Gradient Flow Drifting We present a unified framework in which generative models arise as Wasserstein gradient flows (WGFs) of divergence functionals under KDE approximation. The logical development proceeds in three layers: • Foundation 4.1: Under mild kernel regularity conditions, KDE-level distribution matching is equivalent to matching the original distributions. • Engine 4.2–4.3: General f-divergence WGFs at the KDE level, with energy dissipation and unified identifiability. • Instantiation 4.4–4.6: The Drifting Model, MMD generators, and mixed gradient flows emerge as special cases. The framework extends naturally to Riemannian manifolds 4.7, and is summarized as a complete training pipeline in 4.8. 4.1 Foundation: KDE Smoothing and Distribution Matching The starting point of our framework is that KDE smoothing, under mild kernel regularity, preserves distributional identity and simultaneously provides the smoothness needed for gradient flow analysis. This allows us to work entirely at the KDE level without imposing any regularity on the data distribution p or the generated distribution q. Assumption 4.1 (Kernel regularity; full statement in Appendix A). Let k:ℝd×ℝd→ℝk:R^d×R^d satisfy: K1. Characteristic: the mean embedding μ↦∫k(⋅,)dμ()μ k(·,y)dμ(y) is injective on (ℝd)P(R^d). K2. Uniform gradient bound: Mk:=sup,‖∇k(,)‖<∞M_k:= _x,y\| _xk(x,y)\|<∞. K3. Strict positivity: k(,)>0k(x,y)>0 for all ,x,y. K4. Differentiability: ↦k(,)x k(x,y) is C1C^1 for every y. The Gaussian kernel kh(,)=exp(−‖−‖2/(2h2))k_h(x,y)= (-\|x-y\|^2/(2h^2)) satisfies K1.–K4.; the Laplace kernel used in the original Drifting Model fails K4. (Appendix J). Theorem 4.2 (KDE regularity; proof in Appendix C). Under K2–K4, for any μ∈(ℝd)μ (R^d): (i) μkde∈C1(ℝd) _kde∈ C^1(R^d) with ∇μkde()=∫∇k(,)dμ() _x _kde(x)= _xk(x,y)dμ(y); (i) μkde()>0 _kde(x)>0 for all x; (i) sup‖∇μkde()‖≤Mk _x\|∇ _kde(x)\|≤ M_k. In particular, no moment or smoothness conditions on μ are required: the constant MkM_k serves as a universal dominating function for any probability measure, enabling all subsequent Leibniz interchanges. Proposition 4.3 (KDE injectivity; proof in Appendix B). Under K1, μkde=νkde _kde= _kde pointwise implies μ=νμ=ν. Remark 4.4 (Foundation summary). Under K1.–K4., the KDE-smoothed densities pkdep_kde and qkdeq_kde are strictly positive and C1C^1. In particular, the log-ratio log(pkde/qkde) (p_kde/q_kde) is well-defined and C1C^1, and pkde=qkdep_kde=q_kde if and only if p=qp=q. This means every divergence-minimization argument at the KDE level faithfully transfers to the original distributions. 4.2 Gradient Flows of f-Divergences under KDE Approximation With smoothed densities that are smooth and positive (Theorem 4.2), we can apply the standard Wasserstein gradient flow machinery to f-divergences directly at the KDE level. Recall that for a convex function f:(0,∞)→ℝf:(0,∞) with f(1)=0f(1)=0, the f-divergence is Df(ρ∥π)=∫πf(ρ/π)dD_f(ρ\|π)= π\,f(ρ/π)dx (Definition D.1 in Appendix). The WGF of ℱ[q]=Df(q∥p)F[q]=D_f(q\|p) has first variation δℱδq()=f′(q()/p()) δ q(x)=f (q(x)/p(x)) and particle velocity f()=−∇f′(q()/p())v_f(x)=-∇ f (q(x)/p(x)) (Proposition D.2 in Appendix). Replacing the true densities with their KDE approximations yields the generalized drifting velocity field: f()=−∇f′(q()p()).v_f(x)=-∇ f \! ( q(x)p(x) ). (5) Theorem 4.5 (Energy dissipation). Let f be strictly convex and let qtt≥0\q_t\_t≥ 0 be smooth positive densities evolving according to the continuity equation with velocity (5). Under appropriate boundary conditions (Appendix D, Remark D.3): dtDf(qt∥p)=−∫qt()‖∇f′(qt()p())‖2d≤0. ddtD_f(q_t\|p)=- q_t(x) \|∇ f \! ( q_t(x)p(x) ) \|^2dx≤ 0. (6) On compact Riemannian manifolds without boundary (e.g., d−1S^d-1), the boundary condition is vacuous and (6) holds unconditionally. Table 1 records the specific velocity fields for the divergences of primary interest. Table 1: Generative models as Wasserstein gradient flows of divergences under KDE approximation. All velocity fields are sample-computable via the KDE score formula (Appendix E). f(u)f(u) Divergence f′(u)f (u) KDE velocity field kde()v^kde(x) Model ulogu u Forward KL logu+1 u+1 ∇logpkde−∇logqkde∇ p_kde-∇ q_kde Drifting −logu- u Reverse KL −1/u-1/u pkdeqkde(∇logpkde−∇logqkde) p_kdeq_kde(∇ p_kde-∇ q_kde) – 12(u−1)2 12(u-1)^2 χ2χ^2 (u−1)(u-1) qkdepkde(∇logpkde−∇logqkde) q_kdep_kde(∇ p_kde-∇ q_kde) – 12‖mkp−mkq‖ℋk2 12\|m_k^p-m_k^q\|_H_k^2 ∇∫k(,)d(p−q)()=∇(pkde−qkde) _x k(x,y)d(p-q)(y)=∇(p_kde-q_kde) MMD Remark 4.6 (Factored velocity structure). All f-divergence velocities in Table 1 share the common factor (∇logpkde−∇logqkde)(∇ p_kde-∇ q_kde), modulated by a density-ratio weight w()w(x): w≡1w≡ 1 (forward KL), w=pkde/qkdew=p_kde/q_kde (reverse KL), w=qkde/pkdew=q_kde/p_kde (χ2χ^2). This weight governs the local emphasis: forward KL treats all regions equally, reverse KL up-weights regions of high data density (encouraging precision), and χ2χ^2 up-weights regions of high generated density (penalizing spurious mass). 4.3 Unified Identifiability Combining the distribution-matching foundation 4.1 with the gradient flow machinery 4.2, we obtain a unified identifiability result. Theorem 4.7 (Unified identifiability (Proof in Appendix F.1)). Let k satisfy K1.–K4. and f be strictly convex with f(1)=0f(1)=0. If the generalized drifting velocity (5) vanishes identically, fkde≡v_f^kde 0, then p=qp=q. Corollary 4.8 (Loss landscape). The KDE-level f-divergence Df(qkde∥pkde)D_f(q_kde\|p_kde) satisfies: 1. Df≥0D_f≥ 0, with equality if and only if q=pq=p (identifiability); 2. dtDf≤0 ddtD_f≤ 0 along the Wasserstein gradient flow (energy dissipation); 3. the only equilibrium of the flow is q=pq=p. The unique global optimum of the KDE-level divergence is p=qp=q, and the energy is monotonically non-increasing along the flow. 4.4 The Drifting Model as Forward KL Gradient Flow We now show that the Drifting Model of Deng et al. (2026) is a special case of our framework, corresponding to the forward KL divergence f(u)=uloguf(u)=u u. Theorem 4.9 (Core equivalence; proof in Appendix G). Let kh(,)=exp(−‖−‖2/(2h2))k_h(x,y)= (-\|x-y\|^2/(2h^2)) with h>0h>0, and let p,q∈(ℝd)p,q (R^d). Then the drifting field (4) satisfies p,q()=h2(∇logpkde()−∇logqkde())for all ∈ℝd.V_p,q(x)=h^2 (∇ p_kde(x)-∇ q_kde(x) ) all x ^d. (7) The proof is a direct computation: the Gaussian kernel satisfies ∇kh(,)=−h2kh(,) _xk_h(x,y)= y-xh^2k_h(x,y), and substituting into the KDE score formula (Appendix E) gives exactly the mean-shift vectors p+V_p^+ and q−V_q^- from (4). Corollary 4.10 (Drifting = Forward KL Wasserstein gradient flow + KDE). The right-hand side of (7) is precisely h2KLkde()h^2v_KL^kde(x), the forward KL row of Table 1 scaled by h2h^2. Hence the Drifting Model’s velocity fields correspond to the Wasserstein-2 gradient flow of KL(qkde∥pkde)KL(q_kde\|p_kde), up to a time rescaling by h2h^2. This identification immediately imports the convergence and identifiability results of 4.3 to the Drifting Model. The identifiability proof, in particular, reduces to: p,q≡⇒Thm. 4.9∇logpkde=∇logqkde⇒Thm. 4.7p=q.V_p,q 0\; Thm.~ thm:core-equivalence\;∇ p_kde=∇ q_kde\; Thm.~ thm:identifiability\;p=q. 4.5 MMD Generators as ℒ2L^2 Gradient Flows The squared MMD functional ℱ[q]=12MMDk2(q,p)=12‖mkq−mkp‖ℋk2F[q]= 12MMD_k^2(q,p)= 12\|m_k^q-m_k^p\|_H_k^2 is not an f-divergence, but fits naturally into our framework. Proposition 4.11 (MMD gradient flow velocity; proof in Appendix H). Under K1.–K4., the WGF velocity of 12MMDk2(q,p) 12MMD_k^2(q,p) is MMD()=∇(pkde()−qkde())=∫∇k(,)d(p−q)().v_MMD(x)=∇ (p_kde(x)-q_kde(x) )= _xk(x,y)d(p-q)(y). (8) Note that MMDv_MMD is the gradient of the ℒ2L^2 density difference, while the f-divergence velocities in Table 1 involve the gradient of a nonlinear function of the density ratio. Both families are sample-computable via the KDE score formula, and the same identifiability argument applies (Remark in 4.3). 4.6 Mixed Gradient Flows Different divergences induce complementary failure modes. We propose mixing gradient flows to combine their strengths. Theorem 4.12 (Legitimacy of mixed gradient flows; proof in Appendix I). Let D1,D2D_1,D_2 be divergences (Di(q∥p)≥0D_i(q\|p)≥ 0, with equality iff q=pq=p), and α,β>0α,β>0 with α+β=1α+β=1. Define Dmix=αD1+βD2D_mix=α D_1+β D_2. Then: (a) DmixD_mix is a valid divergence; (b) its WGF velocity is mix=α1+β2v_mix=α\,v_1+β\,v_2; (c) dtDmix[qt]≤0 ddtD_mix[q_t]≤ 0 along the flow. Practical mixed drifting field. We propose combining the reverse KL and χ2χ^2 velocity fields: mix()=α⋅pkdeqkde(∇logpkde−∇logqkde)+β⋅qkdepkde(∇logpkde−∇logqkde),V_mix(x)=α· p_kdeq_kde(∇ p_kde-∇ q_kde)+β· q_kdep_kde(∇ p_kde-∇ q_kde), (9) corresponding to Dmix=αKLkde(p∥q)+βχkde2(q∥p)D_mix=α\,KL_kde(p\|q)+β\,χ^2_kde(q\|p). Referring to Remark 4.6, the reverse KL weight pkde/qkdep_kde/q_kde provides strong attraction toward high-density regions of p (precision-forcing, avoiding mode blurring), while the χ2χ^2 weight qkde/pkdeq_kde/p_kde penalizes spurious generated mass (coverage-forcing, avoiding mode collapse). Their combination reconciles mode-seeking and mode-covering behaviors. Experiments in 5 confirm this qualitative picture. 4.7 Extension to Riemannian Manifolds The Drifting Model Deng et al. (2026) trains in a semantic feature space that is empirically close to a hypersphere. This motivates extending our framework to Riemannian manifolds M. Two benefits emerge: 1. Vacuous boundary conditions. On compact manifolds without boundary (e.g., d−1S^d-1), the energy dissipation inequality (6) holds unconditionally (Theorem 4.5), eliminating the tail-decay assumptions required on ℝdR^d. 2. Richer kernel design. The adapted spherical assumptions K1S–K4S (Appendix J.2) admit kernels with qualitatively different weighting profiles. For example, the von Mises–Fisher (vMF) kernel kκ(,)=exp(κ⊤)k_κ(x,y)= (κ\,x y) provides a spherical analog of the Gaussian kernel, while the spherical logarithmic kernel (Appendix J.2, Proposition J.5) produces polynomial (inverse-distance) weighting, analogous to the Euclidean IMQ kernel, offering heavier tails and better global mode coverage. All results of 4.2–4.6—velocity fields, energy dissipation, identifiability, and mixed flows—extend to the Riemannian setting by replacing Euclidean gradients with Riemannian gradients and requiring the manifold analogues of K1.–K4.. Details and kernel verifications are given in Appendix J.2. 4.8 Training Pipeline Algorithm 1 summarizes the full training procedure of Gradient Flow Drifting. The framework is modular: one selects a divergence (or mixture), a kernel satisfying K1.–K4., and trains a one-step generator via the stop-gradient loss. Algorithm 1 Gradient Flow Drifting: training algorithm 0: Generator fθf_θ, data distribution p, source distribution pϵp_ ε 1: Divergence selection: Choose divergence(s) achieving distributional convergence (e.g., reverse KL, forward KL, χ2χ^2, or mixture thereof). 2: Velocity field: Derive the WGF velocity from the chosen divergence. 3: Kernel design: Select a kernel k satisfying Assumption 4.1 K1.–K4., or their Riemannian analogues. 4: for each training iteration do 5: Sample ϵ∼pϵ ε p_ ε; compute =fθ(ϵ)x=f_θ( ε) (generated samples). 6: Sample +∼py^+ p (data samples). 7: Mini-batch KDE velocity estimation: Compute kde()v^kde(x) over +\y^+\ and \x\. 8: Update: ℒ(θ)=ϵ[‖fθ(ϵ)−sg(fθ(ϵ)+kde(fθ(ϵ)))‖2]L(θ)=E_ ε [\|f_θ( ε)-sg(f_θ( ε)+v^kde(f_θ( ε)))\|^2 ]; θ←θ−η∇θℒθ←θ-η _θL. 9: end for 5 Experiments 5.1 Synthetic 2D Benchmarks We visualized the particle evolution under the velocity field of gradient flow using different implementations of divergence and kernel functions. As shown in Fig.1, the original drifting model and ℒ2L^2 flow drifting (with the same gradient flow with MMD) show a mode-covering training process as harsh punishment from divergence. We can easily find that they both have blur situations. The reverse KL divergence + χ2χ^2 divergence mixture flow drifting shows a totally different evolving process. This model almost only generate precise samples, but not struggles in mode collapse, it quickly explored all the modes. The original drifting model uses laplace kernel which may have some issues in high probability area. Since the Laplace kernel violates the assumptionK4., the gradient flow derived from it is mathematically only "weakly" defined, and during the convergence stage, it causes numerical instability (jittering) of particles near the data manifold. We can observe this phenomenon on the center of the swiss-roll distribution, the generation distribution has weird distortion, while RBF kernel version not. While the drifting model achieved empirical success due to the uniformity of semantic distribution, we can make it much more stable through the design of kernel function. (a) Reverse KL divergence + χ2χ^2 divergence (RBF kernel) (b) ℒ2L^2 distance(RBF kernel) (c) Forward KL divergence (Laplace kernel, original drifting model) (d) Forward KL divergence (RBF kernel) Figure 1: Training results with the velocity field of gradient flow under different implementations of divergence and kernel function on 2D-toy dataset. 6 Conclusion and Discussion We have found a new family of generative models, and given a mathematical equivalence between the Drifting Model and the Wasserstein gradient flow of the KL divergence under KDE approximation as a special case of our method, Gradient Flow Drifting. We have proved that under the aid of a finely designed kernel function, matching the pushforward distribution of KDE can achieve an approximation of the original distribution, and then extended this method onto Riemannian manifold which loosens the constraints on the kernel function, and make this method more suitable for the semantic space which is used in Drifting Model Deng et al. (2026). Besides, we did some preliminary experiments on synthetic benchmarks to validate the framework. Limitations. Our approach utilizes the convergence of the KDE distribution to induce the convergence of the original distribution. However, in practice, we can only approximate this KDE distribution using minibatches. As the dimension increases, the variance of the minibatch estimation will gradually increase, which will seriously affect the stability of the training and the final convergence effect, as other kernel-based methods suffer. Future work. Future work includes extending Gradient Flow Drifting to large-scale, high-dimensional datasets and diverse generation tasks, such as conditional generation and multi-modal generation. We plan to conduct comprehensive ablation experiments to evaluate the contribution of the combined reverse KL and χ2χ^2 gradient flows, explore the impact of different kernel functions and bandwidth choices and investigate acceleration techniques such as mini-batch particle updates and kernel approximation to improve computational efficiency and practical application capabilities. Meanwhile, we will follow the engineering techniques used in drifting model Deng et al. (2026), like training the model in semantic space and using multiple bandwidths. Furthermore, as theoretical analysis indicates, the Riemannian manifold is highly suitable for our approach. We will employ the hyperspherical semantic space constructed by JEPA, use ViT-based instead of a CNN-based architecture to achieve high computation efficiency, and make this type of model more scalable. References L. Ambrosio, N. Gigli, and G. Savaré (2005) Gradient flows: in metric spaces and in the space of probability measures. Springer. Cited by: Appendix D, §2. M. Arbel, A. Korba, A. Salim, and A. Gretton (2019) Maximum mean discrepancy gradient flow. Advances in neural information processing systems 32. Cited by: §2. A. Belhadji, D. Sharp, and Y. Marzouk (2025) Weighted quantization using mmd: from mean field to mean shift via gradient flows. arXiv preprint arXiv:2502.10600. Cited by: §1, §2. Y. Cheng (1995) Mean shift, mode seeking, and clustering. IEEE transactions on pattern analysis and machine intelligence 17 (8), p. 790–799. Cited by: §2. L. Chizat, M. Colombo, R. Colombo, and X. Fernández-Real (2026) Quantitative convergence of wasserstein gradient flows of kernel mean discrepancies. arXiv preprint arXiv:2603.01977. Cited by: §2. J. Choi, J. Choi, and M. Kang (2024) Scalable wasserstein gradient flow for generative modeling through unbalanced optimal transport. In International Conference on Machine Learning, p. 8629–8650. Cited by: §2. D. Comaniciu and P. Meer (2002) Mean shift: a robust approach toward feature space analysis. IEEE Transactions on pattern analysis and machine intelligence 24 (5), p. 603–619. Cited by: §2. M. Deng, H. Li, T. Li, Y. Du, and K. He (2026) Generative modeling via drifting. arXiv preprint arXiv:2602.04770. Cited by: §1, §1, §2, §3.5, §4.4, §4.7, §6, §6. G. K. Dziugaite, D. M. Roy, and Z. Ghahramani (2015) Training generative neural networks via maximum mean discrepancy optimization. arXiv preprint arXiv:1505.03906. Cited by: §2. R. Jordan, D. Kinderlehrer, and F. Otto (1998) The variational formulation of the fokker–planck equation. SIAM journal on mathematical analysis 29 (1), p. 1–17. Cited by: §2. Y. Li, K. Swersky, and R. Zemel (2015) Generative moment matching networks. In International conference on machine learning, p. 1718–1727. Cited by: §2. Z. Li and B. Zhu (2026) A long-short flow-map perspective for drifting models. arXiv preprint arXiv:2602.20463. Cited by: §1, §2. X. Nguyen, M. J. Wainwright, and M. I. Jordan (2010) Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Transactions on Information Theory 56 (11), p. 5847–5861. Cited by: §2. S. Nowozin, B. Cseke, and R. Tomioka (2016) F-gan: training generative neural samplers using variational divergence minimization. Advances in neural information processing systems 29. Cited by: §2. F. Santambrogio (2015) Optimal transport for applied mathematicians: calculus of variations, pdes, and modeling. Progress in Nonlinear Differential Equations and Their Applications, Vol. 87, Birkhäuser. External Links: Document Cited by: §2. B. K. Sriperumbudur, A. Gretton, K. Fukumizu, B. Schölkopf, and G. R. Lanckriet (2010) Hilbert space embeddings and metrics on probability measures. The Journal of Machine Learning Research 11, p. 1517–1561. Cited by: §J.1.1. M. Yi, Z. Zhu, and S. Liu (2023) Monoflow: rethinking divergence gans via the perspective of wasserstein gradient flows. In International Conference on Machine Learning, p. 39984–40000. Cited by: §2, §2. L. Zhou, S. Ermon, and J. Song (2025) Inductive moment matching. arXiv preprint arXiv:2503.07565. Cited by: §2. Appendix Appendix A Definitions and Standing Assumptions Probability measures. Let (ℝd)P(R^d) denote the set of Borel probability measures on ℝdR^d. For μ∈(ℝd)μ (R^d) and a measurable function g, we write μ[g]=∫ℝdg()dμ()E_μ[g]= _R^dg(y)dμ(y) whenever the integral is well-defined. Assumption A.1 (Kernel Regularity K1–K4). Let k:ℝd×ℝd→ℝk:R^d×R^d be a kernel satisfying: K1. Characteristic. The kernel k is positive definite, and the mean embedding μ↦∫k(⋅,)dμ()μ k(·,y)dμ(y) is injective on (ℝd)P(R^d). K2. Uniform gradient bound. Mk:=sup,∈ℝd‖∇k(,)‖<∞ M_k:= _x,y ^d \| _xk(x,y) \|<∞. K3. Strict positivity. k(,)>0k(x,y)>0 for all ,∈ℝdx,y ^d. K4. Differentiability. For every ∈ℝdy ^d, the map ↦k(,)x k(x,y) is continuously differentiable. Remark A.2 (Normalized KDE). For a translation-invariant kernel k(,)=φ(−)k(x,y)= (x-y) with φ∈L1(ℝd) ∈ L^1(R^d), ∫μkde()d=∥φ∥L1=:Zk _kde(x)dx=\| \|_L^1=:Z_k for every μ∈(ℝd)μ (R^d). The normalized density μ¯h:=μkde/Zk∈(ℝd) μ_h:= _kde/Z_k (R^d) satisfies ∇logμ¯h=∇logμkde∇ μ_h=∇ _kde, so all score-based formulas are unaffected by the normalization constant. Appendix B Injectivity of the KDE Operator We show that a characteristic kernel allows the recovery of the original measure from its KDE. Proposition B.1 (KDE Injectivity). Let k satisfy K1.. If μkde()=νkde() _kde(x)= _kde(x) for all ∈ℝdx ^d, then μ=νμ=ν. Proof. Let ℋkH_k be the RKHS of k and denote the mean embeddings mμ:=∫k(⋅,)dμ()m_μ:= k(·,y)dμ(y), mν:=∫k(⋅,)dν()∈ℋkm_ν:= k(·,y)dν(y) _k. By the reproducing property, mμ()=⟨mμ,k(⋅,)⟩ℋk=μkde()m_μ(x)= m_μ,k(·,x) _H_k= _kde(x) and similarly for ν. Hence μkde=νkde _kde= _kde pointwise implies ⟨mμ−mν,k(⋅,)⟩ℋk=0 m_μ-m_ν,k(·,x) _H_k=0 for all x. Since k(⋅,):∈ℝd\k(·,x):x ^d\ spans a dense subset of ℋkH_k, we obtain mμ=mνm_μ=m_ν in ℋkH_k. By K1. (injectivity of the mean embedding), μ=νμ=ν. ∎ Appendix C Regularity of KDE-Smoothed Densities We establish that the smoothness of μkde _kde is inherited entirely from the kernel, regardless of the regularity of μ. Theorem C.1 (KDE Regularity). Let k satisfy K2.–K4. and let μ∈(ℝd)μ (R^d) be arbitrary. Then: (i) μkde∈C1(ℝd) _kde∈ C^1(R^d) and differentiation commutes with integration: ∇μkde()=∫ℝd∇k(,)dμ(). _x\, _kde(x)= _R^d _xk(x,y)dμ(y). (10) (i) μkde()>0 _kde(x)>0 for all ∈ℝdx ^d. (i) sup‖∇μkde()‖≤Mk<∞ _x\|∇ _kde(x)\|≤ M_k<∞. Proof. (i): By K4., ↦k(,)x k(x,y) is C1C^1 for each y. By K2., ‖∇k(,)‖≤Mk\| _xk(x,y)\|≤ M_k for all ,x,y. Since the constant MkM_k is trivially μ-integrable (∫Mkdμ=Mk<∞ M_kdμ=M_k<∞ for any probability measure μ), the Leibniz integral rule yields (10) and continuity of the derivative. (i): By K3., k(,)>0k(x,y)>0 for all ,x,y. Since μ is a nonzero positive measure, μkde()=∫k(,)dμ()>0 _kde(x)= k(x,y)dμ(y)>0. (i): ‖∇μkde()‖=‖∫∇k(,)dμ()‖≤∫‖∇k(,)‖dμ()≤Mk\|∇ _kde(x)\|= \| _xk(x,y)dμ(y) \|≤ \| _xk(x,y)\|dμ(y)≤ M_k. ∎ Corollary C.2. For any p,q∈(ℝd)p,q (R^d), the KDE densities pkde,qkdep_kde,q_kde are strictly positive and C1C^1. In particular, the log-ratio log(pkde/qkde) (p_kde/q_kde) is well-defined and C1C^1, and the f-divergence machinery of the next section applies to pkde,qkdep_kde,q_kde with no additional assumptions on p,qp,q. Appendix D Wasserstein Gradient Flows of f-Divergences We recall the Wasserstein gradient flow (WGF) framework Ambrosio et al. [2005] for f-divergences between smooth positive densities. Definition D.1 (f-divergence). Let f:(0,∞)→ℝf:(0,∞) be convex with f(1)=0f(1)=0. For positive densities ρ,πρ,π on ℝdR^d, Df(ρ∥π):=∫ℝdπ()f(ρ()π())d.D_f(ρ\|π):= _R^dπ(x)\,f\! ( ρ(x)π(x) )dx. (11) Proposition D.2 (First Variation, Velocity, and Energy Dissipation). Let ρ,π∈C1(ℝd)ρ,π∈ C^1(R^d) with ρ,π>0ρ,π>0 everywhere and Df(ρ∥π)<∞D_f(ρ\|π)<∞. Then: (i) The first variation of Df(ρ∥π)D_f(ρ\|π) with respect to ρ is δDf(ρ∥π)δρ()=f′(ρ()π()). δ D_f(ρ\|π)δρ(x)=f \! ( ρ(x)π(x) ). (12) (i) The WGF particle velocity is f()=−∇f′(ρ()π()).v_f(x)=-∇ f \! ( ρ(x)π(x) ). (13) (i) Suppose in addition that the boundary terms arising from integration by parts vanish: limR→∞∫‖=Rρtf′(ρtπ)∇f′(ρtπ)⋅^S=0. _R→∞ _\|x\|=R _t\,f \! ( _tπ )\,∇ f \! ( _tπ )· n\,dS=0. (14) Then along any smooth WGF solution (ρt)t≥0( _t)_t≥ 0: dtDf(ρt∥π)=−∫ℝdρt()‖∇f′(ρt()π())‖2d≤0. ddtD_f( _t\|π)=- _R^d _t(x) \|∇ f \! ( _t(x)π(x) ) \|^2dx≤ 0. (15) Proof. (i) Let η be a smooth, compactly-supported perturbation with ∫ηd=0 =0. Set u:=ρ/πu:=ρ/π. Then dϵ|ϵ=0Df(ρ+ϵη∥π)=∫πf′(u)ηπd=∫f′(u)ηd, . ddε |_ε=0D_f(ρ+εη\|π)= π\,f (u)\, ηπdx= f (u)\, , giving (12). (i) Immediate from Definition 3.4: f=−∇δℱδρ=−∇f′(u)v_f=-∇ δρ=-∇ f (u). (i) Write Φ:=f′(ρt/π)=δℱδρ :=f ( _t/π)= δρ. Using the continuity equation (3): dtDf(ρt∥π)=∫ℝdΦ∂tρtd=∫ℝdΦ∇⋅(ρt∇Φ)d. ddtD_f( _t\|π)= _R^d \, _t _tdx= _R^d \,∇·\! ( _t∇ )dx. The product rule gives Φ∇⋅(ρt∇Φ)=∇⋅(Φρt∇Φ)−ρt‖∇Φ‖2 \,∇·( _t∇ )=∇·( \, _t∇ )- _t\|∇ \|^2. Integrating the divergence term over BRB_R and applying the divergence theorem yields the boundary integral ∫‖=RρtΦ(∇Φ⋅^)S, _\|x\|=R _t\, \,(∇ · n)\,dS, which vanishes as R→∞R→∞ by (14). Hence dtDf(ρt∥π)=−∫ℝdρt‖∇Φ‖2d≤0.∎ ddtD_f( _t\|π)=- _R^d _t\|∇ \|^2dx≤ 0. Remark D.3 (Boundary condition (14)). We distinguish two settings and don’t assume the vanishing-boundary condition (14). Within ℝdR^d, in our context we assume the original p,q∈2(ℝd)p,q _2(R^d) and the induced KDE distributions can be easily satisfy the vanishing-boundary condition. ℝdR^d. Condition (14) is an assumption on the joint tail behavior of ρt _t and π. It holds whenever ρt _t, π, their ratio, and its gradient have sufficient decay at infinity. In the context of this paper, ρt _t and π are KDE-smoothed densities, whose tails are governed by the kernel. For any kernel with exponential or faster decay (e.g., Gaussian, Matérn with ν>1ν>1, Pseudo-Huber), the KDE-smoothed densities inherit exponential decay and the condition is satisfied for all p,q∈2(ℝd)p,q _2(R^d) with finite second moments. For kernels with only polynomial decay (e.g., IMQ with exponent β), the condition requires β>(d−2)/2β>(d-2)/2. d−1S^d-1. If the ambient space is a compact Riemannian manifold M without boundary, then condition (14) is vacuous. Indeed, the divergence theorem on M reads ∫M∇g⋅Xvolg=∫∂Mg(X,^)S=0 _M _g· X\,dvol_g= _∂ Mg(X, n)\,dS=0 since ∂M=∅∂ M= . Consequently, the energy dissipation inequality (15) holds unconditionally for any f-divergence, any kernel satisfying the manifold analogues of K1.–K4., and any p,q∈(M)p,q (M). Proposition D.4 (Identifiability for f-divergence gradient flows). Let M be a connected Riemannian manifold (e.g., ℝdR^d or d−1S^d-1). Let ρ,π∈C1(M)ρ,π∈ C^1(M) with ρ,π>0ρ,π>0 and ∫Mρvol=∫Mπvol _Mρ\,dvol= _Mπ\,dvol. If f is strictly convex and the WGF velocity vanishes identically, f≡v_f 0, then ρ=πρ=π. Proof. By (13), f≡v_f 0 implies ∇f′(ρ/π)≡∇ f (ρ/π) 0 on M. Since M is connected and f′(ρ/π)∈C0(M)f (ρ/π)∈ C^0(M), we have f′(ρ/π)≡cf (ρ/π)≡ c for some constant c. Strict convexity of f implies f′f is strictly monotone, so ρ/π≡(f′)−1(c)=:λ>0ρ/π≡(f )^-1(c)=:λ>0. Integrating: ∫ρ=λ∫π ρ=λ π, hence λ=1λ=1 and ρ=πρ=π. ∎ Remark D.5 (Specific f-divergences). Writing u:=ρ()/π()u:=ρ(x)/π(x), we record the velocities for the divergences of primary interest: f(u)f(u) Divergence f′(u)f (u) Velocity f()v_f(x) ulogu u KL(ρ∥π)KL(ρ\|π) 1+logu1+ u ∇logπ−∇logρ∇ π-∇ ρ −logu- u KL(π∥ρ)KL(π\|ρ) −1/u-1/u πρ(∇logπ−∇logρ) πρ(∇ π-∇ ρ) 12(u−1)2 12(u-1)^2 χ2(ρ∥π)χ^2(ρ\|π) (u−1)(u-1) ρπ(∇logπ−∇logρ) ρπ(∇ π-∇ ρ) Appendix E KDE Score Formula We express the score of μkde _kde as an expectation under μ, establishing sample computability. Proposition E.1 (KDE Score). Let k satisfy K2.–K4. and μ∈(ℝd)μ (R^d). Then for all ∈ℝdx ^d: ∇logμkde()=∫∇k(,)dμ()∫k(,)dμ().∇ _kde(x)= _xk(x,y)dμ(y) k(x,y)dμ(y). (16) Proof. By Theorem C.1(i) and (i), μkde∈C1 _kde∈ C^1 and μkde>0 _kde>0, so ∇logμkde=∇μkde/μkde∇ _kde=∇ _kde/ _kde. Substituting (10) gives (16). ∎ Corollary E.2 (Gaussian Kernel Score). For kh(,)=exp(−‖−‖2/(2h2))k_h(x,y)= \! (-\|x-y\|^2/(2h^2) ): ∇logμkde()=1h2(μ,kh[∣]−),∇ _kde(x)= 1h^2 (E_μ,k_h[y ]-x ), (17) where μ,kh[∣]:=∫kh(,)dμ()/∫kh(,)dμ()E_μ,k_h[y ]:= \,k_h(x,y)dμ(y) / k_h(x,y)dμ(y). Proof. For the Gaussian kernel, ∇kh(,)=−h2kh(,) _xk_h(x,y)= y-xh^2k_h(x,y). Substituting into (16): ∇logμkde()=∫−h2kh(,)dμ()∫kh(,)dμ()=1h2(∫kh(,)dμ()∫kh(,)dμ()−).∇ _kde(x)= y-xh^2k_h(x,y)dμ(y) k_h(x,y)dμ(y)= 1h^2\! (\! \,k_h(x,y)dμ(y) k_h(x,y)dμ(y)-x\! ). ∎ Appendix F Identifiability Throughout this section, π:=pkdeπ:=p_kde and ρ0:=q0,kde _0:=q_0,kde denote the KDE-smoothed densities of the target distribution p and the initial generated distribution q0q_0, respectively. By Theorem C.1, both are strictly positive and C1C^1 under Assumptions K1.–K4., with no regularity conditions on p or q0q_0. We consider the Wasserstein-2 gradient flow of the f-divergence ℱ[ρ]=Df(ρ∥π)F[ρ]=D_f(ρ\|π), i.e., the continuity equation ∂tρt=∇⋅(ρt∇Φt),Φt:=f′(ρtπ), _t _t=∇· ( _t∇ _t ), _t:=f \! ( _tπ ), (18) with initial condition ρ0>0 _0>0. Remark F.1 (Density-level vs. particle-level flow). Equation (18) describes the evolution of a smooth density ρt _t in the Wasserstein-2 metric. In the practical training algorithm 4.8, one evolves a finite collection of particles whose empirical distribution is qtq_t, and estimates the velocity using the KDE density qt,kdeq_t,kde from a mini-batch. The convergence theorems below apply to the idealized density-level flow; the particle system is viewed as a consistent approximation that converges to this flow in the large-sample limit. F.1 Proof of Identifiability (Theorem 4.7) Proof. The proof proceeds in two steps. Step 1: Constant density ratio. By definition, fkde()=−∇f′(qkde()/pkde())v_f^kde(x)=-∇ f \! (q_kde(x)/p_kde(x) ). The hypothesis fkde≡v_f^kde 0 thus implies ∇f′(qkde()pkde())=for all .∇ f \! ( q_kde(x)p_kde(x) )=0 all x. (19) Let M be the ambient space (ℝdR^d or a connected Riemannian manifold). By Theorem C.1, pkdep_kde and qkdeq_kde are C1C^1 with pkde,qkde>0p_kde,q_kde>0, so the ratio u():=qkde()/pkde()u(x):=q_kde(x)/p_kde(x) is C1C^1 and strictly positive. By the composition rule, Φ:=f′∘u∈C1(M) :=f u∈ C^1(M). Since M is connected and ∇Φ≡∇ 0, Φ is a constant: f′(u())=cf (u(x))=c for some c∈ℝc . Strict convexity of f implies that f′f is strictly monotone, hence injective. Therefore u()=(f′)−1(c)=:λ>0u(x)=(f )^-1(c)=:λ>0 for all x, i.e., qkde=λpkdeq_kde=λ\,p_kde everywhere. Step 2: KDE matching implies distribution matching. Integrating both sides over M: ∫Mqkded=λ∫Mpkded. _Mq_kde\,dx=λ _Mp_kde\,dx. For a translation-invariant kernel, ∫qkde=∫pkde=Zk q_kde= p_kde=Z_k (Remark A.2); on a compact manifold the same equality holds by symmetry. Hence λ=1λ=1 and qkde=pkdeq_kde=p_kde. By the injectivity of the KDE operator under characteristic kernels (Proposition B.1), q=pq=p. ∎ Remark F.2 (MMD identifiability). For the MMD functional (Proposition H.2), MMD≡v_MMD 0 implies ∇(pkde−qkde)≡∇(p_kde-q_kde) 0. By connectedness, pkde−qkde=cp_kde-q_kde=c. Integrating gives c=0c=0, so pkde=qkdep_kde=q_kde, and Proposition B.1 yields p=qp=q. Appendix G Core Equivalence: Drifting as KL Gradient Flow Theorem G.1 (Core Equivalence). Let khk_h be the Gaussian kernel and p,q∈(ℝd)p,q (R^d). Then the drifting field (4) satisfies p,q()=h2(∇logpkde()−∇logqkde())=h2KL()|ρ=qkde,π=pkde, V_p,q(x)=h^2 (∇ p_kde(x)-∇ q_kde(x) )=h^2\,v_KL(x) |_ρ=q_kde,\;π=p_kde, (20) where KL=∇logπ−∇logρv_KL=∇ π-∇ ρ is the WGF velocity of the KL divergence DKL(ρ∥π)D_KL(ρ\|π) (Remark D.5). Proof. Apply Corollary E.2 to p and q respectively: h2∇logpkde() h^2∇ p_kde(x) =p,kh[∣]−, =E_p,k_h[y ]-x, (21) h2∇logqkde() h^2∇ q_kde(x) =q,kh[∣]−. =E_q,k_h[y ]-x. (22) Subtracting, the x terms cancel: h2(∇logpkde()−∇logqkde())=p,kh[∣]−q,kh[∣]=p,q(),h^2 (∇ p_kde(x)-∇ q_kde(x) )=E_p,k_h[y ]-E_q,k_h[y ]=V_p,q(x), (23) The identification with KLv_KL follows from Remark D.5 with ρ=qkdeρ=q_kde, π=pkdeπ=p_kde (both smooth and positive by Corollary C.2). ∎ Appendix H MMD Gradient Flow Definition H.1 (Squared MMD). For a positive-definite kernel k with RKHS ℋkH_k, the squared MMD between p,q∈(ℝd)p,q (R^d) is MMDk2(q,p):=12‖mq−mp‖ℋk2,MMD_k^2(q,p):= 12\|m_q-m_p\|_H_k^2, (24) where mμ:=∫k(⋅,)dμ()m_μ:= k(·,y)dμ(y) is the mean embedding. Proposition H.2 (MMD Flow Velocity). Let k satisfy K1.–K4.. The WGF of ℱ[q]:=12MMDk2(q,p)F[q]:= 12MMD_k^2(q,p) has particle velocity MMD()=∫∇k(,)d(p−q)()=∇(pkde()−qkde()).v_MMD(x)= _xk(x,y)d(p-q)(y)=∇ (p_kde(x)-q_kde(x) ). (25) Proof. Perturb q→(1−ϵ)q+ϵδq→(1-ε)q+ε _x, so mqϵ=mq+ϵ(k(⋅,)−mq)m_q_ε=m_q+ε(k(·,x)-m_q). Expanding and differentiating at ϵ=0ε=0: δℱδq()=⟨mq−mp,k(⋅,)⟩ℋk=mq()−mp()=qkde()−pkde(), δ q(x)= m_q-m_p,\,k(·,x) _H_k=m_q(x)-m_p(x)=q_kde(x)-p_kde(x), using the reproducing property. The velocity is MMD=−∇δℱδq=∇(pkde−qkde)v_MMD=-∇ δ q=∇(p_kde-q_kde). By Theorem C.1(i), differentiation under the integral gives the first equality in (25). ∎ Appendix I Superposition of Gradient Flows Proposition I.1 (Superposition). Let ℱ1,ℱ2:(ℝd)→ℝF_1,F_2:P(R^d) be functionals with well-defined first variations. For any α,β≥0α,β≥ 0, the mixed functional ℱmix:=αℱ1+βℱ2F_mix:= _1+ _2 has WGF velocity mix()=α1()+β2(),v_mix(x)=α\,v_1(x)+β\,v_2(x), (26) where i=−∇δℱiδρv_i=-∇ _iδρ. Proof. By linearity of the first variation and the gradient: mix=−∇δℱmixδρ=−∇(αδℱ1δρ+βδℱ2δρ)=α1+β2.v_mix=-∇ _mixδρ=-∇\! (α _1δρ+β _2δρ )= _1+ _2. ∎ Unified framework. Table 2 summarizes the KDE-based gradient flow velocities, all sample-computable via the score formula (16). Table 2: Unified KDE-based gradient flow framework. All velocities are expressed via pkde,qkdep_kde,q_kde and are sample-computable. Functional WGF velocity ()v(x) Model DKL(q¯h∥p¯h)D_KL( q_h\| p_h) ∇logpkde−∇logqkde∇ p_kde-∇ q_kde Drifting (/h2/h^2) DKL(p¯h∥q¯h)D_KL( p_h\| q_h) ∇(pkde/qkde)∇(p_kde/q_kde) - χ2(q¯h∥p¯h)χ^2( q_h\| p_h) −∇(qkde/pkde)-∇(q_kde/p_kde) - MMDkh2(q,p)MMD_k_h^2(q,p) ∇(pkde−qkde)∇(p_kde-q_kde) MMD generators Appendix J Analysis of Specific Kernel Families We verify Assumptions K1.–K4. for several kernel families. For each kernel, we also compute the score weight w(r)w(r) defined by ∇logk(,)=−w(r)(−) _x k(x,y)=-w(r)(x-y) where r:=‖−‖r:=\|x-y\|. J.1 Euclidean Kernels J.1.1 Gaussian (RBF) Kernel Proposition J.1. The Gaussian kernel kh(,):=exp(−‖−‖2/(2h2))k_h(x,y):= \! (-\|x-y\|^2/(2h^2) ) satisfies K1.–K4.. Its score weight is w(r)=1/h2w(r)=1/h^2 (constant). If khk_h is the Gaussian kernel, one additionally obtains μkde∈C∞(ℝd) _kde∈ C^∞(R^d) with every derivative uniformly bounded for any μ∈(ℝd)μ (R^d). Proof. K1.: The Fourier transform φ^()=(2πh2)d/2exp(−h2‖2/2)>0 ( ω)=(2π h^2)^d/2 (-h^2\| ω\|^2/2)>0 for all ω. By Sriperumbudur et al. [2010, Theorem 9], a translation-invariant kernel with strictly positive Fourier transform is characteristic. K2.: ‖∇kh‖=rh2e−r2/(2h2)\| _xk_h\|= rh^2e^-r^2/(2h^2) where r=‖−‖r=\|x-y\|. Maximizing over r≥0r≥ 0: the maximum occurs at r=hr=h, giving Mk=1he−1/2=1he<∞M_k= 1he^-1/2= 1h e<∞. K3.: exp(⋅)>0 (·)>0. K4.: kh∈C∞(ℝd×ℝd)k_h∈ C^∞(R^d×R^d). ∎ J.1.2 Matérn-ν Kernel Proposition J.2. The Matérn kernel with smoothness ν>0ν>0 and length scale ℓ>0 >0, kν,ℓ(,)=21−νΓ(ν)(2νrℓ)νKν(2νrℓ),r:=‖−‖,k_ν, (x,y)= 2^1-ν (ν) ( 2ν\,r )^\!νK_ν\! ( 2ν\,r ), r:=\|x-y\|, where KνK_ν is the modified Bessel function of the second kind, satisfies all four assumptions if and only if ν>1ν>1. Proof. K1.: The Fourier transform φ^()∝(2ν/ℓ2+‖2)−(ν+d/2)>0 ( ω) (2ν/ ^2+\| ω\|^2)^-(ν+d/2)>0 for all ω and ν>0ν>0, so kν,ℓk_ν, is characteristic. K3.: kν,ℓ>0k_ν, >0 since Kν(z)>0K_ν(z)>0 for z>0z>0 and kν,ℓ(,)=1k_ν, (x,x)=1. K4.: The sample path regularity theory of Matérn processes shows that kν,ℓk_ν, is C⌈ν⌉−1C ν -1 as a function of r. When ν≤1ν≤ 1, a cusp at r=0r=0 (Laplace-like) violates K4.. When ν>1ν>1, at least C1C^1 regularity is guaranteed. K2.: When ν>1ν>1, ‖∇kν,ℓ‖\| _xk_ν, \| is continuous (by K4.) and decays exponentially as r→∞r→∞, hence bounded. When ν≤1ν≤ 1, the gradient diverges at r=0r=0. ∎ J.1.3 Summary of Euclidean Kernels Table 3: Verification of K1.–K4. for Euclidean kernels. Kernel K1. K2. K3. K4. Status Gaussian ✓ ✓ ✓ ✓ ✓ IMQ ✓ ✓ ✓ ✓ ✓ Pseudo-Huber ✓ ✓ ✓ ✓ ✓ Matérn (ν>1ν>1) ✓ ✓ ✓ ✓ ✓ Laplace ✓ ✓ ✓ ✗ ✗ ✓: verified; ✗: fails. J.2 Spherical Kernels On the unit sphere d−1:=∈ℝd:‖=1S^d-1:=\x ^d:\|x\|=1\ with the round metric, we adapt the kernel assumptions. Adapted assumptions. With ∇ _S denoting the Riemannian gradient on d−1S^d-1: K1S. k is characteristic on (d−1)P(S^d-1). K2S. sup,∈d−1‖∇,k(,)‖<∞ _x,y ^d-1\| _S,xk(x,y)\|<∞. K3S. k(,)>0k(x,y)>0 for all ,∈d−1x,y ^d-1. K4S. ↦k(,)x k(x,y) is C1C^1 on d−1S^d-1. J.2.1 von Mises–Fisher (vMF) Kernel Proposition J.3. The vMF kernel kκ(,):=exp(κ⊤)k_κ(x,y):= (κ\,x y) with κ>0κ>0 satisfies K1S–K4S. Proof. K1S: The Mercer expansion kκ(,)=∑ℓ=0∞aℓ(κ)∑mYℓm()Yℓm()¯k_κ(x,y)= _ =0^∞a_ (κ) _mY_ ^m(x) Y_ ^m(y) has coefficients aℓ(κ)>0a_ (κ)>0 for all ℓ≥0 ≥ 0. Since all eigenvalues are positive, kκk_κ is universal and hence characteristic. K3S: exp(κ⊤)>0 (κ\,x y)>0. K4S: kκ∈C∞k_κ∈ C^∞. K2S: The Riemannian gradient is ∇,kκ=κkκProjT() _S,xk_κ=κ\,k_κ\,Proj_T_xS(y) where ProjT()=−(⊤)Proj_T_xS(y)=y-(x y)x. Since kκ≤eκk_κ≤ e^κ and ‖ProjT()‖≤1\|Proj_T_xS(y)\|≤ 1: ‖∇,kκ‖≤κeκ<∞\| _S,xk_κ\|≤κ e^κ<∞. ∎ Proposition J.4 (Spherical Core Equivalence). Let p,q∈(d−1)p,q (S^d-1) and kκk_κ be the vMF kernel. Define μkde():=∫d−1kκ(,)dμ() _kde(x):= _S^d-1k_κ(x,y)dμ(y). (i) The spherical KDE score is ∇logμkde()=κProjT(μ,kκ[∣]). _S _kde(x)=κ\,Proj_T_xS\! (E_μ,k_κ[y ] ). (27) (i) The spherical drifting field p,q():=ProjT(p,kκ[∣]−q,kκ[∣])V_p,q^S(x):=Proj_T_xS\! (E_p,k_κ[y ]-E_q,k_κ[y ] ) satisfies p,q()=1κ(∇logpkde()−∇logqkde()).V_p,q^S(x)= 1κ ( _S p_kde(x)- _S q_kde(x) ). (28) Proof. (i) The ambient gradient of kκ(,)k_κ(x,y) with respect to x is κkκ(,) \,k_κ(x,y). Hence the ambient gradient of logμkde _kde is κ∫kκdμ/∫kκdμ=κμ,kκ[∣]κ \,k_κdμ/ k_κdμ=κ\,E_μ,k_κ[y ]. The Riemannian gradient is its tangential projection, giving (27). (i) By linearity of ProjTProj_T_xS: 1κ(∇logpkde−∇logqkde)=ProjT(p,kκ[∣]−q,kκ[∣])=p,q(). 1κ ( _S p_kde- _S q_kde )=Proj_T_xS\! (E_p,k_κ[y ]-E_q,k_κ[y ] )=V_p,q^S(x). ∎ J.2.2 Spherical Logarithmic Kernel Proposition J.5. Let c>0c>0 and 0<α<12+c0<α< 12+c. The spherical logarithmic kernel defined by kc,α(,):=−log(α(1−⊤+c))k_c,α(x,y):=- (α(1-x y+c) ) (29) satisfies all spherical assumptions K1S–K4S. Proof. Let z:=⊤z:=x y. Since ,∈d−1x,y ^d-1, we have z∈[−1,1]z∈[-1,1]. Let the argument of the logarithm be denoted as g(z):=α(1−z+c)g(z):=α(1-z+c). K3S (Strict positivity): We analyze the bounds of g(z)g(z). Since z∈[−1,1]z∈[-1,1], the term 1−z+c1-z+c achieves its minimum at z=1z=1 (yielding c) and its maximum at z=−1z=-1 (yielding 2+c2+c). Therefore, 0<α⋅c≤g(z)≤α(2+c).0<α· c≤ g(z)≤α(2+c). (30) By the given condition α<12+cα< 12+c, we strictly have g(z)<1g(z)<1. Thus, 0<g(z)<10<g(z)<1 for all ,∈d−1x,y ^d-1. Consequently, kc,α(,)=−log(g(z))>0k_c,α(x,y)=- (g(z))>0 everywhere. K4S (Continuous differentiability): Since g(z)≥α⋅c>0g(z)≥α· c>0, the argument of the logarithm is strictly bounded away from zero. The functions z=⊤z=x y and t↦−log(t)t - (t) (for t>0t>0) are smooth (C∞C^∞). Therefore, their composition kc,α∈C∞(d−1×d−1)k_c,α∈ C^∞(S^d-1×S^d-1), which trivially implies C1C^1. K2S (Uniform gradient bound): The ambient gradient of the kernel with respect to x is: ∇kc,α(,)=dz[−log(α(1−z+c))]∇(⊤)=11−z+c. _xk_c,α(x,y)= ddz [- (α(1-z+c)) ] _x(x y)= 11-z+cy. (31) The Riemannian gradient is the projection onto the tangent space Td−1T_xS^d-1: ∇,kc,α(,)=ProjT(∇kc,α)=11−z+c(−(⊤)). _S,xk_c,α(x,y)=Proj_T_xS( _xk_c,α)= 11-z+c (y-(x y)x ). (32) We compute its norm. Since x and y are unit vectors, ‖−z‖2=‖2−2z(⊤)+z2‖2=1−z2\|y-zx\|^2=\|y\|^2-2z(x y)+z^2\|x\|^2=1-z^2. Thus: ‖∇,kc,α(,)‖=1−z21−z+c.\| _S,xk_c,α(x,y)\|= 1-z^21-z+c. (33) Since 1−z2≤1 1-z^2≤ 1 and 1−z+c≥c>01-z+c≥ c>0, we have the uniform bound ‖∇,kc,α‖≤1c<∞\| _S,xk_c,α\|≤ 1c<∞. K1S (Characteristic): We expand the kernel into a Taylor series: kc,α(,) k_c,α(x,y) =−logα−log((1+c)−z) =- α- ((1+c)-z ) (34) =−logα−log(1+c)−log(1−z1+c). =- α- (1+c)- (1- z1+c ). (35) Using the Maclaurin series for −log(1−x)=∑n=1∞xnn- (1-x)= _n=1^∞ x^nn (which converges absolutely since |z1+c|≤11+c<1 | z1+c |≤ 11+c<1), we obtain: kc,α(,)=−log(α(1+c))⏟a0+∑n=1∞1n(1+c)n⏟an(⊤)n.k_c,α(x,y)= - (α(1+c) )_a_0+ _n=1^∞ 1n(1+c)^n_a_n(x y)^n. (36) From the user condition α<12+c<11+cα< 12+c< 11+c, we have α(1+c)<1α(1+c)<1, which strictly implies a0>0a_0>0. For all n≥1n≥ 1, since 1+c>11+c>1, we clearly have an>0a_n>0. Because the kernel kc,αk_c,α can be expressed as a power series f(⊤)=∑n=0∞an(⊤)nf(x y)= _n=0^∞a_n(x y)^n where an>0a_n>0 for all n≥0n≥ 0, Schoenberg’s theorem guarantees that it is strictly positive definite on d−1S^d-1 for any dimension d≥2d≥ 2. Thus, it is a universal (and therefore characteristic) kernel. ∎ Remark J.6 (Spherical Score for Logarithmic Kernel). Following the same logic as Proposition J.4, the spherical KDE score for this logarithmic kernel can be explicitly derived. Define the pairwise weight function Wc(,):=11−⊤+cW_c(x,y):= 11-x y+c. The ambient gradient of μkde _kde is ∫Wc(,)dμ() W_c(x,y)ydμ(y). Consequently, the Riemannian score is: ∇logμkde()=ProjT(∫Wc(,)dμ()∫kc,α(,)dμ()). _S _kde(x)=Proj_T_xS ( W_c(x,y)ydμ(y) k_c,α(x,y)dμ(y) ). (37) Unlike the vMF kernel, the weighting factor Wc(,)W_c(x,y) here is polynomial (specifically, inversely proportional to distance squared, analogous to the Euclidean IMQ kernel), which typically produces heavier tails and better global mode coverage.