Paper deep dive
Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies
Ziheng Cheng, Xin Guo, Huyên Pham, Yufei Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/18/2026, 1:49:47 PM
Summary
This paper proposes a model-free reinforcement learning framework for continuous-time extended mean field control (MFC) problems using deterministic feedback policies. The authors derive a deterministic policy gradient formula based on an advantage-rate function on the Wasserstein space, incorporating both action derivatives and measure-derivative terms. They introduce a martingale-based learning principle and propose a continuous-time deep deterministic policy gradient algorithm that combines particle approximations, measure-dependent neural networks, and temporal-difference learning. Numerical experiments on stochastic Cucker-Smale consensus control and optimal liquidation demonstrate the method's efficiency, stability, and robustness.
Entities (10)
Relation Signals (8)
McKean-Vlasov Dynamics → governs → Extended Mean Field Control
confidence 97% · the controlled state of a (representative) agent is governed by a McKean-Vlasov system
Deterministic Policy Gradient Formula → expressedthrough → Advantage-Rate Function
confidence 96% · derive a deterministic policy gradient formula expressed through an advantage-rate function on the Wasserstein space.
Advantage-Rate Function → definedon → Wasserstein Space
confidence 95% · advantage-rate function on the Wasserstein space
Extended Mean Field Control → solvedby → Deterministic Feedback Policies
confidence 95% · This paper develops a model-free reinforcement learning framework for continuous-time extended mean field control problems... We adopt deterministic feedback policies
Deep Deterministic Policy Gradient Algorithm → appliedto → Cucker-Smale Consensus Control
confidence 93% · Numerical experiments on stochastic Cucker–Smale consensus control
Deep Deterministic Policy Gradient Algorithm → appliedto → Optimal Liquidation
confidence 93% · optimal liquidation with trade crowding demonstrate the efficiency
Deep Deterministic Policy Gradient Algorithm → uses → Temporal-Difference Learning
confidence 92% · combining particle approximations, measure-dependent neural networks, temporal-difference learning
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper develops a model-free reinforcement learning framework for continuous--time extended mean field control problems, where both the dynamics and reward may depend on the joint distribution of states and controls. We adopt deterministic feedback policies, under which the state--action distribution is induced directly as a push--forward of the state law. This avoids optimization over stochastic kernels and bypasses key limitations of existing approaches in extended mean field settings. We first establish a model--free sensitivity formula for parameterized McKean--Vlasov dynamics and use it to derive a deterministic policy gradient formula expressed through an advantage--rate function on the Wasserstein space. We then refine this formula by introducing local value and advantage--rate representations that depend on the state, action, and joint state--action distribution, yielding a policy gradient that includes both action derivatives and measure--derivative terms with respect to the control distribution. These characterizations lead to a martingale--based learning principle and motivate a continuous--time deep deterministic policy gradient algorithm combining particle approximations, measure--dependent neural networks, temporal--difference learning, and exploration in either action or parameter space. Numerical experiments on stochastic Cucker--Smale consensus control and optimal liquidation with trade crowding demonstrate the efficiency, stability, and robustness of the proposed method, including problems with explicit dependence on the control distribution.
Tags
Links
- Source: https://arxiv.org/abs/2607.11005v1
- Canonical: https://arxiv.org/abs/2607.11005v1
Trouble viewing inline? Open PDF directly →
Full Text
120,208 characters extracted from source content.
Expand or collapse full text
Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies Ziheng Cheng University of California, Berkeley. Email: ziheng_cheng@berkeley.edu Xin Guo University of California, Berkeley. Email: xinguo@berkeley.edu Huyên Pham Ecole Polytechnique, CMAP, Email: huyen.pham@polytechnique.edu Yufei Zhang Imperial College London. Email: yufei.zhang@imperial.ac.uk Abstract This paper develops a model-free reinforcement learning framework for continuous-time extended mean field control problems, where both the dynamics and reward may depend on the joint distribution of states and controls. We adopt deterministic feedback policies, under which the state–action distribution is induced directly as a push-forward of the state law. This avoids optimization over stochastic kernels and bypasses key limitations of existing approaches in extended mean field settings. We first establish a model-free sensitivity formula for parameterized McKean–Vlasov dynamics and use it to derive a deterministic policy gradient formula expressed through an advantage-rate function on the Wasserstein space. We then refine this formula by introducing local value and advantage-rate representations that depend on the state, action, and joint state–action distribution, yielding a policy gradient that includes both action derivatives and measure-derivative terms with respect to the control distribution. These characterizations lead to a martingale-based learning principle and motivate a continuous-time deep deterministic policy-gradient algorithm combining particle approximations, measure-dependent neural networks, temporal-difference learning, and exploration in either action or parameter space. Numerical experiments on stochastic Cucker–Smale consensus control and optimal liquidation with trade crowding demonstrate the efficiency, stability, and robustness of the proposed method, including problems with explicit dependence on the control distribution. 1 Introduction (Extended) MFC. (Extended) mean field control (MFC) provides a tractable framework for large population stochastic games, by optimizing the collective behavior of homogeneous and interacting agents through a central planner [1]. In the mean-field limit, the controlled state of a (representative) agent is governed by a McKean-Vlasov system dXsα=b(s,Xsα,αs,ℙ(Xsα,αs))ds+σ(s,Xsα,αs,ℙ(Xsα,αs))dWs, splitdX^α_s=b(s,X^α_s,α_s,P_(X^α_s,α_s))ds+σ(s,X^α_s,α_s,P_(X^α_s,α_s))dW_s, split (1.1) where both the drift and diffusion coefficients depend on the distribution of the state Xα X^α and the control α α. Although this limit removes the explicit dependence on the number of agents, the associated extended MFC is intrinsically an infinite-dimensional control problem: the value function and the optimal feedback generally depend not only on time and the individual state, but also on the state distribution. The most fundamental theoretical ingredients in the development of MFC theory are the dynamic programming principle for the lifted value function defined on the space of probability measures, and the invariance property of this value function along the deterministic state flow (see e.g., [28, 8]). Continuous-time RL with deterministic policy. In extended mean field control, both the dynamics and reward may depend on the joint distribution of states and controls. To address this challenge, we adopt deterministic feedback policies. A deterministic policy φ(t,x,μ) (t,x,μ) selects an action directly based on the current time, state and state law, so that the state law together with the feedback map completely determines the associated state–action distribution through the push-forward measure Γtφ=(id,φ(t,⋅,μt))#μt, _t = (id, (t,·, _t) )_\# _t, with μt _t being the current state distribution. For continuous-time reinforcement learning (RL), recent analysis using deterministic policies [7, 15] have shown superiority over both the discretization approach and the stochastic policies. Analytically, the state–action distribution is induced directly as a push-forward of the state law. Hence the policy optimization over controls is reduced to an optimization over a parameterized family of deterministic policies φθ _θ (with respect to parameter θ θ). Numerically, deterministic policies have also been shown to offer improved computational efficiency and greater stability [7]. Learning for discrete-time MFC. In the RL framework of MFCs, the agent does not know the state coefficients b b and σ σ, nor the reward functions, and instead, learns by interacting with the controlled system. In discrete time, dynamic programming principles have been established for both state- and state-action value functions, leading to value-based algorithms including Q-learning and integrated Q-functions based on local policies [25, 6, 13]. More recently, model-free policy-gradient methods have also been proposed in [24], culminating in the MF-REINFORCE algorithm with theoretical guarantees on the bias and variance of the policy-gradient estimator. Our work for continuous-time extended MFC. This paper develops an RL framework for continuous-time extended MFC, adopting the deterministic policy approach. We first develop a model-free deterministic policy gradient formula for continuous-time extended MFC (Theorem 3.1). This representation expresses the gradient of value function with respect to the policy parameter as an integral involving the derivative of an advantage-rate function along the flow of state distributions. This advantage-rate function is lifted to a function on the space of probability measures and is characterized jointly with the value function through an invariance principle along the flow of state distributions. These results are derived using a model-free sensitivity formula for parametric McKean–Vlasov dynamics (Theorems 2.1 and 2.3), which relies on a performance difference lemma (Lemma 6.2) and the robustness of the state distribution with respect to model parameters. These results for continuous-time extended MFC are consistent with their discrete-time counterpart regarding lifted (state–action) value functions and Q-functions [25, 6, 13]. We then refine the deterministic policy gradient formula by explicitly accounting for its dependence on the joint state–action distribution (Theorem 3.2). Specifically, we decompose the lifted value function and the lifted advantage-rate function into integrals of local functions of the state, action, and the joint state–action distribution. The resulting policy gradient contains both the usual derivative of the local advantage-rate function with respect to the individual action and an additional measure-derivative term with respect to control distribution. We further establish a martingale characterization to learn these local representations directly from observed state and control trajectories. Just as the incorporation of local policies facilitates efficient learning in the discrete-time MFC [13], this local advantage-rate function allows more informative learning than directly learning from the flow of state laws (Section 4). This approach applies to general extended MFC problems, and removes the restrictive assumption in [27, Remark 3.6] which requires a separable structure on the coefficients and explicit and known dependence on the control variable. Finally, we propose a continuous-time deep deterministic policy-gradient algorithm for extended MFC. The algorithm combines measure-dependent neural networks, particle approximations of the McKean-Vlasov dynamics, temporal-difference learning based on the aforementioned martingale characterization, and action- or parameter-space exploration. The use of finite-dimensional distribution embeddings is consistent with the neural approximation principles developed in [23, 36], while the deterministic actor avoids direct optimization over a space of stochastic kernels. Numerical experiments for consensus control in a stochastic Cucker-Smale system and optimal liquidation with mean-field market impact demonstrate the efficiency and robustness of the proposed method, including in settings where the dynamics depend explicitly on the distribution of the control. Earlier work for RL in continuous-time-space MFC. Many earlier studies on continuous-time MFC reinforcement-learning have been built on stochastic, or exploratory, policies; and focuses on the setting where the coefficients depend only on the law of the state process. In this exploratory formulation, the admissible policy class is enlarged to stochastic policies that map the time, state, and state distribution to a probability measure over the action space. However, RL algorithms under the exploratory control formulation appear too restrictive for the infinite-dimensional nature of MFC. For example, in [27], the state coefficients are required to possess a separable structure, and the explicit dependence on the control variable must be known in order to characterize the mean-field q q-function. In [38, 39], algorithms require learning two q q-functions: an integrated q q-function identified via weak martingale conditions by testing against all stochastic policies, and an essential q q-function used for policy improvement. In [33], the exploratory HJB equation contains an additional nonlinear functional of the policy, and the optimal policy is characterized by a two-layer fixed-point problem over the space of stochastic policies. The corresponding algorithms in [34] then face the issue that data generated by the relaxed-control formulation are not directly observable, requiring additional approximation by discretely sampled exploratory actions, as in [37, 17]. To see analytically the fundamental obstacle of stochastic-policy methods for learning extended MFC, consider a stochastic Markov policy π(da|t,x,μ) π(da|t,x,μ) and its induced state-action distribution Γtπ(dx,da)=(id,π(da|t,⋅,μt))#μt _t^π(dx,da)= (id,π(da|t,·, _t) )_\# _t. In a standard MFC problem whose coefficients depend only on the state distribution, stochastic policies produce an auxiliary generator by integrating the original state generator over the action space with respect to the stochastic policy π π, which is linear in the stochastic policy. This linearity is central for stochastic-policy based algorithms [11, 27, 3]. When applying the exploratory control formulation to extended MFC, randomizing actions through relaxed controls necessarily alters the control distribution ℙαt _ _t in (1.1), so that the action randomization can no longer be separated from the interaction induced by the control distribution. Thus, one cannot derive an auxiliary dynamics whose generator depends linearly on the stochastic policy. In fact, stochastic policies face algorithmic challenges even in the single-agent setting. They require sampling random actions at very high frequency, leading to irregular control trajectories that may be impractical in real-world applications [37, 17, 34]. Moreover, the Bellman equation under stochastic policies involves integration over continuous action spaces. Enforcing this condition requires Monte Carlo approximation, which is computationally expensive and often leads to unstable and slow convergence. In contrast, the benefit of deterministic policies for MFC is clear: the state–action distribution is induced directly as a push-forward of the state law. It avoids optimization over stochastic kernels, helps bypass key limitations of exploratory-policy approaches in extended mean field settings, and facilitates policy updates by differentiating through both the selected actions and the induced state–action distributions. As demonstrated in [7] for single-agent RL problems and the numerical experiments in this paper, deterministic policies offer significant computational advantages in continuous-time-space settings, leading to model-free algorithms with improved stability and faster convergence compared with stochastic-policy methods. Notation. We denote by x⋅y x· y the scalar product between two vectors x∈ℝm x∈R^m and y∈ℝm y∈R^m, and by M:N=tr(M⊤N) M N= *tr(M N) the inner product between two matrices M∈ℝm×n M∈R^m× n and N∈ℝm×n N∈R^m× n. Given T>0 T>0, a filtered probability space (Ω,ℱ,=(ℱt)t∈[0,T],ℙ) ( ,F,F=(F_t)_t∈[0,T],P), a normed space (E,|⋅|) (E,|·|), and an E E-valued random variable X X on (Ω,ℱ,ℙ) ( ,F,P), we denote by ℙX _X its probability law under ℙ , by X∼μ X μ the fact that X X follows the distribution μ μ, by ξ∼μ[f(ξ)]=∫Ef(x)μ(dx) E_ξ μ[f(ξ)]= _Ef(x)μ(dx) for a measurable function f:E→ℝ f:E→R, and by L2(ℱt;E) L^2(F_t;E) the set of square integrable E E-valued random variables on (Ω,ℱt,ℙ) ( ,F_t,P) for all t∈[0,T] t∈[0,T]. We denote by 2(E) _2(E) the set of probability measures μ μ on E E that are square integrable, i.e., M2(μ)≔(∫E|x|2μ(dx))1/2<∞. M_2(μ) ( _E|x|^2μ(dx) )^1/2<∞. We shall assume without loss of generality that ℱ0 _0 is rich enough to carry E E-valued random variables with any square integrable distribution, i.e., 2(E)=ℙξ∣ξ∈L2(ℱ0;E). _2(E)=\P_ξ ξ∈ L^2(F_0;E)\. We equip 2(E) _2(E) with the 2-Wasserstein distance 2 _2 defined by 2(μ,μ′) _2(μ,μ ) =inf(∫E×E|x−y|2π(dx,dy))1/2|π∈2(E×E) with marginals μ and μ′ = \ ( _E× E|x-y|^2\,π(dx,dy) )^1/2\, |\,π _2(E× E) with marginals μ and μ \ =inf([|ξ−ξ′|2])1/2|ξ,ξ′∈L2(ℱ0;E),ℙξ=μ,ℙξ′=μ′. = \ (E [|ξ-ξ |^2 ] )^1/2\, |\,ξ,ξ ∈ L^2(F_0;E),\,P_ξ=μ,\,P_ξ =μ \. We denote by C1,2([0,T]×2(ℝn)) C^1,2([0,T]×P_2(R^n)) the space of continuous functions w:[0,T]×2(ℝn)→ℝ w:[0,T]×P_2(R^n) such that for all μ∈2(ℝn) μ _2(R^n), w(⋅,μ) w(·,μ) is C1([0,T]) C^1([0,T]), with ∂tw _tw being continuous in all variables; for all t∈[0,T] t∈[0,T], w(t,⋅) w(t,·) is L-differentiable, and (t,μ,v)↦∂μw(t,μ)(v) (t,μ,v) _μw(t,μ)(v) is continuous at all (t,μ,v) (t,μ,v) with v∈supp(μ) v (μ); for all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), ∂μw(t,μ)(⋅) _μw(t,μ)(·) is C1 C^1, with (t,μ,v)↦∂v∂μw(t,μ)(v) (t,μ,v) _v _μw(t,μ)(v) being continuous at all (t,μ,v) (t,μ,v) with v∈supp(μ) v (μ); and for all compact sets ⊂2(ℝn) _2(R^n), sup(t,μ)∈[0,T]×[∫ℝn|∂μw(t,μ)(ξ)|2μ(dξ)+‖∂v∂μw(t,μ)‖Lμ∞]<∞. _(t,μ)∈[0,T]×K [ _R^n| _μw(t,μ)(ξ)|^2μ(dξ)+\| _v _μw(t,μ)\|_L^∞_μ ]<∞. 2 Model-free sensitivity formula of McKean-Vlasov dynamics This section considers a multidimensional McKean–Vlasov dynamics parameterized by θ θ and derives model-free representations of the gradient of the value functional with respect to θ θ. Consistent with the existing literature on MFCs, we consider the lifted value function defined on the Wasserstein space of probability measures. The derivation of these representations exploits both the invariance property of the value function and the robustness of the state distribution with respect to the model parameters. These representations will be useful for developing learning algorithms for the controlled McKean-Vlasov dynamics in Section 3. Let T>0 T>0 be a given terminal time, and (Ω,ℱ,ℙ) ( ,F,P) be a probability space on which an m m-dimensional Brownian motion W=(Wt)0≤t≤T W=(W_t)_0≤ t≤ T is defined. We denote by =(ℱt)0≤t≤T =(F_t)_0≤ t≤ T the filtration generated by W W, augmented with ℙ -null sets, and assume that the initial σ σ-algebra ℱ0 _0 is rich enough. Let β≥0 β≥ 0 be a discount factor, and consider continuous functions b:[0,T]×ℝn×2(ℝn)×ℝk→ℝn b:[0,T]×R^n×P_2(R^n)×R^k→R^n, σ:[0,T]×ℝn×2(ℝn)×ℝk→ℝn×m σ:[0,T]×R^n×P_2(R^n)×R^k→R^n× m, r:[0,T]×ℝn×2(ℝn)×ℝk→ℝ r:[0,T]×R^n×P_2(R^n)×R^k→R and g:ℝn×2(ℝn)→ℝ g:R^n×P_2(R^n)→R satisfying the following regularity and growth conditions: Assumption 2.1. There exists a locally bounded function ω:[0,∞)→[0,∞) ω:[0,∞)→[0,∞) such that for all (t,θ)∈[0,T]×ℝk (t,θ)∈[0,T]×R^k, (x,x′)∈ℝn (x,x )∈R^n and μ,μ′∈2(ℝn) μ,μ _2(R^n), |b(t,x,μ,θ)−b(t,x′,μ′,θ)|+|σ(t,x,μ,θ)−σ(t,x′,μ′,θ)| |b(t,x,μ,θ)-b(t,x ,μ ,θ)|+|σ(t,x,μ,θ)-σ(t,x ,μ ,θ)| ≤ω(|θ|)(|x−x′|+2(μ,μ′)), ≤ω(|θ|)(|x-x |+W_2(μ,μ )), |b(t,0,δ0,θ)|+|σ(t,0,δ0,θ)|+|r(t,x,μ,θ)|+|g(x,μ)|1+|x|2+M2(μ)2 |b(t,0, _0,θ)|+|σ(t,0, _0,θ)|+ |r(t,x,μ,θ)|+|g(x,μ)|1+|x|^2+M_2(μ)^2 ≤ω(|θ|). ≤ω(|θ|). For each t∈[0,T] t∈[0,T], θ∈ℝk θ∈R^k, and ξ∈L2(ℱt;ℝn) ξ∈ L^2(F_t;R^n), consider the following state dynamics: dXst,ξ,θ dX^t,ξ,θ_s =b(s,Xst,ξ,θ,ℙXst,ξ,θ,θ)ds+σ(s,Xst,ξ,θ,ℙXst,ξ,θ,θ)dWs,s∈[t,T];Xtt,ξ,θ=ξ. =b(s,X^t,ξ,θ_s,P_X^t,ξ,θ_s,θ)ds+σ(s,X^t,ξ,θ_s,P_X^t,ξ,θ_s,θ)dW_s, s∈[t,T]; X^t,ξ,θ_t=ξ. (2.1) Under Assumption 2.1, (2.1) has a unique square-integrable strong solution (Xst,ξ,θ)s∈[t,T] (X^t,ξ,θ_s)_s∈[t,T] (see e.g., [28] and the references therein). Moreover, due to the weak uniqueness of (2.1), the law of (Xst,ξ,θ)s∈[t,T] (X^t,ξ,θ_s)_s∈[t,T] depends on ξ ξ only through its law μ=ℙξ μ=P_ξ, and hence we can define ℙst,μ,θ≔ℙXst,ξ,θ,s∈[t,T].P^t,μ,θ_s _X^t,ξ,θ_s, s∈[t,T]. Define the following value function for a given β≥0 β≥ 0: V(t,μ,θ) V(t,μ,θ) ≔[∫tTe−β(s−t)r(s,Xst,ξ,θ,ℙXst,ξ,θ,θ)ds+e−β(T−t)g(XTt,ξ,θ,ℙXTt,ξ,θ)] E [ _t^Te^-β(s-t)r (s,X^t,ξ,θ_s,P_X^t,ξ,θ_s,θ )ds+e^-β(T-t)g (X^t,ξ,θ_T,P_X^t,ξ,θ_T ) ] (2.2) =∫tTe−β(s−t)r¯(s,ℙst,μ,θ,θ)ds+e−β(T−t)g¯(ℙTt,μ,θ), = _t^Te^-β(s-t) r (s,P^t,μ,θ_s,θ )ds+e^-β(T-t) g(P^t,μ,θ_T), (2.3) where the function r¯:[0,T]×2(ℝn)×ℝk→ℝ r:[0,T]×P_2(R^n)×R^k→R and g¯:2(ℝn)→ℝ g:P_2(R^n)→R are defined by r¯(t,μ,θ)≔∫ℝnr(t,x,μ,θ)μ(dx),g¯(μ)≔∫ℝng(x,μ)μ(dx). r(t,μ,θ) _R^nr(t,x,μ,θ)μ(dx), g(μ) _R^ng(x,μ)μ(dx). (2.4) To characterize the gradient of V V in θ θ, for all w∈C1,2([0,T]×2(ℝn)) w∈ C^1,2([0,T]×P_2(R^n)), define the function A[w]:[0,T]×2(ℝn)×ℝk→ℝ A[w]:[0,T]×P_2(R^n)×R^k→R such that for all (t,μ,θ)∈[0,T]×2(ℝn)×ℝk (t,μ,θ)∈[0,T]×P_2(R^n)×R^k, A[w](t,μ,θ)≔ℒθ[w](t,μ)+r¯(t,μ,θ),A[w](t,μ,θ) ^θ[w](t,μ)+ r(t,μ,θ), (2.5) where ℒθ ^θ is the generator of (2.1) given by ℒθ[w](t,μ)≔∂tw(t,μ)−βw(t,μ)+ξ∼μ[b(t,ξ,μ,θ)⋅∂μw(t,μ)(ξ)+12(σ⊤)(t,ξ,μ,θ):∂v∂μw(t,μ)(ξ)]. splitL^θ[w](t,μ)& _tw(t,μ)-β w(t,μ)\\ & +E_ξ μ [b(t,ξ,μ,θ)· _μw(t,μ)(ξ)+ 12(σ )(t,ξ,μ,θ) _v _μw(t,μ)(ξ) ]. split (2.6) We impose the following regularity conditions on the value function Vθ V^θ and the function A[w] A[w]. Assumption 2.2. (1) For all θ∈ℝk θ∈R^k, the function V(⋅,⋅,θ) V(·,·,θ) defined by (2.2) is in C1,2([0,T]×2(ℝn)) C^1,2([0,T]×P_2(R^n)). (2) For all w∈C1,2([0,T]×2(ℝn)) w∈ C^1,2([0,T]×P_2(R^n)), the function A[w] A[w] defined by (2.5) is differentiable with respect to θ θ, and the derivative ∂θA[w]:[0,T]×2(ℝn)×ℝk→ℝk _θA[w]:[0,T]×P_2(R^n)×R^k→R^k is continuous. Under Assumptions 2.1 and 2.2, the following theorem shows that the gradient ∂θV _θV can be expressed as the integral of ∂θA _θA in (2.5) along the flow of state laws. Theorem 2.1. Suppose Assumptions 2.1 and 2.2 hold. For all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n) and θ∈ℝk θ∈R^k, ∂θV(t,μ,θ)=∫tTe−β(s−t)∂θA[V(⋅,⋅,θ)](s,ℙst,μ,θ,θ)ds. split& _θV(t,μ,θ)= _t^Te^-β(s-t) _θA[V(·,·,θ)](s,P^t,μ,θ_s,θ)\,ds. split (2.7) Theorem 2.1 expresses the gradient ∂θV _θV directly in terms of ∂θA[V] _θA[V], without reference to the model coefficients. The proof leverages a precise characterization of the difference between the value functions corresponding to two different parameters (Proposition 6.2). Such a result extends the performance-difference lemmas in [7, Lemma 6.1] and [35, Lemma 3.2] from classical control problems to the more general setting of McKean–Vlasov dynamics. Compared with existing approaches in [18, 11], our method avoids differentiating the PDE for the value function V(⋅,⋅,θ) V(·,·,θ) with respect to θ θ, and therefore requires weaker regularity assumptions on the coefficients. In particular, we do not require differentiability in θ θ of the first- and second-order time and spatial derivatives of V V, which is imposed in [11, Appendix A2]. To characterize the advantage rate function, we first show that the value function V V and the advantage rate function A[V] A[V] must satisfy a Feynman-Kac formula and an invariance property along the deterministic state flow. Theorem 2.2. Suppose Assumptions 2.1 and 2.2 hold. Let θ∈ℝk θ∈R^k, Vθ≔V(⋅,⋅,θ)∈C1,2([0,T]×2(ℝn)) V^θ V(·,·,θ)∈ C^1,2([0,T]×P_2(R^n)) be defined by (2.2), and qθ=A[Vθ]∈C([0,T]×2(ℝn)×ℝk) q^θ=A[V^θ]∈ C([0,T]×P_2(R^n)×R^k) be defined by (2.5). For all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), Vθ(T,μ)=g¯(μ),qθ(t,μ,θ)=0,V^θ(T,μ)= g(μ), q^θ(t,μ,θ)=0, (2.8) and for all θ′∈ℝk θ ∈R^k and s∈[t,T] s∈[t,T], e−βsVθ(s,ℙst,μ,θ′)−e−βtVθ(t,μ)+∫tse−βu[r¯(u,ℙut,μ,θ′,θ′)−qθ(u,ℙut,μ,θ′,θ′)]du=0,e^-β sV^θ(s,P^t,μ,θ _s)-e^-β tV^θ(t,μ)+ _t^se^-β u[ r(u,P^t,μ,θ _u,θ )-q^θ(u,P^t,μ,θ _u,θ )]\,du=0, (2.9) where ℙrt,μ,θ′=ℙXrt,ξ,θ′ ^t,μ,θ _r=P_X^t,ξ,θ _r, and Xt,ξ,θ′ X^t,ξ,θ satisfies for all s∈[t,T] s∈[t,T], dXst,ξ,θ′ dX^t,ξ,θ _s =b(s,Xst,ξ,θ′,ℙXst,ξ,θ′,θ′)ds+σ(s,Xst,ξ,θ′,ℙXst,ξ,θ′,θ′)dWs,Xtt,ξ,θ′=ξ, =b(s,X^t,ξ,θ _s,P_X^t,ξ,θ _s,θ )ds+σ(s,X^t,ξ,θ _s,P_X^t,ξ,θ _s,θ )dW_s, X^t,ξ,θ _t=ξ, with some ξ∈L2(ℱt;ℝn) ξ∈ L^2(F_t;R^n) having the law μ μ. Conversely, we show that the conditions (2.8) and (2.9) are sufficient to jointly characterize both the value function V V and the advantage rate function A[V] A[V]. This derivation is based on two key properties in general MFC: the first being the invariance property of the value function for any given policy along the deterministic state flow in a neighborhood of θ θ, and the second being the continuity of the law of McKean–Vlasov systems with respect to the perturbations of the initial condition and model parameters. Theorem 2.3. Suppose Assumptions 2.1 and 2.2 hold. Let θ∈ℝk θ∈R^k, V^∈C1,2([0,T]×2(ℝn)) V∈ C^1,2([0,T]×P_2(R^n)), and q^∈C([0,T]×2(ℝn)×ℝk) q∈ C([0,T]×P_2(R^n)×R^k) satisfy the following conditions: for all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), V^(T,μ)=g¯(μ),q^(t,μ,θ)=0, V(T,μ)= g(μ), q(t,μ,θ)=0, (2.10) and there exists a neighborhood t,μ(θ)⊂ℝk _t,μ(θ)⊂R^k of θ θ such that for all θ′∈t,μ(θ) θ _t,μ(θ) and s∈[t,T] s∈[t,T], e−βsV^(s,ℙst,μ,θ′)−e−βtV^(t,μ)+∫tse−βu[r¯(u,ℙut,μ,θ′,θ′)−q^(u,ℙut,μ,θ′,θ′)]du=0,e^-β s V(s,P^t,μ,θ _s)-e^-β t V(t,μ)+ _t^se^-β u[ r(u,P^t,μ,θ _u,θ )- q(u,P^t,μ,θ _u,θ )]\,du=0, (2.11) where (ℙrt,μ,θ′)r∈[t,T] (P^t,μ,θ _r)_r∈[t,T] is defined as in Theorem 2.2. Then for all (t,μ,θ′)∈[0,T]×2(ℝn)×t,μ(θ) (t,μ,θ )∈[0,T]×P_2(R^n)×O_t,μ(θ), V^(t,μ)=V(t,μ,θ),q^(t,μ,θ′)=A[V(⋅,⋅θ)](t,μ,θ′). V(t,μ)=V(t,μ,θ), q(t,μ,θ )=A[V(·,·θ)](t,μ,θ ). For any given θ θ, Theorem 2.3 characterizes the associated value function and advantage rate function in a neighborhood of θ θ, via the Feynman–Kac formula (2.10) and the invariance principle (2.11) along deterministic state flows. This provides a unified framework for characterizing value and advantage functions under arbitrary policies, and applies to both deterministic policies as in Section 3 and stochastic policies as in [27, 39]. 3 Deterministic Policy Gradient for Extended MFC Problems In this section, we apply the sensitivity formula in Section 2 to derive a policy gradient formula for extended mean field control (MFC) problems studied in [1, 29]. In extended MFC problems, both the cost functional and the state dynamics depend on the joint distribution of the controlled state and control processes. This formula provides the foundation for developing a model-free deterministic actor–critic algorithm in Section 4 for solving extended MFC problems. 3.1 Problem formulation Extended MFC problems. Let T>0 T>0 be a given terminal time and (Ω,ℱ,ℙ) ( ,F,P) be a complete probability space which supports an m m-dimensional Brownian motion W W and an independent square-integrable random variable ξ0 _0. We denote by =(ℱt)t≥0 =(F_t)_t≥ 0 the filtration generated by W W and ξ0 _0 augmented by the ℙ -null sets. Let A⊂ℝd A ^d be a measurable set representing the agent’s action space, and let be a space of all -adapted square-integrable processes α:[0,T]×Ω→A α:[0,T]× → A representing the agent’s admissible control space. For each α∈ α , consider the associated state process Xα X^α governed by the following controlled McKean-Vlasov dynamics: dXsα=b(s,Xsα,αs,ℙ(Xsα,αs))ds+σ(s,Xsα,αs,ℙ(Xsα,αs))dWs,X0α=ξ0, splitdX^α_s=b(s,X^α_s,α_s,P_(X^α_s,α_s))ds+σ(s,X^α_s,α_s,P_(X^α_s,α_s))dW_s, X^α_0= _0, split (3.1) where b:[0,T]×ℝn×ℝd×2(ℝn×ℝd)→ℝn b:[0,T]×R^n×R^d×P_2(R^n×R^d)→R^n and σ:[0,T]×ℝn×ℝd×2(ℝn×ℝd)→ℝn×m σ:[0,T]×R^n×R^d×P_2(R^n×R^d)→R^n× m are sufficiently regular functions such that (3.1) has a unique square-integrable strong solution Xα X^α. The agent aims to maximize the following reward functional J(α)≔[∫0Te−βsr(s,Xsα,αs,ℙ(Xsα,αs))ds+e−βTg(XTα,ℙXTα)] J(α) E [ _0^Te^-β sr(s,X^α_s,α_s,P_(X^α_s,α_s))ds+e^-β Tg(X^α_T,P_X^α_T) ] (3.2) over all α∈ α , where β≥0 β≥ 0, r:[0,T]×ℝn×ℝd×2(ℝn×ℝd)→ℝ r:[0,T]×R^n×R^d×P_2(R^n×R^d)→R and g:ℝn×2(ℝn)→ℝ g:R^n×P_2(R^n)→R are continuous functions with at most quadratic growth. It is known that under suitable regularity conditions (see e.g., [28]), it suffices to optimize (3.2) over control processes αφ=(αtφ)t∈[0,T] α =(α _t)_t∈[0,T] given in closed loop (or feedback) form: αtφ=φ(t,Xtφ,ℙXtφ),t∈[0,T],α _t= (t,X _t,P_X _t), t∈[0,T], (3.3) where φ:[0,T]×ℝn×2(ℝn)→A :[0,T]×R^n×P_2(R^n)→ A is a measurable function, called a Markov policy, and Xφ X is a solution to the following controlled McKean-Vlasov dynamics: dXsφ=b(s,Xsφ,φ(s,Xsφ,ℙXsφ),ℙ(Xsφ,φ(s,Xsφ,ℙXsφ)))ds+σ(s,Xsφ,φ(s,Xsφ,ℙXsφ),ℙ(Xsφ,φ(s,Xsφ,ℙXsφ)))dWs,X0φ=ξ0, splitdX _s=&b(s,X _s, (s,X _s,P_X _s),P_(X _s, (s,X _s,P_X _s)))ds\\ &+σ(s,X _s, (s,X _s,P_X _s),P_(X _s, (s,X _s,P_X _s)))dW_s, X _0= _0, split (3.4) Consequently, the goal of the agent is to maximize the following objective J(φ)≔[∫0Te−βsr(s,Xsφ,φ(s,Xsφ,ℙXsφ),ℙ(Xsφ,φ(s,Xsφ,ℙXsφ)))ds+e−βTg(XTφ,ℙXTφ)] J( ) E [ _0^Te^-β sr(s,X _s, (s,X _s,P_X _s),P_(X _s, (s,X _s,P_X _s)))ds+e^-β Tg(X _T,P_X _T) ] (3.5) over all admissible Markov policies φ . RL with deterministic policies. In the RL framework, the agent does not know the coefficients b,σ,r b,σ,r and g g. Instead, the agent interacts directly with the system (3.4) with different actions and updates her policies based on the observed state and reward trajectories. As in the classical McKean-Vlasov control framework, we restrict attention to suitable Markov policies φ:[0,T]×ℝn×2(ℝn)→A :[0,T]×R^n×P_2(R^n)→ A for optimizing (3.5). In the RL literature, such policies are referred to as deterministic policies, since they map each time-state-measure triple directly to an action. A deterministic feedback policy φ(t,x,μ) (t,x,μ) selects an action directly based on the time t t, state x x and law μ μ, and the associated state–action distribution is the push-forward measure Γtφ=(id,φ(t,⋅,μ))#μ _t = (id, (t,·,μ) )_\#μ defined in (3.7). The joint distribution is therefore completely determined by the state law and the feedback map. Although the resulting dynamics may still depend nonlinearly on Γtφ _t , policy optimization can be reduced to an optimization over a parameterized family φθ _θ. More precisely, let Lip([0,T]×ℝn×2(ℝn);A) ([0,T]×R^n×P_2(R^n);A) be the set of continuous functions φ:[0,T]×ℝn×2(ℝn)→A :[0,T]×R^n×P_2(R^n)→ A that are Lipschitz continuous in (x,μ) (x,μ), uniformly on t∈[0,T] t∈[0,T]. Given a class of parameterized policies Θ≔φθ∈Lip([0,T]×ℝn×2(ℝn);A)∣θ∈ℝk P_ \ _θ ([0,T]×R^n×P_2(R^n);A) θ∈R^k\, we maximize the following functional over θ∈ℝk θ∈R^k: J(θ)≔[∫0Te−βsr(s,Xsφθ,φθ(s,Xsφθ,ℙXsφθ),ℙ(Xsφθ,φθ(s,Xsφθ,ℙXsφθ)))ds+e−βTg(XTφθ,ℙXTφθ)], J(θ) E [ _0^Te^-β sr(s,X _θ_s, _θ(s,X _θ_s,P_X _θ_s),P_(X _θ_s, _θ(s,X _θ_s,P_X _θ_s)))ds+e^-β Tg(X _θ_T,P_X _θ_T) ], (3.6) where we write J(θ)=J(φθ) J(θ)=J( _θ) with a slight abuse of notation. In the sequel, we propose model-free deterministic policy gradient (DPG) algorithms that optimize (3.6) via a gradient ascent algorithm. 3.2 Deterministic policy gradient formula Reformulation as parametric McKean-Vlasov dynamics. To compute ∇θJ(φθ) _θJ( _θ) using Theorems 2.1 and 2.3, we rewrite (3.6) as a special case of the parametric McKean-Vlasov dynamics analyzed in Section 2. As in [28], for each policy φ∈Θ ∈ P_ and (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), let (id,φ(t,⋅,μ))♯μ∈2(ℝn×ℝd) ( * id, (t,·,μ))_ μ _2(R^n×R^d) be the induced state-action measure given by [(id,φ(t,⋅,μ))♯μ](B) [( * id, (t,·,μ))_ μ](B) ≔μ(x∈ℝn∣(x,φ(t,x,μ))∈B),∀B∈ℬ(ℝn×ℝd). μ(\x∈R^n (x, (t,x,μ))∈ B\), ∀ B (R^n×R^d). (3.7) Note that for any control αφ α of the form (3.3), the joint state-control distribution satisfies ℙ(Xtφ,αtφ)=(id,φ(t,⋅,ℙXtφ))♯ℙXtφ,t∈[0,T].P_(X _t,α _t)=( * id, (t,·,P_X _t))_ P_X _t, t∈[0,T]. Define the following controlled coefficients associated with Θ P_ : for all ℓ∈b,σ,r ∈\b,σ,r\, ℓφ(t,x,μ,θ) (t,x,μ,θ) ≔ℓ(t,x,φθ(t,x,μ),(id,φθ(t,⋅,μ))♯μ). (t,x, _θ(t,x,μ),( * id, _θ(t,·,μ))_ μ). (3.8) The dynamics (3.4) with φ=φθ = _θ can then be rewritten as dXs=bφ(s,Xs,ℙXs,θ)ds+σφ(s,Xs,ℙXs,θ)dWs. splitdX_s=&b (s,X_s,P_X_s,θ)ds+σ (s,X_s,P_X_s,θ)dW_s. split (3.9) Consider the dynamic version of (3.6) defined as follows (cf. (2.2)) : for all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), let ξ∈L2(ℱt;ℝn) ξ∈ L^2(F_t;R^n) with ℙξ=μ _ξ=μ and Xt,ξ,θ X^t,ξ,θ satisfy (3.9) on [t,T] [t,T] with Xtt,ξ,θ=ξ X^t,ξ,θ_t=ξ, and define Vφ(t,μ,θ) V (t,μ,θ) ≔∫tTe−β(s−t)r¯φ(s,ℙst,μ,θ,θ)ds+e−β(T−t)g¯(ℙTt,μ,θ), _t^Te^-β(s-t) r (s,P^t,μ,θ_s,θ )ds+e^-β(T-t) g(P^t,μ,θ_T), (3.10) where ℙst,μ,θ=ℙXst,ξ,θ ^t,μ,θ_s=P_X^t,ξ,θ_s, and r¯φ(t,μ,θ)≔∫ℝnrφ(t,x,μ,θ)μ(dx),g¯(μ)≔∫ℝng(x,μ)μ(dx). r (t,μ,θ) _R^nr (t,x,μ,θ)μ(dx), g(μ) _R^ng(x,μ)μ(dx). (3.11) For each w∈C1,2([0,T]×2(ℝn)) w∈ C^1,2([0,T]×P_2(R^n)), define Aφ[w](t,μ,θ)=ℒφ[w](t,μ,θ)+r¯φ(t,μ,θ),A [w](t,μ,θ)=L [w](t,μ,θ)+ r (t,μ,θ), (3.12) where ℒφ is the generator of (3.9) and is defined as in (2.6) with (b,σ)=(bφ,σφ) (b,σ)=(b ,σ ). The function Aφ[w] A [w] can be viewed as a mean-field analogue of the advantage rate function in continuous-time RL problems driven by classical diffusion processes [7]. In the sequel, we assume that the model coefficients b,σ,r,g \b,σ,r,g\ and the policies in Θ P_ are sufficiently regular such that the induced coefficients given in (3.8) satisfy Assumption 2.1, and the associated value and advantage rate functions satisfy Assumption 2.2. Assumption 3.1. The functions bφ,σφ,rφ,g \b ,σ ,r ,g\ satisfy Assumption 2.1, and Vφ V in (3.10) and Aφ A in (3.12) satisfy Assumption 2.2. Characterizations of DPG. We now present the first characterization of the deterministic policy gradient (DPG), which combines Theorems 2.1 and 2.3. Theorem 3.1. Suppose Assumption 3.1 holds, and let φθ∈Θ _θ∈ P_ for a given θ∈ℝk θ∈R^k. Let V^∈C1,2([0,T]×2(ℝn)) V∈ C^1,2([0,T]×P_2(R^n)) and q^∈C([0,T]×2(ℝn)×ℝk) q∈ C([0,T]×P_2(R^n)×R^k) satisfy the following conditions: for all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), (i) V^(T,μ)=g¯(μ) V(T,μ)= g(μ) and q^(t,μ,θ)=0 q(t,μ,θ)=0, (i) there exists a neighborhood t,μ(θ)⊂ℝk _t,μ(θ)⊂R^k of θ θ such that for all θ′∈t,μ(θ) θ _t,μ(θ) and s∈[t,T] s∈[t,T], e−βsV^(s,ℙst,μ,θ′)−e−βtV^(t,μ)+∫tse−βu[r¯φ(u,ℙut,μ,θ′,θ′)−q^(u,ℙut,μ,θ′,θ′)]du=0,e^-β s V(s,P^t,μ,θ _s)-e^-β t V(t,μ)+ _t^se^-β u[ r (u,P^t,μ,θ _u,θ )- q(u,P^t,μ,θ _u,θ )]\,du=0, (3.13) where ℙrt,μ,θ′=ℙXrt,ξ,θ′ ^t,μ,θ _r=P_X^t,ξ,θ _r, and Xt,ξ,θ′ X^t,ξ,θ satisfies for all s∈[t,T] s∈[t,T], dXst,ξ,θ′ dX^t,ξ,θ _s =bφ(s,Xst,ξ,θ′,ℙXst,ξ,θ′,θ′)ds+σφ(s,Xst,ξ,θ′,ℙXst,ξ,θ′,θ′)dWs,Xtt,ξ,θ′=ξ, =b (s,X^t,ξ,θ _s,P_X^t,ξ,θ _s,θ )ds+σ (s,X^t,ξ,θ _s,P_X^t,ξ,θ _s,θ )dW_s, X^t,ξ,θ _t=ξ, (3.14) with some ξ∈L2(ℱt;ℝn) ξ∈ L^2(F_t;R^n) having the law μ μ. Then for all (t,μ,θ′)∈[0,T]×2(ℝn)×t,μ(θ) (t,μ,θ )∈[0,T]×P_2(R^n)×O_t,μ(θ), V^(t,μ)=Vφ(t,μ,θ),q^(t,μ,θ′)=Aφ[Vφ(⋅,θ)](t,μ,θ′). V(t,μ)=V (t,μ,θ), q(t,μ,θ )=A [V (·,θ)](t,μ,θ ). Moreover, let ℙsθ=ℙXsφθ ^θ_s=P_X _θ_s with Xφθ X _θ defined by (3.4), it holds that ∇θJ(φθ)=∫0Te−βs(∂θq^)(s,ℙsθ,θ)ds. split& _θJ( _θ)= _0^Te^-β s( _θ q)(s,P^θ_s,θ)\,ds. split (3.15) Theorem 3.1 expresses the DPG in terms of the gradient of the advantage rate function with respect to the policy parameter θ θ. Both the value function and the advantage rate function are lifted to functions on the space of probability measures, and Theorem 3.1 simultaneously characterizes them through an invariance principle along the flow of the state process (3.4) controlled by φθ _θ. This theorem is the continuous analogue of [25, 13] for discrete-time extended MFC problems. In particular, [13] shows that the classical Q-function defined on the original action space cannot be characterized via the dynamic programming principle (DPP). To ensure the DPP, one instead needs to construct a lifted Q-function called IQ-function on the space of policies defined by integrating the classical Q-function over the state–action distribution induced by a given policy. The next theorem refines Theorem 3.1 by decomposing the value function and the advantage rate function into integrals of local functions of the state, control, and their associated measures, and by characterizing these local functions through suitable martingale conditions. We impose the following regularity assumptions on these local functions: Assumption 3.2. Let V^D∈C([0,T]×ℝn×2(ℝn)) V_D∈ C([0,T]×R^n×P_2(R^n)) be such that V^:[0,T]×2(ℝn)→ℝ V:[0,T]×P_2(R^n)→R given by V^(t,μ)≔ξ∼μ[V^D(t,ξ,μ)] V(t,μ) E_ξ μ[ V_D(t,ξ,μ)] (3.16) is well-defined and in the space C1,2([0,T]×2(ℝn)) C^1,2([0,T]×P_2(R^n)). Let q^∈C([0,T]×ℝn×ℝd×2(ℝn×ℝd)) q∈ C([0,T]×R^n×R^d×P_2(R^n×R^d)) and the policy class Θ P_ be such that q^:[0,T]×2(ℝn)×ℝk→ℝ q:[0,T]×P_2(R^n)×R^k→R given by q^(t,μ,θ)≔ξ∼μ[q^D(t,ξ,φθ(t,ξ,μ),(id,φθ(t,⋅,μ))♯μ)] q(t,μ,θ) E_ξ μ[ q_D(t,ξ, _θ(t,ξ,μ),( * id, _θ(t,·,μ))_ μ)] (3.17) is well-defined and continuous. Moreover, ∂θq _θ q exists, is continuous, and satisfies the chain rule: (∂θq^)(t,μ,θ)=[∂θφθ(t,ξ,μ)⊤(∂aq^D)(t,ξ,φθ(t,ξ,μ),(id,φθ(t,⋅,μ))♯μ)+~[(∂νq^D)(t,ξ~,φθ(t,ξ~,μ),(id,φθ(t,⋅,μ))♯μ)](ξ,φθ(t,ξ,μ))], split( _θ q)(t,μ,θ)&=E [ _θ _θ(t,ξ,μ) \( _a q_D)(t,ξ, _θ(t,ξ,μ),( * id, _θ(t,·,μ))_ μ)\\ & + E [( _ν q_D) (t, ξ, _θ(t, ξ,μ),( * id, _θ(t,·,μ))_ μ ) ](ξ, _θ(t,ξ,μ)) \ ], split (3.18) where ξ ξ and ξ~ ξ are independent random variables with law μ μ, ~ E denotes the expectation with respect to ξ~ ξ, and ∂νq^D(⋅,Γ) _ν q_D(·, ) is the partial L-derivative of q^D q_D with respect to the second marginal of Γ∈2(ℝn×ℝd) _2(R^n×R^d). Theorem 3.2. Suppose Assumption 3.1 holds, and V^D∈C([0,T]×ℝn×2(ℝn)) V_D∈ C([0,T]×R^n×P_2(R^n)), q^D∈C([0,T]×ℝn×ℝd×2(ℝn×ℝd)) q_D∈ C([0,T]×R^n×R^d×P_2(R^n×R^d)) and the policy class Θ P_ satisfy Assumption 3.2. Let θ∈ℝk θ∈R^k, and assume that V^D V_D and q^D q_D satisfy the following conditions: (i) For all (t,x,μ)∈[0,T]×ℝn×2(ℝn) (t,x,μ)∈[0,T]×R^n×P_2(R^n), V^D(T,x,μ)=g(x,μ),ξ∼μ[q^D(t,ξ,φθ(t,ξ,μ),(id,φθ(t,⋅,μ))♯μ)]=0. V_D(T,x,μ)=g(x,μ), E_ξ μ[ q_D(t,ξ, _θ(t,ξ,μ),( * id, _θ(t,·,μ))_ μ)]=0. (3.19) (i) For all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), there exists a neighborhood t,μ(θ)⊂ℝk _t,μ(θ)⊂R^k of θ θ such that for all θ′∈t,μ(θ) θ _t,μ(θ), the following process (e−βsV^D(s,Xst,ξ,θ′,ℙst,μ,θ′)+∫tse−βu(r−q^D)(u,Xut,ξ,θ′,φθ′(u,Xut,ξ,θ′,ℙut,μ,θ′),(id,φθ′(u,⋅,ℙut,μ,θ′))♯ℙut,μ,θ′)du)s∈[t,T] split& (e^-β s V_D(s,X_s^t,ξ,θ ,P_s^t,μ,θ )\\ & + _t^se^-β u(r- q_D)(u,X_u^t,ξ,θ , _θ (u,X_u^t,ξ,θ ,P_u^t,μ,θ ),( * id, _θ (u,·,P_u^t,μ,θ ))_ P_u^t,μ,θ )du )_s∈[t,T] split (3.20) is an F-martingale, where Xt,ξ,θ′ X^t,ξ,θ and ℙt,μ,θ′ ^t,μ,θ are defined as in Theorem 3.1. Then the functions V V and q q defined by (3.16) and (3.17), respectively, satisfy for all (t,μ,θ′)∈[0,T]×2(ℝn)×t,μ(θ) (t,μ,θ )∈[0,T]×P_2(R^n)×O_t,μ(θ), V^(t,μ)=Vφ(t,μ,θ),q^(t,μ,θ′)=Aφ[Vφ(⋅,θ)](t,μ,θ′). V(t,μ)=V (t,μ,θ), q(t,μ,θ )=A [V (·,θ)](t,μ,θ ). (3.21) Moreover, ∇θJ(φθ)=[∫0Te−βs∂θφθ(s,Xsθ,ℙsθ)⊤(∂aq^D)(s,Xsθ,φθ(s,Xsθ,ℙsθ),(id,φθ(s,⋅,ℙsθ))♯ℙsθ)+~[(∂νq^D)(s,X~sθ,φθ(s,X~sθ,ℙsθ),(id,φθ(s,⋅,ℙsθ))♯ℙsθ)](Xsθ,φθ(s,Xsθ,ℙsθ))ds], split _θJ( _θ)&=E [ _0^Te^-β s _θ _θ(s,X_s^θ,P_s^θ) \( _a q_D)(s,X_s^θ, _θ(s,X_s^θ,P_s^θ),( * id, _θ(s,·,P_s^θ))_ P_s^θ)\\ & + E [( _ν q_D) (s, X_s^θ, _θ(s, X_s^θ,P_s^θ),( * id, _θ(s,·,P_s^θ))_ P_s^θ ) ](X_s^θ, _θ(s,X_s^θ,P_s^θ)) \ds ], split (3.22) where (Xsθ,ℙsθ)=(X~s0,ℙξ0,θ,ℙs0,ℙξ0,θ) (X_s^θ,P_s^θ)=( X_s^0,P_ _0,θ,P_s^0,P_ _0,θ), X~sθ X_s^θ is an independent copy of Xsθ X_s^θ, and ∂νq^D(⋅,Γ) _ν q_D(·, ) is the partial L-derivative of q^D q_D with respect to the second marginal of Γ∈2(ℝn×ℝd) _2(R^n×R^d). Analytically, Theorem 3.2 is the continuous-time-space analogue of the discrete-time MFC with learning [13] where the lifted Q-function (called IQ-function) defined on the space of policies is represented as a local Q-function and a local policy h:→() h:S (A). These local value functions are analogous to the decomposed value functions studied in the MFC literature (see, e.g., [4]). This decomposition is also consistent with the idea of “centralized training with decentralized execution” [10, 21], which has been adopted to improve the computational efficiency of RL algorithms for MFC [14, 2]. In particular, it enables both the value function and the reward to be decomposed additively across individual observations. Specifically, Theorem 3.2 expresses the functions V V and q q as integrals of a local value function V^D V_D and a local advantage rate function q^D q_D, respectively, as given in (3.16) and (3.17). Such local representations allow learning V^D V_D and q^D q_D from observed state and control trajectories, yielding a more informative learning signal than directly learning V V and q q from the flow of state laws as in (3.13); see Section 4. Note, however, only the integrated value and advantage functions V V and q q, rather than their local counterparts, can be uniquely determined, as shown in (3.21). Importantly, these integrated quantities are exactly those needed for policy updates. 3.3 Existence of V^D V_D and q^D q_D in Theorem 3.2 The function V^D V_D can be taken as the (decomposed) value function for a control problem with decoupled state dynamics studied as in [4, 11], while the function q^D q_D can be taken as the integrated Hamiltonian over the joint state–action distribution. To see it, assume without loss of generality that β=0 β=0. For all (t,x,μ)∈[0,T]×ℝn×2(ℝn) (t,x,μ)∈[0,T]×R^n×P_2(R^n), let ξ∈L2(ℱt;ℝn) ξ∈ L^2(F_t;R^n) with law μ μ, and define the following value function associated with the policy φθ _θ: VDθ(t,x,μ)≔[∫tTrφ(s,Xst,x,μ,θ,ℙXst,ξ,θ,θ)ds+g(XTt,x,μ,θ,ℙXTt,ξ,θ)], V_D^θ(t,x,μ) E [ _t^Tr (s,X^t,x,μ,θ_s,P_X^t,ξ,θ_s,θ)ds+g(X^t,x,μ,θ_T,P_X^t,ξ,θ_T) ], (3.23) where Xt,ξ,θ X^t,ξ,θ and Xt,x,μ,θ X^t,x,μ,θ satisfy the following decoupled state dynamics: for all s∈[t,T] s∈[t,T], dXst,ξ,θ=bφ(s,Xst,ξ,θ,ℙXst,ξ,θ,θ)ds+σφ(s,Xst,ξ,θ,ℙXst,ξ,θ,θ)dWs,Xtt,ξ,θ=ξ,dXst,x,μ,θ=bφ(s,Xst,x,μ,θ,ℙXst,ξ,θ,θ)ds+σφ(s,Xst,x,μ,θ,ℙXst,ξ,θ,θ)dWs,Xtt,x,μ,θ=x. splitdX^t,ξ,θ_s&=b (s,X^t,ξ,θ_s,P_X^t,ξ,θ_s,θ)ds+σ (s,X^t,ξ,θ_s,P_X^t,ξ,θ_s,θ)dW_s, X^t,ξ,θ_t=ξ,\\ dX^t,x,μ,θ_s&=b (s,X^t,x,μ,θ_s,P_X^t,ξ,θ_s,θ)ds+σ (s,X^t,x,μ,θ_s,P_X^t,ξ,θ_s,θ)dW_s, X^t,x,μ,θ_t=x. split (3.24) Note that by the weak uniqueness of (3.24), the measures (ℙXst,ξ,θ)s∈[t,T] (P_X^t,ξ,θ_s)_s∈[t,T] depend on ξ ξ only through its law μ μ, which allows us to regard Xt,x,μ,θ X^t,x,μ,θ as a function of μ μ without specifying the choice of ξ ξ. Suppose that V^D V_D is sufficiently regular, by Itô’s formula [5, Proposition 5.102], [t,T]∋s↦Ms≔VDθ(s,Xst,ξ,θ′,ℙst,μ,θ′) [t,T] s M_s V^θ_D(s,X_s^t,ξ,θ ,P_s^t,μ,θ ) −∫tsℒDφ[VDθ](r,Xrt,ξ,θ′,ℙrt,μ,θ′,θ′)dr - _t^sL _D[V^θ_D](r,X_r^t,ξ,θ ,P_r^t,μ,θ ,θ )dr (3.25) is an F-martingale, where ℒDφ _D is the generator of (3.14) satisfying ℒDφ[w](t,x,μ,θ)≔∂tw(t,x,μ)+bφ(t,x,μ,θ)⋅∂xw(t,x,μ)+12σφ[σφ]⊤(t,x,μ,θ):∂x2w(t,x,μ)+ξ∼μ[bφ(t,ξ,μ,θ)⋅∂μw(t,x,μ)(ξ)+12[σφ(σφ)⊤](t,ξ,μ,θ):∂v∂μw(t,x,μ)(ξ)]. split&L _D[w](t,x,μ,θ)\\ & _tw(t,x,μ)+b (t,x,μ,θ)· _xw(t,x,μ)+ 12σ [σ ] (t,x,μ,θ) ∂^2_xw(t,x,μ)\\ & +E_ξ μ [b (t,ξ,μ,θ)· _μw(t,x,μ)(ξ)+ 12[σ (σ ) ](t,ξ,μ,θ) _v _μw(t,x,μ)(ξ) ]. split (3.26) Define for all (t,x,a,Γ)∈[0,T]×ℝn×ℝd×2(ℝn×ℝd) (t,x,a, )∈[0,T]×R^n×R^d×P_2(R^n×R^d), μ(dx)≔Γ(dx,ℝd) μ(dx) (dx,R^d), and qD(t,x,a,Γ) q_D(t,x,a, ) (3.27) ≔∂tVDθ(t,x,μ)+b(t,x,a,Γ)⋅∂xVDθ(t,x,μ)+12[σσ⊤](t,x,a,Γ):∂x2VDθ(t,x,μ) _tV_D^θ(t,x,μ)+b(t,x,a, )· _xV_D^θ(t,x,μ)+ 12[σ ](t,x,a, ) ∂^2_xV_D^θ(t,x,μ) +(ξ,α)∼Γ[b(t,ξ,α,Γ)⋅∂μVDθ(t,x,μ)(ξ)+12[σ⊤](t,ξ,α,Γ):∂v∂μVDθ(t,x,μ)(ξ)] +E_(ξ,α) [b(t,ξ,α, )· _μV_D^θ(t,x,μ)(ξ)+ 12[σ ](t,ξ,α, ) _v _μV_D^θ(t,x,μ)(ξ) ] +r(t,x,a,Γ). +r(t,x,a, ). Since hφ(t,x,μ,θ)=h(t,x,φθ(t,x,μ),(id,φθ(t,⋅,μ))♯μ) h (t,x,μ,θ)=h(t,x, _θ(t,x,μ),( * id, _θ(t,·,μ))_ μ) for all h∈b,σ,r h∈\b,σ,r\, for all (t,x,μ)∈[0,T]×ℝn×2(ℝn) (t,x,μ)∈[0,T]×R^n×P_2(R^n) and θ′∈ℝk θ ∈R^k, ℒDφ[VDθ](t,x,μ,θ′)=[qD−r](t,x,φθ′(t,x,μ),(id,φθ′(t,⋅,μ))♯μ), _D [V_D^θ](t,x,μ,θ )=[q_D-r](t,x, _θ (t,x,μ),( * id, _θ (t,·,μ))_ μ), which along with (3.25) implies that qD q_D satisfies the martingale condition (3.20). Applying Itô’s formula to s↦[VDθ(s,Xst,ξ,θ,ℙst,μ,θ)] s E[V^θ_D(s,X_s^t,ξ,θ,P_s^t,μ,θ)] further shows that VD V_D and qD q_D satisfy (3.19), which proves that VD V_D and qD q_D are candidate functions satisfying Theorem 3.2. We emphasize that in (3.27), it is essential to consider a function qD q_D on a lifted space 2(ℝn×ℝd) _2(R^n×R^d), even when all coefficients are independent of the control distribution as in classical MFC problems. This is because (3.26) requires simultaneously integrating the coefficients b b, σ σ, and the policy φθ _θ over the state component. The characterization of the advantage function thus necessarily involves integrating with respect to the joint state–action law induced by the policy φθ _θ. 4 Model-free Advantage Actor-Critic Algorithm Building on Theorem 3.2, we now develop a model-free advantage actor–critic RL algorithm for extended MFC problems based on neural network (N) parameterizations. We first outline the key components of the framework and then provide implementation details. In the sequel, we denote by Vϕ,qψ V_φ,q_ψ and φθ _θ generic N approximations of the value function, advantage rate function, and policy, respectively. 4.1 Algorithmic design Neural network architecture with measure inputs. Although both Theorem 3.1 and Theorem 3.2 characterize the DPG for a given policy, Theorem 3.2 provides a more effective representation of the associated value and advantage rate functions than Theorem 3.1, thereby facilitating more efficient learning. Specifically, Theorem 3.1 represents the value function V V as a function of time and the state measure, and the advantage rate function q q as a function of time, the state measure, and the policy parameter. Such measure-dependent functions can be approximated by cylinder functions based on finite-dimensional features of the distribution, namely V^(t,μ)≈NNϕ(t,ξ∼μ[Ψ(ξ)]),q^(t,μ,θ)≈NNψ(t,ξ∼μ[Ψ~(ξ)],θ), V(t,μ)≈ N_φ(t,E_ξ μ[ (ξ)]), q(t,μ,θ)≈ N_ψ(t,E_ξ μ[ (ξ)],θ), (4.1) where NNϕ N_φ and NNψ N_ψ denote generic NNs, and Ψ,Ψ~:ℝn→ℝm , :R^n ^m are some prescribed or learnable feature maps [12, 23, 36]. However, updating the value and advantage rate functions according to (3.13) requires sampling multiple state trajectories to estimate the state distribution, which is not sample efficient. Moreover, since N policies typically involve high-dimensional parameter spaces and complex architectures, it is often challenging to design and learn q q as a function that directly maps policy parameters to the value of the advantage rate function. In contrast, Theorem 3.2 introduces a stronger martingale condition to identify the value function V V and advantage rate function q q through the local value function V^D V_D, and the local advantage rate function q^D q_D. This motivates parameterizing V^D V_D and q^D q_D directly as follows: V^D(t,x,μ)≈NNϕ(t,x,ξ∼μ[Ψ(ξ)]),q^D(t,x,a,Γ)≈NNψ(t,x,a,(ξ,α)∼Γ[Ψ~(ξ,α)]). V_D(t,x,μ)≈ N_φ(t,x,E_ξ μ[ (ξ)]), q_D(t,x,a, )≈ N_ψ(t,x,a,E_(ξ,α) [ (ξ,α)]). (4.2) These neural networks can be viewed as feature extractors, enabling accurate approximations of V V and q q even with simple feature maps Ψ and Ψ~ , such as polynomials. Moreover, the parameterization of q^D q_D in (4.2) takes action as an input variable rather than the policy parameter as in (4.1). Since the action space is typically much lower-dimensional than the policy parameter space, and more directly related to the state dynamics, parameterizing the map from action to the value of the advantage rate function leads to a simpler approximation problem and improves learning efficiency. Finally, the martingale criterion (3.20) also yields more informative temporal-difference updates of the NNs in (4.2) directly based on observed state and control trajectories, thereby improving learning efficiency, as shown in Section 5. Based on the above observations, we adopt (4.2) to parameterize V^D V_D and q^D q_D, and denote them by Vϕ V_φ and qψ q_ψ, respectively. Learning critics. We train Vϕ V_φ and qψ q_ψ so that they satisfy the conditions (3.19) and (3.20). Following similar arguments as in [19, 39, 7, 15], the martingale condition (3.20) can be enforced by updating Vϕ V_φ and qψ q_ψ through the following temporal-difference learning scheme: ϕ←ϕ−η[∂ϕVϕ(t,xt,μt)(Vϕ(t,xt,μt)−[r−qψ](t,xt,at,Γt)h−e−βhVϕ(t+h,xt+h,μt+h))],φ←φ- [ _φV_φ(t,x_t, _t) (V_φ(t,x_t, _t)-[r-q_ψ](t,x_t,a_t, _t)h-e^-β hV_φ(t+h,x_t+h, _t+h) ) ], (4.3) ψ←ψ−η[∂ψqψ(t,xt,at,Γt)(Vϕ(t,xt,μt)−[r−qψ](t,xt,at,Γt)h−e−βhVϕ(t+h,xt+h,μt+h))],ψ←ψ- [ _ψq_ψ(t,x_t,a_t, _t) (V_φ(t,x_t, _t)-[r-q_ψ](t,x_t,a_t, _t)h-e^-β hV_φ(t+h,x_t+h, _t+h) ) ], (4.4) where h>0 h>0 is the time step size, and (xt,μt) (x_t, _t) represents the state process and its distribution at time t t. To meet the Bellman constraint (3.19), we re-parameterize the advantage rate function qψ(t,x,a,Γ)≔q¯ψ(t,x,a,Γ)−ξ∼μq¯ψ(t,ξ,φθ(t,ξ,μ),(id,φθ(t,⋅,μ))♯μ)),q_ψ (t,x,a, ) q_ψ (t,x,a, )-E_ξ μ q_ψ (t,ξ, _θ(t,ξ,μ),( * id, _θ(t,·,μ))_ μ) ), (4.5) where q¯ q is an N of the form (4.2), and φθ _θ represents the current deterministic policy. To enforce the terminal condition in (3.19), we introduce a penalty term (Vϕ(T,xT,μT)−g(xT,μT))2,E(V_φ(T,x_T, _T)-g(x_T, _T))^2, (4.6) and minimize it with respect to ϕ φ via stochastic gradient descent based on the observed xT,μT x_T, _T and g(xT,μT) g(x_T, _T). Exploration. The use of a deterministic actor does not eliminate exploration. During training, exploration can be introduced either by perturbing the actions generated by the current policy (see e.g., [37, 7]) or by perturbing the parameters of the policy network [30]. The essential distinction is that randomness is used as a data-collection mechanism rather than being encoded into the policy class itself. The learned policy remains a deterministic feedback control, while exploration can be adjusted according to the geometry of the action and parameter spaces. In our implementation, given σepl>0 _epl>0, we consider the following two exploration mechanisms: • Approach I (action space exploration): Perform at∼(φθ(t,xt,μt),σepl2I) a_t ( _θ(t,x_t, _t), _epl^2I). • Approach I (parameter space exploration): Perform at=φθ′(t,xt,μt) a_t= _θ (t,x_t, _t) where θ′∼(θ,σepl2I) θ (θ, _epl^2I). Approach I explores the action space by perturbing deterministic policies with noise, which is the standard approach in single-agent RL (see e.g., [37, 7]). In general, this approach may not fully explore the parameter space, especially when the Jacobian ∂θφθ _θ _θ is non-invertible. Approach I explores the parameter space and is more consistent with Theorem 3.2. However, as shown in [30], its performance is more sensitive to the choice of noise scale σepl _epl than Approach I. This is because the effect of perturbing the parameters on the induced actions depends on the geometry of the policy parameterization, which is locally governed by the Jacobian of the policy network. Consequently, depending on the network architecture, parameter-space noise can lead to unstable and unpredictable behavior in the action space, necessitating careful tuning of the exploration noise scale. Our experiments suggest that, with appropriate tuning of the perturbation scale in Approach I, the two approaches achieve comparable performance in standard mean field control problems. However, in certain extended MFC problems, the dependence on the control law can make Approach I less effective, resulting in slower convergence. See Section 5.2 for details. 4.2 Implementation details McKean-Vlasov dynamics simulator. We construct a black-box simulator of the McKean–Vlasov dynamics (3.4) using a standard particle approximation with Euler–Maruyama discretization [32], and observe the resulting state and control trajectories as well as the associated rewards. Specifically, let h>0 h>0 be the time step size, and M∈ℕ M be the number of homogeneous agents (particles). Let (xkh,j,akh,j) (x_kh,j,a_kh,j) be the state-action of the j j-th agent at time kh kh, for j=1,…,M j=1,…,M, define the empirical measures μ^kh≔1M∑j=1Mδxkh,j μ_kh 1M _j=1^M _x_kh,j and Γ^kh≔1M∑j=1Mδ(xkh,j,akh,j) _kh 1M _j=1^M _(x_kh,j,a_kh,j), and simulate the states at time (k+1)h (k+1)h by x(k+1)h,j≔xkh,j+b(kh,xkh,j,akh,j,Γ^kh)h+σ(kh,xkh,j,akh,j,Γ^kh)hZkh,j,1≤j≤M,x_(k+1)h,j x_kh,j+b(kh,x_kh,j,a_kh,j, _kh)h+σ(kh,x_kh,j,a_kh,j, _kh) hZ_kh,j, 1≤ j≤ M, (4.7) where Zkh,jk=0,…,⌈T/h⌉,j=1,…,M \Z_kh,j\_k=0,…, T/h ,j=1,…,M are independent m m-dimensional standard normal random vectors. The observed instantaneous reward at time kh kh is given by rkh,j=r(kh,xkh,j,akh,j,Γ^kh) r_kh,j=r(kh,x_kh,j,a_kh,j, _kh), and the reward at time T T is given by rT,j=g(xT,j,μ^T) r_T,j=g(x_T,j, μ_T), for all j=1,…,M j=1,…,M. Critic update. Due to the off-policy nature of Algorithm 1, we store the observed transition (kh,xkh,j,akh,j,rkh,j,x(k+1)h,j)j=1M (kh,x_kh,j,a_kh,j,r_kh,j,x_(k+1)h,j )_j=1^M in a replay buffer ℛ , and re-sample them later to update actor and critic. We also store the terminal state and reward (Kh,xKh,j,rKh,j)j=1M (Kh,x_Kh,j,r_Kh,j)_j=1^M in ℛ . For brevity, we denote by x~t x_t the concatenation of time and state (t,xt) (t,x_t). We also employ a target value network Vϕtgt V_φ^tgt, whose parameters are updated as an exponential moving average of the value network parameters. This technique is widely used in modern deep RL algorithms, including DDPG [20] and SAC [16], to improve training stability. During each episode, after every m≥1 m≥ 1 steps of simulation, we sample a batch of transitions (i)i=1B \D^(i)\_i=1^B from replay buffer ℛ where (i)=(x~kih,j(i),akih,j(i),rkih,j(i),x~(ki+1)h,j(i))j=1M ^(i)= ( x^(i)_k_ih,j,a^(i)_k_ih,j,r^(i)_k_ih,j, x^(i)_(k_i+1)h,j )_j=1^M, corresponding to different trajectories and time indices, and define the martingale loss by ℒℳ≔1BM∑i=1B∑j=1M(Vϕ(x~kih,j(i),μ^kih(i))−[rkih,j(i)−qψ(x~kih,j(i),akih,j(i),Γ^kih(i))]h−e−βhVϕtgt(x~(ki+1)h,j(i),μ^(ki+1)h(i)))2, split&L^M 1BM _i=1^B _j=1^M (V_φ( x^(i)_k_ih,j, μ^(i)_k_ih)-[r^(i)_k_ih,j-q_ψ( x^(i)_k_ih,j,a^(i)_k_ih,j, ^(i)_k_ih)]h-e^-β hV_φ^tgt( x^(i)_(k_i+1)h,j, μ^(i)_(k_i+1)h) )^2, split (4.8) where qψ q_ψ is given by (4.5). Note that ∂ϕℒℳ _φL^M and ∂ψℒℳ _ψL^M are stochastic estimates of the semi-gradients in (4.3) and (4.4), respectively. To enforce the terminal constraints, we also sample a batch of terminal states ter(i)i=1B \D_ter^(i)\_i=1^B from ℛ from different trajectories and compute the terminal loss ℒ=1BM∑i=1B∑j=1M(Vϕ(x~Kh,j(i),μ^Kh(i))−rKh,j(i))2.L^T= 1BM _i=1^B _j=1^M(V_φ( x^(i)_Kh,j, μ^(i)_Kh)-r^(i)_Kh,j)^2. (4.9) The losses ℒℳ ^M and ℒ ^T are combined to update the critics Vϕ V_φ and qψ q_ψ via gradient descent. Actor update. To implement the exact policy gradient of the current policy φθ _θ given in (3.22), we sample a batch (i)i=1B \D^(i)\_i=1^B from ℛ , and define the actor objective by ℒ=1MB∑j=1M∑i=1Bqψ(x~kih,j,φθ(x~kih,j,μ^kih),Γkihθ)h,Γkihθ≔1M∑j=1Mδ(xkih,j,φθ(x~kih,j,μkih)).L^A= 1MB _j=1^M _i=1^Bq_ψ ( x_k_ih,j, _θ( x_k_ih,j, μ_k_ih), ^θ_k_ih )h, ^θ_k_ih 1M _j=1^M _(x_k_ih,j, _θ( x_k_ih,j, _k_ih)). (4.10) Note that ∂θℒ _θL^A is a stochastic estimator of policy gradient (3.22). Full algorithm. Algorithm 1 Continuous Time Deep Deterministic Policy Gradient Inputs: Discretization stepsize h h, horizon K=T/h K=T/h, number of episodes N N, number of agents M M, policy net φθ _θ, advantage-rate net q¯ψ q_ψ, value net Vϕ V_φ, update frequency m m, exploration noise σepl _epl, soft update parameter τ τ, learning rate η η, batch size B B, terminal constraint weight w w Initialization: ϕ,ψ,θ φ,ψ,θ, target ϕtgt=ϕ φ^tgt=φ, and replay buffer ℛ=∅ = for n=1,⋯,N n=1,·s,N do Observe states (x~0,j)j=1M ( x_0,j )_j=1^M for k=0,⋯,K−1 k=0,·s,K-1 do ⊳ action Option I. Perform akh,j∼(φθ(x~kh,j,μ^kh),σepl2I) a_kh,j ( _θ( x_kh,j, μ_kh), _epl^2I) for each j∈[M] j∈[M] Option I. Perform akh,j=φθ′(x~kh,j,μ^kh) a_kh,j= _θ ( x_kh,j, μ_kh) where θ′∼(θ,σepl2I) θ (θ, _epl^2I) for each j∈[M] j∈[M] Observe =(x~kh,j,akh,j,rkh,j,x~(k+1)h,j)j=1M = ( x_kh,j,a_kh,j,r_kh,j, x_(k+1)h,j )_j=1^M and store them in ℛ if k≡0 mod m k≡ 0 mod m then ⊳ critics Sample (i)i=1B \D^(i)\_i=1^B from ℛ and compute ℒℳ ^M as in (4.8) Sample terminal states ter(i)i=1B \D_ter^(i)\_i=1^B from ℛ and compute ℒ ^T as in (4.9) Update the critics: ψ←ψ−η∂ψℒM ψ←ψ-η _ψL^M, ϕ←ϕ−η∂ϕ(ℒM+wℒ) φ←φ-η _φ(L^M+wL^T) Update the target: ϕtgt←τϕ+(1−τ)ϕtgt φ^tgt←τφ+(1-τ)φ^tgt ⊳ actor Sample (i)i=1B \D^(i)\_i=1^B from ℛ and compute the policy loss ℒ ^A as in (4.10) Update the actor: θ←θ+η∂θℒ θ←θ+η _θL^A end if end for end for The full procedure of Continuous Time Deep Deterministic Policy Gradient (CT-DDPG) for extended MFC problems is summarized in Algorithm 1. 5 Numerical Experiments In this section, we illustrate the efficiency of the proposed CT-DDPG algorithm within the general learning MFC framework. Unless otherwise specified, CT-DDPG refers to the implementation with action space exploration (Option I in Algorithm 1). 5.1 Consensus control of Cucker-Smale model We consider the consensus control of multi-dimensional stochastic Cucker-Smale (C-S) models studied in [26, 31]. In this problem, the agent aims to enforce consensus emergence of an interactive particle system via external intervention. Let T>0 T>0, and consider the following 2n 2n-dimensional controlled McKean-Vlasov dynamics, which can be viewed as the large population limit of the finite-particle model studied in [9]: for all t∈[0,T] t∈[0,T], dzt=vtdt,dvt=(αt+∫ℝn×ℝnκ(zt,vt,z′,v′)ℙ(zt,vt)(dz′,dv′))dt+σdWt,dz_t=v_tdt, dv_t= ( _t+ _R^n×R^nκ(z_t,v_t,z ,v )P_(z_t,v_t)(dz ,dv ) )dt+σdW_t, (5.1) with an initial state (z0,v0)∈L2(ℱ0;ℝn×ℝn) (z_0,v_0)∈ L^2(F_0;R^n×R^n), where σ∈ℝn×n σ ^n× n, W W is an n n-dimensional Brownian motion, and the interaction kernel κ:ℝn×ℝn×ℝn×ℝn→ℝn κ:R^n×R^n×R^n×R^n ^n is defined by κ(z,v,z′,v′)=C(v′−v)(1+‖z′−z‖2)γ,with some C,γ≥0.κ(z,v,z ,v )= C(v -v)(1+\|z -z\|^2)^γ, \ with some C,γ≥ 0. (5.2) The variables z z and v v represent the position and velocity of a representative particle, respectively. The dynamics (5.1) models the self-organization behavior of a crowd, where the interaction kernel κ κ captures the decay of influence between agents as their distance increases. It is known that, in the absence of external control, flocking behavior, namely, the convergence of the velocity trajectories v v to a common limit as t→∞ t→∞, occurs only when the parameter γ γ is sufficiently small. The aim of the agent is to either induce consensus on models that would otherwise diverge, or to accelerate the self-organization behavior. Specifically, given c>0 c>0, we consider maximizing the following reward functional over all adapted controls α α: J(α)=−[∫0T(|vt−[vt]|2+c|αt|2)dt+|vT−[vT]|2]J(α)=-E [ _0^T(|v_t-E[v_t]|^2+c| _t|^2)dt+|v_T-E[v_T]|^2 ] (5.3) Note that in the special case with γ=0 γ=0, it is a linear-quadratic (LQ) mean field control problem. In this case, the optimal policy is linear and given explicitly by (see e.g., [40]): φ∗(t,z,v,μ)=−P(t)c(v−v′∼μ[v′]), ^*(t,z,v,μ)=- P(t)c(v-E_v μ[v ]), (5.4) where dtP(t)−1cP(t)2−2CP(t)+1=0 ddtP(t)- 1cP(t)^2-2CP(t)+1=0 with P(T)=1 P(T)=1. Baselines. We compare our CT-DDPG with existing mean-field RL algorithms, including Actor-Critic (AC) [11], which requires the action dependence of dynamics to be explicitly known, and q-Learning [39], which relies on a closed-form policy given by state-dependent Gibbs measures. Due to the model-aware nature of these two methods, and especially their incompatibility with deep RL framework, we evaluate them only in the LQ case. In addition to the default procedures described in Algorithm 1, we also consider a variant CT-DDPG-Flow, which learns the value function V V directly based on the deterministic flow characterization in Theorem 3.1, while still using the decoupled advantage rate estimator q^D q_D to represent q q. Model architecture. Across all experiments, the policy, advantage network, and value network in CT-DDPG and its variants are implemented as three-layer fully connected MLPs with ReLU activations and hidden width 400 400. To incorporate temporal information, we augment the environment observations with a sinusoidal embedding, yielding x~t=(cos(2πtT),sin(2πtT),zt,vt) x_t=( ( 2π tT), ( 2π tT),z_t,v_t), where T T denotes the maximum horizon. To embed the mean-field distribution, the default parameterization of CT-DDPG is given by Vϕ(x~t,μt)=NNϕ1(x~t,x′∼μtNNϕ2(x′)),V_φ( x_t, _t)=N_ _1 ( x_t,E_x _tN_ _2(x ) ), (CT-DDPG) where ϕ=(ϕ1,ϕ2) φ=( _1, _2) and similarly for φθ _θ and q¯ψ q_ψ. Here we use a learnable feature map to preserve the model-agnostic nature of our method. We also consider a variant based on prescribed moment features of the distribution, referred to as CT-DDPG-Moment: Vϕ(x~t,μt)=NNϕ(x~t,x′∼μt[x′],Varx′∼μt[x′]),V_φ( x_t, _t)=N_φ ( x_t,E_x _t[x ],Var_x _t[x ] ), (CT-DDPG-Moment) As a benchmark, we further evaluate a variant that incorporates prior knowledge of the dynamics through a hand-crafted feature map, which we denote by CT-DDPG-Kernel: Vϕ(x~t,μt)=NNϕ(x~t,(z′,v′)∼μt[v′−vt(1+‖z′−zt‖2)γ]),V_φ( x_t, _t)=N_φ ( x_t,E_(z ,v ) _t [ v -v_t(1+\|z -z_t\|^2)^γ ] ), (CT-DDPG-Kernel) This variant requires the exponent γ γ to be specified a priori, and is therefore less model-agnostic than the default parameterization. Nevertheless, all these variants remain model-free in the sense that they do not exploit closed-form expressions for the optimal policy or value function, in contrast to previous works such as [11, 39]. For AC and q-Learning, we follow [11, 39] respectively and adopt the model-aware LQ parameterization. Specifically, the value function is parameterized as a quadratic function of the state and the policy as a linear function of the state with time dependence incorporated through a three-layer MLP based embedding. Training hyperparameters. We set dimension n=3 n=3 and simulate M=50 M=50 homogeneous agents to approximate the McKean-Vlasov dynamics. To accelerate training, we run 8 environments in parallel, i.e., collecting 8 trajectories per episode. All neural networks are optimized by Adam with a learning rate of 3×10−4 3× 10^-4 and a batch size of B=256 B=256. The update frequency is m=1 m=1. The soft target update parameter is τ=0.1 τ=0.1. The weight for the terminal value constraint is w=0.002 w=0.002. And the default exploration noise scale is σepl=0.1 _epl=0.1. The entropy regularization coefficient in AC and q-Learning is 0.01 0.01. Figure 1: Results for C-S model with γ=0 γ=0. AC and q-Learning exploit the LQ structure and corresponding parameterization, while CT-DDPG uses neural network parameterized policy. Figure 2: Results for C-S model with γ=1 γ=1. Different designs of mean-field distribution embeddings in CT-DDPG lead to different convergence rates. Figure 3: Results for C-S model with γ=1,h=0.0005 γ=1,h=0.0005, where CT-DDPG is implemented with different exploration strategies. Results for LQ case. We first set γ=0,C=1,c=0.1,σ=0.3,T=1 γ=0,C=1,c=0.1,σ=0.3,T=1 in C-S dynamics and the initial state distribution is μ0=Unif([0,1]n)×Unif([−1,1]n) _0=Unif([0,1]^n)×Unif([-1,1]^n). The optimal policy is given explicitly in (5.4). We shall examine the effect of different discretization step sizes; the corresponding results are reported in Fig. 1. Observe that CT-DDPG achieves the fastest convergence to the optimal value, despite not explicitly exploiting the LQ structure. In contrast, AC and q-Learning adopt an LQ-based parameterization and nevertheless underperform CT-DDPG, further demonstrating the effectiveness of the proposed method. Results for non-LQ case. We further set γ=1 γ=1, which imposes a non-trivial coupling between z z and v v. The results of different distribution embedding methods in CT-DDPG are shown in Fig. 2. We observe that both CT-DDPG and CT-DDPG-Kernel converge rapidly, indicating that a learnable neural network feature map can achieve performance comparable to that obtained when prior information is incorporated. This demonstrates the efficiency and robustness of the proposed approach, as well as its compatibility with the deep RL framework. Although CT-DDPG-Moment uses only generic moment embeddings of the distribution and therefore converges more slowly, it still eventually attains the same optimal return as CT-DDPG. This further illustrates the effectiveness of deep-RL-based methods for MFC problems. For CT-DDPG-Flow, it converges rather slowly, and the performance gap becomes more pronounced as h h decreases. As discussed in Sec. 4, although the value function can be directly characterized via Theorem 3.1, the martingale-based temporal difference loss induced by Theorem 3.2 provides a more informative signal for learning the decoupled value function, thereby substantially accelerating convergence. To compare different exploration strategies, we evaluate CT-DDPG with both action-space and parameter-space exploration under varying exploration noise scales. As shown in Fig. 3, action-space exploration is more robust to the choice of noise scale, both in terms of convergence rate and final performance. In contrast, parameter-space exploration may suffer from severe instability when the perturbation scale is not properly tuned. Overall, with appropriate tuning, the two approaches achieve comparable performance in standard MFC problems. 5.2 Optimal liquidation with trade crowding In this section, we examine the performance of Algorithm 1 in a portfolio liquidation problem with trading crowd, as studied in [1, 32, 29]. In this problem, a large number of market participants aim to liquidate their positions in the same asset by a fixed terminal time T T, while accounting for the permanent price impact induced by their trading actions. The resulting cooperative equilibrium gives rise to an extended MFC problem for a representative agent. Let (αt)t∈[0,T] ( _t)_t∈[0,T] denote the trading speed chosen by the representative agent. The state dynamics of the extended MFC problem are given by: for all t∈[0,T] t∈[0,T], dQt=αtdt,dSt=λ[αt]dt+σdWt.dQ_t= _tdt, dS_t= [ _t]dt+σdW_t. (5.5) Here Q=(Qt)t∈[0,T] Q=(Q_t)_t∈[0,T] denotes the inventory process with a random initial state Q0 Q_0 representing the initial inventories for all participants, S=(St)t∈[0,T] S=(S_t)_t∈[0,T] is the asset price process, the term λ[αt] λE[ _t] with λ≥0 λ≥ 0 captures the permanent market impact on the asset price induced by the aggregate trading of all participants, and W W is a one-dimensional Brownian motion representing exogenous market noise. The objective of the agent is to maximize the following reward functional over all adapted trading speeds: J(α)=[∫0T−(Qt2+αtSt+cαt2)dt+QT(ST−CQT)],J(α)=E [ _0^T-(Q_t^2+ _tS_t+c _t^2)dt+Q_T(S_T-CQ_T) ], (5.6) where the term αt(St+cαt) _t(S_t+c _t) represents the instantaneous liquidation cost with c>0 c>0 representing the linear temporary market impact, Qt2 Q^2_t penalizes inventory risk over time, and QT(ST−CQT) Q_T(S_T-CQ_T) with C>0 C>0 is the liquidation value of the remaining inventory at terminal time. Note that (5.5) and (5.6) constitute an LQ extended MFC problem and the optimal policy is given by (see e.g., [40]): φ∗(t,s,q,μ)=−1c(PC(t)(q−q′∼μ[q′])+PC−λ2(t)q′∼μ[q′]), ^*(t,s,q,μ)=- 1c (P_C(t)(q-E_q μ[q ])+P_C- λ2(t)E_q μ[q ] ), (5.7) where for all r∈ℝ r , (Pr(t))t∈[0,T] (P_r(t))_t∈[0,T] satisfies dtPr(t)−1cPr(t)2+1=0 ddtP_r(t)- 1cP_r(t)^2+1=0 with Pr(T)=r P_r(T)=r. The model architectures and training hyperparameters are the same as described in the previous section. We set λ=0.2,σ=0.3,c=0.1,C=1,T=1,h=0.0005 λ=0.2,σ=0.3,c=0.1,C=1,T=1,h=0.0005 and the initial mean-field distribution is μ0=Unif([0.8,1.2])×Unif([0.8,1.2]) _0=Unif([0.8,1.2])×Unif([0.8,1.2]). The results of CT-DDPG with different exploration mechanisms are shown in Fig. 4. Action-space exploration converges rapidly and robustly to the optimal value under various exploration noise scales, demonstrating its effectiveness for the extended MFC problems. Parameter-space exploration is still sensitive to noise scale, but it exhibits a faster convergence if the noise scale is properly tuned. Due to the specific structure of the liquidation problem, action space exploration does not change the expected control, and therefore provides limited exploration for the price process in (5.5), leading to slightly slower convergence. This illustrates the potential advantage of parameter-space exploration in certain extended MFC problems, provided that the hyperparameters are chosen appropriately. Figure 4: Results of CT-DDPG for liquidation problem with h=0.0005 h=0.0005. In summary, CT-DDPG exhibits superior performance in terms of convergence speed and stability across all environment settings, verifying the efficiency and robustness of our method. 6 Proofs 6.1 Proofs of Technical Results This section establishes several technical results used in the proofs of Theorems 2.1 and 2.3. We first recall that when the function V(⋅,⋅,θ) V(·,·,θ) defined in (2.2) is in C1,2([0,T]×2(ℝn)) C^1,2([0,T]×P_2(R^n)), it is the unique solution to an associated linear PDE. Lemma 6.1. Suppose Assumption 2.1 holds. Let θ∈ℝk θ∈R^k and assume V(⋅,⋅,θ)∈C1,2([0,T]×2(ℝn)) V(·,·,θ)∈ C^1,2([0,T]×P_2(R^n)). Then V(⋅,⋅,θ) V(·,·,θ) is the unique classical solution to the following linear PDE: ℒθ[w](t,μ)+r¯(t,μ,θ)=0,∀(t,μ)∈[0,T)×2(ℝn),w(T,μ)=g¯(μ),∀μ∈2(ℝn). split&L^θ[w](t,μ)+ r(t,μ,θ)=0, ∀(t,μ)∈[0,T)×P_2(R^n),\\ &w(T,μ)= g(μ), ∀μ _2(R^n). split (6.1) Using Lemma 6.1, the following proposition characterizes the performance difference between two value functions V(⋅,⋅,θ) V(·,·,θ) and V(⋅,⋅,θ′) V(·,·,θ ). It extends the performance-difference lemmas in [7, Lemma 6.1] and [35, Lemma 3.2] from classical control problems to the more general setting of McKean–Vlasov dynamics. Proposition 6.2. Suppose Assumption 2.1 holds. Let θ∈ℝk θ∈R^k and assume Vθ≔V(⋅,⋅,θ)∈C1,2([0,T]×2(ℝn)) V^θ V(·,·,θ)∈ C^1,2([0,T]×P_2(R^n)). For all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n) and θ′∈ℝk θ ∈R^k, Vθ′(t,μ)−Vθ(t,μ)=∫tTe−β(s−t)((ℒθ′[Vθ]−ℒθ[Vθ])(s,ℙst,μ,θ′)+r¯(s,ℙst,μ,θ′,θ′)−r¯(s,ℙst,μ,θ′,θ))ds. split&V^θ (t,μ)-V^θ(t,μ)\\ & = _t^Te^-β(s-t) ((L^θ [V^θ]-L^θ[V^θ])(s,P^t,μ,θ _s)+ r (s,P^t,μ,θ _s,θ )- r (s,P^t,μ,θ _s,θ ) )ds. split (6.2) Proof. Fix θ,θ′∈ℝk θ,θ ^k, (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), and ξ∈L2(ℱt,ℝn) ξ∈ L^2(F_t,R^n) with ξ∼μ ξ μ. Let Xt,ξ,θ′ X^t,ξ,θ be the solution to (2.1) with the parameter θ′ θ and recall ℙst,μ,θ′=ℙXst,ξ,θ′ ^t,μ,θ _s=P_X^t,ξ,θ _s for all s∈[t,T] s∈[t,T]. By the definition (2.2) of Vθ′(t,μ) V^θ (t,μ), Vθ′(t,μ)−Vθ(t,μ)=∫tTe−β(s−t)r¯(s,ℙst,μ,θ′,θ′)ds+e−β(T−t)g¯(ℙTt,μ,θ′)−Vθ(t,μ)=e−β(T−t)Vθ(T,ℙTt,μ,θ′)−Vθ(t,ℙt,μ,θ′)+∫tTe−β(s−t)r¯(s,ℙst,μ,θ′,θ′)ds, split&V^θ (t,μ)-V^θ(t,μ)\\ &= _t^Te^-β(s-t) r (s,P^t,μ,θ _s,θ )ds+e^-β(T-t) g(P^t,μ,θ _T)-V^θ(t,μ)\\ &=e^-β(T-t)V^θ (T,P^t,μ,θ _T )-V^θ(t,P^t,μ,θ _t)+ _t^Te^-β(s-t) r (s,P^t,μ,θ _s,θ )ds, split (6.3) where the last identity used the fact that Vθ(T,μ)=g¯(μ) V^θ(T,μ)= g(μ) and ℙt,μ,θ′=μ ^t,μ,θ _t=μ. As Vθ∈C1,2([0,T]×2(ℝn)) V^θ∈ C^1,2([0,T]×P_2(R^n)), applying Itô’s formula (e.g. [5, Proposition 5.102]) to s↦e−β(s−t)Vθ(s,ℙst,μ,θ′) s e^-β(s-t)V^θ(s,P^t,μ,θ _s) yields e−β(T−t)Vθ(T,ℙTt,μ,θ′)−Vθ(t,ℙt,μ,θ′)=∫tTe−β(s−t)ℒθ′[Vθ](s,ℙst,μ,θ′)ds e^-β(T-t)V^θ (T,P^t,μ,θ _T )-V^θ(t,P^t,μ,θ _t)= _t^Te^-β(s-t)L^θ [V^θ](s,P^t,μ,θ _s)ds =∫tTe−β(s−t)(ℒθ′[Vθ]−ℒθ[Vθ]+ℒθ[Vθ])(s,ℙst,μ,θ′)ds = _t^Te^-β(s-t)(L^θ [V^θ]-L^θ[V^θ]+L^θ[V^θ])(s,P^t,μ,θ _s)ds =∫tTe−β(s−t)((ℒθ′[Vθ]−ℒθ[Vθ])(s,ℙst,μ,θ′)−r¯(s,ℙst,μ,θ′,θ))ds, = _t^Te^-β(s-t) ((L^θ [V^θ]-L^θ[V^θ])(s,P^t,μ,θ _s)- r (s,P^t,μ,θ _s,θ ) )ds, where ℒθ′[Vθ] ^θ [V^θ] and ℒθ[Vθ] ^θ[V^θ] are defined in (2.6), and the last identity used the PDE (6.1). This along with (6.3) proves the desired result. ∎ We further prove the joint continuity of the state law with respect to time and the parameter. Lemma 6.3. Suppose Assumption 2.1 holds. For all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), the map [t,T]×ℝk∋(s,θ)↦ℙst,μ,θ∈2(ℝn) [t,T]×R^k (s,θ) ^t,μ,θ_s _2(R^n) is continuous. Proof. Fix (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n). By Assumption 2.1, for all θ θ in a compact set, the coefficients b b and σ σ are Lipschitz continuous and of linear growth, and hence Assumptions (H1)-(H4) in [22] hold (the Lyapunov condition (H2) holds with the function V(x,μ)=1+|x|2+M2(μ)2) V(x,μ)=1+|x|^2+M_2(μ)^2). Thus by [22, Theorem 3.1], θ↦ℙst,μ,θ θ ^t,μ,θ_s is continuous for all s∈[t,T] s∈[t,T]. We now claim that for any bounded set K⊂ℝk K⊂R^k, there exists a constant CK≥0 C_K≥ 0 such that for all s,r∈[t,T] s,r∈[t,T] and θ∈K θ∈ K, W2(ℙrt,μ,θ,ℙst,μ,θ)≤CK|r−s|1/2.W_2(P^t,μ,θ_r,P^t,μ,θ_s)≤ C_K|r-s|^1/2. (6.4) This along with the continuity of θ↦ℙst,μ,θ θ ^t,μ,θ_s implies the joint continuity of (θ,s)↦ℙst,μ,θ (θ,s) ^t,μ,θ_s. To show (6.4), assume without loss of generality that t≤s<r≤T t≤ s<r≤ T, let ξ∈L2(ℱt;ℝn) ξ∈ L^2(F_t;R^n) with μ=ℙξ μ=P_ξ, and let C≥0 C≥ 0 be a generic constant independent of s,r s,r and θ θ. Then by (2.1), W22(ℙrt,μ,θ,ℙst,μ,θ)≤[|Xrt,ξ,θ−Xst,ξ,θ|2] W_2^2(P^t,μ,θ_r,P^t,μ,θ_s)≤E[|X^t,ξ,θ_r-X^t,ξ,θ_s|^2] =[|∫srb(u,Xut,ξ,θ,ℙXut,ξ,θ,θ)du+∫srσ(u,Xut,ξ,θ,ℙXut,ξ,θ,θ)dWu|2] =E [ | _s^rb(u,X^t,ξ,θ_u,P_X^t,ξ,θ_u,θ)du+ _s^rσ(u,X^t,ξ,θ_u,P_X^t,ξ,θ_u,θ)dW_u |^2 ] ≤C([∫sr|b(u,Xut,ξ,θ,ℙXut,ξ,θ,θ)|2du]+[∫sr|σ(u,Xut,ξ,θ,ℙXut,ξ,θ,θ)|2du]) ≤ C (E [ _s^r|b(u,X^t,ξ,θ_u,P_X^t,ξ,θ_u,θ)|^2du ]+E [ _s^r|σ(u,X^t,ξ,θ_u,P_X^t,ξ,θ_u,θ)|^2du ] ) ≤C(r−s)(1+supu∈[s,r][|Xut,ξ,θ|2])≤C(r−s), ≤ C(r-s) (1+ _u∈[s,r]E[|X^t,ξ,θ_u|^2] )≤ C(r-s), where we have used the linear growth of b b and σ σ in 2.1 and the fact that [|Xst,ξ,θ|2]≤C E[|X^t,ξ,θ_s|^2]≤ C, uniformly with respect to s∈[t,T] s∈[t,T] and θ∈K θ∈ K. This proves the desired estimate (6.4). ∎ 6.2 Proofs of Theorems 2.1, 2.2, 2.3 and 3.2 We now prove Theorem 2.1 based on Proposition 6.2 and Lemma 6.3. Proof of Theorem 2.1. To simplify the notation, we write Vθ≔V(⋅,⋅,θ) V^θ V(·,·,θ) for all θ∈ℝk θ∈R^k. Since ∂θVθ(t,μ)=(∂θ1Vθ(t,μ),…,∂θkVθ(t,μ))⊤ _θV^θ(t,μ)=( _ _1V^θ(t,μ),…, _ _kV^θ(t,μ)) , it suffices to prove that for any fixed θ,θ′∈ℝk θ,θ ^k and (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n), dϵVθ+ϵθ′(t,μ)|ϵ=0=∫tTe−β(s−t)∂θA[Vθ](s,ℙst,μ,θ,θ)ds⋅θ′. split& ddεV^θ+εθ (t,μ) |_ε=0= _t^Te^-β(s-t) _θA[V^θ](s,P^t,μ,θ_s,θ)\,ds·θ . split (6.5) To this end, let ξ∈L2(ℱt,ℝn) ξ∈ L^2(F_t,R^n), and for all ϵ∈[−1,1] ε∈[-1,1], let (Xst,ξ,θ+ϵθ′)s∈[t,T] (X^t,ξ,θ+εθ _s)_s∈[t,T] be the solution to (2.1) with the parameter θ+ϵθ′ θ+εθ , and let ℙst,μ,θ+ϵθ′=ℙXst,ξ,θ+ϵθ′ ^t,μ,θ+εθ _s=P_X^t,ξ,θ+εθ _s. For all ϵ∈[−1,1] ε∈[-1,1], by Proposition 6.2 and the fundamental theorem of calculus, 1ϵ(Vθ+ϵθ′(t,μ)−Vθ(t,μ))=1ϵ∫tTe−β(s−t)(A[Vθ](s,ℙst,μ,θ+ϵθ′,θ+ϵθ′)−A[Vθ](s,ℙst,μ,θ+ϵθ′,θ))ds=∫tTe−β(s−t)(∫01∂θA[Vθ](s,ℙst,μ,θ+ϵθ′,θ+uϵθ′)du)ds⋅θ′. split& 1ε (V^θ+εθ (t,μ)-V^θ(t,μ) )\\ &= 1ε _t^Te^-β(s-t) (A[V^θ](s,P^t,μ,θ+εθ _s,θ+εθ )-A[V^θ](s,P^t,μ,θ+εθ _s,θ) )ds\\ &= _t^Te^-β(s-t) ( _0^1 _θA[V^θ](s,P^t,μ,θ+εθ _s,θ+uεθ )du )ds·θ . split (6.6) Observe that for all s∈[t,T] s∈[t,T] and u∈[0,1] u∈[0,1], by Lemma 6.3, limϵ→0ℙst,μ,θ+ϵθ′=ℙst,μ,θ _ε→ 0P^t,μ,θ+εθ _s=P^t,μ,θ_s for all s∈[0,T] s∈[0,T], and hence by Assumption 2.2, limϵ→0∂θA[Vθ](s,ℙst,μ,θ+ϵθ′,θ+uϵθ′)=∂θA[Vθ](s,ℙst,μ,θ,θ). _ε→ 0 _θA[V^θ](s,P^t,μ,θ+εθ _s,θ+uεθ )= _θA[V^θ](s,P^t,μ,θ_s,θ). Moreover, the continuity of (s,θ)↦ℙst,μ,θ (s,θ) ^t,μ,θ_s and the compactness of [t,T]×[−1,1] [t,T]×[-1,1] imply that ℙst,μ,θ+ϵθ′∣(s,ϵ)∈[t,T]×[−1,1] \P^t,μ,θ+εθ _s (s,ε)∈[t,T]×[-1,1]\ is compact in 2(ℝn) _2(R^n). Hence using the dominated convergence theorem and passing ϵ→0 ε→ 0 in (6.6) yields the desired identity. ∎ We further prove Theorems 2.2 and 2.3. Proof of Theorem 2.2. Vθ V^θ and qθ q^θ satisfy (2.8) due to Lemma 6.1 and the definition of qθ=A[Vθ] q^θ=A[V^θ] in (2.5). The functions Vθ V^θ and qθ q^θ satisfy (2.9) due to Itô’s formula (e.g. [5, Proposition 5.102]). ∎ Proof of Theorem 2.3. For all (t,μ,θ′)∈[0,T]×2(ℝn)×t,μ(θ) (t,μ,θ )∈[0,T]×P_2(R^n)×O_t,μ(θ) and s∈[t,T] s∈[t,T], by Itô’s formula (e.g. [5, Proposition 5.102]), e−βsV^(s,ℙst,μ,θ′)−e−βtV^(t,ℙt,μ,θ′)+∫tse−βur^(u,ℙut,μ,θ′,θ′)du e^-β s V(s,P^t,μ,θ _s)-e^-β t V(t,P^t,μ,θ _t)+ _t^se^-β u r(u,P^t,μ,θ _u,θ )\,du =∫tse−βu∂tV^(u,ℙut,μ,θ′)−βV^(u,ℙut,μ,θ′)+ξ∼ℙut,μ,θ′[b(u,ξ,ℙut,μ,θ′,θ′)⋅∂μV^(u,ℙut,μ,θ′)(ξ) = _t^se^-β u \ _t V(u,P^t,μ,θ _u)-β V(u,P^t,μ,θ _u)+E_ξ ^t,μ,θ _u [b(u,ξ,P^t,μ,θ _u,θ )· _μ V(u,P^t,μ,θ _u)(ξ) +12(σ⊤)(u,ξ,ℙut,μ,θ′,θ′):∂v∂μV^(u,ℙut,μ,θ′)(ξ)]+r¯(u,ℙut,μ,θ′,θ′)du + 12(σ )(u,ξ,P^t,μ,θ _u,θ ): _v _μ V(u,P^t,μ,θ _u)(ξ) ]+ r(u,P^t,μ,θ _u,θ ) \\,du =∫tse−βuA[V^](u,ℙut,μ,θ′,θ′)du. = _t^se^-β uA[ V](u,P^t,μ,θ _u,θ )\,du. This along with (2.11) implies ∫tse−βu(A[V^](u,ℙut,μ,θ′,θ′)−q^(u,ℙut,μ,θ′,θ′))du=0,∀s∈[t,T]. _t^se^-β u(A[ V](u,P^t,μ,θ _u,θ )- q(u,P^t,μ,θ _u,θ ))\,du=0, ∀ s∈[t,T]. (6.7) By Assumption 2.2 and Lemma 6.3, u↦e−βu(A[V^](u,ℙut,μ,θ′,θ′)−q^(u,ℙut,μ,θ′,θ′)) u e^-β u(A[ V](u,P^t,μ,θ _u,θ )- q(u,P^t,μ,θ _u,θ )) is continuous, and hence by (6.7), A[V^](u,ℙut,μ,θ′,θ′)=q^(u,ℙut,μ,θ′,θ′) A[ V](u,P^t,μ,θ _u,θ )= q(u,P^t,μ,θ _u,θ ) for all u∈[t,T] u∈[t,T]. Setting u=t u=t implies A[V^](t,μ,θ′)=q^(t,μ,θ′),∀(t,μ,θ′)∈[0,T]×2(ℝn)×t,μ(θ).A[ V](t,μ,θ )= q(t,μ,θ ), ∀(t,μ,θ )∈[0,T]×P_2(R^n)×O_t,μ(θ). Taking θ′=θ θ =θ and using (2.10) shows that V V is a classical solution to the following linear PDE: for all (t,μ)∈[0,T)×2(ℝn) (t,μ)∈[0,T)×P_2(R^n), A[V^](t,μ,θ)=ℒθ[V^](t,μ)+r¯(t,μ,θ)=0, A[ V](t,μ,θ)=L^θ[ V](t,μ)+ r(t,μ,θ)=0, and V^(T,μ)=g¯(μ) V(T,μ)= g(μ). By Lemma 6.1, V^(t,μ)=V(t,μ,θ) V(t,μ)=V(t,μ,θ) for all (t,μ)∈[0,T]×2(ℝn) (t,μ)∈[0,T]×P_2(R^n). This further implies that q^(t,μ,θ′)=A[V(⋅,⋅,θ)](t,μ,θ′) q(t,μ,θ )=A[V(·,·,θ)](t,μ,θ ) for all θ′∈t,μ(θ) θ _t,μ(θ). ∎ Proof of Theorem 3.2. It suffices to prove that (3.20) implies that V V given in (3.16) satisfies (3.13) in Theorem 3.1. To see it, by (3.20), e−βtV^D(t,Xtt,ξ,θ′,ℙt,μ,θ′)=[e−βsV^D(s,Xst,ξ,θ′,ℙst,μ,θ′)∣ℱt]+[∫tse−βu(r−q^D)(u,Xut,ξ,θ′,φθ′(u,Xut,ξ,θ′,ℙut,μ,θ′),(id,φθ′(u,⋅,ℙut,μ,θ′))♯ℙut,μ,θ′)du|ℱt], split&e^-β t V_D(t,X_t^t,ξ,θ ,P_t^t,μ,θ )=E[e^-β s V_D(s,X_s^t,ξ,θ ,P_s^t,μ,θ ) _t]\\ & +E [ _t^se^-β u(r- q_D)(u,X_u^t,ξ,θ , _θ (u,X_u^t,ξ,θ ,P_u^t,μ,θ ),( * id, _θ (u,·,P_u^t,μ,θ ))_ P_u^t,μ,θ )du\, |\,F_t ], split from which by taking expectations on both sides yields e−βtV^(t,μ)=[e−βsV^D(s,Xst,ξ,θ′,ℙst,μ,θ′)]+[∫tse−βu(r−q^D)(u,Xut,ξ,θ′,φθ′(u,Xut,ξ,θ′,ℙut,μ,θ′),(id,φθ′(u,⋅,ℙut,μ,θ′))♯ℙut,μ,θ′)du]. split&e^-β t V(t,μ)=E[e^-β s V_D(s,X_s^t,ξ,θ ,P_s^t,μ,θ )]\\ & +E [ _t^se^-β u(r- q_D)(u,X_u^t,ξ,θ , _θ (u,X_u^t,ξ,θ ,P_u^t,μ,θ ),( * id, _θ (u,·,P_u^t,μ,θ ))_ P_u^t,μ,θ )du ]. split This along with Fubini’s theorem, ℙXst,ξ,θ′=ℙst,μ,θ′ _X_s^t,ξ,θ =P_s^t,μ,θ and the definitions of V V, q q and rφ r yields (3.13). The desired conclusion then follows from Theorem 3.1 and the chain rule (3.18). ∎ Acknowledgements and Disclosure of Funding Huyên Pham is supported by the Société Générale Chair “Risques Financiers”, FiME (Laboratory of Finance and Energy Markets), and the EDF–CACIB Chair “Finance and Sustainable Development”. Yufei Zhang’s research is supported by CNRS-Imperial “Abraham de Moivre” International Research Lab in Mathematics. References [1] B. Acciaio, J. Backhoff-Veraguas, and R. Carmona (2019) Extended mean field control problems: stochastic maximum principle and transport perspective. SIAM journal on Control and Optimization 57 (6), p. 3666–3693. Cited by: §1, §3, §5.2. [2] E. Bayraktar, N. Bäuerle, and A. D. Kara (2025) Finite approximations for mean-field type multi-agent control and their near optimality. Applied Mathematics & Optimization 92 (1), p. 7. Cited by: §3.2. [3] E. Bayraktar, M. Hernandez, Q. Yan, and Y. Zhu (2026) Policy gradient for continuous-time mean-field control. arXiv preprint arXiv:2605.20718. Cited by: §1. [4] R. Buckdahn, J. Li, S. Peng, and C. Rainer (2017) Mean-field stochastic differential equations and associated pdes. Annals of Probability: An official journal of the Institute of Mathematical Statistics 45 (2), p. 824–878. Cited by: §3.2, §3.3. [5] R. Carmona and F. Delarue (2018) Probabilistic theory of Mean Field Games with applications i: Mean Field FBSDEs, control, and games. Vol. 83, Springer. Cited by: §3.3, §6.1, §6.2, §6.2. [6] R. Carmona, M. Laurière, and Z. Tan (2023) Model-free mean-field reinforcement learning: mean-field MDP and mean-field Q-learning. The Annals of Applied Probability 33 (6B), p. 5334–5381. Cited by: §1, §1. [7] Z. Cheng, X. Guo, and Y. Zhang (2025) Deterministic policy gradient for reinforcement learning with continuous time and state. arXiv preprint arXiv:2509.23711. Cited by: §1, §1, §2, §3.2, §4.1, §4.1, §4.1, §6.1. [8] A. Cosso, F. Gozzi, I. Kharroubi, H. Pham, and M. Rosestolato (2023) Optimal control of path-dependent McKean–Vlasov SDEs in infinite-dimension. The Annals of Applied Probability 33 (4), p. 2863–2918. Cited by: §1. [9] F. Cucker and S. Smale (2007) Emergent behavior in flocks. IEEE Transactions on automatic control 52 (5), p. 852–862. Cited by: §5.1. [10] J. Foerster, I. A. Assael, N. De Freitas, and S. Whiteson (2016) Learning to communicate with deep multi-agent reinforcement learning. Advances in neural information processing systems 29. Cited by: §3.2. [11] N. Frikha, M. Germain, M. Laurière, H. Pham, and X. Song (2025) Actor-critic learning for mean-field control in continuous time. Journal of Machine Learning Research 26 (127), p. 1–42. Cited by: §1, §2, §3.3, §5.1, §5.1, §5.1. [12] M. Germain, M. Laurière, H. Pham, and X. Warin (2022) DeepSets and their derivative networks for solving symmetric PDEs. Journal of Scientific Computing 91 (2), p. 63. Cited by: §4.1. [13] H. Gu, X. Guo, X. Wei, and R. Xu (2023) Dynamic programming principles for mean-field controls with learning. Operations Research 71 (4), p. 1040–1054. Cited by: §1, §1, §1, §3.2, §3.2. [14] H. Gu, X. Guo, X. Wei, and R. Xu (2025) Mean-field multiagent reinforcement learning: a decentralized network approach. Mathematics of Operations Research 50 (1), p. 506–536. Cited by: §3.2. [15] X. Guo, Y. Huang, and X. Yu (2026) Deterministic policy gradient for learning equilibrium in time-inconsistent control problems. arXiv preprint arXiv:2606.11798. Cited by: §1, §4.1. [16] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine (2018) Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, p. 1861–1870. Cited by: §4.2. [17] Y. Jia, D. Ouyang, and Y. Zhang (2026) Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning. SIAM Journal on Control and Optimization 64 (3), p. 1889–1929. Cited by: §1, §1. [18] Y. Jia and X. Y. Zhou (2022) Policy gradient and actor-critic learning in continuous time and space: theory and algorithms. Journal of Machine Learning Research 23 (275), p. 1–50. Cited by: §2. [19] Y. Jia and X. Y. Zhou (2023) Q-learning in continuous time. Journal of Machine Learning Research 24 (161), p. 1–61. Cited by: §4.1. [20] T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra (2015) Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971. Cited by: §4.2. [21] R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. Pieter Abbeel, and I. Mordatch (2017) Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in neural information processing systems 30. Cited by: §3.2. [22] J. Ma and Z. Liu (2025) Continuous dependence for McKean-Vlasov SDEs under distribution-dependent Lyapunov conditions. Discrete and Continuous Dynamical Systems-S 18 (11), p. 3282–3301. Cited by: §6.1. [23] S. Mekkaoui, H. Pham, and X. Warin (2026) Learning operators on labelled conditional distributions with applications to mean field control of non exchangeable systems. arXiv preprint arXiv:2603.21683. Cited by: §1, §4.1. [24] M. Meunier, H. Pham, and C. Reisinger (2026) Model-free policy gradient for discrete-time mean-field control. arXiv preprint arXiv:2601.11217. Cited by: §1. [25] M. Motte and H. Pham (2022) Mean-field markov decision processes with common noise and open-loop controls. The Annals of Applied Probability 32 (2), p. 1421–1458. Cited by: §1, §1, §3.2. [26] M. Nourian, P. E. Caines, and R. P. Malhamé (2011) Mean field analysis of controlled cucker-smale type flocking: linear analysis and perturbation equations. IFAC Proceedings Volumes 44 (1), p. 4471–4476. Cited by: §5.1. [27] H. Pham and X. Warin (2025) Actor-critic learning algorithms for mean-field control with moment neural networks. Methodology and Computing in Applied Probability 27 (1), p. 13. Cited by: §1, §1, §1, §2. [28] H. Pham and X. Wei (2018) Bellman equation and viscosity solutions for mean-field stochastic control problem. ESAIM: Control, Optimisation and Calculus of Variations 24 (1), p. 437–461. Cited by: §1, §2, §3.1, §3.2. [29] A. Picarelli, M. Scaratti, and J. Tam (2025) Extended mean field control: a global numerical solution via finite-dimensional approximation. arXiv preprint arXiv:2503.20510. Cited by: §3, §5.2. [30] M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz (2017) Parameter space noise for exploration. arXiv preprint arXiv:1706.01905. Cited by: §4.1, §4.1. [31] C. Reisinger, W. Stockinger, M. O. Tsianni, and Y. Zhang (2025) Convergence rates of time discretization in extended mean field control. arXiv preprint arXiv:2509.00904. Cited by: §5.1. [32] C. Reisinger, W. Stockinger, and Y. Zhang (2024) A fast iterative pde-based algorithm for feedback controls of nonsmooth mean-field control problems. SIAM Journal on Scientific Computing 46 (4), p. A2737–A2773. Cited by: §4.2, §5.2. [33] Z. Ren, X. Wei, X. Yu, and X. Y. Zhou (2026) Continuous-time q-learning for mean-field control with common noise, part-i: theoretical foundations. arXiv preprint arXiv:2604.27372. Cited by: §1. [34] Z. Ren, X. Wei, X. Yu, and X. Y. Zhou (2026) Continuous-time q-learning for mean-field control with common noise, part-i: q-learning algorithms. arXiv preprint arXiv:2604.27378. Cited by: §1, §1. [35] D. Sethi, D. Šiška, and Y. Zhang (2025) Entropy annealing for policy mirror descent in continuous time and space. SIAM Journal on Control and Optimization 63 (4), p. 3006–3041. Cited by: §2, §6.1. [36] H. M. Soner, J. Teichmann, and Q. Yan (2025) Learning algorithms for mean field optimal control. arXiv preprint arXiv:2503.17869. Cited by: §1, §4.1. [37] L. Szpruch, T. Treetanthiploet, and Y. Zhang (2024) Optimal scheduling of entropy regularizer for continuous-time linear-quadratic reinforcement learning. SIAM Journal on Control and Optimization 62 (1), p. 135–166. Cited by: §1, §1, §4.1, §4.1. [38] X. Wei, X. Yu, and F. Yuan (2024) Unified continuous-time q-learning for mean-field game and mean-field control problems. arXiv preprint arXiv:2407.04521. Cited by: §1. [39] X. Wei and X. Yu (2025) Continuous time q-learning for mean-field control problems. Applied Mathematics & Optimization 91 (1), p. 10. Cited by: §1, §2, §4.1, §5.1, §5.1, §5.1. [40] J. Yong (2017) Linear-quadratic optimal control problems for mean-field stochastic differential equations—time-consistent solutions. Transactions of the American Mathematical Society 369 (8), p. 5467–5523. Cited by: §5.1, §5.2.