Paper deep dive
Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations
Zhenjie Ren, Xiaoli Wei, Xiang Yu, Xun Yu Zhou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/8/2026, 7:19:26 AM
Summary
This paper develops the theoretical foundations for continuous-time q-learning in entropy-regularized mean-field control (MFC) with controlled common noise. It introduces the integrated q-function (Iq-function), derives the exploratory Hamilton-Jacobi-Bellman (HJB) equation, and characterizes the optimal policy as a two-layer fixed point, explicitly showing it follows a Gaussian distribution in linear-quadratic settings.
Entities (7)
Relation Signals (6)
Mean Field Control → incorporates → Common Noise
confidence 95% · incorporate the Brownian common noise in the mean-field system that affects the entire population simultaneously
Optimal Policy → characterizedas → Two-layer fixed point
confidence 90% · an optimal policy is identified as a two-layer fixed point to the argmax operator of the Iq-function
Integrated q-function → definedon → State distribution and policy
confidence 90% · introduce the integrated q-function (Iq-function) defined on the state distribution and the policy
Common Noise → induces → Nonlinear functional of policy
confidence 90% · the controlled common noise gives rise to an additional nonlinear functional of policy, rendering the policy iteration intricate
Two-layer fixed point → takesformof → Gaussian Distribution
confidence 90% · provide the explicit characterization of an optimal policy as a Gaussian distribution in the general linear-quadratic (LQ) setting
Policy Iteration → verifiedvia → Entropy-regularized optimization
confidence 85% · The policy improvement at each iteration is verified by relating to an entropy-regularized optimization problem over the space of policies
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common noise, coined as q-function by Jia and Zhou (2023) in the single agent's model. We first show that, under discretely sampled actions, the value function in the exploratory formulation converges to the one in the relaxed control formulation as the time grid refines. Leveraging the relaxed control formulation, we derive the exploratory Hamilton-Jacobi-Bellman (HJB) equation, in which the controlled common noise gives rise to an additional nonlinear functional of policy, rendering the policy iteration intricate. Under certain concavity condition, we establish the existence and uniqueness of the optimal one-step policy iteration via a first-order condition using the partial linear functional derivative with respect to policy. The policy improvement at each iteration is verified by relating to an entropy-regularized optimization problem over the space of policies. In the mean-field setting, we introduce the integrated q-function (Iq-function) defined on the state distribution and the policy, and it is shown that an optimal policy is identified as a two-layer fixed point to the argmax operator of the Iq-function. Finally, we provide the explicit characterization of an optimal policy as a Gaussian distribution in the general linear-quadratic (LQ) setting.
Tags
Links
- Source: https://arxiv.org/abs/2604.27372v1
- Canonical: https://arxiv.org/abs/2604.27372v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
132,786 characters extracted from source content.
Expand or collapse full text
Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations Zhenjie Ren Email: zhenjie.ren@univ-evry.fr, LaMME, Université Évry Paris-Saclay, Évry, France. Xiaoli Wei Email: tyswxl@gmail.com. Xiang Yu Email: xiang.yu@polyu.edu.hk, Department of Applied Mathematics, The Hong Kong Polytechnic University, Kowloon, Hong Kong. Xun Yu Zhou Email: xz2574@columbia.edu, Department of Industrial Engineering and Operations Research, Columbia University, New York, USA. (This version: April 30, 2026) Abstract This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common noise, coined as q-function by Jia and Zhou (2023) in the single agent’s model. We first show that, under discretely sampled actions, the value function in the exploratory formulation converges to the one in the relaxed control formulation as the time grid refines. Leveraging the relaxed control formulation, we derive the exploratory Hamilton-Jacobi-Bellman (HJB) equation, in which the controlled common noise gives rise to an additional nonlinear functional of policy, rendering the policy iteration intricate. Under certain concavity condition, we establish the existence and uniqueness of the optimal one-step policy iteration via a first-order condition using the partial linear functional derivative with respect to policy. The policy improvement at each iteration is verified by relating to an entropy-regularized optimization problem over the space of policies. In the mean-field setting, we introduce the integrated q-function (Iq-function) defined on the state distribution and the policy, and it is shown that an optimal policy is identified as a two-layer fixed point to the argmax operator of the Iq-function. Finally, we provide the explicit characterization of an optimal policy as a Gaussian distribution in the general linear-quadratic (LQ) setting. Keywords: Continuous-time reinforcement learning, mean-field control, common noise, policy improvement, integrated q-function, two-layer fixed point 1 Introduction Decision making for a large population system with interacting agents in a competitive or cooperative manner has wide applications across finance, systemic risk control, epidemic control, robot swarms, traffic management, among others. The main challenge in the large system is to understand the coupled influence of decision making on the behavior of all agents. To overcome this complexity and the curse of dimensionality in numerical implementations, the mean-field approximation of the population’s state, proposed independently by Lasry and Lions (2007) and Huang et al. (2006), has been extensively studied over the past decades. The mean-field formulation allows one to study the weak interaction between one representative agent and the population rather than the coupled interactions between agents. Mean-field game (MFG) and mean-field control (MFC) problems have been proposed and developed to model the competitive and cooperative interactions, respectively, see Carmona et al. (2013) for the discussion on their relationship and distinctions. See also Carmona and Delarue (2018a, b) for an extensive overview of existing studies in these two types of problems. In the present paper, we are interested in MFC problems where a social planner coordinates all agents to optimize the aggregated reward function that leads to the social optimum of the population. More importantly, we incorporate the Brownian common noise in the mean-field system that affects the entire population simultaneously. The study of common noise in large population system has recently attracted a lot of attention and spurred new advances in mean-field theories, which can effectively describe the exogenous random risk acting towards the whole system. Unlike the idiosyncratic noise that only affects the specific state dynamics of an individual agent, the presence of common noise leads to the mean-field interaction via the conditional population distribution given the common noise. Thereby, in MFC problems, we encounter the conditional population distribution as a measure-valued process that calls for Itô’s formula along conditional measure flows and the stochastic Fokker-Planck equation, giving rise to many new technical challenges, see Graber (2016), Pham and Wei (2017), Buckdahn et al. (2021), Djete et al. (2022), Motte and Pham (2022), Cheung et al. (2023), Zhou et al. (2024) for some recent developments in MFC problems with common noise. Conventional methods to solve the classical stochastic control and MFC problems typically assume the full knowledge or precise estimations of the underlying model. However, in reality, the agent or the social planner may only have little or no information of the environment. The limited information of unknown environment may cause huge errors or inefficiency in implementing the theoretical solutions. This motivated an upsurge of interests in studying the reinforcement learning (RL) algorithms from the classical single agent’s model to the large stochastic systems. Based on the trial-and-error interactions with the unknown environment, the decision maker can gradually learn to select best actions from the procedure of exploration and exploitation. Albeit the substantial success of RL algorithms in wide applications, the theoretical study of RL has been predominantly limited to discrete time models. However, many real-world applications, particularly in finance and engineering, evolve continuously through time. Recently, for continuous-time stochastic control problems by a single agent, Wang et al. (2020), Jia and Zhou (2022a, b, 2023) have laid the theoretical foundation for RL with entropy regularization with continuous state space and action space. To learn an optimal policy in a continuous-time setting requires the shift from the discrete Bellman equation in conventional RL studies to its continuous-time counterpart, the Hamilton-Jacobi-Bellman (HJB) equation. In particular, Wang et al. (2020) studied the optimal policy in an entropy-regularized exploratory RL framework for diffusion processes. Jia and Zhou (2022a) examined the policy evaluation problem by establishing a martingale condition of the value function. Jia and Zhou (2022b) studied the policy gradient algorithm by connecting it to the martingale approach in Jia and Zhou (2022a). Jia and Zhou (2023) developed the continuous-time q-learning theory by introducing the q-function as the first order time derivative of the advantage function and establishing a joint martingale characterization of the q-function and the value function. This continuous-time entropy-regularized RL method has been rapidly generalized in various context of single agent’s control problems. For example, Wang et al. (2023) proposed an actor-critic RL algorithm for optimal execution in continuous-time Almgren-Chriss model; Han et al. (2023) considered the Choquet entropy regularization for the RL exploration and explored the distribution of the optimal policy in the LQ framework; Dong (2024) examined the entropy-regularized RL method for optimal stopping problems; Dai et al. (2023a) studied the policy iteration algorithms to learn the time-consistent equilibrium policy for mean-variance portfolio optimization problems; Bo et al. (2025) generalized the q-learning algorithm for reflected diffusion dynamics and applied the algorithm in solving the optimal tracking portfolio problem; Dai et al. (2023b) considered the recursive entropy regularization and developed the RL algorithm to learn Merton’s optimal strategy in an incomplete market model; Huang et al. (2025) examined the continuous-time reinforcement learning approach for optimal switching over multiple regimes; Jia (2026) investigated the continuous-time risk-sensitive reinforcement learning using the quadratic variation penalty. Comparing with the large volume of studies in continuous-time RL for single agent’s control problems, the investigations of continuous-time RL for MFG and MFC are relatively underdeveloped. In the LQ-MFG, Guo et al. (2022) examined the theoretical justification that entropy regularization helps stabilizing and accelerating the convergence to the Nash equilibrium. Frikha et al. (2025) generalized the policy gradient algorithm in Jia and Zhou (2022b) to continuous time MFC problems and devised actor-critic algorithms based on a gradient expectation representation of the value function. Liang et al. (2024) similarly generalized the actor-critic policy gradient algorithms to continuous-time MFG problems together with fictitious play to update the population distribution. As a first attempt to generalize the continuous-time q-learning to MFC problems, Wei and Yu (2025) recently proposed the integrated q-function (Iq-function) and the essential q-function, which facilitate the design of q-learning algorithms for MFC without common noise from the social planner’s perspective. Wei et al. (2024) further discussed the proper Iq-function in decoupled form and proposed unified q-learning algorithms for both MFG and MFC problems without common noise from the representative agent’s perspective. Similar to Wei and Yu (2025) in the setting without common noise, it is assumed in the present paper that the social planner is responsible to learn a policy that maximizes the collective profit of the population. That is, the social planner assigns randomized policies to agents who interact with the environment based on their own current states and population’s state distribution. Based on agents’ interactions, the social planner collects the population’s distribution and agent’s individual rewards to generate the population’s aggregated reward and iterate the policies accordingly to learn the optimal policy. A typical example of such setting with learning social planner could be the centralized traffic management system while the population of agents could refer to the drivers and public transit users. The unknown environment is the city’s road network and the learning task of the social planner is to dynamically adjust the timing and phasing of all interconnected traffic lights; set speed limits; or implement dynamic congestion pricing. See also Appendix A for a brief description of the interactions of N players in the RL setting that motivates the RL formulation of the MFC problem. 1.1 Our contributions Discrete-time single-agent Q-learning: Q-function (Watkins and Dayan (1992)) Continuous-time single-agent q-learning: q-function (Jia and Zhou (2023)) Discrete-time mean-field Q-learning: IQ-function (Gu et al. (2021, 2023); Carmona et al (2023)) Continuous-time mean-field q-learning without common noise: Iq-function and essential q-function (Wei and Yu (2025)) Continuous-time mean-field q-learning with controlled common noise. Iq-function is complicated; essential q-function does not exist. How to relate the optimal policy to the Iq-function? Figure 1: Conceptual relationship to the literature The goal of the present paper is to investigate the correct form of Iq-function and develop some q-learning theories in the presence of controlled common noise, see the illustration of our research motivation in Figure 1. The design of q-learning algorithms as well as establishing some supporting theories, such as the martingale condition of the Iq-function and value function, are studied in our accompany paper Ren et al. (2026). The controlled common noise creates many new difficulties in the learning procedure. First, identifying an appropriate relaxed control formulation suitable for theoretical analysis in the presence of controlled common noise remains an open problem. In particular, we need to introduce additional Brownian motions in the relaxed control formulation (see (2.7)) so that the joint law of the state and the common noise ℒ(X,B) L(X,B) coincides with that of the exploratory formulation. This ensures the consistency of the conditional state distributions ℒ(X|B) L(X|B) in two formulations. Second, Itô’s formula on the flow of conditional probability measure yields a more sophisticated exploratory HJB equation with an additional nonlinear functional of policy, see the PDE (3.8). As a result, the one-step policy iteration by the first-order condition no longer admits an explicit form, in sharp contrast to the explicit Gibbs measure iteration rule in Jia and Zhou (2023); Wei and Yu (2025). This nonlinear functional of policy pins down how common noise may fundamentally complicate the iteration in learning. In fact, whether the policy improvement under this implicit iteration operator holds or not deserves some careful investigations. By some heuristic computations, we shall expect to see that the definition of the Iq-function in the present paper also involves this nonlinear functional of policy, see Definition 4.1. As the optimal policy is no longer explicitly related to the Hamiltonian in (3.9), how to learn the optimal policy via the learnt Iq-function is a key issue to address in devising some RL algorithms. We summarize the main contributions of the present paper as follows: (i) To cope with the additional nonlinear functional of policy in the exploratory HJB equation caused by controlled common noise, we establish a rigorous characterization of the unique optimal policy as a two-layer fixed point to an iteration operator using the notion of the partial linear functional derivative with respect to the policy, see Theorem 3.10 and Corollary 3.13. Moreover, the policy iteration operator is implicit as for a given policy, to exercise the optimal one-step policy iteration, one needs to solve a fixed point problem stemming from the first order condition equation in (3.14), which differs significantly from previous results in Jia and Zhou (2023); Wei and Yu (2025). Thanks to some proper concavity condition, we can also prove the policy improvement result for this implicit policy iteration operator, see Theorem 3.12. (i) Using the definition of IQ-function in the discrete-time MFC model and Itô’s lemma along the flow of conditional probability measures, we derive the proper definition of the Iq-function in Definition 4.1, which also carries the nonlinear functional of policy. It then follows from the previous first-order condition that the optimal policy is related to a two-layer fixed point of the argmax operator of the Iq-function. Equivalently, we show that it also corresponds to a two-layer fixed point of the operator in the Gibbs measure form using the partial linear functional derivative of the unregularized Iq-function with respect to the policy (see the expression (4.2) in Corollary 4.2). (i) In the general LQ setting, we pioneer the explicit characterization of the optimal policy as a two-layer fixed point of the implicit policy iteration operator, which is shown to be a Gaussian distribution; see Theorem 5.1 and the optimal policy in (5.4). This justifies the ad hoc choice of Gaussian policy in learning tasks even in the presence of common noise. The remainder of the paper is organized as follows. Section 2 reviews the classical MFC problem with common noise and introduces its relaxed control and exploratory formulations. Section 3 gives the characterization of an optimal policy as the two-layer fixed point to an implicit iteration operator and establishes the policy improvement result. Section 4 proposes the proper definition of the continuous-time Iq-function and characterizes an optimal policy as the two-layer fixed point to the argmax operator of the Iq-function. Section 5 investigates the general LQ-MFC problem with controlled common noise and establishes the explicit characterization of the two-layer fixed point as a Gaussian policy. Finally, Appendix A briefly discusses the continuous-time RL for the cooperative N-player game and Appendix B presents the heuristic derivation of the relaxed control formulation. Notations Given two Polish spaces (S,)(S, S) and (T,)(T, T), for any p>0p>0, we denote by p(S) P_p(S) the space of all probability measures with finite p-th moment on S and equip p(S) P_p(S) with the p-Wasserstein metric p W_p defined by p(μ,ν)=inf(∫S×S|x−y|pγ(dx,dy))1/p:γ∈p(S×S)has marginalsμandν. W_p(μ,ν)= \ ( _S× S|x-y|^pγ(dx,dy) )^1/p:γ∈ P_p(S× S)\; has marginals\;μ\; and\;ν \. For μ∈p(ℝd)μ∈ P_p(R^d), denote ‖μ‖p:=(∫ℝd|x|pμ(dx))1/p\|μ\|_p:= ( _R^d|x|^pμ(dx) )^1/p. Let ac(S) P_ac(S) be the space of probability measures on S that are absolutely continuous with respect to the Lebesgue measure. We denote by (T|S) P(T|S) (resp. ac(T|S) P_ac(T|S)) the space of all probability kernels from S to (T) P(T) (resp. ac(T) P_ac(T)). L2(Ω,ℱ,ℙ;ℝd)L^2( , F,P;R^d) denotes the space of all ℱ F-adapted ℝdR^d-valued square-integrable random variables on the probability space (Ω,ℱ,ℙ)( , F,P). For any measurable function g:S→ℝkg:S ^k, we denote ∫Sg(x)μ(dx) _Sg(x)μ(dx) by ⟨g,μ⟩ g,μ . For a functional F:2(S)→ℝF: P_2(S) , ∂μF(μ)(x) _μF(μ)(x), ∂x∂μF(μ)(x) _x _μF(μ)(x) and ∂μ2F(μ)(x,x′) _μ^2F(μ)(x,x ) stand for the L derivative in μ, the mixed second-order derivative with respect to μ and x, and the second-order derivative in measure μ, respectively (see Definition 5.22 in Carmona and Delarue (2018a)). We use (μ,Σ) N(μ, ) and ([p,q]) U([p,q]) for a Gaussian distribution with mean μ and covariance Σ and a uniform distribution on [p,q][p,q], respectively. 2 Problem Formulation 2.1 Classical MFC problem with common noise Let (Ω,ℱ,ℙ)( , F,P) be a complete probability space with a product structure (Ω0×Ω1,ℱ0⊗ℱ1,ℙ0⊗ℙ1)( ^0× ^1, F^0 F^1,P^0 ^1), where (Ω1,ℱ1,ℙ1)( ^1, F^1,P^1) supports a m-dimensional Brownian motion W=(Ws)s∈[0,T]W=(W_s)_s∈[0,T] and (Ω0,ℱ0,ℙ0)( ^0, F^0,P^0) supports a n-dimensional Brownian motion B=(Bs)s∈[0,T]B=(B_s)_s∈[0,T] with B serving as common noise. We consider two filtrations W,B=(ℱtW,B)0≤t≤TF^W,B=( F_t^W,B)_0≤ t≤ T and =(t)0≤t≤TG=( G_t)_0≤ t≤ T defined by ℱtW,B:=σ(Ws,Bs:s∈[0,t]) F_t^W,B:=σ(W_s,B_s:s∈[0,t]) and t:=σ(Bs:s∈[0,t]) G_t:=σ(B_s:s∈[0,t]), respectively. It is assumed that there exists a sub-algebra ℋ H of ℱ1 F^1 such that ℋ H is independent of W,BF^W,B and it is “rich enough” in the sense that for any μ∈2(ℝd)μ∈ P_2(R^d), there exists an ℋ H-measurable random variable ξ on (Ω1,ℙ1)( ^1,P^1) such that ℙ1∘ξ−1=μP^1 ξ^-1=μ. Let =(ℱs)s≥0F=( F_s)_s≥ 0 be the filtration ℱs=ℱsW,B∨ℋ F_s= F_s^W,B H. We consider a MFC problem by the social planner, for which the representative agent’s state process Xss≥t\X_s\_s≥ t, taking values in ℝdR^d, is described by the controlled conditional McKean-Vlasov SDE that dXs dX_s =b(s,Xs,μs,as)ds+σ(s,Xs,μs,as)dWs+σo(s,Xs,μs,as)dBs, =b(s,X_s, _s,a_s)ds+σ(s,X_s, _s,a_s)dW_s+ _o(s,X_s, _s,a_s)dB_s, (2.1) where ξ∈L2(Ω1,ℋ,ℙ1;ℝd)ξ∈ L^2( ^1, H,P^1;R^d) such that ℒ(ξ|t)=μ∈2(ℝd) L(ξ| G_t)=μ∈ P_2(R^d), and b, σ and σo _o are measurable functions valued in ℝdR^d, ℝd×mR^d× m and ℝd×nR^d× n. Recall that, W stands for the idiosyncratic noise for each representative agent and B plays the role of common noise affecting the entire population. The conditional law μs:=ℒ(Xs|s) _s:= L(X_s| G_s) denotes the conditional regular probability distribution of XsX_s given s G_s that μs(ω0)=ℙ1∘Xs−1(ω0,⋅) _s(ω^0)=P^1 X_s^-1(ω^0,·) for every ω0∈Ω0ω^0∈ ^0. The goal of the social planner in the MFC problem is to find an optimal F-progressively measurable control ast≤s≤T\a_s\_t≤ s≤ T valued in the space A, a closed subset of ℝmR^m, which maximizes the expected discounted total reward that [∫tTe−β(s−t)r(s,Xs,μs,as)s+e−β(T−t)g(XT,μT)|Xt=ξ]. [ _t^Te^-β(s-t)r(s,X_s, _s,a_s)ds+e^-β(T-t)g(X_T, _T) |X_t=ξ ]. (2.2) To ensure the wellposedness of (2.1)-(2.2), we impose the following assumptions. Assumption 2.1. The following conditions for the state dynamics and reward functions hold: (i) b, σ, σo _o, r are jointly continuous in (t,x,μ,a)∈[0,T]×ℝd×2(ℝd)×(t,x,μ,a)∈[0,T]×R^d× P_2(R^d)× A, and g is jointly continuous in (x,μ)∈ℝd×2(ℝd)(x,μ) ^d× P_2(R^d); (i) There exists a constant C>0C>0 such that for f∈b,σ,σof∈\b,σ, _o\, and all t,t′∈[0,T]t,t ∈[0,T], x,x′∈ℝdx,x ^d, μ,μ′∈2(ℝd)μ,μ ∈ P_2(R^d), a∈a∈ A, it holds that |f(t,x,μ,a)−f(t′,x′,μ′,a)|≤C(|t−t′|+|x−x′|+2(μ,μ′)), |f(t,x,μ,a)-f(t ,x ,μ ,a)|≤ C (|t-t |+|x-x |+ W_2(μ,μ ) ), |r(t,x,μ,a)−r(t′,x′,μ,a)|≤C(1+|x|+|x′|+‖μ‖2+‖μ′‖2)(|t−t′|+|x−x′|+2(μ,μ′)). |r(t,x,μ,a)-r(t ,x ,μ,a)|≤ C (1+|x|+|x |+\|μ\|_2+\|μ \|_2 ) (|t-t |+|x-x |+ W_2(μ,μ ) ). (i) There exists some constant C>0C>0 such that for each (t,x,μ,a)∈[0,T]×ℝd×2(ℝd)×(t,x,μ,a)∈[0,T]×R^d× P_2(R^d)× A, it holds that |b(t,x,μ,a)| |b(t,x,μ,a)| ≤C(1+|x|+‖μ‖2+|a|), ≤ C (1+|x|+\|μ\|_2+|a| ), |(σσ⊺+σoσo⊺)(t,x,μ,a)| |(σ + _o _o )(t,x,μ,a) | ≤C(1+|x|2+‖μ‖22+|a|2), ≤ C (1+|x|^2+\|μ\|_2^2+|a|^2 ), |r(t,x,μ,a)|+|g(x,μ)| |r(t,x,μ,a)|+|g(x,μ)| ≤C(1+|x|2+‖μ‖22+|a|2). ≤ C (1+|x|^2+\|μ\|_2^2+|a|^2 ). 2.2 Relaxed control formulation of MFC with common noise It is assumed that the mean-field model is unknown, i.e., we do not have the exact information of the model coefficients b, σ and σo _o in state dynamics (2.1) and the reward function r in (2.2). To determine the optimal control in an unknown model, we choose to apply the RL approach based on the principle of trial-and-error recovery. To capture the exploration step in RL, we randomize the action and consider its distribution as a relaxed control. Therefore, it is necessary to extend the original filtered probability space (Ω,ℱ,,ℙ)( , F,F,P) for the purpose of the action randomization and consider another atomless probability space (Ω2,ℱ2,ℙ2)( ^2, F^2,P^2) that supports an ℱ2 F^2-measurable random variables U0U_0 with uniform distribution on [0,1][0,1]. By standard separation of the decimals of U0U_0 (c.f. Lemma 2.21 in Kallenberg (2002)), there exists an i.i.d. sequence of ℱ2F^2-adapted uniform random variables (Ui)i∈ℕ(U_i)_i , independent of ξ, W and B. Denote (Ωe,ℱe,e,ℙe):=(Ω×Ω2,ℱ⊗ℱ2,ℙ⊗ℙ2)( ^e, F^e,F^e,P^e):=( × ^2, F F^2,P ^2) where e=(ℱte)0≤t≤TF^e=( F^e_t)_0≤ t≤ T and ℱte=ℱt∨σ(Ui,ti≤t) F_t^e= F_t σ(U_i,t_i≤ t). Accordingly, for an element ωe∈Ωeω^e∈ ^e, we write it as ωe=(ω,ω2)∈Ω×Ω2ω^e=(ω,ω^2)∈ × ^2, and we extend canonically W and B on Ω by setting W(ωe):=W(ω)W(ω^e):=W(ω) and B(ωe):=B(ω)B(ω^e):=B(ω). Any random variable on Ω can be extended similarly to that on Ωe ^e. eE^e stands for the expectation under ℙeP^e. We first introduce the relaxed control formulation for theoretical analysis. Let Π stand for the set of admissible policies satisfying the following definition. Definition 2.2. A policy π is called admissible if (i) (⋅|t,x,μ)∈ac() π(·|t,x,μ)∈ P_ac( A), supp(⋅|t,x,μ)= supp π(·|t,x,μ)= A for every (t,x,μ)∈[0,T]×ℝd×2(ℝd)(t,x,μ)∈[0,T]×R^d× P_2(R^d), and π is jointly measurable with respect to (t,x,μ)∈[0,T]×ℝd×2(ℝd)(t,x,μ)∈[0,T]×R^d× P_2(R^d). (i) There exits a constant C>0C>0 independent of (t,a)(t,a) such that for any ∈Π π∈ and any x,x′∈ℝdx,x ^d and μ,μ′∈2(ℝd)μ,μ ∈ P_2(R^d) ∫|(a|t,x,μ)−(a|t,x′,μ′)|da≤C(|x−x′|+2(μ,μ′)). _ A| π(a|t,x,μ)- π(a|t,x ,μ )|da≤ C (|x-x |+ W_2(μ,μ ) ). (i) There exists some C>0C>0 and δ>0δ>0 independent of (t,a)(t,a) such that for every ∈Π π∈ ∫|a|2+δ(a|t,x,μ)a≤C(1+|x|2+‖μ‖22). _ A|a|^2+δ π(a|t,x,μ)da≤ C(1+|x|^2+\|μ\|_2^2). (iv) There exists a constant C>0C>0 such that for any ∈Π π∈ |E(t,x,μ)|≤ |E_ π(t,x,μ) |≤ C(1+|x|2+‖μ‖22) C (1+|x|^2+\|μ\|_2^2 ) |E(t,x,μ)−E(t′,x′,μ′)|≤ |E_ π(t,x,μ)-E_ π(t ,x ,μ ) |≤ C(|t−t′|+|x−x′|+2(μ,μ′)), C (|t-t |+|x-x |+ W_2(μ,μ ) ), where the Shannon entropy E_ π is defined by E(t,x,μ)=−∫log(a|t,x,μ)(a|t,x,μ)a, E_ π(t,x,μ)=- _ A π(a|t,x,μ) π(a|t,x,μ)da, (2.3) For f∈b,σ,σo,rf∈\b,σ, _o,r\, we denote by f_ π the mean of f with respect to ∈Π π∈ that f(t,x,μ) f_ π(t,x,μ) :=∫f(t,x,μ,a)(a|t,x,μ)a, := _ Af(t,x,μ,a) π(a|t,x,μ)da, (2.4) and we denote by cov(f) cov_ π(f) and std(f) std_ π(f) the covariance and standard deviation of f with respect to ∈Π π∈ that cov(f)(t,x,μ):= cov_ π(f)(t,x,μ):= ∫ff⊺(t,x,μ,a)(a|t,x,μ)a−ff⊺(t,x,μ), _ Af (t,x,μ,a) π(a|t,x,μ)da-f_ πf_ π (t,x,μ), (2.5) std(f):= std_ π(f):= cov(f)1/2. cov_ π(f)^1/2. (2.6) To save notations, we only present the relaxed control formulation for d=1d=1 here (see (B.4) in Appendix B for the case d>1d>1). dXs dX_s π =b(s,Xs,μs)ds+σ(s,Xs,μs)dWs+σo,(s,Xs,μs)dBt =b_ π(s,X_s π, _s π)ds+ _ π(s,X_s π, _s π)dW_s+ _o, π(s,X_s π, _s π)dB_t +std(σ)(s,Xs,μs)dW¯s+std(σo)(s,Xs,μs)dB¯s, \;\;\;+ std_ π(σ)(s,X_s π, _s π)d W_s+ std_ π( _o)(s,X_s π, _s π)d B_s, (2.7) where μs=ℒ(Xs|s) _s π= L(X_s π| G_s), and W¯ W and B¯ B are extra one-dimensional Brownian motions independent of W and B. Under Assumption 2.1 and Definition 2.2, all coefficients in (2.7) are Lipschitz continuous in arguments x and μ. Therefore the SDE (2.7) admits a unique strong solution, see Theorem 5.1.1 in Stroock and Varadhan (1997) or Appendix A in Djete et al. (2022). Remark 2.3. Due to the controlled common noise, the above relaxed control formulation is different from that in Jia and Zhou (2023); Wei and Yu (2025). We need to introduce auxiliary Brownian motions arising from the action randomiziation in the relaxed formulation; see more details in Appendix B. To encourage the exploration in the continuous-time framework, we adopt the Shannon differential entropy as in Wang et al. (2020) and we consider the value function in the relaxed control formulation by J~(t,μ;) J(t,μ; π) =e[∫tTe−β(s−t)(r^(s,μs)+γℰ(s,μs,))s+e−β(T−t)g^(μT)], =E^e [ _t^Te^-β(s-t) ( r_ π(s, _s π)+ (s, _s π, π) )ds+e^-β(T-t) g( _T π) ], (2.8) where we denote ℰ(t,μ,): (t,μ, π): =∫ℝdE(t,x,μ)μ(dx),g^(t,μ):=∫ℝdg(x,μ)μ(dx), = _R^dE_ π(t,x,μ)μ(dx),\; g(t,μ):= _R^dg(x,μ)μ(dx), r^(t,μ,): r(t,μ, π): =∫ℝd×r(t,x,μ,a)(a|t,x,μ)aμ(dx). = _R^d× Ar(t,x,μ,a) π(a|t,x,μ)daμ(dx). The optimal value function is given by J~∗(t,μ)=sup∈ΠJ~(t,μ;). J^*(t,μ)= _ π∈ J(t,μ; π). (2.9) 2.3 Exploratory formulation under discretely sampled actions Although the relaxed control formulation is convenient for the theoretical analysis such as deriving the HJB equation, it cannnot be directly observed in practice. As a result, to develop some implementable algorithms, we need to consider the continuous-time exploratory formulation with sampling. However, as discussed in recent studies among Szpruch et al. (2024); Bender and Thuan (2024); Jia et al. (2025); Carmona and Laurière (2025), continuously sampling procedure requires continuum independent draws from a distribution and may cause some measure-theoretical issues. As a remedy, we next proceed to consider the discretely sampled action processes as in Jia et al. (2025) in our mean-field setting. Fix ∈Π π∈ and an initial pair (t,ξ)∈[0,T]×L2(Ω1,ℋ,ℙ1;ℝd)(t,ξ)∈[0,T]× L^2( ^1, H,P^1;R^d), we then consider the exploratory discretely sampled state process for all 0≤i≤n−10≤ i≤ n-1 and s∈[si,si+1]s∈[s_i,s_i+1] that Xs, X_s D, π =Xsi,+∫sisb(u,Xu,,μu,,asi)u+∫sisσ(u,Xu,,μu,,asi)Wu =X_s_i D, π+ _s_i^sb(u,X_u D, π, _u D, π,a_s_i π)du+ _s_i^sσ(u,X_u D, π, _u D, π,a_s_i π)dW_u (2.10) +∫sisσo(u,Xu,,μu,,asi)Bu, \;\;\;\;\;+ _s_i^s _o(u,X_u D, π, _u D, π,a_s_i π)dB_u, where μs,=ℒ(Xs,|s) _s D, π= L(X_s D, π| G_s), and asi=ϕ(si,Xsi,,μsi,,Ui)a_s_i π= _ π(s_i,X_s_i D, π, _s_i D, π,U_i) ∼(⋅|si,Xsi,,μsi,) π(·|s_i,X_s_i D, π, _s_i D, π) for some measurable function ϕ:[0,T]×ℝn×2(ℝn)×[0,1]→ _ π:[0,T]×R^n× P_2(R^n)×[0,1]→ A stands for the sampled actions. Under Assumption 2.1, the SDE (2.10) is well-posed. We may also rewrite (2.10) as dXs, dX_s D, π =b(s,Xs,,μs,,aδ(s),)ds+σ(s,Xs,,μs,,aδ(s))dWs =b(s,X_s D, π, _s D, π,a_δ(s) D, π)ds+σ(s,X_s D, π, _s D, π,a_δ(s) π)dW_s (2.11) +σo(s,Xs,,μs,,aδ(s),)dBs, \;\;\;\;\;+ _o(s,X_s D, π, _s D, π,a_δ(s) D, π)dB_s, where δ(s):=siδ(s):=s_i for s∈[si,si+1)s∈[s_i,s_i+1). The procedure of RL for MFC is then described as follows. On an arbitrary time gird =t=s0<s1<…<sn=T D=\t=s_0<s_1<…<s_n=T\, the representative agent at the current state Xs,X_s D, π observes the current population’s conditional state distribution μs, _s D, π and takes the sequence of actions (asi)si∈(a_s_i)_s_i∈ D only at the time grid in D according to a policy ∈Π π∈ assigned by the social planner. The representative agent will receive a stream of running individual rewards and her state evolves according to the sampled SDE in (2.10). Based on the representative agent’s interactions with the unknown environment, the social planner coordinates the population by assigning policies to the representative agent and collecting the population’s conditional state distribution and the aggregated reward. The value function under the time grid D and the policy π is given by J(t,ξ;) J D(t,ξ; π) =e[∫tTe−β(s−t)(r(s,Xs,,μs,,aδ(s),)+γE(δ(s),Xδ(s),,μδ(s),))ds =E^e [ _t^Te^-β(s-t) (r(s,X_s D, π, _s D, π,a_δ(s) D, π)+γ E_ π(δ(s),X_δ(s) D, π, _δ(s) D, π) )ds +e−β(T−t)g(XT,,μT,)|Xt,=ξ]. \;\;\;\;\;\;+e^-β(T-t)g(X_T D, π, _T D, π) |X_t D, π=ξ ]. (2.12) In view of the arbitrariness of D, we define the value function by taking the limit over all time grids J(t,ξ;)=lim||→0J(t,ξ;)J(t,ξ; π)= _| D|→ 0J D(t,ξ; π), where ||:=max0≤i≤n−1|si+1−si|| D|:= _0≤ i≤ n-1|s_i+1-s_i|. The goal of the social planner is to maximize J(t,ξ;)J(t,ξ; π) that J∗(t,ξ)=sup∈ΠJ(t,ξ;). J^*(t,ξ)= _ π∈ J(t,ξ; π). (2.13) 3 Policy Iteration and Policy Improvement 3.1 Relationship between two formulations In this subsection, we first investigate the connection between the relaxed control formulation and the limit of the exploratory formulation using discretely sampled actions. We first recall the following definition that is frequently used in the rest of the paper. Definition 3.1. We say that V:[0,T]×2(ℝd)→ℝV:[0,T]× P_2(R^d) belongs to 1,2([0,T]×2(ℝd)) C^1,2([0,T]× P_2(R^d)) if • ∂V∂t(t,μ) ∂ V∂ t(t,μ) exists and is jointly continuous in t,μt,μ; • ∂μV(t,μ)(x) _μV(t,μ)(x), ∂x∂μV(t,μ)(x) _x _μV(t,μ)(x) and ∂μ2V(t,μ)(x,x′) _μ^2V(t,μ)(x,x ) exist for any (t,μ,x,x′)∈[0,T]×2(ℝd)×ℝd×ℝd(t,μ,x,x )∈[0,T]× P_2(R^d)×R^d×R^d; • ∂μV(t,μ)(x) _μV(t,μ)(x), ∂x∂μV(t,μ)(x) _x _μV(t,μ)(x) and ∂μ2V(t,μ)(x,x′) _μ^2V(t,μ)(x,x ) are Lipschitz continuous with respect to all entries and satisfy that, for any (t,μ,x,x′)∈[0,T]×2(ℝd)×ℝd×ℝd(t,μ,x,x )∈[0,T]× P_2(R^d)×R^d×R^d, |∂μV(t,μ)(x)|≤C(1+|x|+‖μ‖2),|∂x∂μV(t,μ)(x)|+|∂μ2V(t,μ)(x,x′)|≤C, | _μV(t,μ)(x)|≤ C(1+|x|+\|μ\|_2),\;| _x _μV(t,μ)(x)|+| _μ^2V(t,μ)(x,x )|≤ C, for some constant C>0C>0. Assumption 3.2. b,σb,σ, and σ0 _0 are sufficiently regular such that for f=g^,r^f= g, r_ π or ℰ(⋅,) E(·, π), the PDE ∂V∂t(t,μ)+V(t,μ)=0 ∂ V∂ t(t,μ)+ T πV(t,μ)=0, t∈[0,t′]t∈[0,t ], with the terminal condition V(t′,μ)=f(t′,μ)V(t ,μ)=f(t ,μ) has a classical solution Vf∈C1,2([0,t′]×2(ℝd))V_f∈ C^1,2([0,t ]× P_2(R^d)) satisfying |∂Vf∂t(t,μ)−∂Vf∂t(t,μ′)|+|∂μVf(t,μ)(x)−∂μVf(t,μ′)(x′)| | ∂ V_f∂ t(t,μ)- ∂ V_f∂ t(t,μ )|+| _μV_f(t,μ)(x)- _μV_f(t,μ )(x )| +|∂x∂μVf(t,μ)(x)−∂x∂μVf(t,μ′)(x′)|≤CΠ(|x−x′|+2(μ,μ′)), \;\;\;+| _x _μV_f(t,μ)(x)- _x _μV_f(t,μ )(x )|≤ C_ (|x-x |+ W_2(μ,μ ) ), |∂μ2Vf(t,μ)(x,y)−∂μ2Vf(t,μ′)(x′,y′)|≤CΠ(|x−x′|+|y−y′|+2(μ,μ′)), | _μ^2V_f(t,μ)(x,y)- _μ^2V_f(t,μ )(x ,y )|≤ C_ (|x-x |+|y-y |+ W_2(μ,μ ) ), where the operator T π is defined by V(t,μ)= T πV(t,μ)= ∫ℝd(b(t,x,μ)⊺∂μV(t,μ)(x)+12Tr(σσ⊺+σo,σo,⊺)(t,x,μ)∂x∂μV(t,μ)(x))μ(dx) _R^d (b_ π(t,x,μ) _μV(t,μ)(x)+ 12 Tr ( _ π _ π + _o, π _o, π )(t,x,μ) _x _μV(t,μ)(x) )μ(dx) +12∫ℝd×ℝdTr(σo,(t,x,μ)σo,(t,x′,μ)⊺∂μ2V(t,μ)(x,x′))μ(dx)⊗μ(dx′). + 12 _R^d×R^d Tr ( _o, π(t,x,μ) _o, π(t,x ,μ) _μ^2V(t,μ)(x,x ) )μ(dx) μ(dx ). Proposition 3.3. Under Assumptions 2.1 and 3.2, we have |J(t,μ;)−J~(t,μ;)|≤C||1/2|J D(t,μ; π)- J(t,μ; π)|≤ C| D|^1/2, where the constant C depends on b,σ,σo,rb,σ, _o,r, and Π . Consequently, J(t,μ;)=J~(t,μ;)J(t,μ; π)= J(t,μ; π). The proof of Proposition 3.3 relies on the following lemma. Lemma 3.4. Let Assumptions 2.1, and 3.2 hold. Then for f in Assumption 3.2, there exists a constant C (depending only on T, t, γ, b, σ, σ0 _0, Π and f) such that for all grids D, sups∈[t,T]|e[f(μs,)−f(μs)]|≤C||1/2. _s∈[t,T] |E^e[f( _s D, π)-f( _s π)] |≤ C| D|^1/2. (3.1) Proof of Lemma 3.4. Without loss of generality, let us assume that t′=sit =s_i. As π is fixed throughout the proof, we do not write the superscript π in Xs,X_s D, π, μs, _s D, π and μs _s π. In view that VfV_f is the solution of ∂V∂t(t,μ)+V(t,μ)=0 ∂ V∂ t(t,μ)+ T πV(t,μ)=0, t∈[0,si]t∈[0,s_i], V(si,μ)=f(μ)V(s_i,μ)=f(μ), we have e[f(μsi)]=Vf(t,μ)E^e[f( _s_i)]=V_f(t,μ) by the Feynman-Kac’s formula. It follows that e[f(μsi)−f(μsi)]=e[Vf(si,μsi)−Vf(t,μ)]=∑j=0i−1e[Vf(sj+1,μsj+1)−Vf(sj,μsj)]=∑j=0i−1ej. ^e[f( _s_i D)-f( _s_i)]=E^e[V_f(s_i, _s_i D)-V_f(t,μ)]= _j=0^i-1E^e[V_f(s_j+1, _s_j+1 D)-V_f(s_j, _s_j D)]= _j=0^i-1e_j. It thus suffices to estimate the term eje_j. Applying Itô’s lemma to Vf(s,μs)V_f(s, _s D) between sjs_j and sj+1s_j+1 and taking the expectation on both sides, we get that ej= e_j= e[∫sjsj+1(∂Vf∂t(s,μs)+b(s,Xs,μs,asj)⊺∂μVf(s,μs)(Xs) ^e [ _s_j^s_j+1 ( ∂ V_f∂ t(s, _s D)+b(s,X_s D, _s D,a_s_j) _μV_f(s, _s D)(X_s D) +12Tr(σ+σoσo⊺)(s,Xs,μs,asj)∂x∂μVf(s,μs)(Xs)))ds] + 12 Tr (σ+ _o _o )(s,X_s D, _s D,a_s_j) _x _μV_f(s, _s D)(X_s D)) )ds ] +12e¯e[∫sjsj+1Tr(σ0(s,Xs,μs,asj)σ0⊺(s,X¯s,μs,a¯sj)∂μ2Vf(s,μs)(Xs,X¯s))s]. + 12E^e E^e [ _s_j^s_j+1 Tr ( _0(s,X_s D, _s D,a_s_j) _0 (s, X_s D, _s D, a_s_j) _μ^2V_f(s, _s D)(X_s D, X_s D) )ds ]. On the other hand, (∂Vf∂t(sj,μsj)+Vf(sj,μsj))⋅(sj+1−sj)=0 ( ∂ V_f∂ t(s_j, _s_j D)+ T πV_f(s_j, _s_j D) )·(s_j+1-s_j)=0, for 0≤j≤i−10≤ j≤ i-1. Subtracting this term on both sides of the above equation, we get that ej= e_j= e[∫sjsj+1(∂Vf∂t(s,μs)−∂Vf∂t(sj,μsj)+b(s,Xs,μs,asj)⊺∂μVf(s,μs)(Xs) ^e [ _s_j^s_j+1 ( ∂ V_f∂ t(s, _s D)- ∂ V_f∂ t(s_j, _s_j D)+b(s,X_s D, _s D,a_s_j) _μV_f(s, _s D)(X_s D) −b(sj,Xsj,μsj,asj)⊺∂μVf(sj,μsj)(Xsj)+12Tr(σσ+σoσo⊺)(s,Xs,μs,asj)∂x∂μVf(s,μs)(Xs) -b(s_j,X_s_j D, _s_j D,a_s_j) _μV_f(s_j, _s_j D)(X_s_j D)+ 12 Tr (σ+ _o _o )(s,X_s D, _s D,a_s_j) _x _μV_f(s, _s D)(X_s D) −12Tr(σ+σoσo⊺)(sj,Xsj,μsj,asj)∂x∂μVf(sj,μsj)(Xsj))ds] - 12 Tr (σ+ _o _o )(s_j,X_s_j D, _s_j D,a_s_j) _x _μV_f(s_j, _s_j D)(X_s_j D) )ds ] +12e¯e[∫sjsj+1(Tr(σo(s,Xs,μs,asj)σo⊺(s,X¯s,μs,a¯sj)∂μ2Vf(s,μs)(Xs,X¯s)) + 12E^e E^e [ _s_j^s_j+1 ( Tr ( _o(s,X_s D, _s D,a_s_j) _o (s, X_s D, _s D, a_s_j) _μ^2V_f(s, _s D)(X_s D, X_s D) ) −Tr(σo(sj,Xsj,μsj,asj)σo⊺(sj,X¯sj,μsj,a¯sj)∂μ2Vf(sj,μsj)(Xsj,X¯sj)))ds]. - Tr ( _o(s_j,X_s_j D, _s_j D,a_s_j) _o (s_j, X_s_j D, _s_j D, a_s_j) _μ^2V_f(s_j, _s_j D)(X_s_j D, X_s_j D) ) )ds ]. Thanks to the continuity of XsX_s D, we have e[22(μs,μs′)]≤e[|Xs−Xs′|2]≤C|s−s′|E^e [ W_2^2( _s D, _s D) ] ^e [|X_s D-X_s D|^2 ]≤ C|s-s |. Combining it with the Lipschitz continuity on b, σ, σ0 _0, ∂Vf∂t ∂ V_f∂ t, ∂μVf _μV_f, ∂x∂μVf _x _μV_f and ∂μ2Vf _μ^2V_f, we conclude that |ej|≤C(sj+1−sj)||1/2|e_j|≤ C(s_j+1-s_j)| D|^1/2. We then get the desired result that ∑j=0i−1|ej|≤C(T−t)||1/2 _j=0^i-1|e_j|≤ C(T-t)| D|^1/2. ∎ We next proceed the proof of Proposition 3.3. Proof of Proposition 3.3. Note that J(t,ξ;)−J~(t,μ;) J D(t,ξ; π)- J(t,μ; π) =e[e−β(T−t)(g^(μT)−g^(μT))]+e[∑i=0n−1∫sisi+1e−β(s−t)(r(s,Xs,μs,asi)−r^(s,μs))s] =E^e [e^-β(T-t) ( g( _T D)- g( _T) ) ]+E^e [ _i=0^n-1 _s_i^s_i+1e^-β(s-t) (r(s,X_s D, _s D,a_s_i)- r_ π(s, _s) )ds ] +γ∑i=0n−1e[∫sisi+1e−β(s−t)(ℰ(si,μsi,)−ℰ(s,μs,))ds]=:I+∑i=0n−1IIi+γ∑i=0n−1IIIi. \;+γ _i=0^n-1E^e [ _s_i^s_i+1e^-β(s-t) (E(s_i, _s_i D, π)-E(s, _s, π) )ds ]=:I+ _i=0^n-1I^i+γ _i=0^n-1I^i. By Lemma 3.4 with f=g^f= g, we have that |I|≤C||1/2. |I|≤ C| D|^1/2. (3.2) For the term ∑i=0n−1IIi _i=0^n-1I_i, it holds that IIi= I^i= e[∫sisi+1e−β(s−t)(r(s,Xs,μs,asi)−r(si,Xsi,μsi,asi))s] ^e [ _s_i^s_i+1e^-β(s-t) (r(s,X_s D, _s D,a_s_i)-r(s_i,X_s_i D, _s_i D,a_s_i) )ds ] +e[∫sisi+1e−β(s−t)(r^(si,μsi)−r^(si,μs))s] +E^e [ _s_i^s_i+1e^-β(s-t) ( r_ π(s_i, _s_i D)- r_ π(s_i, _s) )ds ] (3.3) +e[∫sisi+1e−β(s−t)(r^(si,μsi)−r^(s,μs))ds]=:I1i+I2i+I3i, +E^e [ _s_i^s_i+1e^-β(s-t) ( r_ π(s_i, _s_i)- r_ π(s, _s) )ds ]=:I_1^i+I_2^i+I_3^i, where in the term II2iII_2^i, we have e[r(si,Xsi,μsi,asi)]=e[r^(si,μsi)]E^e[r(s_i,X_s_i D, _s_i D,a_s_i)]=E^e[ r_ π(s_i, _s_i D)] by the tower property. For the term II1iII_1^i, we obtain by Assumption 2.1 (i) that |II1i|≤ |I_1^i|≤ e[∫sisi+1e−β(s−t)(|Xs|+|Xsi|+∥μs∥2+∥μsi∥2)(|s−si| ^e [ _s_i^s_i+1e^-β(s-t) (|X_s D|+|X_s_i D|+\| _s D\|_2+\| _s_i D\|_2 ) (|s-s_i| (3.4) +|Xs−Xsi|+2(μs,μsi))ds]≤C(si+1−si)⋅||1/2. +|X_s D-X_s_i D|+ W_2( _s D, _s_i D) )ds ]≤ C(s_i+1-s_i)·| D|^1/2. Similarly, by Lipschitz continuity of μs _s in time and local Lipschitz continuity on r r_ π, it holds that |II3i|≤C(si+1−si)⋅||1/2. |I_3^i|≤ C(s_i+1-s_i)·| D|^1/2. (3.5) It follows from Lemma 3.4 with f=r^f= r_ π that |II2i|≤C(si+1−si)⋅||1/2. |I_2^i|≤ C(s_i+1-s_i)·| D|^1/2. (3.6) Finally, for the term IIIiIII^i, we can deduce that |IIIi|≤ |I^i|≤ e[∫sisi+1e−β(s−t)(ℰ(si,μsi,)−ℰ(si,μsi,)+ℰ(si,μsi,)−ℰ(s,μs,))s] ^e [ _s_i^s_i+1e^-β(s-t) (E(s_i, _s_i D, π)-E(s_i, _s_i, π)+E(s_i, _s_i, π)-E(s, _s, π) )ds ] ≤C(si+1−si)⋅||1/2, ≤ C(s_i+1-s_i)·| D|^1/2, (3.7) where in the last inequality, we have used Lemma 3.4, the local Lipschitz continuity of ℰ(⋅,) E(·, π) and the continuity of μs _s in time s. Combining (3.2), (3.4), (3.5), (3.6), and (3.1), we conclude the result. ∎ The equivalence between value functions in Proposition 3.3 allows us to use the relaxed control formulation for the derivation of HJB equation and some theoretical analysis afterwards. 3.2 First-order condition and policy improvement We first give a PDE characterization of J~(⋅,⋅;) J(·,·; π) in (2.8) based on the Feynman-Kac formula. It is assumed that J~(⋅,⋅;)∈1,2([0,T]×2(ℝd)) J(·,·; π)∈ C^1,2([0,T]× P_2(R^d)) in Lemma 3.5 to avoid heavy technicalities. Interested readers can generalize Malliavin calculus arguments in Buckdahn et al. (2021); Chassagneux et al. (2022); Crisan and McMurray (2018) to the common noise setting to investigate sufficient conditions for the regularity of J~(⋅,⋅;) J(·,·; π). Lemma 3.5. Assume that the function J~(⋅,⋅;) J(·,·; π) belongs to 1,2([0,T]×2(ℝd)) C^1,2([0,T]× P_2(R^d)). Then it satisfies the following PDE ∂J~∂t(t,μ;)+∫ℝd×H(t,x,μ,a,∂μJ~(t,μ;)(x),∂x∂μJ~(t,μ;)(x))(a|t,x,μ)aμ(dx) ∂ J∂ t(t,μ; π)+ _R^d× AH (t,x,μ,a, _μ J(t,μ; π)(x), _x _μ J(t,μ; π)(x) ) π(a|t,x,μ)daμ(dx) +12∫ℝd×ℝdTr(σo,(t,x,μ)σo,(t,x′,μ)⊺∂μ2J~(t,μ;)(x,x′))μ(dx)⊗μ(dx′) + 12 _R^d×R^d Tr ( _o, π(t,x,μ) _o, π(t,x ,μ) _μ^2 J(t,μ; π)(x,x ) )μ(dx) μ(dx ) (3.8) −βJ~(t,μ;)+γℰ(t,μ,)=0, -β J(t,μ; π)+ (t,μ, π)=0, where the Hamiltonian operator H is defined by H(t,x,μ,a,p,q) H(t,x,μ,a,p,q) :=b(t,x,μ,a)⊺p+12Tr((σσ⊺+σoσo⊺)(t,x,μ,a)q)+r(t,x,μ,a). :=b(t,x,μ,a) p+ 12 Tr ( (σ + _o _o )(t,x,μ,a)q )+r(t,x,μ,a). (3.9) Proof. From the flow property of μst−h,μ=μst,μt−h,μ _s^t-h,μ= _s^t, _t^t-h,μ for any 0≤h≤t0≤ h≤ t J~(t−h,μ;) J(t-h,μ; π) =e−βhe[∫t−hte−β(s−t)(r(s,μs)+γℰ(s,μs,))s]+e−βhe[J~(t,μt;)], =e^-β hE^e [ _t-h^te^-β(s-t) (r_ π(s, _s)+γ E(s, _s, π) )ds ]+e^-β hE^e[ J(t, _t; π)], In view that J~(⋅,⋅;)∈1,2([0,T]×2(ℝd)) J(·,·; π)∈ C^1,2([0,T]× P_2(R^d)), applying Itô’s formula in Carmona and Delarue (2018b) leads to h−1(J~(t−h,μ;)−J~(t,μ;)) h^-1 ( J(t-h,μ; π)- J(t,μ; π) ) (3.10) = = h−1e−βhe[∫t−hte−β(s−t)(r^(s,μs)+γℰ(s,μs,))s] h^-1e^-β hE^e [ _t-h^te^-β(s-t) ( r_ π(s, _s)+γ E(s, _s, π) )ds ] +h−1e−βhe[J~(t,μt;)−J~(t,μ;)]+h−1(e−βh−1)J~(t,μ;) +h^-1e^-β hE^e [ J(t, _t; π)- J(t,μ; π) ]+h^-1(e^-β h-1) J(t,μ; π) = = h−1e−βhe[∫t−htℋγ(s,μs,;)s]+h−1(e−βh−1)J~(t,μ;). h^-1e^-β hE^e [ _t-h^t H^γ(s, _s, π; π)ds ]+h^-1(e^-β h-1) J(t,μ; π). Letting h→0h→ 0, we obtain (3.8). ∎ Under proper conditions, the optimal value function J~∗ J^* in (2.9) satisfies the HJB equation ∂J~∗∂t(t,μ)+sup∈(|ℝd)∫ℝd×H(t,x,μ,a,∂μJ~∗(t,μ)(x),∂x∂μJ~∗(t,μ)(x))(a|x)daμ(dx) ∂ J^*∂ t(t,μ)+ _ π∈ P( A|R^d) \ _R^d× AH (t,x,μ,a, _μ J^*(t,μ)(x), _x _μ J^*(t,μ)(x) ) π(a|x)daμ(dx) +12∫ℝd×ℝdTr(σo,(t,x,μ)σo,(t,x′,μ)⊺∂μ2J~∗(t,μ)(x,x′))μ(dx)⊗μ(dx′) + 12 _R^d×R^d Tr ( _o, π(t,x,μ) _o, π(t,x ,μ) _μ^2 J^*(t,μ)(x,x ) )μ(dx) μ(dx ) (3.11) +γℰ(t,μ,)−βJ~∗(t,μ)=0. + (t,μ, π) \-β J^*(t,μ)=0. Here, for fixed (t,μ)∈[0,T]×2(ℝd)(t,μ)∈[0,T]× P_2(R^d), we will frequently identify a policy ∈Π π∈ with the probability transition kernel ∈ac(|ℝd) π∈ P_ac( A|R^d) by abuse of notations, and hereafter we will not distinguish Π and ac(|ℝd) P_ac( A|R^d) when there is no confusion. Note that the sup operator in (3.11) for the optimal value function leads to a fully nonlinear PDE in the Wasserstein space. The conventional policy iteration is to linearize (3.11) from a given policy and hope to iterate the policy to reach the optimal one. This motivates us to consider the functional ℋ:[0,T]×2(ℝd)×ac(|ℝd)→ℝ H:[0,T]× P_2(R^d)× P_ac( A|R^d) as the integrated Hamiltonian under the fixed policy π defined by ℋ(t,μ,;) H(t,μ, h; π) :=∫ℝd×H(t,x,μ,a,∂μJ~(t,μ;)(x),∂x∂μJ~(t,μ;)(x))(a|x)aμ(dx) := _R^d× AH (t,x,μ,a, _μ J(t,μ; π)(x), _x _μ J(t,μ; π)(x) ) h(a|x)daμ(dx) (3.12) +12∫ℝd×ℝdTr(σo,(t,x,μ)σo,(t,x′,μ)⊺∂μ2J~(t,μ;)(x,x′))μ(dx)⊗μ(dx′). \;\;\;+ 12 _R^d×R^d Tr ( _o, h(t,x,μ) _o, h(t,x ,μ) _μ^2 J(t,μ; π)(x,x ) )μ(dx) μ(dx ). Due to the dependence of the coefficient σo _o on the action a, ℋ(t,μ,;) H(t,μ, h; π) is nonlinear in h. For the fixed policy π, we consider the following entropy-regularized optimization problem over the space of policies h (as the probability transition kernel ∈ac(|ℝd) h∈ P_ac( A|R^d) with the fixed (t,μ)(t,μ)): sup∈ac(|ℝd)(ℋ(t,μ,;)+γℰ(t,μ,))=:sup∈ac(|ℝd)ℋγ(t,μ,;). _ h∈ P_ac( A|R^d) ( H(t,μ, h; π)+ (t,μ, h) )=: _ h∈ P_ac( A|R^d) H^γ(t,μ, h; π). (3.13) Given the current policy π, to characterize the maximizer in (3.13), called the optimal one-step iterated policy, we adopt two notions of concavity with respect to measures, namely, the classical concavity and the displacement concavity in McCann (1997) and Villani (2009) to our setting with respect to probability transition kernels. Assumption 3.6. For a fixed ∈Π π∈ and (t,μ)∈[0,T]×2(ℝd)(t,μ)∈[0,T]× P_2(R^d), the integrated Hamiltonian ℋ H in (3.12) satisfies either of the following two conditions: for 0,1 h_0, h_1 in ac(|ℝd) P_ac( A|R^d) and any θ∈[0,1]θ∈[0,1] (i) ℋ H is concave in the classical sense ℋ(t,μ,θ0+(1−θ)1;)≥θℋ(t,μ,0;)+(1−θ)ℋ(t,μ,1;), H(t,μ,θ h_0+(1-θ) h_1; π)≥θ H(t,μ, h_0; π)+(1-θ) H(t,μ, h_1; π), (i) ℋ H is concave in ∈ac(|ℝd) h∈ P_ac( A|R^d) in the displacement concave sense if there exists a geodesic θ:=ℙθϕ0(⋅,U)+(1−θ)ϕ1(⋅,U)2 h_θ:=P^2_θ _ h_0(·,U)+(1-θ) _ h_1(·,U) ℋ(t,μ,θ;)≥θℋ(t,μ,0;)+(1−θ)ℋ(t,μ,1;), H(t,μ, h_θ; π)≥θ H(t,μ, h_0; π)+(1-θ) H(t,μ, h_1; π), where for any ∈ac(|ℝd) h∈ P_ac( A|R^d), ϕ(⋅,U) _ h(·,U) is a random variable on (Ω2,ℱ2,ℙ2)( ^2, F^2,P^2) with ℙϕ(x,U)2=(⋅|x)P^2_ _ h(x,U)= h(·|x). Remark 3.7. The difference between the classical concavity and the displacement concavity lies in the means of interpolation. Both notions have their own merits: • When σo _o is not controlled, ℋ H is linear in ∈ac(|ℝd) h∈ P_ac( A|R^d) and hence concave in the classical sense. However, to make sure that ℋ H is displacement concave in ∈ac(|ℝd) h∈ P_ac( A|R^d), we have to assume that H(t,x,μ,a,p,q)H(t,x,μ,a,p,q) is concave in a∈a∈ A. Therefore, it would be better to utilize the classical concavity rather than the displacement concavity in this case. • When σo _o is controlled, one necessary and sufficient condition that ℋ H in (3.12) is concave in h is that ∂μ2J(t,μ;) _μ^2J(t,μ; π) satisfies for any ,′∈ac(|ℝd) h, h ∈ P_ac( A|R^d) ∫ℝd×ℝdTr((σo,−σo,′)(t,x,μ)(σo,−σo,′)(t,x′,μ)⊺∂μ2J(t,μ;)(x,x′))μ(dx)⊗μ(dx′)≤0. _R^d×R^d Tr ( ( _o, h- _o, h )(t,x,μ) ( _o, h- _o, h )(t,x ,μ) _μ^2J(t,μ; π)(x,x ) )μ(dx) μ(dx )≤ 0. However, some classical LQ-MFC problems, such as the mean-variance optimization problems, do not satisfy the above inequality. Instead, one can easily check that the displacement concavity condition holds for LQ-MFC problems; see section 5 for more details. Essentially, we can use the derivative of ℋγ H^γ with respect to the policy to characterize the improved policy. To this end, in the same spirit of the linear functional derivative with respect to the probability measure, see Definition 5.43 Carmona and Delarue (2018a), we consider the following definition of the partial linear functional derivative with respect to the probability transition kernel, see section 2.1 in Conforti et al. (2023). Definition 3.8 (Partial linear functional derivative ). Fix μ∈2(ℝd)μ∈ P_2(R^d). The functional δGδ:ac(|ℝd)×ℝd×→ℝ δ Gδ h: P_ac( A|R^d)×R^d× A is said to be a partial linear functional derivative of G:ac(|ℝd)→ℝG: P_ac( A|R^d) with respect to the probability transition kernel ∈ac(|ℝd) h∈ P_ac( A|R^d) if for any ,′∈ac(|ℝd) h, h ∈ P_ac( A|R^d), G(′)−G()=∫01∫ℝd∫δGδ((1−λ)+λ′)(x,a)(′−)(a|x)aμ(dx)λ. G( h )-G( h)= _0^1 _R^d _ A δ Gδ h((1-λ) h+λ h )(x,a)( h - h)(a|x)daμ(dx)dλ. Moreover, there exists a constant C>0C>0, possibly depending on μ, such that sup∈ac(|ℝd)|δGδ()(x,a)|≤C(1+|x|2+|a|2). _ h∈ P_ac( A|R^d) | δ Gδ h( h)(x,a) |≤ C(1+|x|^2+|a|^2). Remark 3.9. The partial linear functional derivative with respect to the probability transition kernel is unique up to an additive function κ(x,μ)κ(x,μ) satisfying ∫ℝdκ(x,μ)μ(dx)=C _R^dκ(x,μ)μ(dx)=C for some constant. See Buckdahn et al. (2021) for the partial L-derivative with respect to probability transition kernels. Given the current policy ∈Π π∈ , the next result gives the existence, uniqueness and the first-order condition of the optimal one-step iterated policy for the problem (3.13). Theorem 3.10. Let Assumptions 2.1 and 3.6 hold. Assume that J(⋅,⋅;)∈1,2([0,T]×2(ℝd))J(·,·; π)∈ C^1,2([0,T]× P_2(R^d)). Given ∈Π π∈ and (t,μ)∈[0,T]×2+δ(ℝd)(t,μ)∈[0,T]× P_2+δ(R^d) for some δ>0δ>0. The entropy-regularized integrated Hamiltonian ℋγ(t,μ,;) H^γ(t,μ, h; π) in (3.13) has a unique maximizer ∗∈Π h^*∈ if and only if ∗ h^* satisfies δℋδ(t,μ,∗;)(x,a)−γlog∗(a|t,x,μ)=κ(t,x,μ), δ Hδ h(t,μ, h^*; π)(x,a)-γ h^*(a|t,x,μ)=κ(t,x,μ), (3.14) where δℋδ δ Hδ h is given by δℋδ(t,μ,;)(x,a)= δ Hδ h(t,μ, h; π)(x,a)= H(t,x,μ,a,∂μJ(t,μ;)(x),∂x∂μJ(t,μ;)(x)) H (t,x,μ,a, _μJ(t,μ; π)(x), _x _μJ(t,μ; π)(x) ) (3.15) +∫ℝdTr(σo(t,x,μ,a)σo,(t,x′,μ)⊺∂μ2J(t,μ;)(x,x′))μ(dx′). + _R^d Tr ( _o(t,x,μ,a) _o, h(t,x ,μ) _μ^2J(t,μ; π)(x,x ) )μ(dx ). Or equivalently, ∗ h^* is the fixed point of of Φ:Π→Π _ π: → defined by Φ()(a|t,x,μ)=exp1γδℋδ(t,μ,;)(x,a)∫exp1γδℋδ(t,μ,;)(x,a)a. _ π( h)(a|t,x,μ)= \ 1γ δ Hδ h(t,μ, h; π)(x,a) \ _ A \ 1γ δ Hδ h(t,μ, h; π)(x,a) \da. (3.16) Consequently, ∗ h^* is a map of π and we denote by ∗=ℐ() h^*=I( π). Proof. Step-1. We first show the existence and uniqueness of the maximizer ∗ h^* of (3.13). In this step, we change the admissible policy space from ac(|ℝd) P_ac( A|R^d) to (|ℝd) P( A|R^d), which is more convenient for the compactness and does not affect the result. This is because for those policies not in ac(|ℝd) P_ac( A|R^d), we set ℰ(t,μ,)=−∞ E(t,μ, h)=-∞ by convention. Thereby, if a maximizer h exists, it belongs to ac(|ℝd) P_ac( A|R^d). Define μ:=∈2(ℝd×):(⋅,ℝd)=μ V_μ:=\ ν∈ P_2(R^d× A): ν(·,R^d)=μ\ as the space of probability measures on ℝd×R^d× A whose first marginal μ is fixed. Then, the one-to-one correspondence between ac(|ℝd) P_ac( A|R^d) and μ V_μ holds, i.e., for each ∈μ ν∈ V_μ, there exists a (da|x)∈(|ℝd) h(da|x)∈ P( A|R^d) by disintegration such that (dx,da)=(da|x)μ(dx) ν(dx,da)= h(da|x)μ(dx). Conversely, each (da|x)∈(|ℝd) h(da|x)∈ P( A|R^d) induces a probability measure ∈μ ν∈ V_μ. Thanks to the equivalence between μ V_μ and (|ℝd) P( A|R^d), we may rewrite ℋγ(t,μ,;) H^γ(t,μ, h; π) as a functional of ∈μ ν∈ V_μ with a slight abuse of notation and ℰ(t,μ,)=ℰ(t,μ,),ℋ(t,μ,;)=ℋ(t,μ,;), E(t,μ, h)= E(t,μ, ν),\; H(t,μ, ν; π)= H(t,μ, h; π), (3.17) with ℋ(t,μ,;)=∫ℝd×H(t,x,μ,a,∂μJ~(t,μ;)(x),∂x∂μJ~(t,μ;)(x))(dx,da) H(t,μ, ν; π)= _R^d× AH (t,x,μ,a, _μ J(t,μ; π)(x), _x _μ J(t,μ; π)(x) ) ν(dx,da) +12∫ℝ2d×2Tr(σo(t,x,μ,a)σo(t,x′,μ,a′)⊺∂μ2J~(t,μ;)(x,x′))(dx,da)⊗(dx′,da′), + 12 _R^2d× A^2 Tr ( _o(t,x,μ,a) _o(t,x ,μ,a ) _μ^2 J(t,μ; π)(x,x ) ) ν(dx,da) ν(dx ,da ), and whenever ν is absolutely continuous with respect to μ(dx)daμ(dx)da ℰ(t,μ,) E(t,μ, ν) =−∫ℝd×log(dx,da)μ(dx)da(dx,da). =- _R^d× A ν(dx,da)μ(dx)da ν(dx,da). Otherwise, we set ℰ()=−∞ E( ν)=-∞. Hence the optimization problem (3.13) becomes sup∈μℋγ(t,μ,;). _ ν∈ V_μ H^γ(t,μ, ν; π). (3.18) The next step is to reduce the problem (3.18) to maximizing an upper semicontinuous function on a compact set μ S_μ of μ V_μ under the 2 W_2 metric. First, note that there exists some ¯∈μ ν∈ V_μ such that ℋγ(t,μ,¯;)<+∞ H^γ(t,μ, ν; π)<+∞ and let us introduce a subset μ S_μ of μ V_μ μ:=∈μ:γℰ()≥ℋγ(t,μ,¯;)−sup′∈μℋ(t,μ,′;). S_μ:= \ ν∈ V_μ: ( ν)≥ H^γ(t,μ, ν; π)- _ ν ∈ V_μ H(t,μ, ν ; π) \. It follows from the definition of μ S_μ that ℋγ(t,μ,;)≤ℋγ(t,μ,¯;) H^γ(t,μ, ν; π)≤ H^γ(t,μ, ν; π) for any ∉μ ν∉ S_μ, hence sup∈μℋγ(t,μ,;)=sup∈μℋγ(t,μ,;) _ ν∈ V_μ H^γ(t,μ, ν; π)= _ ν∈ S_μ H^γ(t,μ, ν; π). As the sublevel set of −ℰ()-E( ν) is weakly compact by Lemma 1.4.3 in Dupuis and Ellis (2011), S is weakly compact. Furthermore, by Definition 2.2 (i), we have sup∈μ∫ℝd×(|x|2+|a|2)(2+δ)/2(dx,da)≤C∫ℝd(|x|2+δ+sup∈Π∫|a|2+δ(a|t,x,μ)a)μ(dx)<+∞. _ ν∈ S_μ _R^d× A(|x|^2+|a|^2)^(2+δ)/2 ν(dx,da)≤ C _R^d (|x|^2+δ+ _ h∈ _ A|a|^2+δ h(a|t,x,μ)da )μ(dx)<+∞. Theorem 5.5 in Carmona and Delarue (2018a) guarantees that μ S_μ is compact in 2(ℝd×) P_2(R^d× A) under the 2 W_2 metric. By Assumption 2.1, H(t,x,μ,a,∂μJ~(t,μ;)(x),∂x∂μJ~(t,μ;)(x))H (t,x,μ,a, _μ J(t,μ; π)(x), _x _μ J(t,μ; π)(x) ) is continuous in (x,a)(x,a) and square integrable. Similarly, Tr(σo(t,x,μ,a)σo(t,x′,μ,a′)⊺∂μ2J~(t,μ;)(x,x′)) Tr ( _o(t,x,μ,a) _o(t,x ,μ,a ) _μ^2 J(t,μ; π)(x,x ) ) is continuous in (x,x′,a,a′)(x,x ,a,a ) and square integrable. Therefore, by Lemma A.3 in Lacker (2015), ℋ(t,μ,;) H(t,μ, ν; π) is continuous under the 2 W_2 metric. On the other hand, by Lemma 1.4.3 in Dupuis and Ellis (2011), ℰ(t,μ,)E(t,μ, ν) is upper semicontinuous on 2(ℝd×) P_2(R^d× A) under the weak topology and hence continuous under the 2 W_2 metric on 2(ℝd×) P_2(R^d× A). Therefore, ℋγ H^γ is continuous and thus its supremum is attained in μ S_μ, which implies the existence of the maximizer of ℋγ H^γ. In view that ℋγ H^γ is strictly concave or strictly displacement concave, the maximizer ∗∈μ ν^*∈ V_μ is unique, which implies the uniqueness of the maximizer ∗∈(|ℝd) h^*∈ P( A|R^d) by disintegration. Note that ℰ(t,μ,∗)>−∞ E(t,μ, ν^*)>-∞, it holds that ∗∈ac(|ℝd) h^*∈ P_ac( A|R^d). Step-2. We verify the sufficient and necessary condition of the first-order condition under the condition (i) when ℋγ H^γ is strictly concave in ∈ac(|ℝd) h∈ P_ac( A|R^d). Step-2.1 Let us first prove the sufficient condition. Let ∗ h^* satisfy (3.14). Denote θ:=(1−θ)∗+θ h^θ:=(1-θ) h^*+θ h for any ∈ac(|ℝd) h∈ P_ac( A|R^d) and 0<θ≤10<θ≤ 1. As ℋ H is concave in ∈ac(|ℝd) h∈ P_ac( A|R^d), we have 1θ(ℋ(t,μ,θ;)−ℋ(t,μ,∗;))≤ 1θ ( H(t,μ, h^θ; π)- H(t,μ, h^*; π) )≤ 1θ∫ℝd×δℋδ(t,μ,∗)(x,a)(θ−∗)(a|x)aμ(dx) 1θ _R^d× A δ Hδ h(t,μ, h^*)(x,a)( h^θ- h^*)(a|x)daμ(dx) = = ∫ℝd×δℋδ(t,μ,∗)(x,a)(−∗)(a|x)aμ(dx). _R^d× A δ Hδ h(t,μ, h^*)(x,a)( h- h^*)(a|x)daμ(dx). (3.19) Similarly, by the concavity of ℰ(t,μ,)E(t,μ, h) in h, we get that γθ(ℰ(t,μ,θ)−ℰ(t,μ,∗)))≤ γθ (E(t,μ, h^θ)-E(t,μ, h^*)) )≤ γθ∫ℝd×δℰδ(t,μ,∗)(x,a)(θ−∗)(a|x)aμ(dx) γθ _R^d× A δ h(t,μ, h^*)(x,a)( h^θ- h^*)(a|x)daμ(dx) = = −γ∫ℝd×log∗(a|x)(−∗)(a|x)aμ(dx). -γ _R^d× A h^*(a|x)( h- h^*)(a|x)daμ(dx). (3.20) Summing (3.19) and (3.20), we obtain that ℋγ(t,μ,θ;)−ℋγ(t,μ,∗;)θ H^γ(t,μ, h^θ; π)- H^γ(t,μ, h^*; π)θ ≤ ≤ ∫ℝd×(δℋδ(t,μ,∗;)(x,a)−γlog∗(a|x))(−∗)(a|x)aμ(dx)=0. _R^d× A ( δ Hδ h(t,μ, h^*; π)(x,a)-γ h^*(a|x) )( h- h^*)(a|x)daμ(dx)=0. This implies that ℋγ(t,μ,θ;)≤ℋγ(t,μ,∗;) H^γ(t,μ, h^θ; π)≤ H^γ(t,μ, h^*; π) for any ∈Π h∈ if θ=1θ=1. Step-2.2. We then show the necessary condition. Let ∗ h^* be the maximizer of ℋγ(t,μ,;) H^γ(t,μ, h; π). It then holds that ℋγ(t,μ,θ;)−ℋγ(t,μ,∗;)θ H^γ(t,μ, h^θ; π)- H^γ(t,μ, h^*; π)θ = = 1θ(ℋ(t,μ,θ)−ℋ(t,μ,∗)+γ(ℰ(t,μ,θ)−ℰ(t,μ,∗))) 1θ ( H(t,μ, h^θ)- H(t,μ, h^*)+γ (E(t,μ, h^θ)-E(t,μ, h^*) ) ) = = 1θ∫01∫ℝd×(δℋδ(t,μ,λ,θ;)(x,a)−γlogλ,θ(a|x))(θ−)(a|x)aμ(dx)λ 1θ _0^1 _R^d× A ( δ Hδ h(t,μ, h^λ,θ; π)(x,a)-γ h^λ,θ(a|x) )( h^θ- h)(a|x)daμ(dx)dλ = = ∫01∫ℝd×(δℋδ(t,x,μ,λ,θ,a;)−γlogλ,θ(a|x))(−∗)(a|x)aμ(dx), _0^1 _R^d× A ( δ Hδ h(t,x,μ, h^λ,θ,a; π)-γ h^λ,θ(a|x) )( h- h^*)(a|x)daμ(dx), where we denote λ,θ=λθ+(1−λ)∗ h^λ,θ=λ h^θ+(1-λ) h^*. As δℋδ δ Hδ h in (3.15) is continuous and integrable by Assumption 2.1, by dominated convergence theorem, we obtain that limθ→0ℋγ(t,μ,θ;)−ℋγ(t,μ,∗;)θ _θ→ 0 H^γ(t,μ, h^θ; π)- H^γ(t,μ, h^*; π)θ = = ∫ℝd×(δℋδ(t,μ,∗;)(x,a)−γlog∗(a|x))(∗−1)∗(a|x)aμ(dx)≤0, _R^d× A ( δ Hδ h(t,μ, h^*; π)(x,a)-γ h^*(a|x) ) ( h h^*-1 ) h^*(a|x)daμ(dx)≤ 0, for any h. By the arbitrariness of h and Definition 2.2 (i), it holds that δℋδ(t,μ,∗;)(x,a)−γlog∗(a|x)=κ(t,x,μ) δ Hδ h(t,μ, h^*; π)(x,a)-γ h^*(a|x)=κ(t,x,μ). Step-3. When ℋγ H^γ is strictly displacement concave, we verify the necessary and sufficient condition of the first-order condition. By the definition of displacement concavity, [0,1]∋θ↦ℋγ(t,μ,ℙ(1−θ)ϕ∗(⋅,U)+θϕ(⋅,U)2;)[0,1] θ H^γ(t,μ,P^2_(1-θ) _ h^*(·,U)+θ _ h(·,U); π) is concave in θ∈[0,1]θ∈[0,1]. Therefore, ∗ h^* is the maximizer if and only if for any ϕ:ℝd→ _ h:R^d→ A, dθℋγ(t,μ,ℙ(1−θ)ϕ∗(⋅,U)+θϕ(⋅,U)2;)|θ=0=0. ddθ H^γ(t,μ,P^2_(1-θ) _ h^*(·,U)+θ _ h(·,U); π) |_θ=0=0. Take ϕ=∘ϕ∗ _ h= T _ h^* with the map :→ T: A→ A being injective and surjective, and denote ϕθ:=(1−θ)ϕ∗+θ∘ϕ∗ _ h_θ:=(1-θ) _ h^*+θ T _ h^* and θ:=ℙϕθ(⋅,U)2 h_θ:=P^2_ _ h_θ(·,U). It then holds that dθℋγ(t,μ,θ;)|θ=0 ddθ H^γ(t,μ, h_θ; π) |_θ=0 = = limθ→01θ(ℋγ(t,μ,θ;)−ℋγ(t,μ,∗;)) _θ→ 0 1θ ( H^γ(t,μ, h_θ; π)- H^γ(t,μ, h^*; π) ) = = limθ→01θ∫01∫ℝd×δℋγδ(t,μ,θ,λ)(x,a)(θ−∗)(a|x)μ(dx)λ _θ→ 0 1θ _0^1 _R^d× A δ H^γδ h(t,μ, h_θ,λ)(x,a)( h_θ- h^*)(a|x)μ(dx)dλ = = limθ→01θ∫01(∫ℝd×δℋγδ(t,μ,θ,λ)(x,a)θ(a|x)μ(dx)−δℋγδ(t,μ,θ,λ)(x,a)∗(a|x)μ(dx))λ _θ→ 0 1θ _0^1 ( _R^d× A δ H^γδ h(t,μ, h_θ,λ)(x,a) h_θ(a|x)μ(dx)- δ H^γδ h(t,μ, h_θ,λ)(x,a) h^*(a|x)μ(dx) )dλ = = limθ→01θ∫01(e[δℋγδ(t,μ,θ,λ)(ξ,ϕθ(ξ,U))−δℋγδ(t,μ,θ,λ)(ξ,ϕ∗(ξ,U))])λ _θ→ 0 1θ _0^1 (E^e [ δ H^γδ h (t,μ, h_θ,λ)(ξ, _ h_θ(ξ,U) )- δ H^γδ h (t,μ, h_θ,λ)(ξ, _ h^*(ξ,U) ) ] )dλ = = e[∇aδℋγδ(t,μ,∗)(ξ,ϕ∗(ξ,U))⊺(∘ϕ∗(ξ,U)−ϕ∗(ξ,U))] ^e [ _a δ H^γδ h (t,μ, h^* )(ξ, _ h^*(ξ,U)) ( T _ h^*(ξ,U)- _ h^*(ξ,U) ) ] = = ∫ℝd×∇aδℋγδ(t,μ,∗)(x,a)⊺((a)−a)∗(a|x)μ(dx)=0, _R^d× A _a δ H^γδ h (t,μ, h^* )(x,a) ( T(a)-a ) h^*(a|x)μ(dx)=0, where the second equality follows from the definition of partial linear functional derivative and θ,λ=(1−λ)+λθ h_θ,λ=(1-λ) h+λ h_θ. By the arbitrariness of the map T, we deduce that ∇aδℋγδ(t,μ,∗;)(x,a)=∇aδℋδ(t,μ,∗;)(x,a)−γ∇alog∗(a|x)=0, _a δ H^γδ h(t,μ, h^*; π)(x,a)= _a δ Hδ h(t,μ, h^*; π)(x,a)-γ _a h^*(a|x)=0, which yields the desired result. ∎ Remark 3.11. In particular, when there is no common noise, i.e., σo=0 _o=0, or when the common noise is uncontrolled, i.e., σo=σo(t,x,μ) _o= _o(t,x,μ), the fixed point of (3.16) reduces to the conventional Gibbs measure characterization, which is consistent with (2.11) in Wei and Yu (2025). When there is no common noise nor mean-field term, (3.16) becomes (13) in Jia and Zhou (2023). Based on Theorem 3.10, the learning procedure starts with some policy π and produces a new policy ′ π that improves ℋγ(t,μ,;) H^γ(t,μ, h; π). The next result shows that the resulting new policy that improves ℋγ(t,μ,;) H^γ(t,μ, h; π) will also improve the value function, and if the iterated new policy cannot improve the value function any more, it must be an optimal policy. Theorem 3.12 (Policy improvement). For a given ∈Π π∈ , select a new policy ′ π such that ℋγ(s,μ,′;)≥ℋγ(s,μ,;) H^γ(s,μ, π ; π)≥ H^γ(s,μ, π; π) holds for any s∈[t,T]s∈[t,T], we then have J(t,μ;′)≥J(t,μ;)J(t,μ; π )≥ J(t,μ; π). Proof. For two given admissible policies ,′∈Π π, π ∈ , and any 0≤t≤T0≤ t≤ T, by applying Itô’s formula in Carmona and Delarue (2018b) to e−β(s−t)J~(s,μs′;)e^-β(s-t) J(s, _s π ; π) between t and T, we get that e[e−β(T−t)J~(T,μT′;)−J~(t,μt′;)+∫tTe−β(s−t)(r(s,X~s′,μs′,as′)+γℰ(s,μs′,′))ds] ^e [e^-β(T-t) J(T, _T π ; π)- J(t, _t π ; π)+ _t^Te^-β(s-t) (r(s, X_s π , _s π ,a_s π )+ (s, _s π , π ) )ds ] = = e[∫tTe−β(s−t)(∂J~∂t(s,μs′;)−βJ~(s,μs′;)+ℋγ(s,μs′,′;))ds]. ^e [ _t^Te^-β(s-t) ( ∂ J∂ t(s, _s π ; π)-β J(s, _s π ; π)+ H^γ(s, _s π , π ; π) )ds ]. Using μt′=μ _t π =μ and J~(T,μ;)=g^(μ) J(T,μ; π)= g(μ), we rewrite the above equality as J~(t,μ;′)−J~(t,μ;)=e[∫tTe−β(s−t)(∂J~∂t(s,μs′;)−βJ~(s,μs′;)+ℋγ(s,μs′,′;))ds]. J(t,μ; π )- J(t,μ; π)=E^e [ _t^Te^-β(s-t) ( ∂ J∂ t(s, _s π ; π)-β J(s, _s π ; π)+ H^γ(s, _s π , π ; π) )ds ]. (3.21) Therefore, for any (s,μ)∈[t,T]×2(ℝd)(s,μ)∈[t,T]× P_2(R^d), ℋγ(s,μ,′;)≥ℋγ(s,μ,;) H^γ(s,μ, π ; π)≥ H^γ(s,μ, π; π), we obtain that J~(t,μ;′)−J~(t,μ;) J(t,μ; π )- J(t,μ; π) = = e[∫tTe−β(s−t)(∂J~∂t(s,μs′;)−βJ~(s,μs′;)+ℋγ(s,μs′,′;))ds] ^e [ _t^Te^-β(s-t) ( ∂ J∂ t(s, _s π ; π)-β J(s, _s π ; π)+ H^γ(s, _s π , π ; π) )ds ] ≥ ≥ e[∫tTe−β(s−t)(∂J~∂t(s,μs′;)−βJ~(s,μs′;)+ℋγ(s,μs′,;))ds]=0, ^e [ _t^Te^-β(s-t) ( ∂ J∂ t(s, _s π ; π)-β J(s, _s π ; π)+ H^γ(s, _s π , π; π) )ds ]=0, where the last equality holds because of the dynamic programming equation (3.8). ∎ Corollary 3.13. For a given ∈Π π∈ , define ′=ℐ() π =I( π), with ℐI given in Theorem 3.10. Then J(t,μ;′)≥J(t,μ;)J(t,μ; π )≥ J(t,μ; π). Conversely, if there exists some ^∈Π π∈ such that J(t,μ;^′)=J(t,μ;^)J(t,μ; π )=J(t,μ; π) for any (t,μ)∈[0,T]×2(ℝd)(t,μ)∈[0,T]× P_2(R^d), with ^′=ℐ(^) π = I( π), then π is an optimal policy of (2.13). Proof. If we take ′=ℐ() π =I( π), then for any (t,μ)∈[0,T]×2(ℝd)(t,μ)∈[0,T]× P_2(R^d), it holds that ℋγ(s,μ,′;)≥ℋγ(s,μ,;) H^γ(s,μ, π ; π)≥ H^γ(s,μ, π; π). By Theorem 3.12, we deduce that J~(t,μ;′)≥J~(t,μ;) J(t,μ; π )≥ J(t,μ; π). We next prove the second claim. By (3.21) and J~(t,μ;^)=J~(t,μ;^′) J(t,μ; π)= J(t,μ; π ), we get that for any (t,μ)∈[0,T]×2(ℝd)(t,μ)∈[0,T]× P_2(R^d), e[∫tTe−β(s−t)(∂J~∂t(s,μs′;^)−βJ~(s,μs′;)+ℋγ(s,μs′,^′;^))ds]=0. ^e [ _t^Te^-β(s-t) ( ∂ J∂ t(s, _s π ; π)-β J(s, _s π ; π)+ H^γ(s, _s π , π ; π) )ds ]=0. (3.22) Similarly, in view of (3.21) and J~(t+h,μt+h^′;^)=J~(t+h,μt+h^′;ℐ(^)) J(t+h, _t+h π ; π)= J(t+h, _t+h π ;I( π)), ℙ0P^0-a.s. for any h∈[0,T−t)h∈[0,T-t), we have that e[∫t+hTe−β(s−t)(∂J~∂t(s,μs^′;^)−βJ(s,μs^′;^)+ℋγ(s,μs^′,^′;^))ds]=0. ^e [ _t+h^Te^-β(s-t) ( ∂ J∂ t(s, _s π ; π)-β J(s, _s π ; π)+ H^γ(s, _s π , π ; π) )ds ]=0. (3.23) Subtracting (3.23) from (3.22) and dividing h on both sides, we get that 1he[∫t+he−β(s−t)(∂J∂t(s,μs^′;^)−βJ(s,μs^′;^)+ℋγ(s,μs^′,^′;^))ds]=0. 1hE^e [ _t^t+he^-β(s-t) ( ∂ J∂ t(s, _s π ; π)-β J(s, _s π ; π)+ H^γ(s, _s π , π ; π) )ds ]=0. By the continuity of b,σb,σ, r, J~ J and μs^′ _s π , it holds by sending h→0h→ 0 that ∂J~∂t(t,μ;^)−βJ~(t,μ;^)+ℋγ(t,μ,^′;^)=0. ∂ J∂ t(t,μ; π)-β J(t,μ; π)+ H^γ(t,μ, π ; π)=0. (3.24) As J~(t,μ;^) J(t,μ; π) satisfies the equation (3.8), we arrive at ℋγ(t,μ,^′;^)=ℋγ(t,μ,^;^) H^γ(t,μ, π ; π)= H^γ(t,μ, π; π). Recall from Proposition 3.10 that ^′=ℐ(^) π =I( π) is the unique maximizer of ℋγ(t,μ,;^) H^γ(t,μ, h; π), then we have ℐ(^)=^I( π)= π. By a standard verification argument for the entropy regularized MFC problem, it holds that J(t,μ,^)=J∗(t,μ)J(t,μ, π)=J^*(t,μ) and hence π is an optimal policy. ∎ 4 Continuous-Time Integrated q-Function We investigate in this section the proper definition of the continuous-time integrated q-function (Iq-function), which lays the theoretical foundation of q-learning theory for MFC problems. Similar to Wei and Yu (2025), let us consider a “perturbed policy” ¯∈Π π∈ , which takes ∈ac(|ℝd) h∈ P_ac( A|R^d) on [t,t+Δt)[t,t+ t), and then ∈Π π∈ on [t+Δt,T)[t+ t,T). Then Xs,¯X_s D, π on [t,T)[t,T) is governed by dXs,¯ dX_s D, π =b(s,Xs,¯,μs,¯,aδ(s))ds+σ(s,Xst,ξ,¯,μs,¯,aδ(s))dWs,s∈[t,t+Δt), =b(s,X_s D, π, _s D, π,a_δ(s) h)ds+σ(s,X_s^t,ξ, π, _s D, π,a_δ(s) h)dW_s,\;s∈[t,t+ t), +σo(s,Xst,ξ,¯,μs,¯,aδ(s))dBs,Xt,¯=ξ, \;\;\;+ _o(s,X_s^t,ξ, π, _s D, π,a_δ(s) h)dB_s,\;X_t D, π=ξ, dXs,¯ dX_s D, π =b(s,Xs,¯,μs,¯,aδ(s))ds+σ(s,Xst,ξ,¯,μs,¯,aδ(s))dWs,s∈[t+Δt,T), =b(s,X_s D, π, _s D, π,a_δ(s) π)ds+σ(s,X_s^t,ξ, π, _s D, π,a_δ(s) π)dW_s,\;s∈[t+ t,T), +σo(s,Xst,ξ,¯,μs,¯,aδ(s))dBs,Xt+Δt,¯=Xt+Δt,. \;\;\;+ _o(s,X_s^t,ξ, π, _s D, π,a_δ(s) π)dB_s,\;X_t+ t D, π=X_t+ t D, h. We first consider the discrete time IQ-function defined on [0,T]×L2(Ω;ℝd)×ac(|ℝd)[0,T]× L^2( ;R^d)× P_ac( A|R^d) independent of discretely sampling, with the fixed time interval Δt t and the entropy regularizer that QΔt(t,ξ,;)=:lim||→0QΔt(t,ξ,;) Q_ t(t,ξ, h; π)=: _| D|→ 0Q_ t D(t,ξ, h; π) = = lim||→0e[∫t+Δte−β(s−t)(r(s,Xs,¯,μs,¯,aδ(s))+γE(δ(s),X,δ(s)¯,μs,¯))ds _| D|→ 0E^e [ _t^t+ te^-β(s-t) (r(s,X_s D, π, _s D, π,a h_δ(s))+γ E_ h(δ(s),X_ D,δ(s) π, _s D, π) )ds +∫t+ΔtTe−β(s−t)(r(s,Xs,¯,μs,¯,aδ(s))+γE(δ(s),X,δ(s)¯,μs,¯))s \;+ _t+ t^Te^-β(s-t) (r(s,X_s D, π, _s D, π,a π_δ(s))+γ E_ π(δ(s),X_ D,δ(s) π, _s D, π) )ds +e−β(T−t)g(XT,¯,μT,¯)|Xt,¯=ξ]. \;+e^-β(T-t)g(X_T D, π, _T D, π) |X_t D, π=ξ ]. By noting that QΔt(t,ξ,;)=J(t,ξ;¯)Q_ t D(t,ξ, h; π)=J D(t,ξ; π) and the equivalence result in Proposition 3.3, we have that QΔt(t,ξ,;)=lim|Δ|→0J(t,ξ;¯)=J~(t,μ;¯). Q_ t(t,ξ, h; π)= _| D|→ 0J D(t,ξ; π)= J(t,μ; π). Consequently, we can rewrite QΔtQ_ t in terms of the relaxed control formulation QΔ(t,ξ,;) Q_ (t,ξ, h; π) =e[∫t+Δte−β(s−t)(r^(s,μs¯)+γℰ(s,μs¯,))ds =E^e [ _t^t+ te^-β(s-t) ( r_ h(s, _s π)+ (s, _s π, h) )ds +∫t+ΔtTe−β(s−t)(r^(s,μs¯)+γℰ(s,μs¯,))ds+e−β(T−t)g^(μT¯)]. \;+ _t+ t^Te^-β(s-t) ( r_ π(s, _s π)+ (s, _s π, π) )ds+e^-β(T-t) g( _T π) ]. Noting the collapse of IQ-function to the value function as Δt→0 t→ 0, we instead consider the first order derivative of QΔtQ_ t by using the flow property of μs¯ _s π and applying Itô’s formula (see Theorem 4.14 in Carmona and Delarue (2018b)) to e−βsJ~(s,μs;)e^-β s J(s, _s h; π) between t and t+Δt+ t that QΔt(t,ξ,;)= Q_ t(t,ξ, h; π)= J~(t,μ;)+e[∫t+Δte−β(s−t)(∂J~∂t(s,μs;)−βJ~(s,μs;)+ℋγ(s,μs,;))ds], J(t,μ; π)+E^e [ _t^t+ te^-β(s-t) ( ∂ J∂ t(s, _s h; π)-β J(s, _s h; π)+ H^γ(s, _s h, h; π) )ds ], where ℋγ H^γ is defined in (3.12). By the continuity of J~ J and μs _s h with respect to s, it holds that QΔt(t,ξ,;)≈ Q_ t(t,ξ, h; π)≈ J~(t,μ;)+Δt(∂J~∂t(t,μ;)−βJ~(t,μ;)+ℋγ(t,μ,;))+o(Δt). J(t,μ; π)+ t ( ∂ J∂ t(t,μ; π)-β J(t,μ; π)+ H^γ(t,μ, h; π) )+o( t). (4.1) This leads to the next definition of continuous-time Iq-function. Definition 4.1. Given a policy ∈Π π∈ , for any (t,μ,)∈[0,T]×2(ℝd)×ac(|ℝd)(t,μ, h)∈[0,T]× P_2(R^d)× P_ac( A|R^d), we define the continuous-time integrated q-function (Iq-function) by qγ(t,μ,;) q^γ(t,μ, h; π) :=limΔt→0QΔt(t,ξ,;)−J~(t,μ;)Δt=∂J~∂t(t,μ;)−βJ~(t,μ;)+ℋγ(t,μ,;). := _ t→ 0 Q_ t(t,ξ, h; π)- J(t,μ; π) t= ∂ J∂ t(t,μ; π)-β J(t,μ; π)+ H^γ(t,μ, h; π). We also call q0(t,μ,;)=qγ(t,μ,;)−γℰ(t,μ,)q^0(t,μ, h; π)=q^γ(t,μ, h; π)- (t,μ, h) the unregularized Iq-function. It is straightforward to see that qγq^γ (resp. q0q^0) equals to ℋγ H^γ in (3.13) (resp. ℋ H in (3.12)) compensated by the dispersion term ∂J~∂t(t,μ;)−βJ~(t,μ;) ∂ J∂ t(t,μ; π)-β J(t,μ; π). Therefore, we obtain the following corollary, which is an immediate consequence of Theorem 3.10. Corollary 4.2. Let Assumptions 2.1 and 3.6 hold. Given ∈Π π∈ and (t,μ)∈[0,T]×2+δ(ℝd)(t,μ)∈[0,T]× P_2+δ(R^d) for some δ>0δ>0, there exists a unique maximizer to ∗=argmax∈ac(|ℝd)qγ(t,μ,;) h^*= h∈ P_ac( A|R^d)arg\,max\,q^γ(t,μ, h; π) if and only if ∗(a|t,x,μ)=Φ(∗)=exp1γδq0δ(t,μ,∗;)(x,a)∫exp1γδq0δ(t,μ,∗;)(x,a)a. h^*(a|t,x,μ)= _ π( h^*)= \ 1γ δ q^0δ h(t,μ, h^*; π)(x,a) \ _ A \ 1γ δ q^0δ h(t,μ, h^*; π)(x,a) \da. Furthermore, if there exists some ∗∈Π π^*∈ satisfying the two-layer fixed point to ∗=argmax∈ac(|ℝd)qγ,∗(t,μ,) π^*= h∈ P_ac( A|R^d)arg\,max\,q^γ,*(t,μ, h) or equivalently the two-layer fixed point to ∗(a|t,x,μ)=exp1γδq0,∗δ(t,μ,∗)(x,a)∫exp1γδq0,∗δ(t,μ,∗)(x,a)a, π^*(a|t,x,μ)= \ 1γ δ q^0,*δ h(t,μ, π^*)(x,a) \ _ A \ 1γ δ q^0,*δ h(t,μ, π^*)(x,a) \da, (4.2) where qγ,∗(t,μ,):=qγ(t,μ,;∗)q^γ,*(t,μ, h):=q^γ(t,μ, h; π^*) and q0,∗(t,μ,):=q0(t,μ,;∗)q^0,*(t,μ, h):=q^0(t,μ, h; π^*), then ∗ π^* is an optimal policy. Remark 4.3. 0 π^0 ℐ I⋯ ·s ℐ In π^n Φn _ π^nℐ In+1 π^n+1 ℐ I⋯ ·s ℐ I∗ π^*n,1 π^n,1 Φn _ π^n⋯ ·s Φn _ π^nn,ℓ π^n, Φn _ π^nn,ℓ+1 π^n, +1 Φn _ π^n⋯ ·s Φn _ π^n Figure 2: Illustration of two-layer fixed point (4.2) We could search for the optimal policy by considering (4.2) as a two-layer fixed point problem. Specifically, starting with an initial policy 0 π^0, at each iteration n∈ℕn , we derive n+1 π^n+1 by looking for the fixed point of the map Φn _ π^n, which is the inner layer of the two-layer fixed point problem. Recall that ℐ I is a map from n π^n to n+1 π^n+1, that is, n+1=ℐ(n) π^n+1= I( π^n). And then the optimal policy ∗ π^* is a fixed point of the map ℐ I, which is the outer layer of the two-layer fixed point problem. See Figure 2 for the illustration. 5 Linear Quadratic MFC and Gaussian Optimal Policy Let us consider a controlled linear McKean-Vlasov dynamics with =ℝp A=R^p and assume n=m=1n=m=1 for simplicity. The coefficients of the dynamics are given by b(t,x,μ,a) b(t,x,μ,a) =b0(t)+B(t)x+B¯(t)μ¯+C(t)a, =b_0(t)+B(t)x+ B(t) μ+C(t)a, σ(t,x,μ,a) σ(t,x,μ,a) =ϑ(t)+D(t)x+D¯(t)μ¯+F(t)a, = (t)+D(t)x+ D(t) μ+F(t)a, (5.1) σo(t,x,μ,a) _o(t,x,μ,a) =ϑo(t)+Do(t)x+D¯o(t)μ¯+Fo(t)a. = _o(t)+D_o(t)x+ D_o(t) μ+F_o(t)a. The running and terminal reward functions are given by r(t,x,μ,a)=x⊺M(t)x+μ¯⊺M¯(t)μ¯+a⊺R(t)a+x⊺O(t),g(x,μ)=x⊺Px+μ¯⊺P¯μ¯ r(t,x,μ,a)=x M(t)x+ μ M(t) μ+a R(t)a+x O(t),\;\;g(x,μ)=x Px+ μ P μ (5.2) Here, b0(t),ϑ(t)b_0(t), (t), ϑo(t) _o(t) and O(t)O(t) are deterministic functions of t valued in ℝdR^d, B(t)B(t), B¯(t) B(t), D(t)D(t), D¯(t) D(t), Do(t)D_o(t), D¯o(t) D_o(t), M(t)M(t) and M¯(t) M(t) are deterministic functions of t valued in ℝd×dR^d× d, C(t),F(t),Fo(t)C(t),F(t),F_o(t) are deterministic functions of t valued in ℝd×pR^d× p, R(t)R(t) is deterministic matrix functions of t valued in ℝp×pR^p× p, and P and P¯ P are constant matrices in ℝd×dR^d× d. We may assume without loss of generality that M(t),M¯(t),R(t),P,P¯M(t), M(t),R(t),P, P are symmetric matrices and β=0β=0. Denote the mean and variance of μ by μ¯=∫ℝdxμ(dx) μ= _R^dxμ(dx) and Var(μ)(Λ)=∫ℝd(x−μ¯)⊺Λ(x−μ¯)μ(dx) Var(μ)( )= _R^d(x- μ) (x- μ)μ(dx), respectively, for a symmetric matrix Λ∈ℝd×d ^d× d. We obtain explicit expressions of the optimal value function J~∗ J^* and the optimal policy ∗ π^* in the next result. Theorem 5.1. Under the condition ()P⪯0,P+P¯⪯0,M(t)⪯0,M(t)+M¯(t)⪯0,R(t)⪯−δIq, ( H)\;\;P 0,\;P+ P 0,\;M(t) 0,\;M(t)+ M(t) 0,\;R(t) -δ I_q, for some δ>0δ>0, the optimal value function J~∗ J^* takes the quadratic form J~∗(t,μ)=Var(μ)(Λ∗(t))+μ¯⊺Γ∗(t)μ¯+μ¯⊺ζ∗(t)+χ∗(t), J^*(t,μ)= Var(μ)( ^*(t))+ μ ^*(t) μ+ μ ζ^*(t)+χ^*(t), (5.3) and the optimal policy is unique and satisfies the Gaussian type that ∗(⋅|t,x,μ)=(−(Ut∗)−1St∗(x−μ¯)−(Vt∗)−1Zt∗μ¯−12(Vt∗)−1Yt∗,−γ2(Ut∗)−1). π^*(·|t,x,μ)= N (-(U_t^*)^-1S^*_t(x- μ)-(V^*_t)^-1Z^*_t μ- 12(V_t^*)^-1Y_t^*,- γ2(U^*_t)^-1 ). (5.4) Here, we set Ut=U(t,Λ(t))U_t=U(t, (t)), Vt=V(t,Γ(t))V_t=V(t, (t)), St=S(t,Λ(t))S_t=S(t, (t)), Zt=Z(t,Γ(t),Λ(t))Z_t=Z(t, (t), (t)), Yt=Y(ζ(t),Λ(t),Γ(t))Y_t=Y(ζ(t), (t), (t)) such that Ut=F⊺Λ(t)F+Fo⊺Λ(t)Fo+R,Vt=F⊺Λ(t)F+Fo⊺Γ(t)Fo+R,St=C⊺Λ(t)+F⊺Λ(t)D+Fo⊺Λ(t)Do,Zt=C⊺Γ(t)+F⊺Λ(t)(D+D¯)+Fo⊺Γ(t)(Do+D¯o),Yt=C⊺ζ(t)+2F⊺Λ(t)ϑ+2Fo⊺Γ(t)ϑ0, \ array[]lU_t&=F (t)F+F_o (t)F_o+R,\\ V_t&=F (t)F+F_o (t)F_o+R,\\ S_t&=C (t)+F (t)D+F_o (t)D_o,\\ Z_t&=C (t)+F (t) (D+ D )+F_o (t) (D_o+ D_o ),\\ Y_t&=C ζ(t)+2F (t) +2F_o (t) _0, array . and Λ∗(t) ^*(t), Γ∗(t) ^*(t), ζ∗(t)ζ^*(t) and χ∗(t)χ^*(t) satisfy (Λ∗)′(t)+M+D⊺Λ∗(t)D+Do⊺Λ∗(t)Do+B⊺Λ∗(t)+Λ∗(t)B−(St∗)⊺(Ut∗)−1St∗=0,Λ∗(T)=P, \ array[]rcl( ^*) (t)+M+D ^*(t)D+D_o ^*(t)D_o+B ^*(t)+ ^*(t)B-(S_t^*) (U_t^*)^-1S_t^*=0,\\ ^*(T)=P, array . (5.7) (Γ∗)′(t)+M+M¯+(D+D¯)⊺Λ∗(t)(D+D¯)+(D¯o+Do)⊺Γ∗(t)(D¯o+Do)+(B+B¯)⊺Γ∗(t)+Γ∗(t)(B+B¯)−(Zt∗)⊺(Vt∗)−1Zt∗=0,Γ∗(T)=P+P¯. \ array[]rcl( ^*) (t)+M+ M+ (D+ D) ^*(t) (D+ D)+( D_o+D_o) ^*(t)( D_o+D_o)\\ +(B+ B) ^*(t)+ ^*(t)(B+ B)-(Z_t^*) (V_t^*)^-1Z_t^*=0,\\ ^*(T)=P+ P. array . (5.11) (ζ∗)′(t)+(B+B¯)⊺ζ∗(t)+2Γ∗(t)b0+2(D+D¯)⊺Λ∗(t)ϑ+2(D¯o+Do)⊺Γ∗(t)ϑo−(Zt∗)⊺(Vt∗)−1Yt∗+O=0,ζ∗(T)=0, \ array[]rcl(ζ^*) (t)+(B+ B) ζ^*(t)+2 ^*(t)b_0+2(D+ D) ^*(t) +2( D_o+D_o) ^*(t) _o\\ -(Z_t^*) (V_t^*)^-1Y_t^*+O=0,\\ ζ^*(T)=0, array . (5.15) (χ∗)′(t)+ϑ⊺Λ∗(t)ϑ+ϑo⊺Γ∗(t)ϑo+b0⊺ζ∗(t)−14(Yt∗)⊺(Vt∗)−1Yt∗+γ2log((−γπ)pdet((Ut∗)−1))=0,χ∗(T)=0. \ array[]rcl(χ^*) (t)+ ^*(t) + _o ^*(t) _o+b_0 ζ^*(t)- 14(Y_t^*) (V_t^*)^-1Y_t^*\\ + γ2 ((-γπ)^p det((U_t^*)^-1) )=0,\\ χ^*(T)=0. array . (5.19) Remark 5.2. In view of (5.7)-(5.19), when γ tends to zero, the solution (Λ,Γ,ζ,χ)( , ,ζ,χ) of the LQ-MFC problem with entropy regularizer reduces to that of the classical LQ-MFC problem, and the relaxed optimal policy reduces to the optimal strict control in Pham and Wei (2017). Proof of Theorem 5.1. We conjecture that J∗J^* takes the following quadratic form in (5.3). One can easily check that J∗∈1,2([0,T]×2(ℝd))J^*∈ C^1,2([0,T]× P_2(R^d)) with ∂J∗∂t(t,μ) ∂ J^*∂ t(t,μ) =Var(μ)((Λ∗)′(t))+μ¯⊺(Γ∗)′(t)μ¯+μ¯⊺(ζ∗)′(t)+(χ∗)′(t), = Var(μ)(( ^*) (t))+ μ ( ^*) (t) μ+ μ (ζ^*) (t)+(χ^*) (t), ∂μJ∗(t,μ)(x) _μJ^*(t,μ)(x) =2Λ∗(t)(x−μ¯)+2Γ∗(t)μ¯+ζ∗(t), =2 ^*(t)(x- μ)+2 ^*(t) μ+ζ^*(t), ∂x∂μJ∗(t,μ)(x) _x _μJ^*(t,μ)(x) =2Λ∗(t),∂μ2J∗(t,μ)(x,x′)=2(Γ∗(t)−Λ∗(t)). =2 ^*(t),\; _μ^2J^*(t,μ)(x,x )=2( ^*(t)- ^*(t)). For simplicity, we will suppress the time variable t in the rest of the proof. Plugging (5.3) in the first order condition (3.14), together with (3.15), we have that (b0+Bx+B¯μ¯+Ca)⊺(2Λ∗(x−μ¯)+2Γ∗μ¯+ζ)+(ϑ+Dx+D¯μ¯+Fa)⊺Λ∗(ϑ+Dx+D¯μ¯+Fa) (b_0+Bx+ B μ+Ca ) (2 ^*(x- μ)+2 ^* μ+ζ )+ ( +Dx+ D μ+Fa ) ^* ( +Dx+ D μ+Fa ) +(ϑo+Dox+D¯oμ¯+Foa)⊺Λ∗(ϑo+Dox+D¯oμ¯+Foa)+x⊺Mx+μ¯⊺M¯μ¯+a⊺Ra + ( _o+D_ox+ D_o μ+F_oa ) ^* ( _o+D_ox+ D_o μ+F_oa )+x Mx+ μ M μ+a Ra +(ϑo+Dox+D¯oμ¯+Foa)⊺∫ℝd∫ℝp2(Γ∗−Λ∗)(ϑo+Dox′+D¯oμ¯+Foa)∗(a|t,x′,μ)aμ(dx′) + ( _o+D_ox+ D_o μ+F_oa ) _R^d _R^p2( ^*- ^*) ( _o+D_ox + D_o μ+F_oa ) π^*(a|t,x ,μ)daμ(dx ) = = γlog∗(a|t,x,μ)+κ(t,x,μ). γ π^*(a|t,x,μ)+κ(t,x,μ). (5.20) By comparing both sides of the above equality, the fixed point satisfies the form of log∗(a|t,x,μ)=−12(a−m(t,x,μ))⊺Σ−1(t)(a−m(t,x,μ))−12log((2π)pdet(Σ(t))), π^*(a|t,x,μ)=- 12 (a-m(t,x,μ) ) ^-1(t) (a-m(t,x,μ) )- 12 ((2π)^p det( (t)) ), where m(t,x,μ)=K(t)(x−μ¯)+K¯(t)μ¯+K0(t)m(t,x,μ)=K(t)(x- μ)+ K(t) μ+K_0(t). Comparing the coefficients of terms a and a⊺(⋅)a (·)a on both sides of the above equality, we have that Σ(t) (t), K(t)K(t), K¯(t) K(t) and K0(t)K_0(t) satisfy Σ =−γ2Ut∗,K=2ΣγSt∗,K¯=2Σγ(Zt∗+Fo⊺(Γ∗−Λ∗)FoK¯), =- γ2U_t^*,K= 2 γS_t^*,\; K= 2 γ (Z_t^*+F_o ( ^*- ^*)F_o K ), K0 K_0 =Σγ(Yt∗+2Fo⊺(Γ∗−Λ∗)F0K0). = γ (Y_t^*+2F_o ( ^*- ^*)F_0K_0 ). Therefore, we verify that ∗ π^* is Gaussian in the form of (5.4). After straightforward calculations, we see that J∗J^* satisfies the dynamic programming equation (3.8) if and only if Var(μ)((Λ∗)′+D⊺Λ∗D+Do⊺Λ∗Do+B⊺Λ∗+Λ∗B−γ2K⊺Σ−1K+2(St∗)⊺K+M) Var(μ) (( ^*) +D ^*D+D_o ^*D_o+B ^*+ ^*B- γ2K ^-1K+2(S_t^*) K+M ) +μ¯⊺((Γ∗)′+(D+D¯)⊺Λ∗(D+D¯)+(Do+D¯o)⊺Γ∗(Do+D¯o)+(B+B¯)⊺Γ∗+Γ(B+B¯) + μ (( ^*) + (D+ D) ^* (D+ D)+(D_o+ D_o) ^*(D_o+ D_o)+(B+ B) ^*+ (B+ B) −γ2K¯⊺Σ−1K¯+K¯⊺Fo⊺(Γ∗−Λ∗)FoK¯+2(Zt∗)⊺K¯+M+M¯)μ¯ - γ2 K ^-1 K+ K F_o ( ^*- ^*)F_o K+2(Z_t^*) K+M+ M ) μ +μ¯⊺(ζ′+(B+B¯)⊺ζ+2Γ∗b0+2(D+D¯)⊺Λ∗ϑ+2(Do+D¯o)⊺Λ∗ϑo+2(FC+(D+D¯)⊺Λ∗F + μ (ζ +(B+ B) ζ+2 ^*b_0+2(D+ D) ^* +2(D_o+ D_o) ^* _o+2 (FC+(D+ D) ^*F +(Do+D¯o)⊺Γ∗Fo)K0+K¯⊺Yt⊺+K¯⊺(−γΣ−1+Fo⊺(Γ∗−Λ∗)Fo)K0+O) +(D_o+ D_o) ^*F_o )K_0+ K Y_t + K (-γ ^-1+F_o ( ^*- ^*)F_o )K_0+O ) +(χ∗)′+ϑ⊺Λ∗ϑ+ϑo⊺Γ∗ϑo+b0⊺ζ∗+(Yt)⊺K0−γ2K0⊺Σ−1K0+K0⊺Fo⊺(Γ∗−Λ∗)FoK0 +(χ^*) + ^* + _o ^* _o+b_0 ζ^*+ (Y_t ) K_0- γ2K_0 ^-1K_0+K_0 F_o ( ^*- ^*)F_oK_0 +γ2log((2π)pdet(Σ))=0. + γ2 ((2π)^p det( ) )=0. Setting the coefficients of the terms Var(μ)(⋅) Var(μ)(·), μ¯⊺(⋅)μ¯ μ (·) μ, μ¯ μ to be zero, we arrive at the system of ODEs for Λ∗(t),Γ∗(t),ζ∗(t) ^*(t), ^*(t),ζ^*(t) and χ∗(t)χ^*(t). It is known that, see e.g. Wonham (1968), Yong (2013), under the condition (H), for some δ>0δ>0, the matrix Riccati equations (5.7)-(5.11) admit the unique solution (Λ∗,Γ∗)( ^*, ^*) valued in symmetric negative definite matrices. Given the existence of (Λ∗,Γ∗)( ^*, ^*), we also have the existence of solution to the system of linear ODEs for (ζ∗,χ∗)(ζ^*,χ^*). Finally, we verify Assumption 3.6 (i). Denote ¯:= h:= ∫ℝd∫ℝpa(a|x)aμ(dx)=μ,e[a], _R^d _R^pa h(a|x)daμ(dx)=E^e_μ, h[a h], Var()(Λ∗):= Var( h)( ^*):= ∫ℝd∫ℝp(a−¯)⊺Λ∗(a−¯)(a|x)aμ(dx)=μ,e[(a)⊺Λ∗a]−μ,e[a]⊺Λ∗μ,[a]. _R^d _R^p (a- h ) ^* (a- h ) h(a|x)daμ(dx)=E^e_μ, h[(a h) ^*a h]-E^e_μ, h[a h] ^*E_μ, h[a h]. By (3.12), we have that ℋ(t,μ,) H(t,μ, h) =Var()(Ut∗)+¯⊺Vt∗¯+∫ℝd∫ℝpa⊺((St∗)⊺(x−μ¯)+(Zt∗)⊺μ¯+Yt∗)(a|x)aμ(dx) = Var( h)(U_t^*)+ h V_t^* h+ _R^d _R^pa ((S_t^*) (x- μ)+(Z_t^*) μ+Y_t^* ) h(a|x)daμ(dx) +G(t,μ), \;\;\;+G(t,μ), where G is independent of h. For every x∈ℝdx ^d, we denote θ(⋅|x) h_θ(·|x) as the law of the interpolated random variable ϕθ(x,U)=(1−θ)ϕ1(x,U)+θϕ0(x,U) _ h^θ(x,U)=(1-θ) _ h_1(x,U)+θ _ h_0(x,U). Thus ℋ(t,μ,θ) H(t,μ, h_θ) can be written in terms of ϕθ(x,U) _ h^θ(x,U) that ℋ(t,μ,θ) H(t,μ, h_θ) =e[(ϕθ(ξ,U)−e[ϕθ(ξ,U)])⊺Ut∗(ϕθ(ξ,U)−e[ϕθ(ξ,U)])] =E^e [ ( _ h^θ(ξ,U)-E^e[ _ h^θ(ξ,U)] ) U_t^* ( _ h^θ(ξ,U)-E^e[ _ h^θ(ξ,U)] ) ] +e[ϕθ(ξ,U)]⊺Vt∗e[ϕθ(ξ,U)] +E^e [ _ h^θ(ξ,U)] V_t^*E^e[ _ h^θ(ξ,U) ] +e[ϕθ(ξ,U)⊺((St∗)⊺(ξ−μ¯)+(Zt∗)⊺μ¯+Yt∗)]+G(t,μ) +E^e[ _ h^θ(ξ,U) ((S_t^*) (ξ- μ)+(Z_t^*) μ+Y_t^* ) ]+G(t,μ) ≥(1−θ)ℋ(t,μ,0)+θℋ(t,μ,1), ≥(1-θ) H(t,μ, h_0)+θ H(t,μ, h_1), which implies that ℋ(t,μ,) H(t,μ, h) is displacement concave in view of Ut∗⪯0U_t^* 0 and Vt∗⪯0V_t^* 0. By Villani (2009), ℰ(t,μ,) E(t,μ, π) is also displacement concave. We thus conclude that ℋγ(t,μ,) H^γ(t,μ, h) is displacement concave in h and Assumption 3.6 (i) holds. ∎ Acknowledgement: X. Wei is supported by National Natural Science Foundation of China grant under no.12201343 and no.12571509. X. Yu is supported by the Hong Kong RGC General Research Fund (GRF) under grant no. 15211524. Appendix A Cooperative Mean-Field N-Agent Game Here, we recall the cooperative N-agent game that is coordinated by a social planner in the continuous-time entropy regularized setting. The state Xtj,,X_t^j, D, π of agent j∈1,…,Nj∈\1,…,N\ satisfies the SDE dXsj,, dX_s^j, D, π =b(s,Xsj,,,μsN,,,aδ(s)j,,)ds+σ(s,Xsj,,,μsN,,,aδ(s)j,)dWsj =b(s,X_s^j, D, π, _s^N, D, π,a_δ(s)^j, D, π)ds+σ(s,X_s^j, D, π, _s^N, D, π,a_δ(s)^j, π)dW_s^j (A.1) +σo(s,Xsj,,,μsN,,,aδ(s)j,,)dBs,Xtj,,=xj, \;\;\;\;\;+ _o(s,X_s^j, D, π, _s^N, D, π,a_δ(s)^j, D, π)dB_s,\;X_t^j, D, π=x^j, where W1,…,WNW^1,…,W^N are independent Brownian motions, and independent of B, and μsN,,=1N∑j=1NδXsj,, _s^N, D, π= 1N _j=1^N _X_s^j, D, π is the empirical measure, and aδ(s)j,∼(⋅|δ(s),Xδ(s)j,,,μsN,,)a_δ(s)^j, π π(·|δ(s),X_δ(s)^j, D, π, _s^N, D, π) stands for the discretely sampled actions. The expected accumulated reward for the agent j is e[∫tTe−βs(r(Xsj,,,μsN,,,aδ(s)j,)+γE(δ(s),Xj,,δ(s),μδ(s)N,,))s+g(XTj,,,μTN,,)]. ^e [ _t^Te^-β s (r(X_s^j, D, π, _s^N, D, π,a_δ(s)^j, π)+γ E_ π(δ(s),X_j, D,δ(s) π, _δ(s)^N, D, π) )ds+g(X_T^j, D, π, _T^N, D, π) ]. The learning procedure for the cooperative N-agent game is as follows. At each time s∈[t,T]s∈[t,T], each agent j observes the empirical measure μsN,, _s^N, D, π and takes the action aδ(s)j,,a_δ(s)^j, D, π according to the policy :[0,T]×ℝd×2(ℝd)→() π:[0,T]×R^d× P_2(R^d)→ P( A) assigned by the social planner. His state Xsj,,X_s^j, D, π evolves according to (A.1) and he will receive an individual reward r(Xsj,,,μsN,,,aδ(s)j,,)r(X_s^j, D, π, _s^N, D, π,a_δ(s)^j, D, π) via the interaction with the environment. At the social planner’s level, she selects the policy π, assigns it to all agents and observes the evolution of the empirical measure μsN,, _s^N, D, π over time, and obtains the aggregated reward 1N∑j=1Nr(Xsj,,,μsN,,,aδ(s)j,) 1N _j=1^Nr(X_s^j, D, π, _s^N, D, π,a_δ(s)^j, π) so as to ultimately maximize the overall aggregated reward. As N→+∞N→+∞, this formulation leads to an exploratory MFC problem as described in Section 2.3. Appendix B Heuristical Derivation of Relaxed Control Formulation In this section, we discuss the corresponding relaxed formulation of MFC with controlled common noise. Let c∞(ℝd×ℝm×ℝn) C_c^∞(R^d×R^m×R^n) denote the set of infinitely differentiable function ϕ:ℝd×ℝm×ℝn→ℝφ:R^d×R^m×R^n with compact set, and let DϕDφ and D2ϕD^2φ denote gradient and Hessian of ϕφ, respectively. Let σi _i and σo,i _o,i, 1≤i≤d1≤ i≤ d, denote i-th row of σ and σo _o, respectively. Define the infinitesimal generator Lsa,x,μϕ L_s^a,x,μφ =(b(s,x,μ,a)⊺,0m,0n)Dϕ+12Tr((σoIm0m×n0n×mIn)(σ⊺Im0m×nσo⊺0n×mIn)D2ϕ) =(b(s,x,μ,a) ,0_m,0_n)Dφ+ 12 Tr ( ( matrixσ& _o\\ I_m&0_m× n\\ 0_n× m&I_n matrix ) ( matrixσ &I_m&0_m× n\\ _o &0_n× m&I_n matrix )D^2φ ) =b(s,x,μ,a)⊺Dxϕ(x,,)+12Tr((σ⊺+σoσo⊺)(s,x,μ,a)Dxxϕ(x,,)+Dϕ(x,,) =b(s,x,μ,a) D_xφ(x, w, b)+ 12 Tr ( (σ + _o _o )(s,x,μ,a)D_xφ(x, w, b)+D_ w wφ(x, w, b) +Dϕ(x,,)+2σ(s,x,μ,a)Dxϕ(x,,)+2σo(s,x,μ,a)Dxϕ(x,,)). \;\;\;+D_ b bφ(x, w, b)+2σ(s,x,μ,a)D_x wφ(x, w, b)+2 _o(s,x,μ,a)D_x bφ(x, w, b) ). For any ϕ∈c∞(ℝd×ℝm×ℝn)φ∈ C_c^∞(R^d×R^m×R^n), we define the generator associated with π: Ls,x,μϕ(x,,)=∫Lsa,x,μϕ(x,,)(a|s,x,μ)aL π,x,μ_sφ(x, w, b)= _ AL_s^a,x,μφ(x, w, b) π(a|s,x,μ)da. Heuristically, it is the limit of the infinitesimal generator of the dynamics (2.10) because the action is sampled from π independent of W and B. The rest is to construct a triplet (X,W,B)(X,W,B) defined on (Ωe,ℱe,ℙe)( ^e, F^e,P^e) corresponding to the generator L,x,μϕ(x,,)L π,x,μφ(x, w, b), where (X,W,B)(X,W,B) satisfy ⟨dXs,dXs⟩ dX_s,dX_s =∫(σσ⊺+σoσo⊺)(s,Xs,μs,a)(a|s,Xs,μs)a, = _ A (σ + _o _o )(s,X_s, _s,a) π(a|s,X_s, _s)da, (B.1) ⟨dXs,dWs⟩ dX_s,dW_s =σ(s,Xs,μs)ds,⟨dXs,dBs⟩=σo,(s,Xs,μs)ds, = _ π(s,X_s, _s)ds,\; dX_s,dB_s = _o, π(s,X_s, _s)ds, (B.2) ⟨dWs,dWs⟩ dW_s,dW_s =Imds,⟨dBs,dBs⟩=Inds,⟨dWs,Bs⟩=. =I_mds,\; dB_s,dB_s =I_nds,\; dW_s,B_s = 0. (B.3) In addition to B and W, we add two other d-dimensional Brownian motions B¯ B and W¯ W, which are defined on (Ω2,ℱ2,ℙ2)( ^2, F^2,P^2) and are independent of B and W. Recall from (2.5) that cov(σ) cov_ π(σ) and cov(σo) cov_ π( _o) are positive semidefinite. Hence there exist matrices denoted by std(σ) std_ π(σ) and std(σo) std_ π( _o) such that cov(σ)=std(σ)std(σ)⊺ cov_ π(σ)= std_ π(σ) std_ π(σ) and cov(σo)=std(σo)std(σo)⊺ cov_ π( _o)= std_ π( _o) std_ π( _o) , and we have dXs dX_s π =b(s,Xs,μs)ds+σ(s,Xs,μs)dWs+σo,(s,Xs,μs)dBs =b_ π(s,X_s π, _s π)ds+ _ π(s,X_s π, _s π)dW_s+ _o, π(s,X_s π, _s π)dB_s (B.4) +std(σ)(s,Xs,μs)dW¯s+std(σo)(s,Xs,μs)dB¯s. + std_ π(σ)(s,X_s π, _s π)d W_s+ std_ π( _o)(s,X_s π, _s π)d B_s. It is readily seen that (B.4) satisfies (B.1)-(B.3), and hence (B.4) corresponds to the controlled martingale problem with the generator L,x,μϕ(x,,)L π,x,μφ(x, w, b). References Bender and Thuan (2024) C. Bender and N. T. Thuan (2024): On the grid-sampling limit SDE. Preprint, available at arXiv:2410.07778. Bo et al. (2025) L. Bo, Y. Huang and X. Yu (2025): On optimal tracking portfolio in incomplete markets: The reinforcement learning approach. SIAM Journal on Control and Optimization. 63(1), 321-348. Buckdahn et al. (2017) R. Buckdahn, J. Li, S. Peng, C. Rainer. Mean-field stochastic differential equations and associated PDEs. Annals of Probability. 45(2):824-878. Buckdahn et al. (2021) R. Buckdahn, Y. Chen. and J. Li (2021): Partial derivative with respect to the measure and its application to general controlled mean-field systems. Stochastic Processes and their Applications. 134: 265-307. Carmona et al. (2013) R. Carmona, F. Delarue and A. Lachapelle (2013): Control of McKean-Vlasov dynamics versus mean field games. Mathematics and Financial Economics. 7, 131-166. Carmona and Delarue (2018a) R. Carmona and F. Delarue (2018a): Probabilistic Theory of Mean Field Games with Applications, Vol I. Springer. Carmona and Delarue (2018b) R. Carmona and F. Delarue (2018b): Probabilistic Theory of Mean Field Games with Applications, Vol I. Springer. Carmona and Laurière (2025) R. Carmona and M. Laurière (2025): Reconciling Discrete-Time Mixed Policies and Continuous-Time Relaxed Controls in Reinforcement Learning and Stochastic Control. Preprint, available at arXiv:2504.21793. Carmona et al (2023) R. Carmona, M. Laurière. and Z. Tan. (2023): Model-free mean-field reinforcement learning: mean-field MDP and mean-field Q-learning. Annals of Applied Probability. 33(6B), 5334-5381. Chassagneux et al. (2022) J.F. Chassagneux, D. Crisan, and F. Delarue (2022): A probabilistic approach to classical solutions of the master equation for large population equilibria. Memoirs of the AMS,volume 280. Cheung et al. (2023) H. Cheung, J. Qiu and A. Badescu (2023): A viscosity solution theory of stochastic Hamilton-Jacobi-Bellman equations in the Wasserstein space. Preprint, available at arXiv:2310.14446. Conforti et al. (2023) G. Conforti, A. Kazeykina, Z. Ren (2023): Game on random environment, mean-field Langevin system, and neural networks. Mathematics of Operations Research. 48(1):78-99. Crisan and McMurray (2018) D. Crisan and E. McMurray (2018): Smoothing properties of McKean–Vlasov SDEs. Probability Theory and Related Fields, 171:97–148. Dai et al. (2023a) M. Dai, Y. Dong and Y. Jia (2023): Learning equilibrium mean-variance strategy. Mathematical Finance. 33(4), 1166-1212. Dai et al. (2023b) M. Dai, Y. Dong, Y. Jia and X. Y. Zhou (2023): Data-driven Merton’s strategies via policy randomization. Preprint, available at arXiv:2312.11797. Djete et al. (2022) M. F. Djete, D. Possamaï and X. Tan (2022): McKean–Vlasov optimal control: the dynamic programming principle. The Annals of Probability. 50(2):791-833. Dong (2024) Y. Dong (2024): Randomized optimal stopping problem in continuous time and reinforcement learning algorithm. SIAM Journal on Control and Optimization. 62(3), 1590-1614. Dupuis and Ellis (2011) P. Dupuis, R. S. Ellis (2011): A weak convergence approach to the theory of large deviations. John Wiley & Sons. Kallenberg (2002) O. Kallenberg(2002): Foundations of Modern Probability. Probability and its Applications (New York). Springer Verlag, New York, second edition. Lacker (2015) D. Lacker (2015): Mean field games via controlled martingale problems: existence of Markovian equilibria. Stochastic Processes and their Applications. 125(7):2856-2894. Frikha et al. (2025) N. Frikha, M. Germain, M. Laurière, H. Pham. and X. Song (2023). Actor-Critic learning for mean-field control in continuous time. Journal of Machine Learning Research. 26(127):1-42. Graber (2016) P. Graber(2016): Linear quadratic mean field type control and mean field games with common noise, with applications to production of an exhaustible resource. Applied Mathematics &\& Optimization. 74, 459-486. Gu et al. (2021) H. Gu, X. Guo, X. Wei and R. Xu (2021): Mean-field controls with Q-learning for cooperative MARL: Convergence and complexity analysis. SIAM Journal on Mathematics of Data Science. 3(4), 1168-1196. Gu et al. (2023) H. Gu, X. Guo, X. Wei and R. Xu (2023): Dynamic programming principles for mean-field controls with learning. Operations Research. 71(4), 1040-1054. Guo et al. (2022) X. Guo, R. Xu and T. Zariphopoulou (2022): Entropy regularization for mean field games with learning. Mathematics of Operations Research. 47(4), 3239-3260. Han et al. (2023) X. Han, R. Wang and X. Y. Zhou (2023): Choquet regularization for continuous-time reinforcement learning. SIAM Journal on Control and Optimization. 61(5), 2777-2801. Huang et al. (2006) M. Huang, R.P. Malhamé, P. E. Caines (2006): Large population stochastic dynamic games closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle. Communications in Information and Systems. 6(3), 221–252. Huang et al. (2025) Y. Huang, M. Li, X. Yu and Z. Zhou (2025): Continuous-time reinforcement learning for optimal switching over multiple regimes. Preprint, available at arXiv:2512.04697. Jia and Zhou (2022a) Y. Jia and X. Y. Zhou (2022a): Policy gradient and actor-critic learning in continuous time and space: Theory and algorithms. Journal of Machine Learning Research. 23, 1-50. Jia and Zhou (2022b) Y. Jia and X. Y. Zhou (2022b): Policy evaluation and temporal-difference learning in continuous time and space: A martingale approach. Journal of Machine Learning Research. 23, 1-55. Jia and Zhou (2023) Y. Jia and X. Y. Zhou (2023): q-learning in continuous time. Journal of Machine Learning Research. 24, 1-61. Jia et al. (2025) Y. Jia, D. Ouyang and Y. Zhang(2025): Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning. SIAM Journal on Control and Optimization, forthcoming. Jia (2026) Y. Jia (2026): Continuous-time risk-sensitive reinforcement learning via quadratic variation penalty. Applied Mathematics &\& Optimization, forthcoming. Lasry and Lions (2007) J. M. Lasry and P. L. Lions (2007): Mean field games. Japanese Journal of Mathematics. 2(1), 229-260 Liang et al. (2024) H. Liang, Z. Chen and K. Jing (2024): Actor-critic reinforcement learning algorithms for mean field games in continuous time, state and action spaces. Applied Mathematics and Optimization. 89(3): 72. Lions (2006) P. L. Lions (2006): Cours au collège de france: Théorie des jeux à champ moyens. Audio Conference. McCann (1997) R. J. McCann (1997): A convexity principle for interacting gases. Advances in Mathematics. 128(1): 153-179. Motte and Pham (2022) M. Motte and H. Pham (2022): Mean-field Markov decision processes with common noise and open-loop controls. Annals of Applied Probability. 32(2):1421-1458. Pham and Wei (2017) H. Pham. and X. Wei (2017): Dynamic programming for optimal control of stochastic McKean–Vlasov dynamics. SIAM Journal on Control and Optimization. 55(2), 1069-1101. Ren et al. (2026) Z. Ren, X. Wei, X. Yu and X. Y. Zhou (2026): Continuous-time q-learning for mean-field control with common noise, part-I: q-learning algorithms. Working paper. Stroock and Varadhan (1997) D. Stroock and S. Varadhan (1997): Multidimensional diffusion processes, volume 233 of Grundlehren der mathematischen Wissenschaften. Springer–Verlag Berlin Heidelberg. Szpruch et al. (2024) L. Szpruch, T. Treetanthiploet and Y. Zhang (2024): Optimal scheduling of entropy regularizer for continuous-time linear-quadratic reinforcement learning. SIAM Journal on Control and Optimization. 62(1):135–166. Villani (2009) C. Villani (2009): Optimal transport: old and new. Berlin: Springer. Wang et al. (2020) H. Wang, T. Zariphopoulou and X. Y. Zhou (2020): Reinforcement learning in continuous time and space: A stochastic control approach. Journal of Machine Learning Research. 21(1):8145-8178. Wang et al. (2023) Wang, B., X. Gao and L. Li (2023): Reinforcement learning for continuous-time optimal execution: Actor-Critic algorithm and error analysis. Finance and Stochastics, 30, 597-655. Watkins and Dayan (1992) C. Watkins and P. Dayan (1992): Q-learning. Machine Learning. 8(3):279-292. Wei and Yu (2025) X. Wei and X. Yu (2025): Continuous-time q-learning for mean-field control problems. Applied Mathematics and Optimization. 91: 10. Wei et al. (2024) X. Wei, X. Yu and F. Yuan (2024): Unified continuous-time q-learning for mean-field game and mean-field control problems. Preprint, available at arXiv:2407.04521. Wonham (1968) W. Wonham (1968): On a matrix Riccati equation of stochastic control. SIAM Journal on Control and Optimization, 6(4):681-697. Yong (2013) J. Yong (2013): Linear-quadratic optimal control problems for mean-field stochastic differential equations. SIAM journal on Control and Optimization. 51(4):2809-38. Zhou et al. (2024) J. Zhou, N. Touzi, and J. Zhang (2024): Viscosity solutions for HJB equations on the process space: Application to mean field control with common noise. Preprint, available at arXiv:2401.04920. (52)