Paper deep dive
A Controllability Perspective on Steering Follow-the-Regularized-Leader Learners in Games
Heling Zhang, Siqi Du, Roy Dong
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/31/2026, 1:28:43 AM
Summary
This paper investigates the controllability of multi-agent systems where learners follow continuous-time Follow-the-Regularized-Leader (FTRL) dynamics. The authors model the system as a nonlinear control problem on the relative interior of a simplex, where a single controller influences the learners' strategies through their own mixed strategy without modifying the game's payoff structure. They provide necessary and sufficient conditions for controllability in two-player games and sufficient conditions for multi-learner interactions using geometric control theory, Lie-algebraic rank tests, and neutralization strategies.
Entities (5)
Relation Signals (3)
FTRL → modeledas → Nonlinear Control System
confidence 95% · Viewing the learners' dynamics as a nonlinear control system evolving on the relative interior of a simplex
Geometric Control Theory → providestoolsfor → Controllability Analysis
confidence 95% · To analyze state reachability on the probability simplex, we utilize tools from geometric control theory.
Controller → steers → FTRL Learners
confidence 90% · we ask when the controller can steer the learners to a target state
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Follow-the-regularized-leader (FTRL) algorithms have become popular in the context of games, providing easy-to-implement methods for each agent, as well as theoretical guarantees that the strategies of all agents will converge to some equilibrium concept (provided that all agents follow the appropriate dynamics). However, with these methods, each agent ignores the coupling in the game, and treats their payoff vectors as exogenously given. In this paper, we take the perspective of one agent (the controller) deciding their mixed strategies in a finite game, while one or more other agents update their mixed strategies according to continuous-time FTRL. Viewing the learners' dynamics as a nonlinear control system evolving on the relative interior of a simplex or product of simplices, we ask when the controller can steer the learners to a target state, using only its own mixed strategy and without modifying the game's payoff structure. For the two-player case we provide a necessary and sufficient criterion for controllability based on the existence of a fully mixed neutralizing controller strategy and a rank condition on the projected payoff map. For multi-learner interactions we give two sufficient controllability conditions, one based on uniform neutralization and one based on a periodic-drift hypothesis together with a Lie-algebra rank condition. We illustrate these results on canonical examples such as Rock-Paper-Scissors and a construction related to Brockett's integrator.
Tags
Links
- Source: https://arxiv.org/abs/2603.27081v1
- Canonical: https://arxiv.org/abs/2603.27081v1
Trouble viewing inline? Open PDF directly →
Full Text
73,563 characters extracted from source content.
Expand or collapse full text
A Controllability Perspective on Steering Follow-the-Regularized-Leader Learners in Games Heling Zhang, Siqi Du, and Roy Dong This work was supported by the National Science Foundation under Grant CCF 2236484.H. Zhang is with the Department of Electrical and Computer Engineering at Illinois Grainger Engineering, University of Illinois Urbana-Champaign. (email: hzhng120@illinois.edu).S. Du and R. Dong are with the Department of Industrial and Enterprise Systems Engineering at Illinois Grainger Engineering, University of Illinois Urbana-Champaign. (emails: siqidu3,roydong@illinois.edu). Abstract Follow-the-regularized-leader (FTRL) algorithms have become popular in the context of games, providing easy-to-implement methods for each agent, as well as theoretical guarantees that the strategies of all agents will converge to some equilibrium concept (provided that all agents follow the appropriate dynamics). However, with these methods, each agent ignores the coupling in the game, and treats their payoff vectors as exogenously given. In this paper, we take the perspective of one agent (the controller) deciding their mixed strategies in a finite game, while one or more other agents update their mixed strategies according to continuous-time FTRL. Viewing the learners’ dynamics as a nonlinear control system evolving on the relative interior of a simplex or product of simplices, we ask when the controller can steer the learners to a target state, using only its own mixed strategy and without modifying the game’s payoff structure. For the two-player case we provide a necessary and sufficient criterion for controllability based on the existence of a fully mixed neutralizing controller strategy and a rank condition on the projected payoff map. For multi-learner interactions we give two sufficient controllability conditions, one based on uniform neutralization and one based on a periodic-drift hypothesis together with a Lie-algebra rank condition. We illustrate these results on canonical examples such as Rock-Paper-Scissors and a construction related to Brockett’s integrator. I Introduction Follow-the-regularized-leader (FTRL) algorithms have become standard techniques for handling complicated strategic interactions in multi-agent environments [11, 17]. From a practical standpoint, there is a compelling case for deploying FTRL in real-world systems: it provides an easy-to-implement, computationally efficient update rule for each agent that relies strictly on locally observed payoff feedback [3, 30]. Furthermore, it is supported by a rich theoretical literature ensuring that, provided all agents follow appropriate dynamics, the population’s strategies will converge to established equilibrium concepts [16, 19, 8]. As a result, FTRL is frequently utilized to design autonomous agents navigating complex, repeated interactions. However, this decentralized simplicity comes with a fundamental structural assumption. By employing FTRL, each agent inherently ignores the strategic coupling of the underlying game. Rather than recognizing the interaction as a closed-loop feedback system where their own actions influence the future behavior of others, the learner treats their incoming sequence of payoff vectors as exogenously given [26, 29]. To an FTRL agent, the multi-agent environment is perceived merely as a fluctuating, open-loop landscape to be optimized against [12]. If an entire population of interacting agents blindly adopts this uncoupled learning paradigm, it can introduce new systemic vulnerabilities and opportunities for strategic exploitation [6, 14]. Suppose a single, sophisticated agent is aware of this behavioral structure. Recognizing that their opponents are predictably driven by FTRL updates, this model-aware agent no longer needs to myopically optimize their immediate payoff [22, 24]. Instead, they can actively shape the exogenous payoffs observed by the learners, treating the evolving mixed strategies of the population as a dynamical system to be manipulated [7]. The natural question then arises: where could they steer the system? The practical implications of such reachability are significant. For instance, in automated financial markets or algorithmic pricing, a strategic participant could manipulate FTRL-driven competitors to drive the market into profitable, out-of-equilibrium pricing configurations [6]. Similarly, in intelligent infrastructure, an adversarial or central node could steer independent routing protocols to induce targeted congestion or enforce globally optimal traffic flows [32]. These examples motivate a controllability viewpoint: when the game and learning dynamics satisfy suitable conditions, a model-aware strategic player may be able to steer the induced learning dynamics toward selected states. This paper studies this steering problem as a controllability problem. We consider a finite game with a distinguished controller and one or more learners that follow continuous-time FTRL. Interpreting the controller’s mixed strategy as the control input and the learners’ mixed strategies as the state yields a nonlinear control system evolving on the relative interior of the simplex, or a product of simplices if there are multiple learners. The restriction to the relative interior is based on the following consideration: under common learning flows (e.g., replicator dynamics), the boundary is invariant, so interior initial conditions cannot be driven to the boundary in finite time [15]. Closest to our work are recent papers on steering learners in games, but the control channel, objective, and analysis are different. In contrast to work that steers no-regret learners through external payments or dynamic incentives, we keep the game and payoff structure fixed and allow the controller to act only through its own mixed strategy. In contrast to repeated-game formulations that seek to drive a learner toward a Stackelberg outcome while learning unknown payoffs, we study a model-based controllability question for continuous-time FTRL dynamics on the relative interior of the simplex. Our contribution is a control-theoretic formulation of this steering problem, together with an exact two-player controllability criterion and two sufficient controllability conditions for multi-learner interactions based on neutralization, periodic drift, and Lie-algebraic rank tests. The contributions of this paper are threefold. First, we formulate steering continuous-time FTRL learners as a nonlinear controllability problem on the relative interior of a simplex or product of simplices, with the controller’s mixed strategy serving as the admissible control input. Second, in the two-player case we derive an exact controllability criterion based on a fully mixed neutralizing strategy and a projected-payoff rank condition. Third, in the multi-learner case we derive two sufficient controllability conditions, one via uniform neutralization and one via periodic drift together with a Lie-algebra rank condition. We illustrate our results on canonical examples such as Rock-Paper-Scissors and a construction related to Brockett’s integrator. I Related Works I-A No-Regret Learning in Multi-Agent Settings In multi-agent settings, no-regret learning rules like FTRL do not directly consider the coupling between agents [30]. Rather than explicitly modeling their opponents, learners update their strategies based solely on their own historical payoff observations. They effectively treat the game’s endogenous strategic interactions as exogenous environmental signals. Even so, no-regret learning remains attractive because it is decentralized, requires only limited feedback, and gives each agent a clear long-run performance guarantee against the realized behavior of others. The main technical nuance is that these guarantees are typically about average behavior rather than the last iterate: time-averaged play can exhibit meaningful convergence properties even when the period-by-period strategies continue to move. This makes no-regret dynamics useful when long-run average performance is the relevant objective, but less satisfactory when one needs stable pointwise convergence of strategies. To address this non-convergence, an active branch of algorithmic research modifies the learning rules to achieve last-iterate convergence, utilizing methods such as optimistic gradient extrapolation [31], continuous action perturbations [1], or black-box reductions [8]. However, these approaches require the ability to dictate the agents’ internal algorithms. Our work considers an alternative setting: assuming the agents’ standard FTRL dynamics are fixed, we investigate whether a single strategic player can steer the resulting system purely through their own actions. I-B Strategizing Against Learners Because FTRL dynamics do not explicitly account for the closed-loop effects in a multi-agent setting, a large body of literature focuses on how to strategically exploit FTRL agents. This research usually frames the interaction as an asymmetric dynamic game, where a sophisticated optimizer adjusts its strategy to take advantage of a naive learner over time [24, 14]. The ultimate goal here is almost entirely payoff-oriented. Researchers frequently rely on Stackelberg formulations to compute commitment strategies that extract far more cumulative utility than a standard simultaneous-play Nash equilibrium would allow. Recent work also shows that this viewpoint faces strong computational barriers in general settings: unless P=NP, no polynomial-time optimizer can compute a near-optimal strategy against a learner running a standard no-regret algorithm such as multiplicative weights [4]. While this work offers tight bounds on utility extraction and regret minimization, it treats the learner’s evolving state simply as a vehicle for reaching a specific scalar payoff limit. This framework does not consider whether arbitrary coordinates on the non-Euclidean simplex can actually be reached, nor does it offer operational guidance for transient physical control. For instance, if a controller needs a population to temporarily adopt a specific, non-equilibrium mixed strategy just to avoid a catastrophic failure in a physical routing network, scalar payoff bounds are practically useless. By divorcing utility maximization from the mechanics of nonlinear state reachability, this literature largely ignores the underlying dynamics and state of the system. I-C Steering via Mechanism Design Another active area of research investigates how central planners can guide learning populations toward stable or socially optimal states [10]. A common approach is to introduce dynamic incentives or side payments directly into the learners’ utility functions to mitigate the oscillatory nature of FTRL dynamics [32]. Although incentive-based steering is effective, it requires a mediator with the administrative authority to actively modify the underlying payoff structure. In many decentralized or competitive settings, a standard player lacks the ability to unilaterally alter an opponent’s utility function [23]. Thus, while mechanism design demonstrates that a system can be guided by adjusting incentives, our work investigates whether learners can be steered to target states purely through strategic play within a rigidly fixed game topology. I-D Game Dynamics and Geometric Control To analyze state reachability on the probability simplex, we utilize tools from geometric control theory. Specifically, we rely on the Chow-Rashevskii Theorem [13, 28] to establish local controllability, and we use [5, Theorem 1] to establish global controllability. For broader treatments of these standard tools, one can refer to [20] or [2]. The closest mathematical analogues to our approach appear in evolutionary biology and population dynamics. In those fields, continuous-time game equations like the replicator dynamics have been mapped onto the cotangent bundle of the probability simplex, allowing researchers to establish local controllability [27]. However, these prior studies typically model the control input as an exogenous environmental parameter or mutation rate [18], rather than the actions of a single participant. Our work adapts this methodology to matrix games by defining the control input strictly as the mixed strategy of a participating player. By introducing the concept of a neutralizing strategy to counteract the system’s strong autonomous drift, we are able to map the algebraic tools of geometric control directly onto the primal manifold of the multi-agent interaction. In summary, the existing literature influences learning dynamics in three main ways: by optimizing against learners for payoff, by modifying their incentives, or by introducing exogenous control signals. Our setting is different. We ask whether a standard participating player, without changing the learners’ utility functions or internal algorithms, can use only its own mixed strategy to reach arbitrary interior states through the continuous-time FTRL dynamics. This perspective leads naturally to a control theoretic formulation of the problem and to corresponding controllability criteria. I Notation For any N∈ℕN , we use [N][N] to represent 1,…,N\1,…,N\. Let x1,…,xnx_1,…,x_n be a sequence of n objects; we denote by x−ix_-i the sequence with xix_i removed, i.e. x−i=(x1,…,xi−1,xi+1,…,xn)x_-i=(x_1,…,x_i-1,x_i+1,…,x_n). Let Δn:=(x1,…,xn)|xi≥0,∑i=1nxi=1 _n:=\(x_1,…,x_n)|x_i≥ 0, _i=1^nx_i=1\ be the n-dimensional probability simplex. For a set S, we denote its relative interior by S∘S , and its closure by S¯ S. Let n1_n be the n-dimensional all-ones vector, i.e. =(1,1,…,1)∈ℝn1=(1,1,…,1) ^n. We omit subscripts and write 1 when the dimension can be inferred from the context. We denote the concatenation of vectors x∈ℝnx ^n and y∈ℝmy ^m by x⊕yx y. For a function f:→f:X and a set S⊆S , we use f|Sf|_S to represent the restriction of f to S. For a vector field g(x)g(x), we use etg(x0)e^tg(x_0) to denote the time-t flow of g starting from x0x_0. For a set of vector fields (x)=g1(x),…,gn(x)G(x)=\g_1(x),…,g_n(x)\, we denote the Lie algebra generated by (x)G(x) as Lie()Lie(G). These notations will be further elaborated in the next section when we talk about geometric control. For manifolds M,NM,N and k∈0,1,∞k∈\0,1,∞\, we write Ck(M,N)C^k(M,N) for the space of k times continuously differentiable functions from M to N. When N=ℝN=R, we abbreviate this as Ck(M)C^k(M). For q∈Mq∈ M, TqMT_qM denotes the tangent space of M at q. If X∈C1(M,N)X∈ C^1(M,N), then DX(q):TqM→TX(q)NDX(q):T_qM→ T_X(q)N denotes the differential of X at q. IV Preliminaries IV-A Finite Games In this paper, we focus on finite games, i.e., normal-form games with finite action spaces. Formally, such games are defined as follows. Definition 1 (Finite Games). A finite game consists of the following: 1. a finite set of players indexed by [N][N], 2. for each player i, a finite set of actions [ni][n_i], 3. for each player i, a utility function ri:[n1]×⋯×[nN]→ℝr_i:[n_1]×…×[n_N] that quantifies the player’s preference over the joint action profile of all players. The above definition naturally extends to mixed strategies. A mixed strategy of player i is a vector xi∈Δnix_i∈ _n_i representing a probability distribution over [ni][n_i], where xi(a)x_i(a) denotes the probability assigned to action a∈[ni]a∈[n_i]. For convenience, let i:=[ni]A_i:=[n_i] and let −iA_-i denote the joint action set of all players other than i. Given a mixed strategy profile (x1,…,xN)∈Δn1×⋯×ΔnN(x_1,…,x_N)∈ _n_1×…× _n_N, the expected utility of player i is R(xi,x−i) R(x_i,x_-i) =∑ai∈ixi(ai)∑a−i∈−iri(ai,a−i)∏j∈[N]∖ixj(aj) = _a_i _ix_i(a_i) _a_-i _-ir_i(a_i,a_-i) _j∈[N] \i\x_j(a_j) =⟨xi,pi⟩, = x_i,p_i , where pi=(∑a−1∈−1ri(1,a−i)∏j∈[N]∖ixj(aj)⋮∑a−1∈−1ri(ni,a−i)∏j∈[N]∖ixj(aj))p_i= pmatrix _a_-1 _-1r_i(1,a_-i) _j∈[N] \i\x_j(a_j)\\ \\ _a_-1 _-1r_i(n_i,a_-i) _j∈[N] \i\x_j(a_j)\\ pmatrix is the corresponding payoff vector of player i. IV-B Continuous-Time FTRL Learning in games studies how players adapt their strategies over time in response to payoff feedback generated by the evolving play of others. In this paper, we focus on learning rules in the following form, x=f(y),y˙=g(y,p),x=f(y), y=g(y,p), where x is the mixed strategy, p is the payoff vector and y is an auxiliary state that stores historical information about the environment. One of the most studied families of such learning rules is the FTRL class of algorithms, which is given by x=Q(y),y˙=p,x=Q(y), y=p, where Q(y)=argmaxx′∈Δk⟨x′,y⟩−h(x′)Q(y)= *argmax_x ∈ _k x ,y -h(x ) (1) for some regularizer h. In this paper, we make the following assumptions on h: Assumption 1. We assume 1. h:ℝn→ℝh:R^n is smooth in the sense that h∈C∞h∈ C^∞ on Δn∘ _n , and 2. the Hessian ∇2h(x)∇^2h(x) is positive definite for all x∈Δn∘x∈ _n , where Δn∘ _n is the relative interior of the probability simplex. IV-C Geometric Control In this subsection, we present our main theoretic tools from geometric control. Consider the nonlinear control system Σx:x˙=f(x,u),x∈M,u∈U, _x: x=f(x,u), x∈ M, u∈ U, (2) where M is a connected smooth manifold (in our applications, M will be the relative interior of a simplex or a product of simplices) and U is the set of admissible controls. We assume that for each fixed u∈Uu∈ U the map x↦f(x,u)x f(x,u) defines a smooth vector field on M. For a fixed control value u∈Uu∈ U write fu(⋅):=f(⋅,u)f_u(·):=f(·,u) for the corresponding vector field. We denote by etfu(x0)e^tf_u(x_0) the time-t flow of fuf_u starting from x0x_0, i.e., the unique value at time t of the solution x(⋅)x(·) to x˙=fu(x) x=f_u(x) with initial condition x(0)=x0x(0)=x_0 whenever the solution uniquely exists. If the control signal u(⋅)u(·) is piecewise constant, taking values u1,…,uk∈Uu_1,…,u_k∈ U over successive time intervals of lengths t1,…,tk≥0t_1,…,t_k≥ 0, then the resulting trajectory satisfies x(t1+⋯+tk)=etkfuk∘⋯∘et1fu1(x0),x(t_1+…+t_k)=e^t_kf_u_k … e^t_1f_u_1(x_0), which we sometimes abbreviate x=etkfuk…et1fu1x0x=e^t_kf_u_k… e^t_1f_u_1x_0. Let ℱ:=fu(⋅):u∈UF:=\f_u(·):u∈ U\ denote the family of control vector fields. For x0∈Mx_0∈ M and T≥0T≥ 0, define the attainable set within time T by Ax(≤T,x0) A_x(≤ T,x_0) =etkfuk∘⋯∘et1fu1(x0): = \e^t_kf_u_k … e^t_1f_u_1(x_0): k∈ℕ,ui∈U,ti≥0,∑i=1kti≤T, k ,u_i∈ U,t_i≥ 0, _i=1^kt_i≤ T \, and the (forward) attainable set by Ax(x0):=⋃T>0Ax(≤T,x0)A_x(x_0):= _T>0A_x(≤ T,x_0). Based on these attainable sets, we define three forms of controllability for the system (2): 1. Small-time local controllability (STLC): The system is STLC if for every x∈Mx∈ M and every T>0T>0, the state x is in the interior of the attainable set within time T, i.e., x∈int(Ax(≤T,x))x (A_x(≤ T,x)). 2. Local controllability: The system is locally controllable if for every x∈Mx∈ M, the state x is in the interior of the forward attainable set, i.e., x∈int(Ax(x))x (A_x(x)). 3. Global controllability: The system is globally controllable (or simply controllable) on M if the attainable set from any state is the entire manifold, i.e., Ax(x)=MA_x(x)=M for all x∈Mx∈ M. Throughout this paper, we will invoke controllability results for driftless systems under a mild condition on the set of admissible controls. Definition 2 (Proper Control Set). We call a control set U⊂ℝmU ^m proper if it is convex, compact, and contains the origin in its relative interior. This condition guarantees that one can realize small control variations in all directions within the affine subspace generated by U. In particular, if u0∈Δm∘u_0∈ _m is fully mixed, then the translated set V:=Δm−u0V:= _m-u_0 is proper in the affine subspace v∈ℝm:⟨v,⟩=0\v ^m: v,1 =0\ and satisfies 0∈V∘0∈ V . For C1C^1 vector fields X and Y on M, their Lie bracket [X,Y][X,Y] is the C0C^0 vector field defined (in local coordinates) by [X,Y](x)=DY(x)X(x)−DX(x)Y(x).[X,Y](x)=DY(x)X(x)-DX(x)Y(x). The Lie algebra generated by a family of smooth vector fields G, denoted Lie()Lie(G), is the smallest collection of vector fields containing G and closed under finite linear combinations and Lie brackets. Its evaluation at q∈Mq∈ M is the subspace Lieq():=X(q):X∈Lie()⊆TqM.Lie_q(G):=\X(q):X (G)\ T_qM. Definition 3 (Bracket Generating). We say that a family of vector fields G is bracket generating on M if Lieq()=TqMLie_q(G)=T_qM for every q∈Mq∈ M. We will use (without proof) the following classical facts. Proposition 1 (Chow-Rashevskii). Consider a driftless control-affine system x˙=∑i=1muifi(x) x= _i=1^mu_if_i(x) on a connected manifold M with a proper control set U. If the family of vector fields f1,…,fm\f_1,…,f_m\ is bracket generating on M, then the system is small-time locally controllable (STLC) on M. Proposition 2 (Local Controllability Implies Global Controllability). Consider a control system defined on a connected finite-dimensional smooth manifold M. If the system is locally controllable, then it is globally controllable [5, Theorem 1]. Proposition 3 (Krener’s Theorem). Consider (2) and suppose the vector fields in ℱF are real-analytic (in particular, algebraic). For any q∈Mq∈ M, the attainable set Ax(≤T,q)A_x(≤ T,q) has nonempty interior in M for every T>0T>0 (and hence Ax(q)A_x(q) has nonempty interior) if and only if the system is bracket generating at q, i.e., Lieq(ℱ)=TqMLie_q(F)=T_qM. Remark 1. Proposition 1 is the workhorse for our driftless reductions. Proposition 3 is used as a Lie-rank test for the existence of reachable directions under analytic- ity/algebraicity. Because STLC strictly implies local controllability, Propositions 1 and 2 together guarantee that an everywhere STLC system on a connected manifold is globally controllable. We will also use state equivalence to pass controllability between systems. Consider another control system Σy:y˙=g(y,u),y∈N,u∈U, _y: y=g(y,u), y∈ N, u∈ U, (3) with the same control set U. We say that Σx _x and Σy _y are state equivalent if there exists a diffeomorphism Φ:M→N :M→ N such that for every admissible control u(⋅)u(·) and every solution x(⋅)x(·) of Σx _x, the curve y(t):=Φ(x(t))y(t):= (x(t)) is a solution of Σy _y under the same control u(⋅)u(·). Proposition 4 (State equivalence preserves controllability). If Σx _x and Σy _y are state equivalent, then Σx _x is globally controllable if and only if Σy _y is globally controllable. V Problem Formulation We formulate the steering problem on the relative interior of learners’ strategy space and introduce the projected dual coordinates that will be used throughout the paper. V-A Relative interior of the state space and projected mirror coordinates Consider N learners and one controller. Learner i has action set [ni][n_i], the controller has action set [m][m], and learner i uses the FTRL choice map Qi(yi)=argmaxxi′∈Δni⟨xi′,yi⟩−hi(xi′),Q_i(y_i)= *argmax_x_i ∈ _n_i \ x_i ,y_i -h_i(x_i ) \, where hih_i satisfies Assumption 1. Define ∘ :=Δn1∘×⋯×ΔnN∘, = _n_1 ×·s× _n_N , (4) Hi H_i :=zi∈ℝni:⟨zi,ni⟩=0, =\z_i ^n_i: z_i,1_n_i =0\, PHi P_H_i :=Ini−1ninini⊤, =I_n_i- 1n_i1_n_i1_n_i , H H :=H1×⋯×HN, =H_1×·s× H_N, PH P_H :=diag(PH1,…,PHN). =diag(P_H_1,…,P_H_N). For convenience, write Q Q :=Q1⊕⋯⊕QN, =Q_1 ·s Q_N, (5) ∇h(x) ∇ h(x) :=∇h1(x1)⊕⋯⊕∇hN(xN). =∇ h_1(x_1) ·s ∇ h_N(x_N). For each learner define Mi M_i :=PHi∇hi(Δni∘)⊂Hi, =P_H_i∇ h_i( _n_i )⊂ H_i, (6) M M :=M1×⋯×MN⊂H. =M_1×·s× M_N⊂ H. Our formulation is based upon the following crucial lemma. We state this lemma as follows, but we will delay the proof to the next section. Lemma 1. For each i, the restriction Φi:=Qi|Mi:Mi→Δni∘ _i:=Q_i|_M_i:M_i→ _n_i is a diffeomorphism with inverse Φi−1(xi)=PHi∇hi(xi). _i^-1(x_i)=P_H_i∇ h_i(x_i). Consequently, Φ:=Φ1⊕⋯⊕ΦN:M→X∘ := _1 ·s _N:M→ X is a diffeomorphism with inverse Φ−1(x)=PH∇h(x). ^-1(x)=P_H∇ h(x). Because each QiQ_i is invariant under translations along spannispan\1_n_i\, the joint map Q satisfies Q(y)=Q(PHy)Q(y)=Q(P_Hy). Moreover, for every z∈Mz∈ M, DQ(z)PH=DQ(z),imDQ(z)⊂TQ(z)∘=H.DQ(z)P_H=DQ(z), \,DQ(z)⊂ T_Q(z)X =H. These identities allow us to pass freely between primal and projected dual coordinates, as we will make clear in the following subsections. Remark 2 (Why we work on ∘X ). The controllability results in this paper are stated on the relative interior of the state space ∘X . This is the natural domain of the projected mirror coordinates Φ−1 ^-1, and it is forward invariant for important FTRL dynamics such as replicator dynamics. We therefore study steering on the relative interior only. We do not claim boundary inaccessibility for every regularizer covered by Assumption 1. V-B Two-player case We first consider one learner with n actions and one controller with m actions. Let A∈ℝn×mA ^n× m be the payoff matrix mapping the controller mixed strategy u∈Δmu∈ _m to the learner payoff vector AuAu. The raw FTRL dynamics are x=Q(y),y˙=Au,u∈Δm.x=Q(y), y=Au, u∈ _m. (7) Define H H :=z∈ℝn:⟨z,n⟩=0, =\z ^n: z,1_n =0\, (8) PH P_H :=In−1nnn⊤, =I_n- 1n1_n1_n , M M :=PH∇h(Δn∘), =P_H∇ h( _n ), and let Φ:=Q|M:M→Δn∘ :=Q|_M:M→ _n be the diffeomorphism from Lemma 1. With the projected dual state z:=PHyz:=P_Hy, the system becomes Σz2p:z˙=PHAu,z∈M,u∈Δm. _z^2p: z=P_HAu, z∈ M, u∈ _m. (9) Transporting this dynamics to the relative interior of the probability simplex via Φ yields the induced control system Σx2p:x˙ _x^2p: x =DΦ(Φ−1(x))PHAu =D ( ^-1(x))\,P_HAu (10) =DQ(Φ−1(x))Au,x∈Δn∘,u∈Δm. =DQ( ^-1(x))\,Au, x∈ _n ,\ u∈ _m. Systems (9) and (10) are state equivalent. Throughout the paper, controllability in the two-player case means controllability of (10) on Δn∘ _n , or equivalently of (9) on M. In this paper, we identify necessary and sufficient conditions for the controllability of (10), i.e., when can we steer a single FTRL agent to any desired probability distribution through repeated play of the game? V-C Multi-player case Now consider N learners and one controller. Let pi(x−i,u)p_i(x_-i,u) denote the payoff vector of learner i. Since it is linear in the controller strategy, we write pi(x−i,u)=Ai(x−i)u,i∈[N].p_i(x_-i,u)=A_i(x_-i)u, i∈[N]. Stack the learner states and dual variables as x:=x1⊕⋯⊕xN,y:=y1⊕⋯⊕yN,x:=x_1 ·s x_N, y:=y_1 ·s y_N, and define A(x):=[A1(x−1)⋮AN(x−N)]=[a1(x)⋯am(x)].A(x):= bmatrixA_1(x_-1)\\ \\ A_N(x_-N) bmatrix= [a_1(x)\ ·s\ a_m(x) ]. The raw learning dynamics are xi=Qi(yi),y˙i=Ai(x−i)u,i∈[N],u∈Δm.x_i=Q_i(y_i), y_i=A_i(x_-i)u, i∈[N], u∈ _m. (11) Equivalently, x=Q(y),y˙=A(x)u,x∈∘,u∈Δm.x=Q(y), y=A(x)u, x , u∈ _m. (12) With the projected dual state z:=PHyz:=P_Hy and the diffeomorphism Φ:M→∘ :M , the state-equivalent projected dual dynamics are Σz:z˙=PHA(Φ(z))u,z∈M,u∈Δm. _z: z=P_HA( (z))u, z∈ M, u∈ _m. (13) The induced primal dynamics on the product of simplices are Σx:x˙ _x: x =DΦ(Φ−1(x))PHA(x)u =D ( ^-1(x))\,P_HA(x)u (14) =DQ(Φ−1(x))A(x)u,x∈∘,u∈Δm. =DQ( ^-1(x))\,A(x)u, x ,\ u∈ _m. Again, (13) and (14) are state equivalent. All controllability statements for the multi-learner model refer to (14), while the proofs will often be carried out in the projected dual coordinates (13). In this paper, we identify multiple sufficient conditions for the controllability of (14), i.e., when can we simultaneously steer multiple FTRL agents to desired states through repeated play of the game? VI Main Results VI-A The 2-Player Case We first analyze the setting of a single controller interacting with a single learner. Building on the dynamics established in the problem formulation, our goal is to determine exactly when the controller can steer the learner to any target mixed strategy within the relative interior of the simplex. A key technical challenge in analyzing this setup is the learner’s continuous, state-dependent drift. To address this, we explore how the controller can effectively cancel out this autonomous drift. Doing so allows us to transform the original nonlinear dynamics into a simpler driftless control system, which ultimately reduces the complex question of steerability to a straightforward algebraic rank test. Our first result gives a necessary and sufficient condition for controllability of (10) on Δn∘ _n . The condition is based on the concept of a neutralizing strategy: a mixed strategy such that, if the controller plays it, the learner’s payoff vector is constant (and hence independent of the learner’s current action). Formally: Definition 4 (Neutralizing Strategy). A mixed strategy u0∈Δmu_0∈ _m is neutralizing if Au0=k 1Au_0=k\,1 with some k∈ℝk . Intuitively, when the controller plays a neutralizing strategy, the learner is made indifferent among its actions because all actions yield the same expected payoff. In the controllability analysis below, the existence of a fully mixed neutralizing strategy ensures that the projected control directions contain a neighborhood of the origin, which is essential for moving the learner’s state in all directions within the relative interior. Now, let H,PHH,P_H and M be defined as in (8). Our first result can be stated as follows: Theorem 1 (Controllability of the two-player system). The system (10) is controllable on Δn∘ _n if and only if 1. There exists a fully mixed neutralizing strategy u0∈Δm∘u_0∈ _m for the controller, and 2. rank(PHA)=n−1rank(P_HA)=n-1. As discussed briefly in Section V-B, our main proof technique is to transform system (10) into a state-equivalent system with a simpler form, i.e. z=PHAu,z∈PH∇h(Δn∘),u∈U.z=P_HAu, z∈ P_H∇ h( _n ), u∈ U. (15) The state-equivalence between the primal and dual systems is a consequence of Lemma 1, which we stated without proof. We now state the complete two-player version of this lemma as follows, and give a full proof of it. Lemma 2. The restriction of the choice map Φ:=Q|M:M→Δn∘ :=Q|_M:M→ _n is a diffeomorphism with inverse Φ−1(x)=PH∇h(x). ^-1(x)=P_H∇ h(x). Proof. We first confirm that Q is well defined and continuous. For fixed y, the function x↦⟨x,y⟩−h(x)x x,y -h(x) is strictly concave. Since Δn _n is compact and convex, the maximizer Q(y)Q(y) exists and is unique. Continuity of Q as a function of y follows from Berge’s maximum theorem [25]. To prove the lemma, it suffices to show that (i) Q|M:M→Δn∘Q|_M:M→ _n is bijective, and (i) Q|MQ|_M is smooth. To show bijectivity, we examine the optimization problem whose solution is Q(y)Q(y): minimizez∈Δn *minimize_z∈ _n h(z)−⟨z,y⟩ h(z)- z,y (16) s.t. s.t. ⟨,z⟩−1=0, 1,z -1=0, zi≥0. z_i≥ 0. The Lagrangian is ℒ(z,λ,μ)=h(z)−⟨z,y⟩+λ(⟨,z⟩−1)−⟨μ,z⟩,L(z,λ,μ)=h(z)- z,y +λ( 1,z -1)- μ,z , where λ∈ℝλ and μ∈ℝ≥0nμ ^n_≥ 0. By the KKT conditions, the solution to (16), x=Q(y)x=Q(y), must satisfy ∇zℒ(z,λ,μ)=∇h(x)−y+λ−μ=0,⟨1,x⟩−1=0,μi≥0,μixi=0 for i∈[n]. cases _zL(z,λ,μ)=∇ h(x)-y+ 1-μ=0,\\ 1,x -1=0,\\ _i≥ 0,\, _ix_i=0 for i∈[n].\\ cases Since x∈Δn∘x∈ _n , these conditions reduce to ∇h(x)−y+λ=0,∇ h(x)-y+ 1=0, (17) which is also sufficient for optimality. In fact, this also gives us a exact characterization of the preimage Q−1(Δn∘)Q^-1( _n ): Q−1(Δn∘)=∇h(Δn∘)+span.Q^-1( _n )=∇ h( _n )+ span\1\. (18) Further, since ⟨x,⟩ x,1 is constant over Δn _n, we have Q(y)=Q(y+k)Q(y)=Q(y+k1) for any y∈ℝny ^n and k∈ℝk . This implies that Q(∇h(Δn∘)+span)=Q(PH∇h(Δn∘))=Δn∘,Q(∇ h( _n )+ span\1\)=Q(P_H∇ h( _n ))= _n , so Q|PH∇h(Δn∘)Q|_P_H∇ h( _n ) is surjective onto Δn∘ _n . To show injectivity, take any y1,y2∈PH∇h(Δn∘)y_1,y_2∈ P_H∇ h( _n ). If Q(y1)=Q(y2)=xQ(y_1)=Q(y_2)=x, then (17) yields y1−y2=(λ1−λ2),y_1-y_2=( _1- _2)1, which can only be zero since y1,y2∈Hy_1,y_2∈ H. This establishes bijectivity. It remains to show smoothness. Define the map F(x,λ;y):=[∇h(x)−y+λ⟨1,x⟩−1].F(x,λ;y):= bmatrix∇ h(x)-y+ 1\\ 1,x -1 bmatrix. Since h is smooth, F is smooth, and F(x,λ;y)=0F(x,λ;y)=0 corresponds exactly to the KKT condition. The Jacobian of F is J=[∇2h(x)⊺0],J= bmatrix∇^2h(x)&1\\ 1 &0 bmatrix, which is invertible since ∇2h∇^2h is invertible. By the Implicit Function Theorem, for every (x0,λ0,y0)(x_0, _0,y_0) with F(x0,λ0;y0)=0F(x_0, _0;y_0)=0, there exist neighborhoods of y0y_0 and (x0,λ0)(x_0, _0) and a unique smooth mapping (x(⋅),λ(⋅))(x(·),λ(·)) such that F(x(y),λ(y);y)=0F(x(y),λ(y);y)=0 in that neighborhood. When x∈Δn∘x∈ _n , this solution coincides with the unique maximizer of (16), i.e., x(y)=Q(y)x(y)=Q(y). Since this holds for every y∈Q−1(Δn∘)y∈ Q^-1( _n ), we conclude that Q is smooth on Q−1(Δn∘)=∇h(Δn∘)+spanQ^-1( _n )=∇ h( _n )+ span\1\. Since M=PH∇h(Δn∘)⊂∇h(Δn∘)+spanM=P_H∇ h( _n )⊂∇ h( _n )+ span\1\, the claim follows. ∎ We are now ready to prove Theorem 1. Proof of Theorem 1. As discussed in Section V-B, the primal system (10) is state equivalent to the projected dual system (9); it therefore suffices to analyze controllability on M. Step 1: Necessity of 1). Assume that there is no fully mixed neutralizing strategy. Since PHAP_HA is linear and Δm∘ _m is the relative interior of Δm _m, we have PHA(Δm∘)=(PHA(Δm))∘P_HA( _m )= (P_HA( _m) ) . Therefore 0∉(PHA(Δm))∘0∉ (P_HA( _m) ) . By the separating hyperplane theorem, there exists a nonzero vector w∈Hw∈ H such that ⟨w,PHAu⟩≥0 w,P_HAu ≥ 0 for all u∈Δmu∈ _m. Let z(⋅)z(·) be any trajectory of (9). Then dt⟨w,z(t)⟩=⟨w,PHAu(t)⟩≥0 ddt w,z(t) = w,P_HAu(t) ≥ 0, so the scalar function t↦⟨w,z(t)⟩t w,z(t) is non-decreasing along every trajectory. Fix any z0∈Mz_0∈ M. Since w∈H=Tz0Mw∈ H=T_z_0M, there exists a smooth curve γ:(−ε,ε)→Mγ:(- , )→ M such that γ(0)=z0γ(0)=z_0 and γ˙(0)=−w γ(0)=-w. Consequently, ⟨w,γ(s)⟩<⟨w,z0⟩ w,γ(s) < w,z_0 for all sufficiently small s>0s>0. Such points cannot be attained from z0z_0, which shows that (9) is not locally controllable, and hence not globally controllable. Thus 1) is necessary. Assume now that 1) holds. Choose u0∈Δm∘u_0∈ _m such that PHAu0=0P_HAu_0=0, and define V:=Δm−u0V:= _m-u_0. Since u0∈Δm∘u_0∈ _m , the set V is a proper control set in the affine subspace aff(V)=v∈ℝm:⟨v,1m⟩=0aff(V)=\v ^m: v,1_m =0\, and 0∈V∘0∈ V . Writing u=u0+vu=u_0+v, the auxiliary system becomes z˙=PHAv,z∈M,v∈V. z=P_HAv, z∈ M, v∈ V. (19) Step 2: Necessity of 2). Let z(⋅)z(·) be any trajectory of (19) with initial condition z(0)=z0z(0)=z_0. Then for every T≥0T≥ 0, z(T)−z0=∫0TPHAv(t)t∈im(PHA).z(T)-z_0= _0^TP_HAv(t)\,dt (P_HA). Hence the attainable set from z0z_0 satisfies Az(z0)⊂(z0+im(PHA))∩MA_z(z_0)⊂ (z_0+im(P_HA) )∩ M. If rank(PHA)<n−1rank(P_HA)<n-1, then z0+im(PHA)z_0+im(P_HA) is a proper affine subspace of H. Since M is an (n−1)(n-1)-dimensional manifold, it cannot be contained in a proper affine subspace of H. Therefore Az(z0)≠MA_z(z_0)≠ M, so (19) is not controllable. Thus 2) is necessary. Step 3: Sufficiency of 2). Write the columns of PHAP_HA as PHA=[b1⋯bm],bj∈HP_HA=[b_1\ ·s\ b_m], b_j∈ H. Then (19) is the driftless control-affine system z˙=∑j=1mvjbj z= _j=1^mv_jb_j, with v∈Vv∈ V, whose control vector fields are constant on M. Therefore all Lie brackets vanish, and for every z∈Mz∈ M, Liezb1,…,bm=spanb1,…,bm=im(PHA).Lie_z\b_1,…,b_m\=span\b_1,…,b_m\=im(P_HA). If rank(PHA)=n−1rank(P_HA)=n-1, then Liezb1,…,bm=H=TzM,∀z∈M.Lie_z\b_1,…,b_m\=H=T_zM, ∀ z∈ M. Thus the family b1,…,bm\b_1,…,b_m\ is bracket generating on M. Since V is a proper control set, Proposition 1 implies that (19) is small-time locally controllable everywhere on M. By Proposition 2, it follows that (19) is globally controllable on the connected manifold M. Finally, Proposition 4 transfers controllability back to the original system, so the primal system is controllable on Δn∘ _n . This completes the proof. ∎ VI-B The N-player case Moving beyond a single opponent, we now consider the case where the controller interacts with multiple independent learners at the same time. Steering a joint system is naturally more demanding, as any change in the controller’s strategy simultaneously affects the entire population. Because finding a tight necessary and sufficient condition is difficult in this high-dimensional setting, we instead focus on providing reliable sufficient conditions. We outline two distinct approaches. The first generalizes our earlier intuition by asking if the controller can uniformly neutralize the drift for all learners at once. Since this can be a restrictive requirement in larger games, we also provide a complementary condition that relies on the periodicity of the system’s drift. We will start by laying out the first sufficient condition, which is based on the concept of a uniformly neutralizing strategy: Definition 5 (Uniformly Neutralizing Strategy). A mixed strategy u0∈Δmu_0∈ _m is uniformly neutralizing if for every fully mixed strategy profile x, Ai(x)u0=ki 1A_i(x)u_0=k_i\,1 with some ki∈ℝk_i for every i∈[N]i∈[N]. This extends the two-player notion of a neutralizing strategy: a uniformly neutralizing strategy simultaneously neutralizes each learner’s payoff vector across all system states. As a result, it removes the drift term in the projected auxiliary system, reducing the controllability analysis to a driftless control-affine problem. Let Hi,H,PHH_i,H,P_H and ∘X be defined as in (4), let Q and h be defined as in (5), and let M be defined as in (6). Define the vector fields η1,…,ηm−1:∘→∘ _1,…, _m-1:X by ηi(x)=DQ(∇h(x))(ai(x)−am(x)). _i(x)=DQ(∇ h(x)) (a_i(x)-a_m(x) ). (20) Remark 3. Since Φ−1(x)=PH∇h(x) ^-1(x)=P_H∇ h(x) and Q is invariant under blockwise translations along H⟂H , we have DQ(Φ−1(x))=DQ(PH∇h(x))=DQ(∇h(x)).DQ( ^-1(x))=DQ(P_H∇ h(x))=DQ(∇ h(x)). Accordingly, in theorem statements we use ηi(x):=DQ(∇h(x))(ai(x)−am(x)),i=1,…,m−1. _i(x):=DQ(∇ h(x))(a_i(x)-a_m(x)),\ i=1,…,m-1. For a family G of smooth vector fields on X∘X , define Liex():=Y(x):Y∈Lie()⊂TxX∘.Lie_x(G):=\Y(x):Y (G)\⊂ T_xX . We say that Lie(G)Lie(G) has full rank on ∘X if Liex()=Tx∘Lie_x(G)=T_xX for all x∈∘x , or, equivalently, dimLiex()=dim(∘) _x(G)= (X ) for all x∈∘x . Our sufficient condition for controllability of the N-learner system is as follows: Theorem 2 (Sufficient condition for controllability of the N-player system with a uniformly neutralizing strategy). The system (14) is controllable if 1. There exists a fully mixed uniformly neutralizing strategy u0∈Δm∘u_0∈ _m for the controller, and 2. Lieη1,…,ηm−1Lie\ _1,…, _m-1\ has full rank on X∘X . As in the two-player case, we prove this by analyzing the dual system granted by Lemma 1. We have proven the two-player version of Lemma 1 in the previous section, now we state and prove the second part, i.e. the multi-player version. Lemma 3. The restriction of the choice map Φ:=Q|M=Φ1⊕⋯⊕ΦN:M→∘ :=Q|_M= _1 ·s _N:M is a diffeomorphism with inverse Φ−1(x)=PH∇h(x). ^-1(x)=P_H∇ h(x). Proof. Apply the single-learner argument to each block. The KKT conditions show that Qi−1(Δni∘)=∇hi(Δni∘)+spanniQ_i^-1( _n_i )=∇ h_i( _n_i )+span\1_n_i\, and restriction to HiH_i removes the nonuniqueness along the all-ones direction. Smoothness follows from the implicit function theorem exactly as in the one-learner case. ∎ Now we are ready to prove Theorem 2. Proof of Theorem 2. Let d:=∑i=1N(ni−1)=∑i=1Nni−Nd:= _i=1^N(n_i-1)= _i=1^Nn_i-N. As in the two-player case, the primal system (14) and the dual system (13) are state-equivalent due to Lemma 3. We will focus on the analysis of the dual system for simplicity. Recall that the dual system is given by Σz:z˙=PHA(Φ(z))u,z∈M,u∈Δm, _z: z=P_HA( (z))u, z∈ M,\ u∈ _m, and the primal dynamics is given by Σx:x˙=DQ(Φ−1(x))A(x)u,x∈∘,u∈Δm. _x: x=DQ( ^-1(x))A(x)u, x ,\ u∈ _m. Moreover, Σz _z and Σx _x are state equivalent via the diffeomorphism Φ . Indeed, for x=Φ(z)x= (z), DΦ(z)(PHA(Φ(z))u)=DQ(z)PHA(x)u=DQ(z)A(x)u,D (z) (P_HA( (z))u )=DQ(z)P_HA(x)u=DQ(z)A(x)u, where we again used DQ(z)PH=DQ(z)DQ(z)P_H=DQ(z). Therefore, by Proposition 4, controllability of Σx _x on ∘X is equivalent to controllability of Σz _z on M. Step 1: Neutralizing the drift and introducing effective controls. By Condition 1, there exists u0∈Δm∘u_0∈ _m such that for every x∈∘x and every learner i, Ai(x)u0=ki(x)niA_i(x)u_0=k_i(x)1_n_i for some scalar ki(x)∈ℝk_i(x) . Applying PHiP_H_i to each block gives PHiAi(x)u0=0P_H_iA_i(x)u_0=0, and hence, after stacking the blocks, PHA(x)u0=0P_HA(x)u_0=0 for all x∈∘x . Introduce the matrix E:=[e1−em⋯em−1−em]∈ℝm×(m−1).E:= bmatrixe_1-e_m&·s&e_m-1-e_m bmatrix ^m×(m-1). Since im(E)=v∈ℝm:⟨v,m⟩=0im(E)=\v ^m: v,1_m =0\, every u∈Δmu∈ _m can be written uniquely as u=u0+Ewu=u_0+Ew for some w∈ℝm−1w ^m-1. Define W:=w∈ℝm−1:u0+Ew∈ΔmW:=\w ^m-1:u_0+Ew∈ _m\. Because u0∈Δm∘u_0∈ _m , the translated simplex Δm−u0 _m-u_0 is a proper control set in v:⟨v,m⟩=0\v: v,1_m =0\, and since E is a linear isomorphism from ℝm−1R^m-1 onto this hyperplane, W is a proper control set in ℝm−1R^m-1. Substituting u=u0+Ewu=u_0+Ew into Σz _z yields z˙ z =PHA(Φ(z))(u0+Ew) =P_HA( (z))(u_0+Ew) =PHA(Φ(z))Ew =P_HA( (z))Ew =∑k=1m−1wkbk(z),w∈W, = _k=1^m-1w_k\,b_k(z), w∈ W, where bk(z):=PH(ak(Φ(z))−am(Φ(z))),k=1,…,m−1.b_k(z):=P_H (a_k( (z))-a_m( (z)) ), k=1,…,m-1. Thus the projected dual system is a driftless control-affine system on M with control vector fields b1,…,bm−1b_1,…,b_m-1. Step 2: Push-forward of the control vector fields. For x=Φ(z)x= (z), (Φ∗bk)(x)=DΦ(Φ−1(x))bk(Φ−1(x)).( _*b_k)(x)=D ( ^-1(x))\,b_k( ^-1(x)). Since bk(z)=PH(ak(Φ(z))−am(Φ(z)))b_k(z)=P_H (a_k( (z))-a_m( (z)) ), we obtain (Φ∗bk)(x)=DΦ(Φ−1(x))PH(ak(x)−am(x))( _*b_k)(x)=D ( ^-1(x))\,P_H (a_k(x)-a_m(x) ). Using Remark 3 and the identity DQ(ζ)PH=DQ(ζ)DQ(ζ)P_H=DQ(ζ), this becomes (Φ∗bk)(x)=DQ(∇h(x))(ak(x)−am(x))=ηk(x).( _*b_k)(x)=DQ(∇ h(x)) (a_k(x)-a_m(x) )= _k(x). Therefore, Φ∗(Lieb1,…,bm−1)=Lieη1,…,ηm−1. _* (Lie\b_1,…,b_m-1\ )=Lie\ _1,…, _m-1\. Hence, for every z∈Mz∈ M and x=Φ(z)x= (z), DΦ(z)Liezb1,…,bm−1=Liexη1,…,ηm−1.D (z)\,Lie_z\b_1,…,b_m-1\=Lie_x\ _1,…, _m-1\. Step 3: Wrapping up the proof. By Condition 2), Lieη1,…,ηm−1Lie\ _1,…, _m-1\ has full rank on ∘X , i.e., Liexη1,…,ηm−1=Tx∘Lie_x\ _1,…, _m-1\=T_xX for all x∈∘x . Since DΦ(z):TzM→TxX∘D (z):T_zM→ T_xX is a linear isomorphism, it follows that Liezb1,…,bm−1=TzMLie_z\b_1,…,b_m-1\=T_zM for all z∈Mz∈ M. Thus the family b1,…,bm−1\b_1,…,b_m-1\ is bracket generating on M. Since W is a proper control set, Proposition 1 implies that the driftless dual system is small-time locally controllable everywhere on M. The manifold M is connected because it is diffeomorphic to the connected manifold ∘X . Hence, by Proposition 2, the dual system is globally controllable on M. Finally, Proposition 4 transfers controllability back through the diffeomorphism Φ , so the primal system (14) is controllable on ∘X . This completes the proof. ∎ For a concrete example, please see Section VII-B1. Intuitively, the first condition of Theorem 2 ensures that the auxiliary system has no drift term. However, controllability is sometimes possible even with drift. We provide below another sufficient condition that captures one such case, which is due to [21, Chapter 4, Theorem 5], which we present here as Proposition 5. This condition relies on periodicity. Definition 6 (Periodicity of Vector Fields). Let f be a vector field on M. We say f is periodic if t↦etf(x)t e^tf(x) is periodic for every x∈Mx∈ M. Proposition 5. [21, Chapter 4, Theorem 5] Suppose that X0,…,XmX_0,…,X_m are vector fields on M that define a system affine in control. Assume that (a) the drift X0X_0 is periodic, and (b) the origin of ℝmR^m lies in the interior of the convex hull of U. Then the corresponding control system is controllable provided that LiexX0,…,Xm=TxMLie_x\X_0,…,X_m\=T_xM for each x∈Mx∈ M. Further, let u¯:=1mm u:= 1m1_m. Under the change of variables u=u¯+Ewu= u+Ew, the induced primal system (14) can be written as x˙=η0(x)+∑k=1m−1wkηk(x),w∈W, x= _0(x)+ _k=1^m-1w_k _k(x), w∈ W, where η0(x):=1mDQ(∇h(x))A(x)m. _0(x):= 1mDQ(∇ h(x))A(x)1_m. (21) Then our second sufficient condition can be stated as: Theorem 3 (Sufficient condition for controllability of the N-player system with a periodic drift). The system (14) is controllable if 1. the vector field η0 _0 defined in (21) is periodic on ∘X , and 2. Lieη1,…,ηm−1Lie\ _1,…, _m-1\ has full rank on X∘X . where η1,…,ηm−1 _1,…, _m-1 are as defined in (20). Proof of Theorem 3. Under the affine change of variables u=u¯+Ewu= u+Ew, the induced primal system (14) becomes x˙=η0(x)+∑k=1m−1wkηk(x) x= _0(x)+ _k=1^m-1w_k _k(x), with w∈Ww∈ W. Since u¯=1mm∈Δm∘ u= 1m1_m∈ _m and E:ℝm−1→v∈ℝm:⟨v,m⟩=0E:R^m-1→\v ^m: v,1_m =0\ is a linear isomorphism, the set W:=w∈ℝm−1:u¯+Ew∈ΔmW:=\w ^m-1: u+Ew∈ _m\ is convex, compact, and contains 0 in its interior. By Condition 2), the family η1,…,ηm−1\ _1,…, _m-1\ has full rank on ∘X . Hence the larger family η0,η1,…,ηm−1\ _0, _1,…, _m-1\ is bracket generating on ∘X . By Condition 1), every trajectory of the drift system x˙=η0(x) x= _0(x) is periodic; in particular, η0 _0 is Poisson stable. Therefore, we can invoke [21, Chapter 4, Theorem 5], a classical controllability theorem for control-affine systems with Poisson-stable drift, bracket-generating vector fields, and control set whose convex hull is a neighborhood of the origin, and conclude global controllability on ∘X . Since the change of variables u=u¯+Ewu= u+Ew is bijective, the original induced primal system (14) is controllable on ∘X . ∎ Remark 4. A sharper treatment for constrained controls with 0 possibly on the boundary of the control set is available in [9], but we do not use it here. For a concrete example, please see Section VII-B2. VII Examples VII-A Examples for The two-player case VII-A1 Rock-Paper-Scissor Game A canonical steerable two player finite game is the Rock-Paper-Scissor (RPS) game, which is a zero sum game with the following payoff matrix: A=[ϵ−111ϵ−1−11ϵ]A= bmatrixε&-1&1\\ 1&ε&-1\\ -1&1&ε bmatrix with some ϵ∈(−1,1)ε∈(-1,1). An obvious choice for a neutralizing strategy is u0=(1/3,1/3,1/3)u_0=(1/3,1/3,1/3). Further, a straightforward computation yields PHA=[2ϵ3−1−ϵ31−ϵ31−ϵ32ϵ3−1−ϵ3−1−ϵ31−ϵ32ϵ3],P_HA= bmatrix 2ε3&-1- ε3&1- ε3\\ 1- ε3& 2ε3&-1- ε3\\ -1- ε3&1- ε3& 2ε3\\ bmatrix, which has rank 22 for −1≤ϵ≤1-1≤ε≤ 1. By Theorem 1, this game is steerable. VII-A2 Non-Steerable Case: Modified RPS We can break steerability with a small modification of the payoff matrix. Consider A=[0−1110−1−113].A= bmatrix0&-1&1\\ 1&0&-1\\ -1&1&3 bmatrix. Despite PHA=[0−1010−2−112]P_HA= bmatrix0&-1&0\\ 1&0&-2\\ -1&1&2 bmatrix which has rank rank(PHA)=2rank(P_HA)=2, the only neutralizing strategy u0=(2/3,0,1/3)u_0=(2/3,0,1/3) lies on the boundary of Δ3 _3. Therefore, the game is not steerable. Figure 1 illustrates the attainable set of this system from three random initial points, which does not cover Δ3∘ _3 . Figure 1: Approximate attainable sets for the Modified RPS game: We plot the approximate attainable set for three random initial points in Δ3∘ _3 , for two different variants of FTRL. The figures on the left shows the case where the learner adopts the Replicator Dynamics, while the figures on the right shows that of the learner who adopts FTRL with h(x)=1/2‖x‖2h(x)=1/2\|x\|^2. The reachable set is approximated by plotting the union of states generated by constant controls u∈Δ3u∈ _3 sampled on the simplex lattice u=(i,j,k)/50u=(i,j,k)/50, i+j+k=50i+j+k=50, and time horizons t∈[0,12]t∈[0,12] sampled on a uniform grid of 4545 points. VII-B Examples for The multi-player case VII-B1 Brockett’s Integrator Game Consider the game with 3 learners and a controller. Each learner has 2 pure strategies, and the controller has 3 pure strategies. Learners 1 and 2 interact exclusively with the controller, receiving payoff depending on the first two and last two strategies of the controller, respectively. Learner 3 receives payoff depending on all three strategies of the controller, but the exact payoff it receives depends on the strategies of the other two learners. Figure 2 illustrates the structure of this game. Controlleru∈Δ3u∈ ^3Learner 1x1∈Δ2x_1∈ ^2p1(u(1),u(2))p_1(u^(1),u^(2))Learner 3x3∈Δ2x_3∈ ^2p3(x1(1),x2(1),u)p_3(x_1^(1),x_2^(1),u)Learner 2x2∈Δ2x_2∈ ^2p2(u(2),u(3))p_2(u^(2),u^(3))u(1),u(2)u^(1),u^(2)u(1),u(2),u(3)u^(1),u^(2),u^(3)u(2),u(3) u^(2),u^(3)x1(1)→r2x_1^(1)→ r_2x2(1)→r2x_2^(1)→ r_2 Figure 2: Interdependency graph for the Brockett’s Integrator game: This game shows the interdependency of players in the Brockett’s Integrator game presented in Section VII-B1. As shown in the graph, learner 1’s payoff only depend on the probability of the controller playing his first two strategies, and learner 2’s payoff only s on the probability of the controller playing his last two strategies. Learner three’s payoff depends on the probability of learner 1 and 2 playing their first strategies, as well as the entire mixed strategy of the controller. We can model this game in the form of (11) with A1(x2,x3)=[1−10−110],A2(x1,x3)=[01−10−11],A3(x1,x2)=[−x2(1)x1(1)+x2(1)−x1(1)x2(1)−(x1(1)+x2(1))x1(1)], casesA_1(x_2,x_3)= bmatrix1&-1&0\\ -1&1&0 bmatrix,\\ A_2(x_1,x_3)= bmatrix0&1&-1\\ 0&-1&1 bmatrix,\\ A_3(x_1,x_2)= bmatrix-x_2^(1)&x_1^(1)+x_2^(1)&-x_1^(1)\\ x_2^(1)&-(x_1^(1)+x_2^(1))&x_1^(1) bmatrix, cases and x=Q(y),y˙=A(x)u,x∈Δ2×Δ2×Δ2,u∈Δ3x=Q(y), y=A(x)u, x∈ _2× _2× _2,\,u∈ _3 (22) with A(x)=[A1(x1,x2)A2(x1,x3)A3(x2,x2)].A(x)= bmatrixA_1(x_1,x_2)\\ A_2(x_1,x_3)\\ A_3(x_2,x_2) bmatrix. We now verify the two conditions of Theorem 2 using the vector fields η1(x) _1(x) =DQ(∇h(x))(a1(x)−a3(x)), =DQ(∇ h(x))(a_1(x)-a_3(x)), η2(x) _2(x) =DQ(∇h(x))(a2(x)−a3(x)). =DQ(∇ h(x))(a_2(x)-a_3(x)). The first condition is easily satisfied, since u0=(1/3,1/3,1/3)∈Δ3∘u_0=(1/3,1/3,1/3)∈ _3 is uniformly neutralizing. To verify Condition 2, write xi=(xi(1),xi(2)),ξi:=12(xi(1)−xi(2)),i=1,2,3.x_i= (x_i^(1),x_i^(2) ), _i:= 12 (x_i^(1)-x_i^(2) ), i=1,2,3. Since each learner has two pure strategies, each tangent space Hi=(r,−r):r∈ℝH_i=\(r,-r):r \ is one-dimensional. Hence, for each block there exists a smooth positive scalar function λi(ξi)>0 _i( _i)>0 such that PHiDQi(∇hi(xi))[1−1]=λi(ξi)[1−1].P_H_iDQ_i(∇ h_i(x_i)) bmatrix1\\ -1 bmatrix= _i( _i) bmatrix1\\ -1 bmatrix. From the matrices in the example, the three columns of A(x)A(x) are a1(x) a_1(x) =[1−100−x2(1)x2(1)]⊺, = bmatrix1&-1&0&0&-x_2^(1)&x_2^(1) bmatrix , a2(x) a_2(x) =[−111−1x1(1)+x2(1)−(x1(1)+x2(1))]⊺, = bmatrix-1&1&1&-1&x_1^(1)+x_2^(1)&-(x_1^(1)+x_2^(1)) bmatrix , a3(x) a_3(x) =[00−11−x1(1)x1(1)]⊺. = bmatrix0&0&-1&1&-x_1^(1)&x_1^(1) bmatrix . Therefore, in the local coordinates ξ=(ξ1,ξ2,ξ3)ξ=( _1, _2, _3), the push-forward vector fields become η1(ξ)=[λ1(ξ1)λ2(ξ2)(ξ1−ξ2)λ3(ξ3)],η2(ξ)=[−λ1(ξ1)2λ2(ξ2)(2ξ1+ξ2+32)λ3(ξ3)]. _1(ξ)= bmatrix _1( _1)\\[5.69054pt] _2( _2)\\[5.69054pt] ( _1- _2) _3( _3) bmatrix, _2(ξ)= bmatrix- _1( _1)\\[5.69054pt] 2 _2( _2)\\[5.69054pt] (2 _1+ _2+ 32 ) _3( _3) bmatrix. A direct computation gives [η1,η2](ξ)=[003(λ1(ξ1)+λ2(ξ2))λ3(ξ3)].[ _1, _2](ξ)= bmatrix0\\[5.69054pt] 0\\[5.69054pt] 3 ( _1( _1)+ _2( _2) ) _3( _3) bmatrix. Since λi(ξi)>0 _i( _i)>0 for all i, the vectors η1(ξ) _1(ξ), η2(ξ) _2(ξ), and [η1,η2](ξ)[ _1, _2](ξ) are linearly independent for every ξ∈(−1/2,1/2)3ξ∈(-1/2,1/2)^3. Hence rank(Lieη1,η2)=3=∑i=13ni−3rank (Lie\ _1, _2\ )=3= _i=1^3n_i-3. By Theorem 2, induced primal system and thus (22) is steerable on Δ2∘×Δ2∘×Δ2∘ _2 × _2 × _2 . Interestingly, when all learners adopt FTRL with h(x)=1/2‖x‖2h(x)=1/2\|x\|^2, we can adopt the following change of variables: ξ1=12(x1(1)−x1(2)),ξ2=12(x2(1)−x2(2)),ξ3=12(x3(1)−x3(2)),w1=u1−u2,w2=u2−u3. cases _1= 12(x_1^(1)-x_1^(2)),\\ _2= 12(x_2^(1)-x_2^(2)),\\ _3= 12(x_3^(1)-x_3^(2)),\\ w_1=u_1-u_2,\\ w_2=u_2-u_3. cases The resulting system is exactly the Brockett’s Integrator: ξ˙=[1001−ξ2ξ1]w,w∈W, ξ= bmatrix1&0\\ 0&1\\ - _2& _1 bmatrixw, w∈ W, where the control set W is the convex polygon W=(v1,v2)∈ℝ2:1+2v1+v2≥0,1−v1+v2≥0,1−v1−2v2≥0.W= \(v_1,v_2) ^2:\; array[]l1+2v_1+v_2≥ 0,\\[1.0pt] 1-v_1+v_2≥ 0,\\[1.0pt] 1-v_1-2v_2≥ 0 array \. VII-B2 Regulated Matching Pennies Consider a three player game with two learners and one controller, in which the two learners each have two pure actions and the controller has three. For any fixed controller action, the two learners are involved in a two player zero-sum game, with a payoff matrix dependent on the controller’s action. Specifically, the three actions of the controller correspond to the following payoff matrices: B1=[1100],B2=[2−5−32],B3=[0101].B_1= bmatrix1&1\\ 0&0 bmatrix,B_2= bmatrix2&-5\\ -3&2 bmatrix,B_3= bmatrix0&1\\ 0&1 bmatrix. Figure 3 illustrates the structure of this game. Denoting x1=(α,1−α)x_1=(α,1-α) and x2=(β,1−β)x_2=(β,1-β), we can model this game in the form of (11) with A1(x2) A_1(x_2) =[B1x2,B2x2,B3x2], = bmatrixB_1x_2,B_2x_2,B_3x_2 bmatrix, A2(x1) A_2(x_1) =[−B1x1,−B2x1,−B3x1] = bmatrix-B_1x_1,-B_2x_1,-B_3x_1 bmatrix Suppose that the learning agents adopt the replicator dynamics. For u¯=(1/3,1/3,1/3) u=(1/3,1/3,1/3), the drift field is η0(x)=13DQ(∇h(x))A(x)3. _0(x)= 13DQ(∇ h(x))A(x)1_3. Under replicator dynamics and the reduced coordinates (α,β)(α,β), this becomes η0(α,β)=[2α(1−α)(2β−1)2β(1−β)(1−2α)], _0(α,β)= bmatrix2α(1-α)(2β-1)\\ 2β(1-β)(1-2α) bmatrix, which is the classical matching-pennies replicator field and is periodic on (0,1)2(0,1)^2. This satisfies condition 1) of Theorem 3. Controlleru∈Δ3u∈ ^3Matrix selectionBk,k∈1,2,3B^k,\ k∈\1,2,3\Learner 1x1∈Δ2x_1∈ ^2 p1=Bkx2p_1=B_kx_2Learner 2x2∈Δ2x_2∈ ^2 p2=−Bk⊺x1p_2=-B_k x_1uuBkB^kBkB^kx2x_2x1x_1 Figure 3: Interdependency graph for the Regulated Matching Pennies game: This figure shows the interdependency among players in the Regulated Matching Pennies game presented in Section VII-B2. The two learners are involved in a two-player zero-sum game with a payoff matrix chosen by the controller. Again, we use u∈Δ3u∈ _3 to represents the controller’s mixed strategy, xi∈Δ2x_i∈ _2 represents the mixed strategy of learner i, and pip_i represents the payoff vector of learner i. We now verify Condition 2). Taking again η1(x) _1(x) =DQ(∇h(x))(a1(x)−a3(x)), =DQ(∇ h(x))(a_1(x)-a_3(x)), η2(x) _2(x) =DQ(∇h(x))(a2(x)−a3(x)), =DQ(∇ h(x))(a_2(x)-a_3(x)), where a1,a2,a3a_1,a_2,a_3 are the columns of A(x)A(x). A direct computation gives a1(x)=[10−α−α],a2(x)=[7β−52−5β3−5α7α−2],a3(x)=[1−β1−β0−1].a_1(x)= bmatrix1\\ 0\\ -α\\ -α bmatrix,a_2(x)= bmatrix7β-5\\ 2-5β\\ 3-5α\\ 7α-2 bmatrix,a_3(x)= bmatrix1-β\\ 1-β\\ 0\\ -1 bmatrix. For replicator dynamics, in the reduced coordinates (α,β)(α,β) these become η1(α,β)=[α(1−α)−β(1−β)],η2(α,β)=[(12β−7)α(1−α)(4−12α)β(1−β)]. _1(α,β)= bmatrixα(1-α)\\[2.84526pt] -β(1-β) bmatrix, _2(α,β)= bmatrix(12β-7)α(1-α)\\[2.84526pt] (4-12α)β(1-β) bmatrix. Their Lie bracket is [η1,η2](α,β)=−12αβ(1−α)(1−β)[11].[ _1, _2](α,β)=-12αβ(1-α)(1-β) bmatrix1\\ 1 bmatrix. Therefore, det[η1(α,β),[η1,η2](α,β)] [ _1(α,β),[ _1, _2](α,β) ] = = −12αβ(1−α)(1−β)(α(1−α)+β(1−β)). -2αβ(1-α)(1-β) (α(1-α)+β(1-β) ). Since (α,β)∈(0,1)2(α,β)∈(0,1)^2, the right-hand side is never zero. Thus rank(Lieη1,η2)=2=(2−1)+(2−1)rank (Lie\ _1, _2\ )=2=(2-1)+(2-1) for every (α,β)∈Δ2∘×Δ2∘(α,β)∈ _2 × _2 . Hence Theorem 3 applies, and the induced primal system (14) is controllable on Δ2∘×Δ2∘ _2 × _2 . VIII Conclusion and Further Discussions In this paper, we show that steering continuous-time FTRL learners in finite games can be understood as a control problem on the relative interior of a simplex, or a product of simplices. From this perspective, the issue is not only how learners adapt, but when a model-aware player can guide that adaptation to a desired interior strategy configuration through standard strategic play alone, without changing the game’s payoff structure. In the two-player case, we characterize this exactly through the existence of a fully mixed neutralizing strategy and a rank condition on the projected payoff map. In settings with multiple learners, we identify two sufficient routes to controllability: one based on uniform neutralization, and one based on periodic drift together with a Lie-algebra rank condition. These results followed from the insight that steering FTRL dynamics has the structure of a geometric control problem, allowing us to bring ideas of controllability into the analysis of vulnerabilities in strategic learning and identify how the geometry of the game determines the influence an agent can have on the overall learning behavior. References [1] K. Abe, M. Sakamoto, and A. Iwasaki Mutation-driven follow the regularized leader for last-iterate convergence in zero-sum games. In Uncertainty in Artificial Intelligence, Cited by: §I-A. [2] A. A. Agrachev and Yu. L. Sachkov (2004) Control theory from the geometric viewpoint. Springer Berlin, Heidelberg. Cited by: §I-D. [3] S. Arora, E. Hazan, and S. Kale (2012) The multiplicative weights update method: a meta-algorithm and applications. Theory of Computing. Cited by: §I. [4] A. Assos, Y. Dagan, and N. Rajaraman Computational intractability of strategizing against online learners. In Conference on Learning Theory, Cited by: §I-B. [5] U. Boscain, D. Cannarsa, V. Franceschi, and M. Sigalotti Local controllability does imply global controllability. Comptes Rendus. Mathématique. Cited by: §I-D, Proposition 2. [6] M. Braverman, J. Mao, J. Schneider, and M. Weinberg Selling to a no-regret buyer. Cited by: §I, §I. [7] W. Brown, C. Papadimitriou, and T. Roughgarden Online Stackelberg optimization via nonlinear control. Cited by: §I. [8] Y. Cai, H. Luo, C. Wei, and W. Zheng From average-iterate to last-iterate convergence in games: a reduction and its applications. In Neural Information Processing Systems (NeurIPS), Cited by: §I, §I-A. [9] J. Caillau, L. Dell’Elce, A. Herasimenka, and J. Pomet On the controllability of nonlinear systems with a periodic drift. SIAM Journal on Control and Optimization. Cited by: Remark 4. [10] I. Canyakmaz, I. Sakos, W. Lin, A. Varvitsiotis, and G. Piliouras Learning and steering game dynamics towards desirable outcomes. In Learning for Dynamics & Control Conference, Cited by: §I-C. [11] N. Cesa-Bianchi Cited by: §I. [12] Y. K. Cheung and G. Piliouras (2021) Online optimization in games via control theory: connecting regret, passivity and poincaré recurrence. In International Conference on Machine Learning, p. 1855–1865. Cited by: §I. [13] W. Chow (1939) Über die multiplizität der schnittpunkte von hyperflächen. Mathematische Annalen. Cited by: §I-D. [14] Y. Deng, J. Schneider, and B. Sivan Strategizing against no-regret learners. In Neural Information Processing Systems (NeurIPS), Cited by: §I, §I-B. [15] (1995) Evolutionary game theory. MIT Press. Cited by: §I. [16] D. P. Foster and R. V. Vohra (1997) Calibrated learning and correlated equilibrium. Games and Economic Behavior. Cited by: §I. [17] D. Fudenberg and D. K. Levine Learning in games and the interpretation of natural experiments. American Economic Journal: Microeconomics. Cited by: §I. [18] U. Halder, V. Raju, M. Mischiati, B. Dey, and P. Krishnaprasad Flocks, games, and cognition: a geometric approach. Systems & Control Letters. Cited by: §I-D. [19] S. Hart and A. Mas-Colell A simple adaptive procedure leading to correlated equilibrium. Econometrica. Cited by: §I. [20] A. Isidori (1985) Nonlinear control systems: an introduction. Springer. Cited by: §I-D. [21] Velimir. Jurdjevic (1996) Geometric control theory. Cambridge University Press. Cited by: §VI-B, §VI-B, Proposition 5. [22] N. T. Lauffer, M. Ghasemi, A. Hashemi, Y. Savas, and U. Topcu (2022) No-regret learning in dynamic Stackelberg games. IEEE Transactions on Automatic Control. Cited by: §I. [23] D. Manheim Multiparty dynamics and failure modes for machine learning and artificial intelligence. Big Data and Cognitive Computing. Cited by: §I-C. [24] Y. Mansour, M. Mohri, J. Schneider, and B. Sivan Strategizing against learners in Bayesian games. In Conference on Learning Theory, Cited by: §I, §I-B. [25] E.A. Ok (2007) Real Analysis with Economic Applications. Princeton University Press. Cited by: §VI-A. [26] C. H. Papadimitriou and G. Piliouras (2019) Game dynamics as the meaning of a game. SIGecom Exch. 16, p. 53–63. Cited by: §I. [27] V. Raju and P. Krishnaprasad (2020) Lie algebra structure of fitness and replicator control. arXiv:2005.09792. Cited by: §I-D. [28] P.K. Rashevskii (1939) About connecting two points of complete non-holonomic space by admissible curve (in russian). Uch. Zapiski Ped. Inst. Libknexta (2). Cited by: §I-D. [29] S. A. Toonsi and J. S. Shamma (2025) Higher-order uncoupled learning dynamics and nash equilibrium. arXiv:2506.10874. Cited by: §I. [30] S. Toonsi and J. Shamma (2023) Higher-order uncoupled dynamics do not lead to Nash equilibrium–except when they do. In Neural Information Processing Systems (NeurIPS), Cited by: §I, §I-A. [31] T. Tsuchiya and S. Ito In Neural Information Processing Systems (NeurIPS), Cited by: §I-A. [32] B. H. Zhang, G. Farina, I. Anagnostides, F. Cacciamani, S. McAleer, A. Haupt, A. Celli, N. Gatti, V. Conitzer, and T. Sandholm (2024) Steering no-regret learners to a desired equilibrium. Cited by: §I, §I-C. Heling Zhang (Graduate Student Member, IEEE) received his B.S. and M.S. degrees in Electrical and Computer Engineering from the University of Illinois at Urbana-Champaign, Urbana, IL, USA. He is currently pursuing a Ph.D. degree in Electrical and Computer Engineering at the same institution. His research interests include control theory and machine learning. Siqi Du (Graduate Student Member, IEEE) received the B.S. degree in Management Science from Sichuan University and the M.S. degree in Industrial Engineering from the University of Illinois at Urbana-Champaign. She is currently pursuing the Ph.D. degree in the Department of Industrial and Enterprise Systems Engineering at UIUC. Her research interests include operations research and control systems. Roy Dong (Member, IEEE) is an Assistant Professor in the Industrial & Enterprise Systems Engineering department at the University of Illinois at Urbana-Champaign. He received a BS Honors in Computer Engineering and a BS Honors in Economics from Michigan State University in 2010. He received a PhD in Electrical Engineering and Computer Sciences at the University of California, Berkeley in 2017, where he was funded in part by the NSF Graduate Research Fellowship. Prior to his current position, he was a postdoctoral researcher in the Berkeley Energy & Climate Institute, a visiting lecturer in the Industrial Engineering and Operations Research department at UC Berkeley, and a Research Assistant Professor in the Electrical and Computer Engineering department at the University of Illinois at Urbana-Champaign. His research uses tools from control theory, economics, statistics, and optimization to understand the closed-loop effects of machine learning, with applications in cyber-physical systems such as the smart grid, modern transportation networks, and autonomous vehicles.